Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
89
datasets available to search
ShareScore release 0.7.1
Dataset results
89 results for “Urdu”
Offensive content dataset in Urdu language
<p>The archive contains python code and various feature files of offensive language dataset in urdu. The purpose of sharing this archive is to regenerate the results produced by the research article and can extend the findings.</p>
Urdu Handwritten Text Dataset
<p>The dataset contains the images of handwritten text in Urdu language, one of the most widely spoken languages in South-East Asian regions. The native-speaking authors from different social domains were invited to write a pre-written text in their handwritings. The pre-written text is carefully written in a way that it includes almost all the characters, ligatures, diacritics, and dots used in writing the text Urdu script. The disabled persons are also involved to write the text to make the data collection more comprehensive. The demographic data of the authors is also recorded for supporting the research activities like author identification, text-matching etc.</p>
Urdu Text Normalization and Tokenization Dataset
<p>This dataset is made public so researchers can perform NLP tasks.</p>
UHaT Dataset: Urdu Handwritten Text Dataset
<p><strong>UHaT Dataset</strong></p> <p><strong>UHaT: Urdu Handwritten Text Dataset</strong></p> <p>This dataset contains handwritten characters and digits of Urdu language. The samples are written by 900+ individuals.</p> <p><strong>Description and organization</strong></p> <p>Size of images: All the images are stored in 28 by 28 resolution.</p> <p>How many images: The training set per each character contains of 700 images on average. For example, there are 811 train set images for AYN and 697 train set images for ALIF. Similarly, the train set per each contains 700 images on average. For example, there are 678 train set images for digits one. The test set per each character contains 140 images on average. For example, there are 145 test set images for character ALIF. The test set per each digit contains 140 images on average. For example, there are 147 test set images for digit nine.</p> <p>The dataset is organized into four sub-directories. Characters Training set, Characters Test set, Digits training set and digits test set. Each sub-director contains one sub-folder per one character. For example, all the train images for character ALIF are placed in sub-folder Alif.</p> <p>The folder hierarchy is given as:</p> <p>*Data > characters train set > alif</p> <p>Data > characters train set > ayn*</p> <p>And so on….</p> <p><strong>How to load directly?</strong></p> <p>You can also load it directly from the <em>uhat_<em>dataset.npz</em> file. See the kernel </em><strong>load_dataset</strong></p> <p><strong>Acknowledgements</strong></p> <p>Thanks to all volunteers who contributed by providing handwriting samples.</p> <p><strong>Inspiration</strong></p> <p>This is an <em>MNIST</em> style dataset. The machine learning community in general will find it useful for experimentation, demonstration purposes of machine learning models.<br> The dataset will also provide an opportunity to researchers to work on Urdu text recognition.</p>
Ontolex-lemon and TIAD versions of Apertium Urdu-Hindi dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2014, Francis M. Tyers 2010, Eknath Venkataramani 2012, Jim O'Regan 2012, San_ 2014, Kevin Brubeck Unhammer 2014, Sudarsh Rathi
URDU Dataset for Multi-modal Sentiment Analysis
<p>The "Multi-modal Sentiment Analysis Dataset for Urdu Language Opinion Videos" is a valuable resource aimed at advancing research in sentiment analysis, natural language processing, and multimedia content understanding. This dataset is specifically curated to cater to the unique context of Urdu language opinion videos, a dynamic and influential content category in the digital landscape.</p> <p><strong>Dataset Description:</strong></p> <ul> <li><strong>Size and Diversity:</strong> This dataset comprises an extensive collection of Urdu language opinion videos, encompassing a wide spectrum of topics and sentiments. It consists of a total of 214 videos, each of varying lengths, offering a diverse and comprehensive representation of the Urdu language content landscape.</li> <li><strong>Sentiment Annotations:</strong> The dataset is meticulously annotated with sentiment labels, providing information on the emotional tone expressed in each video. The sentiment labels include "positive," "negative," and "neutral," offering a nuanced understanding of the sentiment conveyed in these multimedia opinion pieces.</li> <li><strong>Multi-modal Approach:</strong> A unique feature of this dataset is its multi-modal approach. It combines text, audio, and visual data to enable researchers to delve into the various dimensions of sentiment analysis within the context of opinion videos. The multi-modal annotations encompass the textual content of spoken words, the auditory characteristics of the videos, and the visual cues from the video frames.</li> </ul> <p><strong>Significance and Applications:</strong></p> <p>This dataset holds significant value for both the research community and practical applications:</p> <ul> <li><strong>Research Advancement:</strong> Researchers can employ this dataset to investigate the complex landscape of sentiment analysis within the context of opinion videos. It facilitates inquiries into sentiment trends, the development of sentiment analysis models, and the creation of sentiment-aware multimedia content analysis tools.</li> <li><strong>Content Recommendation:</strong> The dataset can play a pivotal role in the development of content recommendation systems that cater to viewers' emotional preferences. Understanding sentiment in opinion videos is crucial for improving content engagement and user experience.</li> <li><strong>User Engagement Analysis:</strong> The dataset can empower studies on user engagement and interaction with multimedia content. It is an essential resource for researchers aiming to decode the factors influencing viewer reactions and engagement in multimedia.</li> </ul> <p> </p> <p>Researchers are encouraged to explore and utilize this dataset for various academic and commercial purposes, fostering innovation in sentiment analysis and multimedia understanding. The dataset is made available with open access to facilitate collaborative research and to contribute to the broader knowledge in the field.</p>
Urdu Translation of Bimanaual Fine Motor Functional Classification System 2
ClinicalTrials.gov study NCT06484465. IPD Sharing: NO. Countries: 1. Publications: 1.
Translation of Modified Fatigue Impact Scale in Urdu Language
ClinicalTrials.gov study NCT05405517. IPD Sharing: NO. Countries: 1. Publications: 4.
Urdu Translation of Wayfinding Questionnaire for Stroke Patients
ClinicalTrials.gov study NCT05375578. IPD Sharing: NO. Countries: 1. Publications: 2.
Translation and Validation of Rivermead Mobility Index in Urdu
ClinicalTrials.gov study NCT06519227. IPD Sharing: NO. Countries: 1. Publications: 5.
Urdu Translation and Cross Cultural Validation of PedsQL Inventory Infant Scale Parent Report 1 to 12 Months
ClinicalTrials.gov study NCT05721430. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Urdu Translation of Duke Activity Status
ClinicalTrials.gov study NCT04739163. IPD Sharing: NO. Countries: 1. Publications: 4.
Cross Cultural Adaptation and Psychometric Properties of Urdu Version of Fall Risk Questionnaire
ClinicalTrials.gov study NCT06687031. IPD Sharing: NO. Countries: 1. Publications: 6.
Cross Cultural Adaptation of Functional Status Questionnaire in Urdu Language in CABG Patients
ClinicalTrials.gov study NCT05023083. IPD Sharing: NO. Countries: 1. Publications: 10.
Validation of Urdu Version of Leicester Cough Questionnaire (LCQ)
ClinicalTrials.gov study NCT06674395. IPD Sharing: NO. Countries: 1. Publications: 10.
Urdu Version of National Institutes of Health Stroke Scale: Reliability and Validity Study
ClinicalTrials.gov study NCT05203081. IPD Sharing: NO. Countries: 1. Publications: 10.
Urdu Translation of Revised High-Level Mobility Assessment Tool (HiMAT) in Healthy Children
ClinicalTrials.gov study NCT06986720. IPD Sharing: NO. Countries: 1. Publications: 1.
Translation, Cultural Adaptation and Psychometric Properties of Urdu Version of Upper Limb Functional Index Questionnaire in Patients With Upper Limb Musculoskeletal Disorders
ClinicalTrials.gov study NCT05088096. IPD Sharing: NO. Countries: 1. Publications: 1.
Urdu Translation and Validation of Michigan Hand Outcome Questionnaire
ClinicalTrials.gov study NCT05931133. IPD Sharing: NO. Countries: 1. Publications: 3.
Translation and Cross -Cultural Validation of ECOS-16 Questionnaire in Urdu Language
ClinicalTrials.gov study NCT04873960. IPD Sharing: NO. Countries: 1. Publications: 6.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.