Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

89 results for “Urdu”

Learn how ShareScore rates datasets ↗
zenodo40/100

Offensive content dataset in Urdu language

<p>The archive contains python code and various feature files of offensive language dataset in urdu. The purpose of sharing this archive is to regenerate the results produced by the research article and can extend the findings.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Urdu Handwritten Text Dataset

<p>The dataset contains the images of handwritten text in Urdu language, one of the most widely spoken languages in South-East Asian regions. The native-speaking authors from different social domains were invited to write a pre-written text in their handwritings. The pre-written text is carefully written in a way that it includes almost all the characters, ligatures, diacritics, and dots used in writing the text Urdu script. The disabled persons are also involved to write the text to make the data collection more comprehensive. The demographic data of the authors is also recorded for supporting the research activities like author identification, text-matching etc.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Urdu Text Normalization and Tokenization Dataset

<p>This dataset is made public so researchers can perform NLP tasks.</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

UHaT Dataset: Urdu Handwritten Text Dataset

<p><strong>UHaT Dataset</strong></p> <p><strong>UHaT: Urdu Handwritten Text Dataset</strong></p> <p>This dataset contains handwritten characters and digits of Urdu language. The samples are written by 900+ individuals.</p> <p><strong>Description and organization</strong></p> <p>Size of images: All the images are stored in 28 by 28 resolution.</p> <p>How many images: The training set per each character contains of 700 images on average. For example, there are 811 train set images for AYN and 697 train set images for ALIF. Similarly, the train set per each contains 700 images on average. For example, there are 678 train set images for digits one. The test set per each character contains 140 images on average. For example, there are 145 test set images for character ALIF. The test set per each digit contains 140 images on average. For example, there are 147 test set images for digit nine.</p> <p>The dataset is organized into four sub-directories. Characters Training set, Characters Test set, Digits training set and digits test set. Each sub-director contains one sub-folder per one character. For example, all the train images for character ALIF are placed in sub-folder Alif.</p> <p>The folder hierarchy is given as:</p> <p>*Data &gt; characters train set &gt; alif</p> <p>Data &gt; characters train set &gt; ayn*</p> <p>And so on&hellip;.</p> <p><strong>How to load directly?</strong></p> <p>You can also load it directly from the&nbsp;<em>uhat_<em>dataset.npz</em>&nbsp;file. See the kernel </em><strong>load_dataset</strong></p> <p><strong>Acknowledgements</strong></p> <p>Thanks to all volunteers who contributed by providing handwriting samples.</p> <p><strong>Inspiration</strong></p> <p>This is an&nbsp;<em>MNIST</em>&nbsp;style dataset. The machine learning community in general will find it useful for experimentation, demonstration purposes of machine learning models.<br> The dataset will also provide an opportunity to researchers to work on Urdu text recognition.</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Ontolex-lemon and TIAD versions of Apertium Urdu-Hindi dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2014, Francis M. Tyers 2010, Eknath Venkataramani 2012, Jim O'Regan 2012, San_ 2014, Kevin Brubeck Unhammer 2014, Sudarsh Rathi

opengpl-2.0-or-laterMar 2020View details →
zenodo32/100

URDU Dataset for Multi-modal Sentiment Analysis

<p>The "Multi-modal Sentiment Analysis Dataset for Urdu Language Opinion Videos" is a valuable resource aimed at advancing research in sentiment analysis, natural language processing, and multimedia content understanding. This dataset is specifically curated to cater to the unique context of Urdu language opinion videos, a dynamic and influential content category in the digital landscape.</p> <p><strong>Dataset Description:</strong></p> <ul> <li><strong>Size and Diversity:</strong>&nbsp;This dataset comprises an extensive collection of Urdu language opinion videos, encompassing a wide spectrum of topics and sentiments. It consists of a total of 214 videos, each of varying lengths, offering a diverse and comprehensive representation of the Urdu language content landscape.</li> <li><strong>Sentiment Annotations:</strong>&nbsp;The dataset is meticulously annotated with sentiment labels, providing information on the emotional tone expressed in each video. The sentiment labels include "positive," "negative," and "neutral," offering a nuanced understanding of the sentiment conveyed in these multimedia opinion pieces.</li> <li><strong>Multi-modal Approach:</strong>&nbsp;A unique feature of this dataset is its multi-modal approach. It combines text, audio, and visual data to enable researchers to delve into the various dimensions of sentiment analysis within the context of opinion videos. The multi-modal annotations encompass the textual content of spoken words, the auditory characteristics of the videos, and the visual cues from the video frames.</li> </ul> <p><strong>Significance and Applications:</strong></p> <p>This dataset holds significant value for both the research community and practical applications:</p> <ul> <li><strong>Research Advancement:</strong>&nbsp;Researchers can employ this dataset to investigate the complex landscape of sentiment analysis within the context of opinion videos. It facilitates inquiries into sentiment trends, the development of sentiment analysis models, and the creation of sentiment-aware multimedia content analysis tools.</li> <li><strong>Content Recommendation:</strong> The dataset can play a pivotal role in the development of content recommendation systems that cater to viewers' emotional preferences. Understanding sentiment in opinion videos is crucial for improving content engagement and user experience.</li> <li><strong>User Engagement Analysis:</strong> The dataset can empower studies on user engagement and interaction with multimedia content. It is an essential resource for researchers aiming to decode the factors influencing viewer reactions and engagement in multimedia.</li> </ul> <p>&nbsp;</p> <p>Researchers are encouraged to explore and utilize this dataset for various academic and commercial purposes, fostering innovation in sentiment analysis and multimedia understanding. The dataset is made available with open access to facilitate collaborative research and to contribute to the broader knowledge in the field.</p>

opencc-by-4.0Dec 2024View details →
ClinicalTrials.gov32/100

Urdu Translation of Bimanaual Fine Motor Functional Classification System 2

ClinicalTrials.gov study NCT06484465. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Translation of Modified Fatigue Impact Scale in Urdu Language

ClinicalTrials.gov study NCT05405517. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Translation of Wayfinding Questionnaire for Stroke Patients

ClinicalTrials.gov study NCT05375578. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Translation and Validation of Rivermead Mobility Index in Urdu

ClinicalTrials.gov study NCT06519227. IPD Sharing: NO. Countries: 1. Publications: 5.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Translation and Cross Cultural Validation of PedsQL Inventory Infant Scale Parent Report 1 to 12 Months

ClinicalTrials.gov study NCT05721430. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Translation of Duke Activity Status

ClinicalTrials.gov study NCT04739163. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Cross Cultural Adaptation and Psychometric Properties of Urdu Version of Fall Risk Questionnaire

ClinicalTrials.gov study NCT06687031. IPD Sharing: NO. Countries: 1. Publications: 6.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Cross Cultural Adaptation of Functional Status Questionnaire in Urdu Language in CABG Patients

ClinicalTrials.gov study NCT05023083. IPD Sharing: NO. Countries: 1. Publications: 10.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Validation of Urdu Version of Leicester Cough Questionnaire (LCQ)

ClinicalTrials.gov study NCT06674395. IPD Sharing: NO. Countries: 1. Publications: 10.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Version of National Institutes of Health Stroke Scale: Reliability and Validity Study

ClinicalTrials.gov study NCT05203081. IPD Sharing: NO. Countries: 1. Publications: 10.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Translation of Revised High-Level Mobility Assessment Tool (HiMAT) in Healthy Children

ClinicalTrials.gov study NCT06986720. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Translation, Cultural Adaptation and Psychometric Properties of Urdu Version of Upper Limb Functional Index Questionnaire in Patients With Upper Limb Musculoskeletal Disorders

ClinicalTrials.gov study NCT05088096. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Urdu Translation and Validation of Michigan Hand Outcome Questionnaire

ClinicalTrials.gov study NCT05931133. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Translation and Cross -Cultural Validation of ECOS-16 Questionnaire in Urdu Language

ClinicalTrials.gov study NCT04873960. IPD Sharing: NO. Countries: 1. Publications: 6.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record