Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

12

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

12 results for “Hindi”

Learn how ShareScore rates datasets ↗
zenodo44/100

CWID-hi: A Dataset for Complex Word Identification in Hindi Text

<p>This dataset was created by conducting a human intelligence test, wherein native and non-native Hindi speakers annotated words they could not understand in Hindi text. They were then asked to rank the complexity of these words along with their synonyms. A word that received an average rank of &lt;=3 (out of 5) is labeled 1 and the word that received an average rank of &gt;3 is labeled 0. 1 indicates complex and 0 indicates simple.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

CNN for Modeling Sanskrit Originated Bengali and Hindi Language Dataset

<p>Though recent works have focused on modeling high resource languages, the area is still unexplored for low resource languages like Bengali and Hindi. We propose an end-to-end trainable memory efficient CNN architecture named CoCNN to handle specific characteristics such as high inflection, morphological richness, flexible word order and phonetical spelling errors of Bengali and Hindi. In particular, we introduce two learnable convolutional sub-models at word and at sentence level that are end-to-end trainable. We show that state-of-the-art (SOTA) Transformer models including pretrained BERT do not necessarily yield the best performance for Bengali and Hindi. CoCNN outperforms pretrained BERT with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. This is the first study on the effectiveness of different architectures from Convolution, Recurrent, and Transformer neural net paradigm for modeling Bengali and Hindi.</p>

openmit-licenseOct 2022View details →
zenodo36/100

Hindi News Article Text Dataset

<p>The Hindi News Article Dataset (HNAD) comprises over a million meticulously curated news articles in the Hindi language, sourced from diverse online platforms. Covering a wide range of topics, including politics, economics, culture, sports, and more, it offers a comprehensive representation of the Hindi news landscape.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Multilingual Fake News Detection Dataset: Gujarati, Hindi, Marathi, and Telugu

<p>This dataset is designed to support research in fake news detection across four major Indian languages: Gujarati, Hindi, Marathi, and Telugu. The dataset includes a diverse set of news articles collected from various sources, each labeled as either 'fake' or 'real'. The primary goal is to provide a resource that helps in the development and evaluation of natural language processing (NLP) models capable of detecting fake news in these regional languages.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

India (Hindi)

भारत, आधिकारिक तौर पर भारत गणराज्य, दक्षिण एशिया का एक देश है। यह क्षेत्रफल के हिसाब से सातवां सबसे बड़ा देश, दूसरा सबसे अधिक आबादी वाला देश और दुनिया में सबसे अधिक आबादी वाला लोकतंत्र है। Source: Objaverse 1.0 / Sketchfab

opencc-byJun 2022View details →
ClinicalTrials.gov36/100

Examining the Feasibility of Wysa in Hindi

ClinicalTrials.gov study NCT06320756. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
zenodo32/100

Ontolex-lemon and TIAD versions of Apertium Urdu-Hindi dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2014, Francis M. Tyers 2010, Eknath Venkataramani 2012, Jim O'Regan 2012, San_ 2014, Kevin Brubeck Unhammer 2014, Sudarsh Rathi

opengpl-2.0-or-laterMar 2020View details →
zenodo32/100

H-Prop and H-Prop-News Propaganda Datasets in Hindi

<p>The H-Prop dataset contains 28,630 articles created by translating a portion of Proppy Corpus in Hindi. Each article is labeled as either &ldquo;propagandistic&rdquo; (positive class) or &ldquo;non-propagandistic&rdquo; (negative class). The labeling done indirectly in Proppy corpus using a technique known as distant supervision is retained.&nbsp;</p> <p>The H-Prop-News dataset contains 5,500 Hindi News articles collected from 30+ prominent Hindi News websites. Each article is labeled as either &ldquo;propagandistic&rdquo; (positive class) or &ldquo;non-propagandistic&rdquo; (negative class). The labeling was done by human annotators and the inter-annotator agreement using Cohen&rsquo;s Kappa measure observed is 0.81.</p> <p>## Data format</p> <p>We provide the H-Prop dataset in three tsv files, including training, testing and validation partitions. The H-Prop-News dataset is provided in csv files including training, testing and validation partitions.</p> <p>Each line represents one article in H-Prop dataset with the following information:</p> <p>1. article_text: the text of the article translated from Proppy corpus.<br> 2. propaganda_label: label for articles retained from Proppy corpus.</p> <p>Each line represents one article in H-Prop-News dataset with the following information:</p> <p>1. news_website: Name of the news source website<br> 2. article_url: the direct URL for the published article in its source website<br> 3. news_headline: news headline<br> 4. article_text: the text of the article retrieved via parsehub tool<br> 5. propaganda_label: label for articles</p> <p>## About</p> <p>The H-Prop dataset was translated using IBM Watson Language Translator.&nbsp;</p> <p>## Credit</p> <p>Please cite the dataset as:<br> [HProp-News] Deptii Chaudhari, Ambika Pawar, and Alberto Barr&oacute;n-Cede&ntilde;o. 2022. H-Prop and H-Prop-News: Computational Propaganda Datasets in Hindi. doi: 10.5281/zenodo.5828240</p> <p>## Authors</p> <p>Deptii Chaudhari;<br> Ambika Pawar;<br> Alberto Barr&oacute;n-Cedeno</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

PB Hindi ASR dataset

<p>Hindi ASR dataset -&nbsp; released along with the paper "TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR".&nbsp; If you find this useful, please cite as&nbsp;</p> <pre>@article{ravi2024teles, title={TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR}, author={Ravi, Nagarathna, T, Thishyan Raj and Arora, Vipul}, journal={arXiv preprint arXiv:2401.03251}, year={2024} }</pre>

opencc-by-4.0May 2024View details →
zenodo32/100

BHAAV (भाव) - A Text Corpus for Emotion Analysis from Hindi Stories

<p>The first and largest Hindi text corpus, named BHAAV (भाव), which means emotions&nbsp;in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus consists of 20,304 sentences collected from 230 different short stories spanning across 18 genres such as प्रेरणादायक (Inspirational) and रहस्यमयी (Mystery). Each sentence has been annotated into one of the five emotion categories anger, joy, suspense, sad, and neutral)&nbsp;by three native Hindi speakers with at least ten years of formal education in Hindi.</p>

opencc-by-4.0Sep 2019View details →
ClinicalTrials.gov24/100

Validation of Hindi Version Oral Health Impact Profile for Periodontitis

ClinicalTrials.gov study NCT07303491. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov24/100

Responsiveness of Hindi Version of Oral Health Impact Profile for Periodontitis (OHIP-P-HIN)

ClinicalTrials.gov study NCT07304596. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record