Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “Hindi”
CWID-hi: A Dataset for Complex Word Identification in Hindi Text
<p>This dataset was created by conducting a human intelligence test, wherein native and non-native Hindi speakers annotated words they could not understand in Hindi text. They were then asked to rank the complexity of these words along with their synonyms. A word that received an average rank of <=3 (out of 5) is labeled 1 and the word that received an average rank of >3 is labeled 0. 1 indicates complex and 0 indicates simple.</p>
CNN for Modeling Sanskrit Originated Bengali and Hindi Language Dataset
<p>Though recent works have focused on modeling high resource languages, the area is still unexplored for low resource languages like Bengali and Hindi. We propose an end-to-end trainable memory efficient CNN architecture named CoCNN to handle specific characteristics such as high inflection, morphological richness, flexible word order and phonetical spelling errors of Bengali and Hindi. In particular, we introduce two learnable convolutional sub-models at word and at sentence level that are end-to-end trainable. We show that state-of-the-art (SOTA) Transformer models including pretrained BERT do not necessarily yield the best performance for Bengali and Hindi. CoCNN outperforms pretrained BERT with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. This is the first study on the effectiveness of different architectures from Convolution, Recurrent, and Transformer neural net paradigm for modeling Bengali and Hindi.</p>
Hindi News Article Text Dataset
<p>The Hindi News Article Dataset (HNAD) comprises over a million meticulously curated news articles in the Hindi language, sourced from diverse online platforms. Covering a wide range of topics, including politics, economics, culture, sports, and more, it offers a comprehensive representation of the Hindi news landscape. </p>
Multilingual Fake News Detection Dataset: Gujarati, Hindi, Marathi, and Telugu
<p>This dataset is designed to support research in fake news detection across four major Indian languages: Gujarati, Hindi, Marathi, and Telugu. The dataset includes a diverse set of news articles collected from various sources, each labeled as either 'fake' or 'real'. The primary goal is to provide a resource that helps in the development and evaluation of natural language processing (NLP) models capable of detecting fake news in these regional languages.</p>
India (Hindi)
भारत, आधिकारिक तौर पर भारत गणराज्य, दक्षिण एशिया का एक देश है। यह क्षेत्रफल के हिसाब से सातवां सबसे बड़ा देश, दूसरा सबसे अधिक आबादी वाला देश और दुनिया में सबसे अधिक आबादी वाला लोकतंत्र है। Source: Objaverse 1.0 / Sketchfab
Examining the Feasibility of Wysa in Hindi
ClinicalTrials.gov study NCT06320756. IPD Sharing: NO. Countries: 1. Publications: 2.
Ontolex-lemon and TIAD versions of Apertium Urdu-Hindi dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2014, Francis M. Tyers 2010, Eknath Venkataramani 2012, Jim O'Regan 2012, San_ 2014, Kevin Brubeck Unhammer 2014, Sudarsh Rathi
H-Prop and H-Prop-News Propaganda Datasets in Hindi
<p>The H-Prop dataset contains 28,630 articles created by translating a portion of Proppy Corpus in Hindi. Each article is labeled as either “propagandistic” (positive class) or “non-propagandistic” (negative class). The labeling done indirectly in Proppy corpus using a technique known as distant supervision is retained. </p> <p>The H-Prop-News dataset contains 5,500 Hindi News articles collected from 30+ prominent Hindi News websites. Each article is labeled as either “propagandistic” (positive class) or “non-propagandistic” (negative class). The labeling was done by human annotators and the inter-annotator agreement using Cohen’s Kappa measure observed is 0.81.</p> <p>## Data format</p> <p>We provide the H-Prop dataset in three tsv files, including training, testing and validation partitions. The H-Prop-News dataset is provided in csv files including training, testing and validation partitions.</p> <p>Each line represents one article in H-Prop dataset with the following information:</p> <p>1. article_text: the text of the article translated from Proppy corpus.<br> 2. propaganda_label: label for articles retained from Proppy corpus.</p> <p>Each line represents one article in H-Prop-News dataset with the following information:</p> <p>1. news_website: Name of the news source website<br> 2. article_url: the direct URL for the published article in its source website<br> 3. news_headline: news headline<br> 4. article_text: the text of the article retrieved via parsehub tool<br> 5. propaganda_label: label for articles</p> <p>## About</p> <p>The H-Prop dataset was translated using IBM Watson Language Translator. </p> <p>## Credit</p> <p>Please cite the dataset as:<br> [HProp-News] Deptii Chaudhari, Ambika Pawar, and Alberto Barrón-Cedeño. 2022. H-Prop and H-Prop-News: Computational Propaganda Datasets in Hindi. doi: 10.5281/zenodo.5828240</p> <p>## Authors</p> <p>Deptii Chaudhari;<br> Ambika Pawar;<br> Alberto Barrón-Cedeno</p>
PB Hindi ASR dataset
<p>Hindi ASR dataset - released along with the paper "TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR". If you find this useful, please cite as </p> <pre>@article{ravi2024teles, title={TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR}, author={Ravi, Nagarathna, T, Thishyan Raj and Arora, Vipul}, journal={arXiv preprint arXiv:2401.03251}, year={2024} }</pre>
BHAAV (भाव) - A Text Corpus for Emotion Analysis from Hindi Stories
<p>The first and largest Hindi text corpus, named BHAAV (भाव), which means emotions in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus consists of 20,304 sentences collected from 230 different short stories spanning across 18 genres such as प्रेरणादायक (Inspirational) and रहस्यमयी (Mystery). Each sentence has been annotated into one of the five emotion categories anger, joy, suspense, sad, and neutral) by three native Hindi speakers with at least ten years of formal education in Hindi.</p>
Validation of Hindi Version Oral Health Impact Profile for Periodontitis
ClinicalTrials.gov study NCT07303491. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Responsiveness of Hindi Version of Oral Health Impact Profile for Periodontitis (OHIP-P-HIN)
ClinicalTrials.gov study NCT07304596. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.