Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

475

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

475 results for “word”

Learn how ShareScore rates datasets ↗
zenodo40/100

Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes

<p>The dataset was made for the purposes of the author&#39;s master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">&quot;Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo&rsquo;s Core Vocabulary&quot;</a>. The dataset includes two worksheets. The first is titled &quot;Main table&quot;, and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled &quot;Extended table&quot;, includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively).&nbsp;<br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">at&nbsp;this link</a>.&nbsp;&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Word-related parameters and discriminative power of NWRT for bilingual children (Italian-German) with/without DLD

<p>Dataset with word-related parameters, adult monolingual ratings&nbsp;and mean repetition scores by bilingual children with/without (risk of) DLD. Word-related parameters include length, number of syllables, number of clusters,&nbsp; monolingual adults&#39; direct (in-person) repetition rates and direct ratings of Pronounceability (PR), Word-likeness (WLike); online ratings of Pronounceability (avPR), Specificity (avSP, corresponding to&nbsp;percent target language assignment), and children&#39;s repetition rates (averages from children with/without DLD).&nbsp;IT, Italian, GER, German.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Word-to-word transcripts of the interviews/consultations on What do FNS-Cloud Food Researchers Want to Know?

<p>Word to word transcripts of 11&nbsp;semi-structured interviews carried out during the period April-October 2020 on &#39;What do FNS-Cloud Food Researchers Want to Know&#39;. The notes were independently taken by the two interviewers during the interviews&nbsp;and combined to produce almost word-to-word transcripts of the interviews. Included in this Excel are also the questions addressed. The semi-guided interviews were&nbsp;undertaken to ensure the content validity of trainings for the FNS-Cloud community, and to&nbsp;reflect the training needs and preferences of FNS-Cloud project partners.</p> <p>Both interviewers (2 persons) and interviewed (15 persons) were food science professionals participating in the FNS-Cloud project (H2020 No. 863059). The Interviews, which lasted 30 minutes, aimed at identifying training needs and preferences related to open science and the use of the project datasets, tools and services: what partners want to learn, how they prefer to learn, and who are their ideal teachers.&nbsp;</p> <p>Inductive coding of the transcripts was done with the NVivo 12 Pro&copy; software for qualitative analysis, following an iterative approach involving three researchers reviewing interview transcripts, codes, sub-codes, and coded phrases. The results of the qualitative analysis are presented in the paper:&nbsp;Teaching Open Science. What do FNS-Cloud Food Researchers Want to Know? presented at the 8th International Conference on Higher Education Advances (HEAd&rsquo;22),<a href="http://headconf.org/">&nbsp;</a><a href="http://headconf.org/">HEAd&#39;22 | June 14-17, 2022 &middot; Valencia, Spain (headconf.org)</a>&nbsp;and published as a&nbsp;Peer-reviewed article&nbsp;in the: Proceedings of the 8th International Conference on Higher Education Advances (HEAd&rsquo;22) (includes DOI and ISBN)&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Challow 100-word Swadesh list

<p>This repository contains a collection of high-quality (WAV) audio files corresponding to the 100-word Swadesh list&nbsp;in Challow, a Trans-Himalayan (Tibeto-Burman) language spoken in Manipur, North-East India.</p> <p>This repository is part of the Supplementary Materials for the paper: Ivani, Jessica K. <em>to appear.</em>&nbsp;2024. &quot;Some preliminary notes on Challow, a Trans-Himalayan language from Manipur, India&quot;.&nbsp;<em>Languages and Peoples of the Eastern Himalayan Region (LPEHR),&nbsp;</em>Vol.22:2.</p> <p>The material in this file may be freely quoted, copied, or reproduced for non-commercial purposes. If used, a citation is required.</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Key words related to public engagement with science by Society of Freshwater Science journals and conference sessions (1997-2019)

Open the record for dataset details and reuse information.

publicApr 2021View details →
zenodo36/100

Supplementary Material for the paper: Automatic Document Screening of Medical Literature Using Word and Text Embeddings in an Active Learning Setting

<p>This is the dataset used in the paper:&nbsp;Automatic Document Screening of Medical Literature Using Word and Text Embeddings in an Active Learning Setting.&nbsp;</p> <p>It is composed of:&nbsp;</p> <p>- Pre-trained models using active learning for document screening on HealthCLEF and Epistemonikos datasets.&nbsp;</p> <p>- Epistemonikos and HealthCLEF datasets containing medical questions and relevant/non relevant articles.&nbsp;</p> <p>- Embeddings and Document Representations used for experiments on both datasets.&nbsp;</p> <p>Scripts to run experiments can be found at:&nbsp;<a href="https://github.com/afcarvallo/active_learning_document_screening">https://github.com/afcarvallo/active_learning_document_screening</a></p> <p>&nbsp;</p> <p><strong>Paper abstract:</strong></p> <p>Document screening is a fundamental task within Evidence-based Medicine (EBM), a practice that provides scientific evidence to support medical decisions. Several approaches have tried to reduce physicians&#39; workload of screening and labeling vast amounts of documents to answer clinical questions. Previous works tried to semi-automate document screening, reporting promising results, but their evaluation was conducted on small datasets, which hinders generalization. Moreover, recent works in natural language processing have introduced neural language models, but none have compared their performance in EBM. In this paper, we evaluate the impact of several document representations such as TF-IDF along with neural language models (BioBERT, BERT, Word2vec, and GloVe) on an active learning-based setting for document screening in EBM. Our goal is to reduce the number of documents that physicians need to label to answer clinical questions. We evaluate these methods using both a small challenging dataset (HealthCLEF 2017) as well as a larger one but easier to rank (Epistemonikos). Our results indicate that word as well as textual neural embeddings always outperform the traditional TF-IDF representation. When comparing among neural and textual embeddings, in the HealthCLEF dataset the models BERT and BioBERT yielded the best results. On the larger dataset, Epistemonikos, Word2Vec and BERT were the most competitive, showing that BERT was the most consistent model across different corpuses. In term of active learning, an uncertainty sampling strategy combined with logistic regression achieved the best performance overall, above other methods under evaluation, and in fewer iterations.</p>

opencc-by-4.0Mar 2020View details →
dryad36/100

Data from: Shared and modality-specific brain regions that mediate auditory and visual word comprehension

<p>Visual speech carried by lip movements is an integral part of communication. Yet, it remains unclear in how far visual and acoustic speech comprehension are mediated by the same brain regions. Using multivariate classification of full-brain MEG data, we first probed where the brain represents acoustically and visually conveyed word identities. We then tested where these sensory-driven representations are predictive of participants' trial-wise comprehension. The comprehension-relevant representations of auditory and visual speech converged only in anterior angular and inferior frontal regions and were spatially dissociated from those representations that best reflected the sensory-driven word identity. These results provide a neural explanation for the behavioural dissociation of acoustic and visual speech comprehension and suggest that cerebral representations encoding word identities may be more modality-specific than often upheld.</p>

opencc-zeroAug 2020View details →
zenodo36/100

Network Graph showing the PMI results for Akkadian words related to Masculinities

<p>Network Graph displaying the top 10 results of a PMI measurement between Akkadian words which relate to the concept of masculinities.</p> <p>Based on data from ORACC downloaded in April 2020.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Lesley Dill - Wood Word Woman

with Wood Word Pedestal. Oil stick and silver leaf on wood 2011, this 3D scan 2014 Scanned at: deCordova Sculpture Garden Source: Objaverse 1.0 / Sketchfab

opencc-byDec 2014View details →
zenodo36/100

Classification of word levels with usage frequency, expert opinions and machine learning

<p>This dataset includes classification of English words according to CEFR language levels. It can be used in various educational applications including determining levels of text that is appropriate for students learning English.&nbsp;</p> <p>For each word, part-of-speech, the word lemma and usage frequency is provided. For words that have no survey results, a machine learning based methodology is used to predict levels. These predictions are also included as a separate file. This data is released as part of the&nbsp;submission process to British Journal of Educational Technology Special Issue on Open Data.</p> <p>The included readme.pdf file contains a detailed&nbsp;description of data.&nbsp;</p>

opencc-by-4.0Oct 2014View details →
zenodo36/100

The Software Sustainability Institute's Collaborations Workshop 2015 (CW15) attendees computational tools word-cloud

<p>Word cloud representing the computational tools used by those attending the Software Sustainability Institute&#39;s Collaborations Workshop 2015 (CW15).</p> <p>For more information see www.software.ac.uk/cw15</p> <p>Please note there is an ERROR in the diagram for some reason wordle.net did not pickup &#39;R&#39; in the dataset -&nbsp;http://dx.doi.org/10.5281/zenodo.19828&nbsp;- i.e. the usage of R in research software in the people who attended CW15 is not represented in this diagram.</p>

opencc-by-nc-4.0Jul 2015View details →
zenodo36/100

Hacker News lda2vec pretrained word vectors

<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900&nbsp;and https://zenodo.org/record/49902</p>

opencc-zeroApr 2016View details →
zenodo36/100

Hacker News lda2vec model pretrained word vectors

<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900</p>

opencc-zeroApr 2016View details →
zenodo36/100

Russian Distributional Thesaurus (RDT): Word Embeddings

<p>This resource is a part of the Russian Distributional Thesaurus (RDT): see http://russe.nlpub.ru/downloads and http://nlpub.ru/RDT. </p> <p>This dataset contains a large scale word embeddings model for Russian trained using the SGNS model (Mikolov et al., 2013) on a 12.9 billion word collection of books in Russian. According to the results of our participation in the shared task on Russian semantic similarity (Panchenko et al., 2015), this approach scored in the top 5 among 105 submissions (Arefyev et al., 2015). Following our prior experiments (Arefyev et al., 2015) we have selected the following parameters for the model: minimal word frequency – 5, number of dimensions in a word vector – 500, three or five iterations of the learning algorithm over the input corpus, context window size of 1, 2, 3, 5, 7 and 10 words. Parameters of the model are listed below:</p> <ul> <li>Model: skip-gram</li> <li>Corpus: a 150Gb sample of the lib.rus.ec book collection.</li> <li>Context window size: 10 words</li> <li>Number of dimensions: 500</li> <li>Number of iterations: 3</li> <li>Minimal word frequency: 5</li> </ul> <p>References:</p> <ul> <li>Panchenko A., Ustalov D., Arefyev N., Paperno D., Konstantinova N., Loukachevitch N. and Biemann C. (2016): Human and Machine Judgements about Russian Semantic Relatedness. In Proceedings of the 5th Conference on Analysis of Images, Social Networks, and Texts (AIST'2016). Communications in Computer and Information Science (CCIS). Springer-Verlag Berlin Heidelberg</li> </ul> <ul> <li>Panchenko A., Loukachevitch N. V., Ustalov D., Paperno D., Meyer C. M., Konstantinova N. (2015): RUSSE: The First International Workshop on Russian Semantic Similarity. In Proceedings of the 21st International Conference on Computational Linguistics and Intellectual Technologies (Dialogue'2015). Moscow, Russia. RGGU</li> </ul> <ul> <li>Arefyev N., Panchenko A., Lukanin A., Lesota O., Romanov P. (2015): Evaluating Three Corpus-Based Semantic Similarity Systems for Russian. In Proceedings of the 21st International Conference on Computational Linguistics and Intellectual Technologies (Dialogue'2015). Moscow, Russia. RGGU</li> </ul>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation

<p>This dataset contains the models for interpretable Word Sense Disambiguation (WSD) that were employed in Panchenko et al. (2017; the paper can be accessed at https://www.lt.informatik.tu-darmstadt.de/fileadmin/user_upload/Group_LangTech/publications/EACL_Interpretability___FINAL__1_.pdf).</p> <p>The files were computed on a 2015 dump from the English Wikipedia. Their contents:</p> <ul> <li>Induced Sense Inventories: <strong>wp_stanford_sense_inventories.tar.gz</strong><br> This file contains 3 inventories (coarse, medium fine)</li> <li>Language Model (3-gram): <strong>wiki_text.3.arpa.gz</strong><br> This file contains all n-grams up to n=3 and can be loaded into an index</li> <li>Weighted Dependency Features: <strong>wp_stanford_lemma_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000.gz</strong><br> This file contains weighted word--context-feature combinations and includes their count and an LMI significance score</li> <li>Distributional Thesaurus (DT) of Dependency Features: <strong>wp_stanford_lemma_BIM_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000_simsortlimit200_feature expansion.gz</strong><br> This file contains a DT of context features. The context feature similarities can be used for context expansion</li> </ul> <p>For further information, consult the paper and the companion page: http://jobimtext.org/wsd/</p> <p>Panchenko A., Ruppert E., Faralli S., Ponzetto S. P., and Biemann C. (2017): Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL'2017). Valencia, Spain. Association for Computational Linguistics.</p> <p> </p> <p> </p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Word Embedding Data Sets Learned from Tweets and General Data

<p>This includes 10 word embedding data sets learned from about 400 million tweets and 7 billion words from general data. They can be used in tasks involving social media data, especially tweets, and other types of textual data. Users can choose different embedding sets based on their use cases; they can also easily try all of them to see which one provides the best performance for their application.</p> <p>More details about the training data collection, word embedding generation, preprocessing steps, and how to use them can be found from the following paper:</p> <p>Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh, Data Set: Word Embeddings Learned from Tweets and General Data, The 11th International AAAI Conference on Web and Social Media (ICWSM-17).  Montreal, Canada. May 16-18, 2017</p>

opencc-by-4.0May 2017View details →
zenodo36/100

A Dataset for Sanskrit Word Segmentation

<p>The work was accepted in Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, colocated with ACL 2017</p> <p>The last decade saw a surge in digitisation efforts for ancient manuscripts in Sanskrit. Due to various linguistic peculiarities inherent to the language, even the preliminary tasks such as word segmentation are non-trivial in Sanskrit. Elegant models for Word Segmentation in Sanskrit are indispensable for further syntactic and semantic processing of the manuscripts. Current works in word segmentation for Sanskrit, though commendable in their novelty, often have variations in their objective and evaluation criteria. In this work, we set the record straight. We formally define the objectives and the requirements for the word segmentation task. In order to encourage research in the field and to alleviate the time and effort required in pre-processing, we release a dataset of 115,000 sentences for word segmentation. For each sentence in the dataset we include the input character sequence, ground truth segmentation, and additionally lexical and morphological information about all the phonetically possible segments for the given sentence. In this work, we also discuss the linguistic considerations made while generating the candidate space of the possible segments. </p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Low-frequency Cortical Activity Reflects Context-dependent Parsing of Word Sequences

<p>The dataset and code used for the publication of "<a href="https://www.biorxiv.org/content/10.1101/2024.11.12.623335v1" target="_blank" rel="noopener"><strong>Low-frequency Cortical Activity Reflects Context-dependent Parsing of Word Sequences</strong></a>".</p> <div> <div><strong>Data path</strong></div> <div> <ol> <li>derivatives (./data/derivatives/): preprocessed data. <ul> <li>meg_quat-tsss-withEOG_100hz_0d3-40: meg data preprocessed at 100 Hz filtered between 0.3 and 40 Hz.</li> <li>FreeSurfer: anatomy data preprocessed</li> </ul> </li> <li>behavior_data (./behavior_data/)</li> <li>stimuli (./exp_procedure/material/): exp stimuli and procedure</li> <li>features (./data/features/): calculated phonological and linguistic features.</li> </ol> </div> <strong>Code path</strong></div> <div><br> <div>Preprocessing<strong>:</strong></div> <div> <ol> <li>[preprocessing.py]: preprocessing of behavior and MEG data.</li> <li>[recon.py]: reconstruction of T1 .niift data into surface data in FreeSurfer.</li> <li>[analysis_stimuli.ipynb]: preprocessing and feature extraction of stimuli.</li> </ol> </div> Analysis</div> <div> <ul> <li>[analysis_meg.ipynb]: spectral, time course, and RSA analysis.</li> </ul> </div>

opencc-by-4.0Jun 2024View details →
zenodo36/100

JasperD-UGent/LexComSpaL2: LexComSpaL2 corpus enriched with word family information

<p>Dataset presented in the <a href="https://aclanthology.org/2024.lrec-main.912/">Degraeuwe and Goethals (2024) paper</a> presented at LREC-COLING2024 plus the word family-enriched version of the dataset presented in the <a href="https://aclanthology.org/2025.bea-1.24/">Degraeuwe (2025) paper</a> presented at the BEA2025 workshop.</p>

openodc-byOct 2024View details →
zenodo36/100

Timely Dictionary Development: In-person and Virtual Rapid Word Collection for Endangered Indigenous Languages

<p>Timely Dictionary Development: In-person and virtual Rapid Word Collection for endangered Indigenous Languages</p> <p>Dorothea Hoffmann, Wilhelm Meya, Abbie Hantgan-Sonko &amp; Elliot Thornton</p> <p>Presented 5 October 2022 at the Berlin-Brandenburg Academy of Sciences and Humanities Where Do We Need to Go From Here? Language Documentation and Archiving in the International Decade of Indigenous Languages</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record