Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

475

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

475 results for “word”

Learn how ShareScore rates datasets ↗
zenodo36/100

Minga of words and thoughts in a non-formal educational process. Systematization of the CORPOMANIGUA experience

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo36/100

Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"

<h1>Replication Package for the Paper: &ldquo;Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis&rdquo;</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled &ldquo;Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.&rdquo;</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>

openmit-licenseApr 2024View details →
zenodo36/100

Word: O Essencial sobre Documentos Estruturados

<p>Word: O Essencial sobre Documentos Estruturados</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

ND250 as a prediction error signal in orthographic processing: insights from the comparison of handwritten and printed words

<p><span>This dataset contains electroencephalography (EEG) recordings and behavioral data from a study investigating the neural mechanisms of visual word recognition in native Chinese speakers. The study used a color decision task, where participants viewed printed and handwritten Chinese single-character words varying in lexical frequency (high-frequency vs. low-frequency). The primary aim was to examine the N250 ERP component, a 250-ms difference in brain activity observed between certain word types, and determine whether it reflects activation of the orthographic lexicon or a prediction error signal during orthographic processing. The findings suggest that the N250 is related to prediction error, providing support for the Interactive Account of orthographic processing.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

The role of spatial terms in time expressions: A case study of Chinese temporal words

<p>This project provides the research data, namely all the Chinese temporal words collected from the BCC corpus for The role of spatial terms in time expressions: A case study of Chinese temporal words. There is one table here which includes all tokens taken from the BCC corpus. More specifically, it lists 692 unique Chinese temporal words along with their frequencies. All of these temporal words have been classified into three categories: explicitly spatialized temporal words, implicitly spatialized temporal words, and others. All explicitly spatialized temporal words have been annotated based on the type of spatial terms they are marked with, mainly including location/direction, motion, dimension, objects, deixis, and spatial mixtures.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

DiaWUG: Diatopic Word Usage Graphs for Spanish

<p>This data collection contains diatopic Word Usage Graphs (WUGs) for Spanish. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Note:</p> <ul> <li>The date given for each word use does not correspond to the exact date of the document from which the use was sampled but only to the midpoint of the rough time period covered by the Corpus del Espa&ntilde;ol (~2000-2014).</li> <li>The numbers given as grouping for each word use map to Spanish variants in the following way: 0: Spain (ES), 1: Cuba (CU), 2: Colombia (CO), 3: Argentina (AR), 4: Peru (PE), 6: Venezuela (VE).</li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p>Version: 1.1.2, 11.1.2025. Update description. Update Reference. Normalize filenames. Assign noise uses the cluster label '-1' instead of removing them. Update plots. Additional removal of wrongly copied graphs.</p> <h3>Reference</h3> <p>Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde. 2022. <a href="https://aclanthology.org/2022.lrec-1.278/">DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish</a>. Proceedings of the Thirteenth Language Resources and Evaluation Conference.</p>

opencc-by-nd-4.0Sep 2021View details →
zenodo36/100

DWUG ES: Diachronic Word Usage Graphs for Spanish

<p>This data collection contains diachronic Word Usage Graphs (WUGs) for Spanish. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Please find more information on the provided data in the papers referenced below.</p> <p>The annotation was funded by</p> <ul> <li>ANID FONDECYT grant 11200290, U-Inicia VID Project UI-004/20,</li> <li>ANID - Millennium Science Initiative Program - Code ICN17 002 and</li> <li>SemRel Group (DFG Grants SCHU 2580/1 and SCHU 2580/2).</li> </ul> <p>Version: 4.0.2, 7.1.2025. <strong>Full data</strong>. Quoting issues in uses resolved. Target word and target sentence indices corrected. One corrected context for word 'metro'. Judgments anonymized. Annotator 'gecsa' removed. Issues with special characters in filenames resolved. Additional removal of wrongly copied graphs.</p> <h3>Reference</h3> <p>Frank D. Zamora-Reina, Felipe Bravo-Marquez, Dominik Schlechtweg. 2022. <a href="https://aclanthology.org/2022.lchange-1.16/">LSCDiscovery: A shared task on semantic change discovery and detection in Spanish</a>. In Proceedings of the 3rd International Workshop on Computational Approaches to Historical Language Change. Association for Computational Linguistics.</p> <p>Dominik Schlechtweg, Tejaswi Choppa, Wei Zhao, Michael Roth. 2025. <a href="https://aclanthology.org/2025.comedi-1.4/">The CoMeDi Shared Task: Median Judgment Classification &amp; Mean Disagreement Ranking with Ordinal Word-in-Context Judgments</a>. In Proceedings of the 1st Workshop on Context and Meaning--Navigating Disagreements in NLP Annotations.</p>

opencc-by-nd-4.0Mar 2022View details →
zenodo36/100

Support for a novel, simple method for calculating word frequency of output on language production tasks

<p>Pre-recorded presentation for the&nbsp;International Workshop on Language Production 2021</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Replication package for: Taking the Fed at its Word: A New Approach to Estimating Central Bank Objectives using Text Analysis

<p>This zip file contains&nbsp;all of the programs and data necessary to replicate the results, tables, and figures&nbsp;<br> in the following paper:</p> <p>Shapiro, Adam H., and Daniel J. Wilson (2021). &quot;Taking the Fed at its Word: A New Approach to Estimating Central Bank Objectives using Text Analysis,&quot; forthcoming at Review of Economic Studies.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

RefWUG: Diachronic Reference Word Usage Graphs for German

<p>This data collection contains diachronic Word Usage Graphs (WUGs) for German created with reference use sampling. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Please find more information on the provided data in the paper referenced below.</p> <p>Version: 1.1.0, 15.12.2021.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg and Sabine Schulte im Walde. submitted. Clustering Word Usage Graphs: A Flexible Framework to Measure Changes in Contextual Word Meaning.</p>

opencc-by-nd-4.0Sep 2021View details →
zenodo36/100

Armenian: Word order and information structure

<ul> <li> <p><strong>Basic word order: arguments in favour of OV and VO</strong></p> </li> <li> <p><strong>What is information structure: topic, focus</strong></p> </li> <li> <p><strong>Usual position of sentential stress and (EA) auxiliary</strong></p> </li> <li> <p><strong>&lsquo;Topical&rsquo; constituents (agent/experiencer subject, specific object): positions and properties</strong></p> </li> <li> <p><strong>Preverbal and postverbal focus</strong></p> </li> <li> <p><strong>Head-final vs. head-initial constituents</strong></p> </li> <li> <p><strong>Factors favouring VO order (cognitive, areal, typological)</strong></p> </li> </ul> <p>&nbsp;</p> <p>This lecture is part of the lecture series:</p> <p><em>Glottoth&egrave;que: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online </em>(electronic resource). Bamberg, Cambridge, G&ouml;ttingen, Moskow, Nicosia, Paris: LACIM network, at https://spw.uni-goettingen.de/projects/lacim/, edited by Christiane Bulut, Ana&iuml;d Donab&eacute;dian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

The SIAEW Corpus of Spanish Iso-Accented English Words

<p>The SIAEW Corpus (Spanish Iso-Accented English Words) is a collection of monosyllabic English words in which one segment (the &#39;target&#39;) is replaced with its Spanish-accented counterpart, at one of 5 gradations of accentedness. Steps are equally-spaced in accentedness as judged by native listeners. The procedure used to generate the Corpus is described in detail in P&eacute;rez Ram&oacute;n, Garc&iacute;a Lecumberri and Cooke (submitted to Interspeech 2022); for a preprint contact rperez.ram@gmail.com.&nbsp;</p> <p>Please refer to the document SIAEW.pdf for a longer description of the SIAEW Corpus contents.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Webis-Context-sensitive-Word-Search-Queries-2022

<pre>This is the dataset created for Language Models as Context-sensitive Word Search Engines at the In2Writing workshop at ACL22. </pre> <p>&nbsp;</p> <p><strong>Cite</strong>&nbsp;</p> <pre><code>@inproceedings{wiegmann:2022, title = "Language Models as Context-sensitive Word Search Engines", author = "Wiegmann, Matti and V{\"{o}}lske, Michael and Potthast, Martin and Stein, Benno", booktitle = "Proceedings of the 1st Workshop on Intelligent and Interactive Writing Assistants (In2Writing 2022)", month = may, year = "2022", address = "Online", publisher = "Association for Computational Linguistics",</code></pre> <p><strong>Datasets</strong></p> <p>This repository contains two datasets with word search queries. Each&nbsp;word search query consists of a token n-gram with one&nbsp;wildcard token ([MASK]). The answers to each query are the most likely token to replace the mask. All queries originate from wikitext-103 and CLOTH, the respected source is annotated for each query.</p> <p>The <em>original-token</em> dataset lists exactly one top answer for each query. The&nbsp;<em>ranked-answers</em> dataset lists multiple, sorted answers in three relevance categories, where 3 is the most relevant. Please refer to the citation for more details.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Frequency dataset for "Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch."

<p>The dataset is structured according to:</p> <ul> <li>the lexical field (CLOTHING, TRAFFIC, IT, and EMOTION);</li> <li>the part of speech (noun or adjective);</li> <li>the corpus;</li> <li>the concept;</li> <li>the term.</li> </ul> <p>It first gives the absolute frequency as found in the corpus and also after it was disambiguated. The concept frequency and relative frequency is calculated based on the &quot;disambiguated&quot; absolute frequency.</p> <p>More information can be found in this&nbsp;dissertation:</p> <p>Daems, Jocelyne. 2022.&nbsp;<em>Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch. </em>KU Leuven.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Modeling Word Importance in Conversational Transcripts from the Perspective of Deaf and Hard of Hearing Viewers

<ol> <li> <p><strong>MaskedPOSAugmentedData.csv</strong></p> </li> </ol> <p><strong>This file contains word embeddings of 6659 tokens augmented with POS tagging and word importance score. The embedding size for each token is a one dimensional vector of length 768 by 1. Due to augmenting POS tagging, the new feature vector becomes a size of 769 by 1. So now the total feature matrix size is 6659 by 769. In the dataset the final column represents the word importance score.&nbsp;</strong></p> <p><br> &nbsp;</p> <ol> <li> <p><strong>MaskedSentenceWithImportanceTag.csv</strong></p> </li> </ol> <p><strong>This dataset contains the newly generated sentence using standard Masking technique and their corresponding token-wise importance score.</strong><br> &nbsp;</p> <ol> <li> <p><strong>Data cleaning procedure</strong></p> </li> </ol> <p><strong>Using several steps, we have cleaned the &ldquo;switchboard corpus&rdquo; that we are using for this particular study. Here are the steps we followed:</strong></p> <p><strong>&ndash; Convert all the letter into lower case</strong></p> <p><strong>&ndash; Punctuation has be removed</strong></p> <p><strong>&ndash; Numbers or Cardinal values have been converted to text representation&nbsp;</strong></p> <p><strong>&ndash; We use lemmatization techniques to eradicate the possibility of multiple versions of the same word token.</strong><br> <br> &nbsp;</p> <p>&nbsp;</p> <ol> <li> <p><strong>Annotation Instruction</strong></p> </li> </ol> <p><strong>If researchers intend to produce additional masked text for further the size of the dataset, we recommend using the method we described in the paper.</strong></p> <p><strong>However, if someone wants to manually annotate the words or tokens in a dataset, it is important to remember that annotators need to put a score on each word based on its relative importance within a sentence or the information available around that text. It may not be appropriate to allow annotators to read the whole document first and then conduct annotation. Also using multiple annotators is recommended otherwise interrater agreement may not work well.</strong></p> <p>&nbsp;</p> <ol> <li> <p><strong>Test and training data splitting</strong></p> </li> </ol> <p><strong>As described in the paper, during our experiment, we have retained 10% of data as test data and use 90% data to train the models. For proper replication, we recommend reading our paper thoroughly.&nbsp;</strong></p> <p>&nbsp;</p> <ol> <li> <p><strong>We understand that the dataset size is relatively small for training and testing a model that might be reliable. It is important to remember that data annotation with this particular user group might be challenging.&nbsp;</strong></p> </li> </ol> <p>&nbsp;</p> <p><strong>N.B: While augmenting the new feature within the dataset, a portion of data has been excluded during the data curation phase. For validation, we have replicated the previous models so that we can measure how the dataset can perform with this newly formed dataset.&nbsp;</strong></p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

SignBD-Word: Video-Based Bangla Word-Level Sign Language Dataset

<p>Bangla sign language (BdSL) is a complete and independent natural sign language with its own linguistic characteristics. While there exists video datasets for well-known sign languages, there is currently no available dataset for word-level BdSL. In this study, we present a video-based word-level dataset for Bangla sign language, called SignBD-Word, consisting of 6000 sign videos representing 200 unique words. The dataset includes full and upper-body views of the signers, along with 2D body pose information. This dataset can also be used as a benchmark for testing sign video classification algorithms.<br><br>Official Train Test Spllit (for both RGB and bodypose) can be found from the following link:&nbsp;<br>https://sites.google.com/view/signbd-word/dataset<br><br>This dataset is part of the following paper:<br>A. Sams, A. H. Akash and S. M. M. Rahman, "SignBD-Word: Video-Based Bangla Word-Level Sign Language and Pose Translation," 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, 2023, pp. 1-7, doi: 10.1109/ICCCNT56998.2023.10306914.<br><br>Download the corresponding paper from this link:<br>https://asnsams.github.io/Publications.html</p>

opencc-by-sa-4.0Jun 2022View details →
zenodo36/100

Obesity and acute stress modulate appetite and neural responses in food word reactivity task

<p>Data used in &quot;Carnell S, Benson L, Papantoni A, Chen L, Huo Y, Wang Z, Peterson BS, Geliebter A. Obesity and acute stress modulate appetite and neural responses in food word reactivity task. PloS one&quot;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Database of Spoken Multisyllabic English Words (SMEW)

<p>The database of Spoken Multisyllabic English Words (SMEW) contains 4940 words that vary in syllable length from one to five.&nbsp; There are 1027 one-syllable words, 1086 two-syllable words, 982 three-syllable words, 884 four-syllable words, and 961 five-syllable words.&nbsp; Individual words in each number of syllable group were generated from the English Lexicon Project website (<a href="http://elexicon.wustl.edu/">http://elexicon.wustl.edu/</a>) and recorded by a male native speaker of the American English.&nbsp; In addition to the zip file containing all words as individual .wav files, there is an excel file with a listing of all words in each number of syllable group and their lexical characteristics.&nbsp; A PRAAT script used to estimate word onset and offset within a .wav file and word duration is also included.&nbsp; &nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo36/100

Week 3 - How to write a resume CV with Microsoft Word

<p>Video for Week 3 ICT Tourism &amp; Hospitality</p> <p>Step by step guide on how to write a resume CV with Microsoft Word</p>

opencc-by-4.0Aug 2013View details →
zenodo36/100

Munji audio word list

<p>On April 29, 2013, I documented Munji of Hetaowan 核桃湾, Yongli 永利村, Donggan Township 董干镇, Malipo County, Yunnan, China. Recordings were collected from two Munji speakers, who were father and son. The father knew a few words that the son did not know.</p> <p>Munji is spoken by the Flowery Yi of Donggan Town, Malipo County, Yunnan. It is closely related to Mantsi of Vietnam, as the similar-sounding autonyms would imply.</p> <p>Munji data was collected by myself in April 2013 in Hetaowan 核桃湾, Yongli Village 永利村, Donggan Township 董干镇, Malipo County, Yunnan, while Mondzi is from YYFC (1983), as cited in the appendix of Lama (2012). Mongi data is from the appendix of Chen Kang (2010).</p> <p>My Munji informants in Yongli Village 永利村, Donggan Township 董干镇, Malipo County reported that their ancestors had come from the Wuhua Mountains 五华山 of Kunming.</p>

opencc-by-4.0Dec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record