Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

116

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

116 results for “Natural Language”

Learn how ShareScore rates datasets ↗
zenodo40/100

DeepThought DPR: Distributed peer review enhanced with natural language processing and machine learning - Dataset I

<p>This is the anonymized dataset obtained from the DPR Experiment run at ESO in Fall 2018. If this dataset is used both this DOI as well as the main paper need to be cited.&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Bangla Natural Language Image to Text (BNLIT)

<p>We represented a new Bangla dataset with a Hybrid Recurrent Neural Network model which generated Bangla natural language description of images. This dataset achieved by a large number of images with classification and containing natural language process of images. We conducted experiments on our self-made Bangla Natural Language Image to Text (BNLIT) dataset. Our dataset contained 8,743 images. We made this dataset using Bangladesh perspective images. We used one annotation for each image. In our repository, we added two types of pre-processed data which is 224 &times; 224 and 500 &times; 375 respectively alongside annotations of full dataset. We also added CNN features file of whole dataset in our repository which is features.pkl.</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Conditional Driving from Natural Language Instruction

<p>A pre-trained model of the paper &quot;Conditional Driving from Natural Language Instructions,&quot; CoRL 2019.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

MTL-QA : A dataset and multi-task learning approach for knowledge graph and natural language question answering

<p>The dataset used for this project is created by enhancing the publicly available MetaQA (Movie Text Audio QA), which is primarily a KGQA dataset pertaining to movies, an extension of WikiMovies. This involves questions requiring 1, 2, and 3 hops which can be answered by using a MetaQA Knowledge Graph. The questions are available in text and audio format. The text has vanilla (original) and its paraphrased version, and is called ntm.&nbsp;</p> <p>In order to develop a dataset to support NLQA, a series of dataset augmentation steps has been performed.</p> <p>The dataset consists of natural language questions and a tagged topic entity as ground truth. This topic entity is used to retrieve textual information related to the question from Wikipedia. The introduction section of the entity&#39;s page is used as the context that is required for NLQA. Hence, this dataset has information related to both KGQA and NLQA. Certain preliminary checks and validations are done to only retain those data samples whose context can be used to answer a given question.</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing

<p><strong>Objectives</strong></p> <p>There is much interest in utilizing clinical data for developing prediction models for Alzheimer disease (AD) risk, progression, and outcomes. Existing studies have mostly utilized curated research registries, image analysis, and structured Electronic Health Record (EHR) data. However, much critical information resides in relatively inaccessible unstructured clinical notes within the EHR.</p> <p><strong>Materials and Methods</strong></p> <p>We developed a natural language processing (NLP)-based pipeline to extract AD-related clinical phenotypes, documenting strategies for success and assessing the utility of mining unstructured clinical notes. We evaluated the pipeline against gold-standard manual annotations performed by two clinical dementia experts for AD-related clinical phenotypes including medical comorbidities, biomarkers, neurobehavioral test scores, behavioral indicators of cognitive decline, family history, and neuroimaging findings.</p> <p><strong>Results</strong></p> <p>Documentation rates for each phenotype varied in the structured versus unstructured EHR. Inter-annotator agreement was high (Cohen's kappa = 0.72–1) and positively correlated with the NLP-based phenotype extraction pipeline's performance (average F1-score = 0.65-0.99) for each phenotype.</p> <p><strong>Discussion</strong></p> <p>We developed an automated NLP-based pipeline to extract informative phenotypes that may improve the performance of eventual machine-learning predictive models for AD. In the process, we examined documentation practices for each phenotype relevant to the care of AD patients and identified factors for success.</p> <p><strong>Conclusion</strong></p> <p>Success of our NLP-based phenotype extraction pipeline depended on domain-specific knowledge and focus on a specific clinical domain instead of maximizing generalizability. </p>

opencc-zeroFeb 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Intermediary Result

<p>The intermediary result of the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Category Confusion Matrix

<p>Resulting category confusion matrix&nbsp;for the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Result

<p>The result for the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo40/100

NLAS-multi: A Multilingual Corpus of Automatically Generated Natural Language Argumentation Schemes

<p>The multilingual corpus of natural language argumentation schemes (NLAS-multi)&nbsp;consists of 3,810 natural language argumentation schemes&nbsp;of which 1,893 are in English and 1,917 in Spanish. It has a total of 253,516 words distributed in 118,493 words in English and 135,023 words in Spanish.&nbsp;In terms of inferences, our corpus has a total of 7,964 (3,949 in English and 4,015 in Spanish).&nbsp;Furthermore, the NLAS-multi&nbsp;corpus contains a total of 23,781 conflict relations between arguments in the same topic.</p>

opencc-by-nc-sa-4.0Sep 2023View details →
dryad40/100

Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing

Open the record for dataset details and reuse information.

publicFeb 2023View details →
zenodo36/100

Conclusion Stability for Natural Language Based Mining of Design Discussions: Dataset

<p>All the data from raw to processed are included in this package that was used during the experiment.</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

Natural-language species descriptions of Anasillomos as exported from Lucid Builder 3.5 in XML SDD format

Natural-language species descriptions of Anasillomos as exported from Lucid Builder 3.5 in XML SDD format.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Data for: Speech naturalness detection and language representation in the dog brain

<p>Abstract<br> Family dogs are exposed to a continuous flow of human speech throughout their lives. However, the extent of their abilities in speech perception is unknown. Here, we used functional magnetic resonance imaging (fMRI) to test speech detection and language representation in the dog brain. Dogs (n = 18) listened to natural speech and scrambled speech in a familiar and an unfamiliar language. Speech scrambling distorts auditory regularities specific to speech and to a given language, but keeps spectral voice cues intact. We hypothesized that if dogs can extract auditory regularities of speech, and of a familiar language, then there will be distinct patterns of brain activity for natural speech vs. scrambled speech, and also for familiar vs. unfamiliar language. Using multivoxel pattern analysis (MVPA) we found that bilateral auditory cortical regions represented natural speech and scrambled speech differently; with a better classifier performance in longer-headed dogs in a right auditory region. This neural capacity for speech detection was not based on preferential processing for speech but rather on sensitivity to sound naturalness.<br> Furthermore, in case of natural speech, distinct activity patterns were found for the two languages in the secondary auditory cortex and in the precruciate gyrus; with a greater difference in responses to the familiar and unfamiliar languages in older dogs, indicating a role for the amount of language exposure. No regions represented differently the scrambled versions of the two languages, suggesting that the activity difference between languages in natural speech reflected sensitivity to language-specific regularities rather than to spectral voice cues. These findings suggest that separate cortical regions support speech naturalness detection and language representation in the dog brain.<br> <br> This dataset contains</p> <ul> <li>Raw data (four functional runs and Matlab logs n = 18)</li> <li>Dog brain&nbsp;template</li> <li>Stimuli (natural and scrambled speech in Hungarian and Spanish)&nbsp;</li> <li>MVPA main results maps (Speech detection and Language discrimination, n = 18)</li> <li>GLM All sounds &gt; Silence contrast (n = 18)&nbsp;</li> </ul>

opencc-by-3.0Jan 2022View details →
zenodo36/100

Datasets and code for "Mapping the plague through natural language processing"

<p>This project investigates the performance of various NLP libraries and geocoding services for the semi-automated generation of quantitative datasets from narrative texts. We provide the original files, several intermediate data products as well as the final plague datasets.&nbsp;Please note that some of the steps in this process were done manually, thus some of the scripts cannot be run completely.&nbsp;</p> <p>This work is based on two plague treatises:</p> <p>- Sticker, G. 1908 <em>Abhandlungen aus der Seuchengeschichte und Seuchenlehre. Band 1: Die Pest</em>. Giessen, A. T&ouml;pelmann.</p> <p>- Biraben, J.-N. 1975 <em>Les hommes et la peste en France et dans les pays europ&eacute;ens et m&eacute;diterran&eacute;ens</em>. Paris, Mouton.</p> <p>The final geocoded, plague datasets are:</p> <p><strong>- plague_sticker_v1.csv</strong></p> <p><strong>- plague_biraben_v1.csv</strong></p> <p>A data dictionary is available as</p> <p><strong>plague_datadict.xlsx</strong></p> <p>Other files:</p> <table> <tbody> <tr> <td>file name</td> <td>content</td> </tr> <tr> <td>sticker_OCR_orig.txt</td> <td>Original OCR text</td> </tr> <tr> <td>sticker_OCR.txt</td> <td>Original OCR text without parenthesis (author names)</td> </tr> <tr> <td>sticker_textprep.rds</td> <td>Original OCR text with further preparations</td> </tr> <tr> <td>sticker_goldstandard_annotated_1.tsv</td> <td>manual annotations file 1</td> </tr> <tr> <td>sticker_goldstandard_annotated_2.tsv</td> <td>manual annotations file 2</td> </tr> <tr> <td>sticker_goldstandard_annotated_consensus.tsv</td> <td>consenus annotation file</td> </tr> <tr> <td>sticker_standard_toponyms.csv</td> <td>Gold standard for toponym recognition. Contains the tokenization, the start/end character respective to the OCR text (orig and without parenthesis) and whether a token is a location or other</td> </tr> <tr> <td>sticker_comparison_NER.rds</td> <td>Comparison of NER performance</td> </tr> <tr> <td>sticker_comparison_geocoding.rds</td> <td>Comparison of Geocoding performance</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-nc-4.0Apr 2021View details →
zenodo36/100

Deciphering microbial gene function using natural language processing

<p>Dataset for the support of a journal publication.&nbsp;</p> <p>The data include both computed models&nbsp;presented in the paper and all data used for analysis.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Supplementary Material for Requirements documentation containing natural language: A Systematic Tertiary Literature Review

<p>Context: Requirements documentation in natural language&nbsp;has diverse artifacts, but few studies address their suitability to types<br>of requirements or ease of communication.</p> <p>Methods: We conducted a&nbsp;systematic tertiary literature review (STLR) and identified 22 relevant&nbsp;review papers that address natural language artifacts used by practitioners to document software requirements. We also investigated which&nbsp;types of requirements are addressed by artifacts and if there are guidelines for each.</p> <p>Results: A variety of artifacts used for this purpose were&nbsp;identified, of which the most referenced in the literature were diagrams,<br>use cases, conceptual models, user stories, and prototypes. The analysis highlighted that artifacts are applied differently to functional and&nbsp;non-functional requirements. In general, diagrams, use cases, scenarios,&nbsp;and prototypes can be used for both types of requirements, depending&nbsp;on the content (usability, security, etc.). However, user stories and derived artifacts are more recommended for functional requirements and&nbsp;have limitations for non-functional requirements.</p> <p>Conclusion: Furthermore, the study explored different guidelines, structures, and formats used in documentation artifacts, reflecting the diversity in requirements documentation practices in software projects.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language

<p>Autonomous tuning of particle accelerators is an active and challenging field of research with the goal of enabling novel accelerator technologies cutting-edge high-impact applications, such as physics discovery, cancer research and material sciences. A key challenge with autonomous accelerator tuning remains that the most capable algorithms require an expert in optimisation, machine learning or a similar field to implement the algorithm for every new tuning task. In this work, we propose the use of large language models (LLMs) to tune particle accelerators. We demonstrate on a proof-of-principle example the ability of LLMs to successfully and autonomously tune a particle accelerator subsystem based on nothing more than a natural language prompt from the operator, and compare the performance of our LLM-based solution to state-of-the-art optimisation algorithms, such as Bayesian optimisation (BO) and reinforcement learning-trained optimisation (RLO). In doing so, we also show how LLMs can perform numerical optimisation of a highly non-linear real-world objective function. Ultimately, this work represents yet another complex task that LLMs are capable of solving and promises to help accelerate the deployment of autonomous tuning algorithms to the day-to-day operations of particle accelerators.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Natural Language Grounding in TeMoto

<p>An illustrative overview of natural language grounding, via action decomposition, in TeMoto framework.</p> <p><em>a) The operator gives an instruction in NL. </em><br> <em>b) The instruction is mapped to an action tree.<br> c) Actions execute the developer defined code and request resources from the Management Layer.</em></p>

opencc-by-sa-4.0Dec 2017View details →
zenodo36/100

Comparing the Use of Research Resource Identifiers and Natural Language Processing for Citation of Databases, Software and Other Digital Artifacts

<p><strong>The Research Resource Identifier was introduced in biomedicine in 2014 to more precisely identify the reagents and tools used in published biomedical research and to track use of tools across the breadth of the biomedical literature. The current RRID specification covers key biological and digital resources. Authors are instructed to include an RRID after the first mention of any resource used. RRIDs are designed to be easy to find using &nbsp;a full text search search engine. </strong></p> <p><strong>The published data sets were used in our comparative study where comparing the output of our RRID curation workflow with the outputs of automated text mining systems that have been used to identify mentions of resources in the text of publications. All files in tab-separated format (tsv). </strong></p> <p><strong>Scibot.tsv: Records of the RRID curation workflow using SciBot. </strong></p> <p>Each record shows that a resource RRID was identified in paper PMID with curator tags (Tag1, Tag2, both optional)</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag1: Curator tags (optional)</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag2: Additional curator tags (optional)</p> <p><strong>rdwsorted.tsv: Records of the output from RDW, a text mining software. </strong></p> <p>RDW identifies mentions of research resources in papers. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p><strong>rridbyrdw05282019.tsv:&nbsp;Records of the output of the RRID-by-RDW in RDW. </strong></p> <p>RRID-by-RDW is a component in RDW that identifies mentions of research resources in papers by matching patterns of RRID specifications. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Context: Snippet where the RRID was found</p> <p><strong>resource_metadata20190418.tsv: Metadata of RRIDs</strong></p> <p>This file contains metadata of resources and their RRIDs. See file header for column definitions.</p> <p><strong>RRIDCUR-definitions.tsv: Definitions of curator tags used in Scibot.tsv.</strong></p> <p><strong>&nbsp;&nbsp; </strong>tag: Tag name</p> <p>&nbsp;&nbsp;&nbsp; definition: Definition of the tag</p>

openbsd-3-clause-clearJun 2019View details →
zenodo36/100

Yongning Na for Natural Language Processing: a single-speaker audio corpus with transcriptions

<p><em>(fran&ccedil;ais ci-dessous)</em></p> <p>This archive contains a dataset (audio files and transcriptions) of a minority language, Yongning Na (Glottocode: yong1288; closest iso 639-3 code: nru). The archive contains a subset of the Na corpus of the Pangloss Collection: it is a single-speaker corpus, consisting of all the audio resources transcribed, for the main speaker of this corpus (Ms. LATAMI Dashilame).<br> The corpus is versioned, so that the experiments carried out on these resources (for linguistic research or for Natural Language Processing) are fully reproducible. All relevant information is contained in YAML files (.yml extension; one in French, one in English).<br> The data sub-folder contains the converted and demultiplexed audio files, as well as the annotations associated with each channel of the audio files.<br> The summary files contain, among other things, the list of graphemes used in the language (complex graphemes are particularly important), as well as information on the various resources (audio and annotations), such as their identifiers (DOIs) and links to the original files.<br> From a computational point of view, the list of DOIs of the audios and annotations described in this YAML file is sufficient to generate this corpus at a given time. A corpus like the present one can be viewed as the version, at a given time, of a set of documents in the Pangloss collection: a corpus as it stands at a precise version.</p> <p>Further information is available from&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p> <p>---------------</p> <p>Cette archive contient un jeu de donn&eacute;es (audios et transcriptions) d&rsquo;une langue &agrave; tradition orale, le na de Yongning &nbsp;(Glottocode: yong1288; code iso 639-3 le plus proche : nru). L&rsquo;archive contient un sous-ensemble du corpus na de la collection Pangloss : c&rsquo;est un corpus monolocuteur, constitu&eacute; de l&rsquo;int&eacute;gralit&eacute; des ressources audio transcrites pour la locutrice principale de ce corpus (Mme LATAMI Dashilame).<br> Le corpus est versionn&eacute;, de sorte que les exp&eacute;riences men&eacute;es sur ces ressources (pour la linguistique ou pour le Traitement automatique des langues) soient reproductibles de fa&ccedil;on exacte (en pensant bien &agrave; joindre l&rsquo;algorithme : param&egrave;tres, r&eacute;partitions des fichiers dans les diff&eacute;rents ensembles, etc.). Toutes les informations pertinentes se trouvent dans les fichiers YAML (extension .yml ; un en fran&ccedil;ais, un autre en anglais).<br> Le sous-dossier des donn&eacute;es contient d&rsquo;une part les audios convertis et d&eacute;multiplex&eacute;s et d&rsquo;autre part les annotations associ&eacute;es &agrave; chaque canal desdits audios.<br> Les fichiers r&eacute;capitulatifs contiennent notamment la liste des graph&egrave;mes utilis&eacute;s dans cette langue (les graph&egrave;mes complexes sont particuli&egrave;rement importants), ainsi que des informations sur les diff&eacute;rentes ressources (audios et annotations), comme les identifiants (DOI), les liens vers les fichiers originaux, etc.<br> Au plan informatique, la liste des identifiants DOI des audios et annotations d&eacute;crits dans ce fichier YAML suffit pour g&eacute;n&eacute;rer ce corpus &agrave; un instant t. Un corpus comme celui-ci peut &ecirc;tre vu comme la version &agrave; l&rsquo;instant t d&rsquo;un ensemble de documents de la collection Pangloss : un corpus arr&ecirc;t&eacute; &agrave; une version pr&eacute;cise.<br> Pour plus de pr&eacute;cisions :&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p>

openAug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record