Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

55

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

55 results for “natural language processing”

Learn how ShareScore rates datasets ↗
zenodo44/100

Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification

<p>Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as &quot;black boxes&quot;, providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD.</p> <p>This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19

<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities.&nbsp;</span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764&nbsp;398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p>&nbsp;</p> <p>&nbsp;</p>

openJun 2022View details →
zenodo40/100

Technical Debt Classification in Issue Trackers using Natural Language Processing based on Transformers

<p>In order to ensure transparency and reproducibility, we have&nbsp;made everything available publicly here, including the Code, Models, Datasets and more. All the files and their functionality used in this paper are explained clearly in the <strong>README.md</strong> file.</p> <p>Background: &nbsp;Technical Debt (TD) needs to be controlled and tracked during software development. Support to automatically track TD in issue trackers is limited.&nbsp;</p> <p>Aim: We explore the usage of a large dataset of developer-labeled TD issues in combination with cutting-edge Natural Language Processing (NLP) approaches to automatically classify TD in issue trackers.</p> <p>Method: &nbsp;We mine and analyze more than 160GB of textual data from GitHub projects, collecting over 55,600 TD issues and consolidating them into a large dataset (GTD dataset). We use such datasets to train and test Transformer ML models. Then we test the model&#39;s&nbsp;generalization ability by testing them on six unseen projects. Finally, we re-train the models including part of the TD issues from the target project to test their adaptability.&nbsp;</p> <p>Results and Conclusion: (i) We create and release the GTD dataset, a comprehensive dataset including TD issues from 6,401 public repositories with various contexts; (ii) By training Transformers using the GTD dataset, we achieve performance metrics that are promising; (iii) Our results are a significant step forward towards supporting the automatic classification of TD in issue trackers, especially when the models are adapted to the context of unseen projects after fine-tuning.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

DeepThought DPR: Distributed peer review enhanced with natural language processing and machine learning - Dataset I

<p>This is the anonymized dataset obtained from the DPR Experiment run at ESO in Fall 2018. If this dataset is used both this DOI as well as the main paper need to be cited.&nbsp;</p>

opencc-by-4.0Apr 2020View details →
dryad40/100

Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing

<p><strong>Objectives</strong></p> <p>There is much interest in utilizing clinical data for developing prediction models for Alzheimer disease (AD) risk, progression, and outcomes. Existing studies have mostly utilized curated research registries, image analysis, and structured Electronic Health Record (EHR) data. However, much critical information resides in relatively inaccessible unstructured clinical notes within the EHR.</p> <p><strong>Materials and Methods</strong></p> <p>We developed a natural language processing (NLP)-based pipeline to extract AD-related clinical phenotypes, documenting strategies for success and assessing the utility of mining unstructured clinical notes. We evaluated the pipeline against gold-standard manual annotations performed by two clinical dementia experts for AD-related clinical phenotypes including medical comorbidities, biomarkers, neurobehavioral test scores, behavioral indicators of cognitive decline, family history, and neuroimaging findings.</p> <p><strong>Results</strong></p> <p>Documentation rates for each phenotype varied in the structured versus unstructured EHR. Inter-annotator agreement was high (Cohen's kappa = 0.72–1) and positively correlated with the NLP-based phenotype extraction pipeline's performance (average F1-score = 0.65-0.99) for each phenotype.</p> <p><strong>Discussion</strong></p> <p>We developed an automated NLP-based pipeline to extract informative phenotypes that may improve the performance of eventual machine-learning predictive models for AD. In the process, we examined documentation practices for each phenotype relevant to the care of AD patients and identified factors for success.</p> <p><strong>Conclusion</strong></p> <p>Success of our NLP-based phenotype extraction pipeline depended on domain-specific knowledge and focus on a specific clinical domain instead of maximizing generalizability. </p>

opencc-zeroFeb 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Intermediary Result

<p>The intermediary result of the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Category Confusion Matrix

<p>Resulting category confusion matrix&nbsp;for the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo40/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Result

<p>The result for the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
dryad40/100

Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing

Open the record for dataset details and reuse information.

publicFeb 2023View details →
zenodo36/100

Datasets and code for "Mapping the plague through natural language processing"

<p>This project investigates the performance of various NLP libraries and geocoding services for the semi-automated generation of quantitative datasets from narrative texts. We provide the original files, several intermediate data products as well as the final plague datasets.&nbsp;Please note that some of the steps in this process were done manually, thus some of the scripts cannot be run completely.&nbsp;</p> <p>This work is based on two plague treatises:</p> <p>- Sticker, G. 1908 <em>Abhandlungen aus der Seuchengeschichte und Seuchenlehre. Band 1: Die Pest</em>. Giessen, A. T&ouml;pelmann.</p> <p>- Biraben, J.-N. 1975 <em>Les hommes et la peste en France et dans les pays europ&eacute;ens et m&eacute;diterran&eacute;ens</em>. Paris, Mouton.</p> <p>The final geocoded, plague datasets are:</p> <p><strong>- plague_sticker_v1.csv</strong></p> <p><strong>- plague_biraben_v1.csv</strong></p> <p>A data dictionary is available as</p> <p><strong>plague_datadict.xlsx</strong></p> <p>Other files:</p> <table> <tbody> <tr> <td>file name</td> <td>content</td> </tr> <tr> <td>sticker_OCR_orig.txt</td> <td>Original OCR text</td> </tr> <tr> <td>sticker_OCR.txt</td> <td>Original OCR text without parenthesis (author names)</td> </tr> <tr> <td>sticker_textprep.rds</td> <td>Original OCR text with further preparations</td> </tr> <tr> <td>sticker_goldstandard_annotated_1.tsv</td> <td>manual annotations file 1</td> </tr> <tr> <td>sticker_goldstandard_annotated_2.tsv</td> <td>manual annotations file 2</td> </tr> <tr> <td>sticker_goldstandard_annotated_consensus.tsv</td> <td>consenus annotation file</td> </tr> <tr> <td>sticker_standard_toponyms.csv</td> <td>Gold standard for toponym recognition. Contains the tokenization, the start/end character respective to the OCR text (orig and without parenthesis) and whether a token is a location or other</td> </tr> <tr> <td>sticker_comparison_NER.rds</td> <td>Comparison of NER performance</td> </tr> <tr> <td>sticker_comparison_geocoding.rds</td> <td>Comparison of Geocoding performance</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-nc-4.0Apr 2021View details →
zenodo36/100

Deciphering microbial gene function using natural language processing

<p>Dataset for the support of a journal publication.&nbsp;</p> <p>The data include both computed models&nbsp;presented in the paper and all data used for analysis.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Comparing the Use of Research Resource Identifiers and Natural Language Processing for Citation of Databases, Software and Other Digital Artifacts

<p><strong>The Research Resource Identifier was introduced in biomedicine in 2014 to more precisely identify the reagents and tools used in published biomedical research and to track use of tools across the breadth of the biomedical literature. The current RRID specification covers key biological and digital resources. Authors are instructed to include an RRID after the first mention of any resource used. RRIDs are designed to be easy to find using &nbsp;a full text search search engine. </strong></p> <p><strong>The published data sets were used in our comparative study where comparing the output of our RRID curation workflow with the outputs of automated text mining systems that have been used to identify mentions of resources in the text of publications. All files in tab-separated format (tsv). </strong></p> <p><strong>Scibot.tsv: Records of the RRID curation workflow using SciBot. </strong></p> <p>Each record shows that a resource RRID was identified in paper PMID with curator tags (Tag1, Tag2, both optional)</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag1: Curator tags (optional)</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag2: Additional curator tags (optional)</p> <p><strong>rdwsorted.tsv: Records of the output from RDW, a text mining software. </strong></p> <p>RDW identifies mentions of research resources in papers. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p><strong>rridbyrdw05282019.tsv:&nbsp;Records of the output of the RRID-by-RDW in RDW. </strong></p> <p>RRID-by-RDW is a component in RDW that identifies mentions of research resources in papers by matching patterns of RRID specifications. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Context: Snippet where the RRID was found</p> <p><strong>resource_metadata20190418.tsv: Metadata of RRIDs</strong></p> <p>This file contains metadata of resources and their RRIDs. See file header for column definitions.</p> <p><strong>RRIDCUR-definitions.tsv: Definitions of curator tags used in Scibot.tsv.</strong></p> <p><strong>&nbsp;&nbsp; </strong>tag: Tag name</p> <p>&nbsp;&nbsp;&nbsp; definition: Definition of the tag</p>

openbsd-3-clause-clearJun 2019View details →
zenodo36/100

Yongning Na for Natural Language Processing: a single-speaker audio corpus with transcriptions

<p><em>(fran&ccedil;ais ci-dessous)</em></p> <p>This archive contains a dataset (audio files and transcriptions) of a minority language, Yongning Na (Glottocode: yong1288; closest iso 639-3 code: nru). The archive contains a subset of the Na corpus of the Pangloss Collection: it is a single-speaker corpus, consisting of all the audio resources transcribed, for the main speaker of this corpus (Ms. LATAMI Dashilame).<br> The corpus is versioned, so that the experiments carried out on these resources (for linguistic research or for Natural Language Processing) are fully reproducible. All relevant information is contained in YAML files (.yml extension; one in French, one in English).<br> The data sub-folder contains the converted and demultiplexed audio files, as well as the annotations associated with each channel of the audio files.<br> The summary files contain, among other things, the list of graphemes used in the language (complex graphemes are particularly important), as well as information on the various resources (audio and annotations), such as their identifiers (DOIs) and links to the original files.<br> From a computational point of view, the list of DOIs of the audios and annotations described in this YAML file is sufficient to generate this corpus at a given time. A corpus like the present one can be viewed as the version, at a given time, of a set of documents in the Pangloss collection: a corpus as it stands at a precise version.</p> <p>Further information is available from&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p> <p>---------------</p> <p>Cette archive contient un jeu de donn&eacute;es (audios et transcriptions) d&rsquo;une langue &agrave; tradition orale, le na de Yongning &nbsp;(Glottocode: yong1288; code iso 639-3 le plus proche : nru). L&rsquo;archive contient un sous-ensemble du corpus na de la collection Pangloss : c&rsquo;est un corpus monolocuteur, constitu&eacute; de l&rsquo;int&eacute;gralit&eacute; des ressources audio transcrites pour la locutrice principale de ce corpus (Mme LATAMI Dashilame).<br> Le corpus est versionn&eacute;, de sorte que les exp&eacute;riences men&eacute;es sur ces ressources (pour la linguistique ou pour le Traitement automatique des langues) soient reproductibles de fa&ccedil;on exacte (en pensant bien &agrave; joindre l&rsquo;algorithme : param&egrave;tres, r&eacute;partitions des fichiers dans les diff&eacute;rents ensembles, etc.). Toutes les informations pertinentes se trouvent dans les fichiers YAML (extension .yml ; un en fran&ccedil;ais, un autre en anglais).<br> Le sous-dossier des donn&eacute;es contient d&rsquo;une part les audios convertis et d&eacute;multiplex&eacute;s et d&rsquo;autre part les annotations associ&eacute;es &agrave; chaque canal desdits audios.<br> Les fichiers r&eacute;capitulatifs contiennent notamment la liste des graph&egrave;mes utilis&eacute;s dans cette langue (les graph&egrave;mes complexes sont particuli&egrave;rement importants), ainsi que des informations sur les diff&eacute;rentes ressources (audios et annotations), comme les identifiants (DOI), les liens vers les fichiers originaux, etc.<br> Au plan informatique, la liste des identifiants DOI des audios et annotations d&eacute;crits dans ce fichier YAML suffit pour g&eacute;n&eacute;rer ce corpus &agrave; un instant t. Un corpus comme celui-ci peut &ecirc;tre vu comme la version &agrave; l&rsquo;instant t d&rsquo;un ensemble de documents de la collection Pangloss : un corpus arr&ecirc;t&eacute; &agrave; une version pr&eacute;cise.<br> Pour plus de pr&eacute;cisions :&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p>

openAug 2021View details →
dryad36/100

Data from: Natural language processing systems for pathology parsing in limited data environments with uncertainty estimation

Open the record for dataset details and reuse information.

publicJul 2021View details →
zenodo32/100

Processed data for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This is the data used to reproduce the results from &quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Scatter plots for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains the test-score-vs-metric plots generated by the paper&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Generalization metrics for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains all the generalization metrics that can be used to reproduce the results of&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Rank correlation results for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains the rank correlation results from the paper&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Trends in Natural Language Processing

<p>This dataset contains the supporting parsed corpus as described in the publication: "Analyzing a Decade of Evolution: Trends in<br>Natural Language Processing"&nbsp;</p> <p>This dataset contains a single zip containing data between the years 2010 and 2022, for the conferences:</p> <ul> <li>Meeting of the Association for Computational Linguistics (ACL)</li> <li>Conference on Empirical Methods in Natural Language Processing (EMNLP)</li> <li>American Chapter of the Association for Computational Linguistics (NAACL)</li> <li>Conference on Computational Linguistics (COLING)</li> <li>International Conference on Language Resources and Evaluation (LREC)</li> <li>Conference on Computational Natural Language Learning (CoNLL)</li> <li>European Chapter of the Association for Computational Linguistics (EACL)</li> <li>International Joint Conference on Natural Language Processing (IJCNLP)</li> </ul> <p>The data inlcuded is a PDF and a JSON file for each confernce. The JSON file is constucted from using the python packages SciPDF and PyPDF2. PyPDF2 extract all text from a page and is presented in the 'full_text' field, where as the SciPDF parser utilize machine lerarning to create a smart representation of the data, which is presented in the remaining fields.</p> <p>For further details on how this dataset was generated, please see our <a href="https://github.com/ieeta-pt/nlp-trends" target="_blank" rel="noopener">GitHub</a> repository, and our paper.</p> <p>Citation:</p> <p>AWAITING PUBLICATION</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Enhancing Name Entity Recognition Through Hybrid Deep Learning in Natural Language Processing

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record