Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
55
datasets available to search
ShareScore release 0.7.1
Dataset results
55 results for “natural language processing”
Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification
<p>Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as "black boxes", providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD.</p> <p>This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.</p>
Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19
<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities. </span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764 398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p> </p> <p> </p>
Technical Debt Classification in Issue Trackers using Natural Language Processing based on Transformers
<p>In order to ensure transparency and reproducibility, we have made everything available publicly here, including the Code, Models, Datasets and more. All the files and their functionality used in this paper are explained clearly in the <strong>README.md</strong> file.</p> <p>Background: Technical Debt (TD) needs to be controlled and tracked during software development. Support to automatically track TD in issue trackers is limited. </p> <p>Aim: We explore the usage of a large dataset of developer-labeled TD issues in combination with cutting-edge Natural Language Processing (NLP) approaches to automatically classify TD in issue trackers.</p> <p>Method: We mine and analyze more than 160GB of textual data from GitHub projects, collecting over 55,600 TD issues and consolidating them into a large dataset (GTD dataset). We use such datasets to train and test Transformer ML models. Then we test the model's generalization ability by testing them on six unseen projects. Finally, we re-train the models including part of the TD issues from the target project to test their adaptability. </p> <p>Results and Conclusion: (i) We create and release the GTD dataset, a comprehensive dataset including TD issues from 6,401 public repositories with various contexts; (ii) By training Transformers using the GTD dataset, we achieve performance metrics that are promising; (iii) Our results are a significant step forward towards supporting the automatic classification of TD in issue trackers, especially when the models are adapted to the context of unseen projects after fine-tuning.</p>
DeepThought DPR: Distributed peer review enhanced with natural language processing and machine learning - Dataset I
<p>This is the anonymized dataset obtained from the DPR Experiment run at ESO in Fall 2018. If this dataset is used both this DOI as well as the main paper need to be cited. </p>
Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing
<p><strong>Objectives</strong></p> <p>There is much interest in utilizing clinical data for developing prediction models for Alzheimer disease (AD) risk, progression, and outcomes. Existing studies have mostly utilized curated research registries, image analysis, and structured Electronic Health Record (EHR) data. However, much critical information resides in relatively inaccessible unstructured clinical notes within the EHR.</p> <p><strong>Materials and Methods</strong></p> <p>We developed a natural language processing (NLP)-based pipeline to extract AD-related clinical phenotypes, documenting strategies for success and assessing the utility of mining unstructured clinical notes. We evaluated the pipeline against gold-standard manual annotations performed by two clinical dementia experts for AD-related clinical phenotypes including medical comorbidities, biomarkers, neurobehavioral test scores, behavioral indicators of cognitive decline, family history, and neuroimaging findings.</p> <p><strong>Results</strong></p> <p>Documentation rates for each phenotype varied in the structured versus unstructured EHR. Inter-annotator agreement was high (Cohen's kappa = 0.72–1) and positively correlated with the NLP-based phenotype extraction pipeline's performance (average F1-score = 0.65-0.99) for each phenotype.</p> <p><strong>Discussion</strong></p> <p>We developed an automated NLP-based pipeline to extract informative phenotypes that may improve the performance of eventual machine-learning predictive models for AD. In the process, we examined documentation practices for each phenotype relevant to the care of AD patients and identified factors for success.</p> <p><strong>Conclusion</strong></p> <p>Success of our NLP-based phenotype extraction pipeline depended on domain-specific knowledge and focus on a specific clinical domain instead of maximizing generalizability. </p>
Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Intermediary Result
<p>The intermediary result of the experiment "Evaluation of a simple score-based Natural Language Processing (NLP) algorithm".</p>
Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Category Confusion Matrix
<p>Resulting category confusion matrix for the experiment "Evaluation of a simple score-based Natural Language Processing (NLP) algorithm".</p>
Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Result
<p>The result for the experiment "Evaluation of a simple score-based Natural Language Processing (NLP) algorithm".</p>
Extraction of clinical phenotypes for Alzheimer disease dementia from clinical notes using natural language processing
Open the record for dataset details and reuse information.
Datasets and code for "Mapping the plague through natural language processing"
<p>This project investigates the performance of various NLP libraries and geocoding services for the semi-automated generation of quantitative datasets from narrative texts. We provide the original files, several intermediate data products as well as the final plague datasets. Please note that some of the steps in this process were done manually, thus some of the scripts cannot be run completely. </p> <p>This work is based on two plague treatises:</p> <p>- Sticker, G. 1908 <em>Abhandlungen aus der Seuchengeschichte und Seuchenlehre. Band 1: Die Pest</em>. Giessen, A. Töpelmann.</p> <p>- Biraben, J.-N. 1975 <em>Les hommes et la peste en France et dans les pays européens et méditerranéens</em>. Paris, Mouton.</p> <p>The final geocoded, plague datasets are:</p> <p><strong>- plague_sticker_v1.csv</strong></p> <p><strong>- plague_biraben_v1.csv</strong></p> <p>A data dictionary is available as</p> <p><strong>plague_datadict.xlsx</strong></p> <p>Other files:</p> <table> <tbody> <tr> <td>file name</td> <td>content</td> </tr> <tr> <td>sticker_OCR_orig.txt</td> <td>Original OCR text</td> </tr> <tr> <td>sticker_OCR.txt</td> <td>Original OCR text without parenthesis (author names)</td> </tr> <tr> <td>sticker_textprep.rds</td> <td>Original OCR text with further preparations</td> </tr> <tr> <td>sticker_goldstandard_annotated_1.tsv</td> <td>manual annotations file 1</td> </tr> <tr> <td>sticker_goldstandard_annotated_2.tsv</td> <td>manual annotations file 2</td> </tr> <tr> <td>sticker_goldstandard_annotated_consensus.tsv</td> <td>consenus annotation file</td> </tr> <tr> <td>sticker_standard_toponyms.csv</td> <td>Gold standard for toponym recognition. Contains the tokenization, the start/end character respective to the OCR text (orig and without parenthesis) and whether a token is a location or other</td> </tr> <tr> <td>sticker_comparison_NER.rds</td> <td>Comparison of NER performance</td> </tr> <tr> <td>sticker_comparison_geocoding.rds</td> <td>Comparison of Geocoding performance</td> </tr> </tbody> </table> <p> </p>
Deciphering microbial gene function using natural language processing
<p>Dataset for the support of a journal publication. </p> <p>The data include both computed models presented in the paper and all data used for analysis.</p>
Comparing the Use of Research Resource Identifiers and Natural Language Processing for Citation of Databases, Software and Other Digital Artifacts
<p><strong>The Research Resource Identifier was introduced in biomedicine in 2014 to more precisely identify the reagents and tools used in published biomedical research and to track use of tools across the breadth of the biomedical literature. The current RRID specification covers key biological and digital resources. Authors are instructed to include an RRID after the first mention of any resource used. RRIDs are designed to be easy to find using a full text search search engine. </strong></p> <p><strong>The published data sets were used in our comparative study where comparing the output of our RRID curation workflow with the outputs of automated text mining systems that have been used to identify mentions of resources in the text of publications. All files in tab-separated format (tsv). </strong></p> <p><strong>Scibot.tsv: Records of the RRID curation workflow using SciBot. </strong></p> <p>Each record shows that a resource RRID was identified in paper PMID with curator tags (Tag1, Tag2, both optional)</p> <p><strong> </strong>PMID: Pubmed ID</p> <p> RRID: Research Resource Identifier</p> <p> Tag1: Curator tags (optional)</p> <p> Tag2: Additional curator tags (optional)</p> <p><strong>rdwsorted.tsv: Records of the output from RDW, a text mining software. </strong></p> <p>RDW identifies mentions of research resources in papers. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong> </strong>PMID: Pubmed ID</p> <p> RRID: Research Resource Identifier</p> <p><strong>rridbyrdw05282019.tsv: Records of the output of the RRID-by-RDW in RDW. </strong></p> <p>RRID-by-RDW is a component in RDW that identifies mentions of research resources in papers by matching patterns of RRID specifications. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong> </strong>PMID: Pubmed ID</p> <p> RRID: Research Resource Identifier</p> <p> Context: Snippet where the RRID was found</p> <p><strong>resource_metadata20190418.tsv: Metadata of RRIDs</strong></p> <p>This file contains metadata of resources and their RRIDs. See file header for column definitions.</p> <p><strong>RRIDCUR-definitions.tsv: Definitions of curator tags used in Scibot.tsv.</strong></p> <p><strong> </strong>tag: Tag name</p> <p> definition: Definition of the tag</p>
Yongning Na for Natural Language Processing: a single-speaker audio corpus with transcriptions
<p><em>(français ci-dessous)</em></p> <p>This archive contains a dataset (audio files and transcriptions) of a minority language, Yongning Na (Glottocode: yong1288; closest iso 639-3 code: nru). The archive contains a subset of the Na corpus of the Pangloss Collection: it is a single-speaker corpus, consisting of all the audio resources transcribed, for the main speaker of this corpus (Ms. LATAMI Dashilame).<br> The corpus is versioned, so that the experiments carried out on these resources (for linguistic research or for Natural Language Processing) are fully reproducible. All relevant information is contained in YAML files (.yml extension; one in French, one in English).<br> The data sub-folder contains the converted and demultiplexed audio files, as well as the annotations associated with each channel of the audio files.<br> The summary files contain, among other things, the list of graphemes used in the language (complex graphemes are particularly important), as well as information on the various resources (audio and annotations), such as their identifiers (DOIs) and links to the original files.<br> From a computational point of view, the list of DOIs of the audios and annotations described in this YAML file is sufficient to generate this corpus at a given time. A corpus like the present one can be viewed as the version, at a given time, of a set of documents in the Pangloss collection: a corpus as it stands at a precise version.</p> <p>Further information is available from <a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p> <p>---------------</p> <p>Cette archive contient un jeu de données (audios et transcriptions) d’une langue à tradition orale, le na de Yongning (Glottocode: yong1288; code iso 639-3 le plus proche : nru). L’archive contient un sous-ensemble du corpus na de la collection Pangloss : c’est un corpus monolocuteur, constitué de l’intégralité des ressources audio transcrites pour la locutrice principale de ce corpus (Mme LATAMI Dashilame).<br> Le corpus est versionné, de sorte que les expériences menées sur ces ressources (pour la linguistique ou pour le Traitement automatique des langues) soient reproductibles de façon exacte (en pensant bien à joindre l’algorithme : paramètres, répartitions des fichiers dans les différents ensembles, etc.). Toutes les informations pertinentes se trouvent dans les fichiers YAML (extension .yml ; un en français, un autre en anglais).<br> Le sous-dossier des données contient d’une part les audios convertis et démultiplexés et d’autre part les annotations associées à chaque canal desdits audios.<br> Les fichiers récapitulatifs contiennent notamment la liste des graphèmes utilisés dans cette langue (les graphèmes complexes sont particulièrement importants), ainsi que des informations sur les différentes ressources (audios et annotations), comme les identifiants (DOI), les liens vers les fichiers originaux, etc.<br> Au plan informatique, la liste des identifiants DOI des audios et annotations décrits dans ce fichier YAML suffit pour générer ce corpus à un instant t. Un corpus comme celui-ci peut être vu comme la version à l’instant t d’un ensemble de documents de la collection Pangloss : un corpus arrêté à une version précise.<br> Pour plus de précisions : <a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p>
Data from: Natural language processing systems for pathology parsing in limited data environments with uncertainty estimation
Open the record for dataset details and reuse information.
Processed data for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This is the data used to reproduce the results from "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Scatter plots for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains the test-score-vs-metric plots generated by the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Generalization metrics for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains all the generalization metrics that can be used to reproduce the results of "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Rank correlation results for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains the rank correlation results from the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Trends in Natural Language Processing
<p>This dataset contains the supporting parsed corpus as described in the publication: "Analyzing a Decade of Evolution: Trends in<br>Natural Language Processing" </p> <p>This dataset contains a single zip containing data between the years 2010 and 2022, for the conferences:</p> <ul> <li>Meeting of the Association for Computational Linguistics (ACL)</li> <li>Conference on Empirical Methods in Natural Language Processing (EMNLP)</li> <li>American Chapter of the Association for Computational Linguistics (NAACL)</li> <li>Conference on Computational Linguistics (COLING)</li> <li>International Conference on Language Resources and Evaluation (LREC)</li> <li>Conference on Computational Natural Language Learning (CoNLL)</li> <li>European Chapter of the Association for Computational Linguistics (EACL)</li> <li>International Joint Conference on Natural Language Processing (IJCNLP)</li> </ul> <p>The data inlcuded is a PDF and a JSON file for each confernce. The JSON file is constucted from using the python packages SciPDF and PyPDF2. PyPDF2 extract all text from a page and is presented in the 'full_text' field, where as the SciPDF parser utilize machine lerarning to create a smart representation of the data, which is presented in the remaining fields.</p> <p>For further details on how this dataset was generated, please see our <a href="https://github.com/ieeta-pt/nlp-trends" target="_blank" rel="noopener">GitHub</a> repository, and our paper.</p> <p>Citation:</p> <p>AWAITING PUBLICATION</p> <p> </p>
Enhancing Name Entity Recognition Through Hybrid Deep Learning in Natural Language Processing
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.