Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

116

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

116 results for “Natural Language”

Learn how ShareScore rates datasets ↗
zenodo32/100

Supplementary material for: Inconsistency Detection in Natural Language Requirements using ChatGPT: a Preliminary Evaluation

<p>Supplementary material including both data and annotations. For each document we present:</p> <ol> <li>annotated results of the chatGPT answers;</li> <li>annotated results of the manual analysis</li> <li>The requirements and the grund truth pairs. Originals requirements are marked with &quot;1&quot; in the third column, mutants are marked with &quot;0&quot;.&nbsp;</li> </ol>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Datasets and Trained Models for "Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning"

<p>Datasets and Trained models for ICSE 2020 &quot;Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning&quot;</p>

opencc-by-4.0Sep 2020View details →
zenodo32/100

Predicting brain activity from word embeddings during natural language comprehension: Example Data

<p>Example data for the &quot;Predicting brain activity from word embeddings during natural language comprehension&quot; tutorial.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

A Large Dataset of Tweets on the 2023 Presidential Elections in Nigeria for Natural Language Processing Tasks

<p>The dataset contains tweets related to the 2023 presidential elections in Nigeria. The data was retrieved from the social media&nbsp;network, Twitter (Now X) between February 4<sup>th</sup>, 2023 and April 4<sup>th</sup>, 2023. The hashtags from the official handles and other popular hashtags endorsed and/or representing the candidates of each party were considered for retrieving election related tweets using an API from Twitter social media platform. Three major political parties in Nigeria were considered and they have been labelled as Party A, Party L and Party P in this dataset. The party or group called &quot;General&quot; contains tweets from the Independent National Electoral Commission (INEC) hashtags such as <em>@inecnigeria</em> and <em>#2023election</em> which is not directly for any political party.</p> <p>The dataset has been pre-processed lightly to make it very useful to researcher for a wide range of natural language processing tasks like sentiment analysis, topic modelling, fake news detection, emotion detection, election stance, etc.</p> <p>Details of the dataset collection such as hashtags, retrieved tweets, duplicates removed, and the remaining unique tweets is presented in Table 1.&nbsp;</p> <p>&nbsp;</p> <p>Table 1: Tweets collection and duplicates removal</p> <table> <tbody> <tr> <td> <p>S/N</p> </td> <td> <p>Party</p> </td> <td> <p>Hash tags</p> </td> <td> <p>Retrieved tweets</p> </td> <td> <p>Duplicates tweets</p> </td> <td> <p>Unique tweets</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; X</p> </td> <td> <p><em>@inecnigeria</em></p> <p><em>#2023election</em></p> </td> <td> <p>64,496</p> </td> <td> <p>47,275</p> </td> <td> <p>17,195</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; A</p> <p>&nbsp;</p> </td> <td> <p><em>#TinubuIsComing</em></p> <p><em>#emilokan </em></p> <p><em>#jagabanarmy</em></p> <p><em>#RenewedHope </em></p> <p><em>#BATKSM2023</em></p> </td> <td> <p>263,870</p> </td> <td> <p>231,036</p> </td> <td> <p>32,832</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; L</p> </td> <td> <p><em>#VoteLP</em></p> <p><em>#NigeriaMustBeBright</em></p> <p><em>#PeterObiForPresident2023</em></p> <p><em>#ObiDatti2023 </em></p> <p><em>#PeterObi</em></p> </td> <td> <p>664,083</p> </td> <td> <p>310,857</p> </td> <td> <p>353,226</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>&nbsp;</p> <p>P</p> </td> <td> <p><em>#NigeriaDecides</em></p> <p><em>#VotePDP</em></p> <p><em>#AtikuOkowa2023</em></p> <p><em>#FinalPushToVictory</em></p> <p><em>#RecoverNigeria</em></p> </td> <td> <p>387,450</p> </td> <td> <p>318,425</p> </td> <td> <p>66,227</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>1,379,899</p> </td> <td> <p>907,593</p> </td> <td> <p>468,480</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>To encourage NLP tasks, we uploaded in this Version One the following files:</p> <ol> <li>The combined dataset with pre-processed tweets and their meta data but with&nbsp;removed duplicates are in the file labelled &ldquo;Combined Dataset Pre-processed without duplicates.csv&rdquo;</li> <li>General statistics on each corpus is in the file labelled &ldquo;Dataset Statistics.xlsx&rdquo;&nbsp;&nbsp;</li> <li>The preprocessed corpus from the general group with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_GENERAL.xlsx&rdquo;</li> <li>The preprocessed corpus from Party A with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party A.xlsx&rdquo;</li> <li>The preprocessed corpus from Party L with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party L.xlsx&rdquo;</li> <li>The preprocessed corpus from Party P with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party P.xlsx&rdquo;</li> <li>The top 100 frequent tokens are in the file labelled &ldquo;Top 100 Tokens and weights.xlsx&rdquo;</li> <li>The top frequent bigrams and their weights are&nbsp;in the file labelled &ldquo;Top 100 Bigrams and weights.xlsx&rdquo;</li> <li>The top frequent trigrams and their weights are in the file labelled &ldquo;Top 100 Trigrams and weights.xlsx&rdquo;</li> </ol>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov32/100

Neurobiology of Language Recovery in Aphasia: Natural History and Treatment-Induced Recovery

ClinicalTrials.gov study NCT01927302. IPD Sharing: Not stated. Countries: 1. Publications: 64.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing for Screening Opioid Misuse

ClinicalTrials.gov study NCT05745480. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing for Headache Medicine

ClinicalTrials.gov study NCT05377437. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing and Quality Assessment in Primary Care

ClinicalTrials.gov study NCT01023243. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Leveraging Natural Language for Program Search and Abstraction Learning Regex

<p>Program synthesis dataset containing text editing tasks and language annotations (synthetic and human annotated) for the&nbsp;Leveraging Natural Language for Program Search and Abstraction Learning (currently under review at NeurIPS 2020). Will be deanonymized upon review.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Figure 1d from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1d A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Fossilised animal skin (Natural History Museum 2009)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1b from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1b A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Pinned insect specimen (Natural History Museum 2018)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1c from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1c A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Microscope slide (Natural History Museum 2017)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1a from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1a A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Herbarium specimen (Natural History Museum 2007a)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 11 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 11 The distribution of languages across the specimen and herbaria. EN=English, FR=French, LA=Latin, ET=Estonian, DE=German, NL=Dutch, PT=Portuguese, ES=Spanish, SV=Swedish, RU=Russian, FI=Finnish, IT=Italian, ZZ=Unknown. The codes for the contributing herbaria are listed in Table 11 (from Dillen et al. 2019).

opencc-by-4.0Jul 2020View details →
zenodo28/100

Supplementary material 1 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Appendices

opencc-zeroJul 2020View details →
zenodo28/100

Figure 1e from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1e A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Liquid preserved specimen (Natural History Museum 2010)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 2 from: Owen D, Groom Q, Hardisty A, Leegwater T, Livermore L, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e58030. https://doi.org/10.3897/rio.6.e58030

Figure 2 A possible semi-automatic digitisation workflow to extract data from the labels of collection specimens.

opencc-by-4.0Sep 2020View details →
zenodo28/100

PHONOTACTIC PHENOMENON IN THE ENGLISH LANGUAGE AND ITS LINGUISTIC NATURE

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

Combining Natural Language and Images for Garbage Classification: A Public Benchmark

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

Retrieving API knowledge from Tutorials and Stack Overflow based on Natural Language Queries

<p>The replication package of PLAN</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record