Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

87

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

87 results for “LANGUAGE PROCESSING”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Natural language processing systems for pathology parsing in limited data environments with uncertainty estimation

Open the record for dataset details and reuse information.

publicJul 2021View details →
zenodo32/100

Supplementary Data for "Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes"

<p>This package contains supplementary data not directly included in the paper, including per-commit results for each prototype and the prompts used in the proof-of-concept implementations.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Text simplification in second language: process and product data

<p>This folder contains data on text simplification collected from an experimental study with second-language university students. We adopted a pre-test and post-test design, and randomly divided participants into experimental and control group. In the pre-test, participants were given an extract of a corporate report dealing with sustainability and were asked to revise it to make it easier to read for a lay customer. Subsequently, they took part in training. The experimental group received training on both plain language and sustainability, while the control group received training exclusively on the topic of sustainability. In the post-test session (2-3 days after the pre-test), all participants were assigned a second extract of a corporate report dealing with sustainability, and were asked again to make it easier to read for a lay customer by applying what they had learned from their respective training. This design allowed us to examine the impact of plain language training on text simplification (revision) tasks. The texts were in English while the participants were native speakers of other languages (mainly Dutch), so the text simplification took place in their second language.</p> <p>This project (PLanTra) has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 888918.</p> <p>Please see &quot;readme&quot; file for additional information.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Processed data for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This is the data used to reproduce the results from &quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Scatter plots for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains the test-score-vs-metric plots generated by the paper&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Generalization metrics for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains all the generalization metrics that can be used to reproduce the results of&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Rank correlation results for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"

<p>This file contains the rank correlation results from the paper&nbsp;&quot;Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Trends in Natural Language Processing

<p>This dataset contains the supporting parsed corpus as described in the publication: "Analyzing a Decade of Evolution: Trends in<br>Natural Language Processing"&nbsp;</p> <p>This dataset contains a single zip containing data between the years 2010 and 2022, for the conferences:</p> <ul> <li>Meeting of the Association for Computational Linguistics (ACL)</li> <li>Conference on Empirical Methods in Natural Language Processing (EMNLP)</li> <li>American Chapter of the Association for Computational Linguistics (NAACL)</li> <li>Conference on Computational Linguistics (COLING)</li> <li>International Conference on Language Resources and Evaluation (LREC)</li> <li>Conference on Computational Natural Language Learning (CoNLL)</li> <li>European Chapter of the Association for Computational Linguistics (EACL)</li> <li>International Joint Conference on Natural Language Processing (IJCNLP)</li> </ul> <p>The data inlcuded is a PDF and a JSON file for each confernce. The JSON file is constucted from using the python packages SciPDF and PyPDF2. PyPDF2 extract all text from a page and is presented in the 'full_text' field, where as the SciPDF parser utilize machine lerarning to create a smart representation of the data, which is presented in the remaining fields.</p> <p>For further details on how this dataset was generated, please see our <a href="https://github.com/ieeta-pt/nlp-trends" target="_blank" rel="noopener">GitHub</a> repository, and our paper.</p> <p>Citation:</p> <p>AWAITING PUBLICATION</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Enhancing Name Entity Recognition Through Hybrid Deep Learning in Natural Language Processing

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Modeling Compliance Specifications in Linear Temporal Logic, Event Processing Language and Property Specification Patterns

<p>Experimental material &amp; data</p>

opencc-by-4.0May 2018View details →
zenodo32/100

Dataset: Language and semantic processing in blind and sighted individuals. A simulation study

<p>-------------------<br> GENERAL INFORMATION<br> -------------------</p> <p>1. Title of Dataset&nbsp;<br> &nbsp;&nbsp; &nbsp;<br> Dataset_SimulationsOutput_BlindvsSightedModels<br> &nbsp;&nbsp; &nbsp;<br> 2. Author Information</p> <p>This dataset is being made public to act as supplementary data for publication. Running title:&nbsp;</p> <p>R. Tomasello, M. Garagnani, T. Wennekers, and F. Pulverm&uuml;ller, Recruitment of visual cortex for language processing in blind individuals is explained by Hebbian learning.</p> <p>&nbsp; Corresponding author:<br> &nbsp; &nbsp; &nbsp; &nbsp; Name: &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Rosario Tomasello&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; Institution: &nbsp;&nbsp; &nbsp;Brain Language Laboratory, Freie Universit&auml;t Berlin<br> &nbsp; &nbsp; &nbsp; &nbsp; Address: &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Habelschwerdter Allee 45, 14195 Berlin<br> &nbsp; &nbsp; &nbsp; &nbsp; Email: &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Tomasello.r@fu-berlin.de</p> <p>--------------------------<br> RESEARCH PROJECT &amp; METHODOLOGICAL INFORMATION<br> --------------------------</p> <p>The data were generated by a neurobiologically constrained cortex model of the fronto-temporal-occipital lobes applied to simulate word meaning acquisition of object- and action-related words in action and perception system under undeprived and visually deprived conditions. The purpose of this study was to investigate how and why the visual system is recruited for language processing in blind individuals, as documented by neurocognitive empirical experiments.</p> <p>The files listed above shows the cell assemblies (CA) distributions across the different cortical areas spontaneously emerged as a result of Hebbian learning. For more information about the general features of the model see our previous publications:</p> <p>R. Tomasello, M. Garagnani, T. Wennekers, and F. Pulverm&uuml;ller, (2018) A Neurobiologically Constrained Cortex Model of Semantic Grounding With Spiking Neurons and Brain-Like Connectivity, Front. Comput. Neurosci., vol. 12, p. 88.</p> <p>M. Garagnani, G. Lucchese, R. Tomasello, T. Wennekers, and F. Pulverm&uuml;ller, (2017) A Spiking Neurocomputational Model of High-Frequency Oscillatory Brain Responses to Words and Pseudowords, Front. Comput. Neurosci., vol. 10, no. January, pp. 1&ndash;19.</p> <p><br> ---------------------<br> DATA &amp; FILE OVERVIEW<br> ---------------------</p> <p><br> 1. File List<br> &nbsp; &nbsp;<br> &nbsp;&nbsp; &nbsp;Filename: CA_Structure_Blind VS SightedModel.xlsx (size 150KB) &nbsp; &nbsp;&nbsp;<br> &nbsp;&nbsp; &nbsp;<br> The excel file includes the cell assembly (CA) distributions of the learnt action and object words of both sighted and blind models along with their CA structure comparisons and the related figures.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br> &nbsp;&nbsp; &nbsp;Folders: Raw_Data_SightedModel (size 6KB) &amp; Raw_Data_BlindModel (size 6KB)<br> &nbsp; &nbsp; &nbsp;&nbsp;<br> These folders include the raw data output of the CA structure of the 13 sighted and 13 blind simulated models after word meaning acquisition.&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo32/100

Natural Language Processing of Clinical Notes on Chronic Diseases: Systematic Review

<p>Supplementary material outlining complete list of reviewed papers, chronic diseases and their classifications, algorithms used, publication venues, and excluded papers.</p>

opencc-by-4.0Apr 2019View details →
zenodo32/100

Corpus Nummorum - Natural Language Processing Dataset

<p>This Natural Language Processing (NLP) dataset contains a part of the MySQL <a href="https://www.corpus-nummorum.eu/">Corpus Nummorum (CN)</a> database. It covers Greek and Roman coins from ancient Thrace, Moesia Inferior, Troad and Mysia.</p> <p>The dataset contains 7,900 coin descriptions (or designs) created by the members of the CN project. Most of them are actual coin designs which can be linked through our relational database to the matching CN coins, types and their images. However, some of them (about 450) were only created for the training of the NLP model.&nbsp;</p> <p>There are nine different MySQL tables: &nbsp;</p> <ol> <li>data_coins: contains the data of all coins in the CN database</li> <li>data_coins_images: contains data of all images in the CN database</li> <li>data_coins_imagesets: contains the image pairs for the CN coins</li> <li>data_designs: contains every coin description in German, English and Bulgarian</li> <li>data_types: contains the data of alle coin types the Cn database</li> <li>nlp_hierarchy: contains the classes and subclasses of all entity categories&nbsp;</li> <li>nlp_list_entities: contains the data of all nlp entities in the CN database</li> <li>nlp_relation_extraction_en_v2: contains the annotations for the training of our NLP model</li> <li>nlp_training_designs: contains the coin designs used for training our NLP model</li> </ol> <p>Only tables 8 and 9 are important for NLP training, as they contain the descriptions and the corresponding annotations. The other tables (data_...) make it possible to link the coin descriptions with the various coins and types in the CN database. It is therefore also possible to provide the CN image data sets with the appropriate descriptions (<a href="../records/10033993">CN - Coin Image Dataset</a> and <a href="../records/13748799">CN - Object Detection Coin Dataset</a>). The other NLP tables provide information about the entities and relations in the descriptions and are used to create the RDF data for the <a href="https://nomisma.org/datasets">nomisma.org</a> portal. The tables of the relational CN database can be related via the various ID columns using foreign keys.</p> <p>For easier access without MySQL, we have attached two csv files with the descriptions in English and German and the annotations for the English designs. The annotations can be related to the descriptions via the Design_ID column.&nbsp;</p> <p>During the summer semester 2024, we held the "Data Challenge" event at our Department of Computer Science at the Goethe-University. Our students could choose between the Object Detection dataset and a Natural Language Processing dataset as their challenge. We gave the teams that decided to take part in the NLP challenge this dataset with the task of trying out their own ideas. Here are the results:</p> <ul> <li><a href="https://github.com/jasperforth/DataChallenge_LLM_REPipeline">LLM_RE Pipeline</a></li> <li><a href="https://github.com/Axolord/coin-description-embeddings">Coin description embeddings</a></li> <li><a href="https://github.com/calul0/nlp_coin_app/tree/master">NLP coin app</a></li> </ul> <p>Now we would like to invite you to try out your own ideas and models on our coin data.</p> <p>If you have any questions or suggestions, please, feel free to contact us.&nbsp;</p>

openSep 2024View details →
zenodo32/100

A Large Dataset of Tweets on the 2023 Presidential Elections in Nigeria for Natural Language Processing Tasks

<p>The dataset contains tweets related to the 2023 presidential elections in Nigeria. The data was retrieved from the social media&nbsp;network, Twitter (Now X) between February 4<sup>th</sup>, 2023 and April 4<sup>th</sup>, 2023. The hashtags from the official handles and other popular hashtags endorsed and/or representing the candidates of each party were considered for retrieving election related tweets using an API from Twitter social media platform. Three major political parties in Nigeria were considered and they have been labelled as Party A, Party L and Party P in this dataset. The party or group called &quot;General&quot; contains tweets from the Independent National Electoral Commission (INEC) hashtags such as <em>@inecnigeria</em> and <em>#2023election</em> which is not directly for any political party.</p> <p>The dataset has been pre-processed lightly to make it very useful to researcher for a wide range of natural language processing tasks like sentiment analysis, topic modelling, fake news detection, emotion detection, election stance, etc.</p> <p>Details of the dataset collection such as hashtags, retrieved tweets, duplicates removed, and the remaining unique tweets is presented in Table 1.&nbsp;</p> <p>&nbsp;</p> <p>Table 1: Tweets collection and duplicates removal</p> <table> <tbody> <tr> <td> <p>S/N</p> </td> <td> <p>Party</p> </td> <td> <p>Hash tags</p> </td> <td> <p>Retrieved tweets</p> </td> <td> <p>Duplicates tweets</p> </td> <td> <p>Unique tweets</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; X</p> </td> <td> <p><em>@inecnigeria</em></p> <p><em>#2023election</em></p> </td> <td> <p>64,496</p> </td> <td> <p>47,275</p> </td> <td> <p>17,195</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; A</p> <p>&nbsp;</p> </td> <td> <p><em>#TinubuIsComing</em></p> <p><em>#emilokan </em></p> <p><em>#jagabanarmy</em></p> <p><em>#RenewedHope </em></p> <p><em>#BATKSM2023</em></p> </td> <td> <p>263,870</p> </td> <td> <p>231,036</p> </td> <td> <p>32,832</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; L</p> </td> <td> <p><em>#VoteLP</em></p> <p><em>#NigeriaMustBeBright</em></p> <p><em>#PeterObiForPresident2023</em></p> <p><em>#ObiDatti2023 </em></p> <p><em>#PeterObi</em></p> </td> <td> <p>664,083</p> </td> <td> <p>310,857</p> </td> <td> <p>353,226</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>&nbsp;</p> <p>P</p> </td> <td> <p><em>#NigeriaDecides</em></p> <p><em>#VotePDP</em></p> <p><em>#AtikuOkowa2023</em></p> <p><em>#FinalPushToVictory</em></p> <p><em>#RecoverNigeria</em></p> </td> <td> <p>387,450</p> </td> <td> <p>318,425</p> </td> <td> <p>66,227</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>1,379,899</p> </td> <td> <p>907,593</p> </td> <td> <p>468,480</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>To encourage NLP tasks, we uploaded in this Version One the following files:</p> <ol> <li>The combined dataset with pre-processed tweets and their meta data but with&nbsp;removed duplicates are in the file labelled &ldquo;Combined Dataset Pre-processed without duplicates.csv&rdquo;</li> <li>General statistics on each corpus is in the file labelled &ldquo;Dataset Statistics.xlsx&rdquo;&nbsp;&nbsp;</li> <li>The preprocessed corpus from the general group with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_GENERAL.xlsx&rdquo;</li> <li>The preprocessed corpus from Party A with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party A.xlsx&rdquo;</li> <li>The preprocessed corpus from Party L with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party L.xlsx&rdquo;</li> <li>The preprocessed corpus from Party P with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party P.xlsx&rdquo;</li> <li>The top 100 frequent tokens are in the file labelled &ldquo;Top 100 Tokens and weights.xlsx&rdquo;</li> <li>The top frequent bigrams and their weights are&nbsp;in the file labelled &ldquo;Top 100 Bigrams and weights.xlsx&rdquo;</li> <li>The top frequent trigrams and their weights are in the file labelled &ldquo;Top 100 Trigrams and weights.xlsx&rdquo;</li> </ol>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov32/100

Computational Neuroscience of Language Processing in the Human Brain

ClinicalTrials.gov study NCT05222594. IPD Sharing: NO. Countries: 1. Publications: 14.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Infant Diet Effects on Brain Function and Language Processing (fMRI)

ClinicalTrials.gov study NCT00735423. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing for Screening Opioid Misuse

ClinicalTrials.gov study NCT05745480. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Language Processing and TMS

ClinicalTrials.gov study NCT05425615. IPD Sharing: NO. Countries: 1. Publications: 22.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Brain Mechanisms for Language Processing in Adolescents With Autism Spectrum Disorder

ClinicalTrials.gov study NCT02700074. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing for Headache Medicine

ClinicalTrials.gov study NCT05377437. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record