Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
87
datasets available to search
ShareScore release 0.9.0
Dataset results
87 results for “LANGUAGE PROCESSING”
Data from: Natural language processing systems for pathology parsing in limited data environments with uncertainty estimation
Open the record for dataset details and reuse information.
Supplementary Data for "Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes"
<p>This package contains supplementary data not directly included in the paper, including per-commit results for each prototype and the prompts used in the proof-of-concept implementations.</p>
Text simplification in second language: process and product data
<p>This folder contains data on text simplification collected from an experimental study with second-language university students. We adopted a pre-test and post-test design, and randomly divided participants into experimental and control group. In the pre-test, participants were given an extract of a corporate report dealing with sustainability and were asked to revise it to make it easier to read for a lay customer. Subsequently, they took part in training. The experimental group received training on both plain language and sustainability, while the control group received training exclusively on the topic of sustainability. In the post-test session (2-3 days after the pre-test), all participants were assigned a second extract of a corporate report dealing with sustainability, and were asked again to make it easier to read for a lay customer by applying what they had learned from their respective training. This design allowed us to examine the impact of plain language training on text simplification (revision) tasks. The texts were in English while the participants were native speakers of other languages (mainly Dutch), so the text simplification took place in their second language.</p> <p>This project (PLanTra) has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 888918.</p> <p>Please see "readme" file for additional information.</p>
Processed data for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This is the data used to reproduce the results from "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Scatter plots for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains the test-score-vs-metric plots generated by the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Generalization metrics for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains all the generalization metrics that can be used to reproduce the results of "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Rank correlation results for the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data"
<p>This file contains the rank correlation results from the paper "Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data".</p>
Trends in Natural Language Processing
<p>This dataset contains the supporting parsed corpus as described in the publication: "Analyzing a Decade of Evolution: Trends in<br>Natural Language Processing" </p> <p>This dataset contains a single zip containing data between the years 2010 and 2022, for the conferences:</p> <ul> <li>Meeting of the Association for Computational Linguistics (ACL)</li> <li>Conference on Empirical Methods in Natural Language Processing (EMNLP)</li> <li>American Chapter of the Association for Computational Linguistics (NAACL)</li> <li>Conference on Computational Linguistics (COLING)</li> <li>International Conference on Language Resources and Evaluation (LREC)</li> <li>Conference on Computational Natural Language Learning (CoNLL)</li> <li>European Chapter of the Association for Computational Linguistics (EACL)</li> <li>International Joint Conference on Natural Language Processing (IJCNLP)</li> </ul> <p>The data inlcuded is a PDF and a JSON file for each confernce. The JSON file is constucted from using the python packages SciPDF and PyPDF2. PyPDF2 extract all text from a page and is presented in the 'full_text' field, where as the SciPDF parser utilize machine lerarning to create a smart representation of the data, which is presented in the remaining fields.</p> <p>For further details on how this dataset was generated, please see our <a href="https://github.com/ieeta-pt/nlp-trends" target="_blank" rel="noopener">GitHub</a> repository, and our paper.</p> <p>Citation:</p> <p>AWAITING PUBLICATION</p> <p> </p>
Enhancing Name Entity Recognition Through Hybrid Deep Learning in Natural Language Processing
Open the record for dataset details and reuse information.
Modeling Compliance Specifications in Linear Temporal Logic, Event Processing Language and Property Specification Patterns
<p>Experimental material & data</p>
Dataset: Language and semantic processing in blind and sighted individuals. A simulation study
<p>-------------------<br> GENERAL INFORMATION<br> -------------------</p> <p>1. Title of Dataset <br> <br> Dataset_SimulationsOutput_BlindvsSightedModels<br> <br> 2. Author Information</p> <p>This dataset is being made public to act as supplementary data for publication. Running title: </p> <p>R. Tomasello, M. Garagnani, T. Wennekers, and F. Pulvermüller, Recruitment of visual cortex for language processing in blind individuals is explained by Hebbian learning.</p> <p> Corresponding author:<br> Name: Rosario Tomasello <br> Institution: Brain Language Laboratory, Freie Universität Berlin<br> Address: Habelschwerdter Allee 45, 14195 Berlin<br> Email: Tomasello.r@fu-berlin.de</p> <p>--------------------------<br> RESEARCH PROJECT & METHODOLOGICAL INFORMATION<br> --------------------------</p> <p>The data were generated by a neurobiologically constrained cortex model of the fronto-temporal-occipital lobes applied to simulate word meaning acquisition of object- and action-related words in action and perception system under undeprived and visually deprived conditions. The purpose of this study was to investigate how and why the visual system is recruited for language processing in blind individuals, as documented by neurocognitive empirical experiments.</p> <p>The files listed above shows the cell assemblies (CA) distributions across the different cortical areas spontaneously emerged as a result of Hebbian learning. For more information about the general features of the model see our previous publications:</p> <p>R. Tomasello, M. Garagnani, T. Wennekers, and F. Pulvermüller, (2018) A Neurobiologically Constrained Cortex Model of Semantic Grounding With Spiking Neurons and Brain-Like Connectivity, Front. Comput. Neurosci., vol. 12, p. 88.</p> <p>M. Garagnani, G. Lucchese, R. Tomasello, T. Wennekers, and F. Pulvermüller, (2017) A Spiking Neurocomputational Model of High-Frequency Oscillatory Brain Responses to Words and Pseudowords, Front. Comput. Neurosci., vol. 10, no. January, pp. 1–19.</p> <p><br> ---------------------<br> DATA & FILE OVERVIEW<br> ---------------------</p> <p><br> 1. File List<br> <br> Filename: CA_Structure_Blind VS SightedModel.xlsx (size 150KB) <br> <br> The excel file includes the cell assembly (CA) distributions of the learnt action and object words of both sighted and blind models along with their CA structure comparisons and the related figures. <br> <br> Folders: Raw_Data_SightedModel (size 6KB) & Raw_Data_BlindModel (size 6KB)<br> <br> These folders include the raw data output of the CA structure of the 13 sighted and 13 blind simulated models after word meaning acquisition. </p>
Natural Language Processing of Clinical Notes on Chronic Diseases: Systematic Review
<p>Supplementary material outlining complete list of reviewed papers, chronic diseases and their classifications, algorithms used, publication venues, and excluded papers.</p>
Corpus Nummorum - Natural Language Processing Dataset
<p>This Natural Language Processing (NLP) dataset contains a part of the MySQL <a href="https://www.corpus-nummorum.eu/">Corpus Nummorum (CN)</a> database. It covers Greek and Roman coins from ancient Thrace, Moesia Inferior, Troad and Mysia.</p> <p>The dataset contains 7,900 coin descriptions (or designs) created by the members of the CN project. Most of them are actual coin designs which can be linked through our relational database to the matching CN coins, types and their images. However, some of them (about 450) were only created for the training of the NLP model. </p> <p>There are nine different MySQL tables: </p> <ol> <li>data_coins: contains the data of all coins in the CN database</li> <li>data_coins_images: contains data of all images in the CN database</li> <li>data_coins_imagesets: contains the image pairs for the CN coins</li> <li>data_designs: contains every coin description in German, English and Bulgarian</li> <li>data_types: contains the data of alle coin types the Cn database</li> <li>nlp_hierarchy: contains the classes and subclasses of all entity categories </li> <li>nlp_list_entities: contains the data of all nlp entities in the CN database</li> <li>nlp_relation_extraction_en_v2: contains the annotations for the training of our NLP model</li> <li>nlp_training_designs: contains the coin designs used for training our NLP model</li> </ol> <p>Only tables 8 and 9 are important for NLP training, as they contain the descriptions and the corresponding annotations. The other tables (data_...) make it possible to link the coin descriptions with the various coins and types in the CN database. It is therefore also possible to provide the CN image data sets with the appropriate descriptions (<a href="../records/10033993">CN - Coin Image Dataset</a> and <a href="../records/13748799">CN - Object Detection Coin Dataset</a>). The other NLP tables provide information about the entities and relations in the descriptions and are used to create the RDF data for the <a href="https://nomisma.org/datasets">nomisma.org</a> portal. The tables of the relational CN database can be related via the various ID columns using foreign keys.</p> <p>For easier access without MySQL, we have attached two csv files with the descriptions in English and German and the annotations for the English designs. The annotations can be related to the descriptions via the Design_ID column. </p> <p>During the summer semester 2024, we held the "Data Challenge" event at our Department of Computer Science at the Goethe-University. Our students could choose between the Object Detection dataset and a Natural Language Processing dataset as their challenge. We gave the teams that decided to take part in the NLP challenge this dataset with the task of trying out their own ideas. Here are the results:</p> <ul> <li><a href="https://github.com/jasperforth/DataChallenge_LLM_REPipeline">LLM_RE Pipeline</a></li> <li><a href="https://github.com/Axolord/coin-description-embeddings">Coin description embeddings</a></li> <li><a href="https://github.com/calul0/nlp_coin_app/tree/master">NLP coin app</a></li> </ul> <p>Now we would like to invite you to try out your own ideas and models on our coin data.</p> <p>If you have any questions or suggestions, please, feel free to contact us. </p>
A Large Dataset of Tweets on the 2023 Presidential Elections in Nigeria for Natural Language Processing Tasks
<p>The dataset contains tweets related to the 2023 presidential elections in Nigeria. The data was retrieved from the social media network, Twitter (Now X) between February 4<sup>th</sup>, 2023 and April 4<sup>th</sup>, 2023. The hashtags from the official handles and other popular hashtags endorsed and/or representing the candidates of each party were considered for retrieving election related tweets using an API from Twitter social media platform. Three major political parties in Nigeria were considered and they have been labelled as Party A, Party L and Party P in this dataset. The party or group called "General" contains tweets from the Independent National Electoral Commission (INEC) hashtags such as <em>@inecnigeria</em> and <em>#2023election</em> which is not directly for any political party.</p> <p>The dataset has been pre-processed lightly to make it very useful to researcher for a wide range of natural language processing tasks like sentiment analysis, topic modelling, fake news detection, emotion detection, election stance, etc.</p> <p>Details of the dataset collection such as hashtags, retrieved tweets, duplicates removed, and the remaining unique tweets is presented in Table 1. </p> <p> </p> <p>Table 1: Tweets collection and duplicates removal</p> <table> <tbody> <tr> <td> <p>S/N</p> </td> <td> <p>Party</p> </td> <td> <p>Hash tags</p> </td> <td> <p>Retrieved tweets</p> </td> <td> <p>Duplicates tweets</p> </td> <td> <p>Unique tweets</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p> X</p> </td> <td> <p><em>@inecnigeria</em></p> <p><em>#2023election</em></p> </td> <td> <p>64,496</p> </td> <td> <p>47,275</p> </td> <td> <p>17,195</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p> A</p> <p> </p> </td> <td> <p><em>#TinubuIsComing</em></p> <p><em>#emilokan </em></p> <p><em>#jagabanarmy</em></p> <p><em>#RenewedHope </em></p> <p><em>#BATKSM2023</em></p> </td> <td> <p>263,870</p> </td> <td> <p>231,036</p> </td> <td> <p>32,832</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p> </p> <p> L</p> </td> <td> <p><em>#VoteLP</em></p> <p><em>#NigeriaMustBeBright</em></p> <p><em>#PeterObiForPresident2023</em></p> <p><em>#ObiDatti2023 </em></p> <p><em>#PeterObi</em></p> </td> <td> <p>664,083</p> </td> <td> <p>310,857</p> </td> <td> <p>353,226</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p> </p> <p>P</p> </td> <td> <p><em>#NigeriaDecides</em></p> <p><em>#VotePDP</em></p> <p><em>#AtikuOkowa2023</em></p> <p><em>#FinalPushToVictory</em></p> <p><em>#RecoverNigeria</em></p> </td> <td> <p>387,450</p> </td> <td> <p>318,425</p> </td> <td> <p>66,227</p> </td> </tr> <tr> <td> <p> </p> </td> <td> <p> </p> </td> <td> <p> </p> </td> <td> <p>1,379,899</p> </td> <td> <p>907,593</p> </td> <td> <p>468,480</p> </td> </tr> </tbody> </table> <p> </p> <p>To encourage NLP tasks, we uploaded in this Version One the following files:</p> <ol> <li>The combined dataset with pre-processed tweets and their meta data but with removed duplicates are in the file labelled “Combined Dataset Pre-processed without duplicates.csv”</li> <li>General statistics on each corpus is in the file labelled “Dataset Statistics.xlsx” </li> <li>The preprocessed corpus from the general group with the tweet contents only is in file labelled “Preprocessed_Tweet only_GENERAL.xlsx”</li> <li>The preprocessed corpus from Party A with the tweet contents only is in file labelled “Preprocessed_Tweet only_Party A.xlsx”</li> <li>The preprocessed corpus from Party L with the tweet contents only is in file labelled “Preprocessed_Tweet only_Party L.xlsx”</li> <li>The preprocessed corpus from Party P with the tweet contents only is in file labelled “Preprocessed_Tweet only_Party P.xlsx”</li> <li>The top 100 frequent tokens are in the file labelled “Top 100 Tokens and weights.xlsx”</li> <li>The top frequent bigrams and their weights are in the file labelled “Top 100 Bigrams and weights.xlsx”</li> <li>The top frequent trigrams and their weights are in the file labelled “Top 100 Trigrams and weights.xlsx”</li> </ol>
Computational Neuroscience of Language Processing in the Human Brain
ClinicalTrials.gov study NCT05222594. IPD Sharing: NO. Countries: 1. Publications: 14.
Infant Diet Effects on Brain Function and Language Processing (fMRI)
ClinicalTrials.gov study NCT00735423. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Natural Language Processing for Screening Opioid Misuse
ClinicalTrials.gov study NCT05745480. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.
Language Processing and TMS
ClinicalTrials.gov study NCT05425615. IPD Sharing: NO. Countries: 1. Publications: 22.
Brain Mechanisms for Language Processing in Adolescents With Autism Spectrum Disorder
ClinicalTrials.gov study NCT02700074. IPD Sharing: NO. Countries: 1. Publications: 1.
Natural Language Processing for Headache Medicine
ClinicalTrials.gov study NCT05377437. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.