Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

48

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

48 results for “Language tasks”

Learn how ShareScore rates datasets ↗
zenodo40/100

MTL-QA : A dataset and multi-task learning approach for knowledge graph and natural language question answering

<p>The dataset used for this project is created by enhancing the publicly available MetaQA (Movie Text Audio QA), which is primarily a KGQA dataset pertaining to movies, an extension of WikiMovies. This involves questions requiring 1, 2, and 3 hops which can be answered by using a MetaQA Knowledge Graph. The questions are available in text and audio format. The text has vanilla (original) and its paraphrased version, and is called ntm.&nbsp;</p> <p>In order to develop a dataset to support NLQA, a series of dataset augmentation steps has been performed.</p> <p>The dataset consists of natural language questions and a tagged topic entity as ground truth. This topic entity is used to retrieve textual information related to the question from Wikipedia. The introduction section of the entity&#39;s page is used as the context that is required for NLQA. Hence, this dataset has information related to both KGQA and NLQA. Certain preliminary checks and validations are done to only retain those data samples whose context can be used to answer a given question.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

SemEval-2020 Task 5: Modelling Causal Reasoning in Language: Detecting Counterfactuals

<p><strong>SemEval-2020 Task 5</strong></p> <p>&nbsp;</p> <p><strong>Subtask-1:</strong> Recognizing Counterfactual Statements (RCS) -- Determine whether a given sentence is counterfactual or not.</p> <p><strong>Subtask-2: </strong>Detecting Antecedent and Consequent (DAC) -- Extract the antecedent and consequent part in a given counterfactual sentence.</p> <p>&nbsp;</p> <p>The released dataset consists of train/test data of both subtask-1 and subtask-2. In our competition, participants could only use the corresponding dataset in each subtask.</p> <p>&nbsp;</p> <p><strong>Task 5 Codalab Website:</strong> <a href="https://competitions.codalab.org/competitions/21691">https://competitions.codalab.org/competitions/21691</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)

<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab:&nbsp;<a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is&nbsp;<a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Validation Set

<p>This is the <strong>validation set for Task 9, Language-Queried Audio Source Separation (LASS), in DCASE 2024 Challenge</strong>.&nbsp;</p> <p>This validation split is meant to be used for Task 9 at the scientific challenge DCASE 2024. This split is not meant to be used for training LASS methods. This split is meant to be used for evaluating LASS methods during the model development stage.</p> <p>This validation set consists of 1000 audio files sourced from Freesound [1], uploaded between April and October 2023. Each audio file has been manually annotated with three captions. In the annotation guidance, we instructed annotators to describe the content of audio clips using 5-20 words (similar to the caption style in Clotho [3] and AudioCaps [4] datasets). The tags of each audio file were verified and revised according to the FSD50K [2] sound event categories. Each audio file has been chunked into a 10-second clip and downsampled to 16kHz.</p> <p><strong>== Details ==</strong></p> <p>The audio files in the archives:</p> <ul> <li>lass_validation.zip</li> </ul> <p>and the associated metadata (including tags and captions) in the JSON file:</p> <ul> <li>lass_validation.json</li> </ul> <p>Participants will evaluate their LASS models using synthetic mixture data in the development stage. Specifically, given an audio clip A1 and its corresponding caption C, we select an additional audio clip, A2, to serve as background noise, thereby creating a mixed audio, A3. We anticipate that the LASS system, given A3 and C as inputs, will be able to separate the A1 source. We use the revised tags information to ensure that the two audio clips used in each mix do not share overlapping sound source classes. Three thousand synthetic audio mixtures with signal-to-noise ratios (SNR) ranging from -15dB to 15dB will be generated for the validation of LASS model development. These synthetic mixtures can be generated based on the provided CSV file:</p> <ul> <li>lass_synthetic_validation.csv</li> </ul> <p>The evaluation tool can be found at: https://github.com/Audio-AGI/dcase2024_task9_baseline/blob/main/dcase_evaluator.py</p> <p><strong>== References ==</strong></p> <p>[1] Fonseca E, Pons Puig J, Favory X, et al. Freesound datasets: a platform for the creation of open audio datasets. International Society for Music Information Retrieval (ISMIR), 2017.</p> <p>[2] Fonseca E, Favory X, Pons J, et al. FSD50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 30: 829-852.</p> <p>[3] Drossos K, Lipping S, Virtanen T. Clotho: An audio captioning dataset. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2020: 736-740.</p> <p>[4] Kim C D, Kim B, Lee H, et al. AudioCaps: Generating captions for audios in the wild. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2019: 119-132.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Development Set

<p><strong>== Description ==&nbsp;</strong></p> <p>The development set is composed of audio samples from FSD50K [1] and Clotho v2 [2] datasets. FSD50K contains over 51k audio clips (~100 hours) manually labeled using 200 classes drawn from the AudioSet Ontology. For each audio clip in the FSD50K dataset, we generated one automatic caption for each audio clip by prompting ChatGPT (GPT-4) with its sound event tags. All audio files should be converted to mono 16 kHz audio for training LASS models.&nbsp;</p> <p>Clotho v2: <a href="../records/4783391">https://zenodo.org/records/4783391</a></p> <p>FSD50K: <a href="../records/4060432">https://zenodo.org/records/4060432</a></p> <p>Automatic captions generated for FSD50K:</p> <ul> <li>fsd50k_dev_auto_caption.json</li> <li>fsd50k_eval_auto_caption.json</li> </ul> <p>Prompt for generating captions:</p> <blockquote> <p>I will give you a number of lists containing sound events. Please write an one-sentence audio caption to describe these sounds.</p> <p>Make sure you are using grammatical subject-verb-object sentences. Directly describe the sounds and avoid using the word &ldquo;heard&rdquo;. Please don't describe the temporal order of these sound events. The caption should be less than 20 words.</p> </blockquote> <p>In addition to the development set, participants are free to use any external data (including private data) but are not allowed to use audio in Freesound uploaded between April and October 2023. Participants must specify all external resources utilized in their submission in the technical report.</p> <p><strong>== References ==</strong></p> <p>[1] Fonseca E, Favory X, Pons J, et al. FSD50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 30: 829-852.</p> <p>[2] Drossos K, Lipping S, Virtanen T. Clotho: An audio captioning dataset. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2020: 736-740.</p> <p><strong>== Contact ==</strong></p> <p>Xubo Liu, xubo.liu@surrey.ac.uk</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Support for a novel, simple method for calculating word frequency of output on language production tasks

<p>Pre-recorded presentation for the&nbsp;International Workshop on Language Production 2021</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Evaluation Set

<p>This is the&nbsp;<strong>evaluation set for Task 9, Language-Queried Audio Source Separation (LASS), in DCASE 2024 Challenge</strong>.&nbsp;</p> <p>This evaluation set is meant to be used for Task 9 at the scientific challenge DCASE 2024. This split is not meant to be used for training LASS methods. This split is meant to be used for evaluating LASS methods in the final testing &amp; ranking stage. All audio clips are sourced from Freesound, uploaded between April and October 2023. Each audio file has been segmented into 10-second clips and converted to mono 16 kHz.</p> <p>This evaluation set consists of<strong> evaluation set (synth)</strong> and an&nbsp;<strong>evaluation set (real)</strong>.&nbsp;</p> <p><strong>== Evaluation set (synth) ==</strong></p> <p>This evaluation set is created using 1,000 audio clips. Each clip is annotated with three captions describing the content of the clip. We created 3,000 synthetic mixtures with signal-to-noise ratios (SNR) ranging from -15 to 15 dB. Each synthetic mixture includes one natural language query and its corresponding target source. We used annotated tag information to ensure that the two audio clips used in each mix do not share overlapping sound source classes. The original audio files used to create these mixtures are not released. The mixtures and language queries are available for evaluation.</p> <p>The audio files in the archives:</p> <ul> <li>lass_evaluation_synth.zip</li> </ul> <p>and the associated metadata (including audio filename and text queries) in the CSV file:</p> <ul> <li>lass_synthetic_evaluation.csv</li> </ul> <p><strong>== Evaluation set (real) ==</strong></p> <p>This evaluation set consists of 100 audio clips. Each audio clip contains at least two overlapping sound sources. For each audio clip, we manually annotated their component sources using text descriptions, so that each clip can be used as a 'mixture' from which to extract one or more of the component sources based on a text query. Each audio clip in evaluation (real) was labeled with two such text queries.</p> <p>The audio files in the archives:</p> <ul> <li>lass_evaluation_real.zip</li> </ul> <p>and the associated metadata (including audio filename and text queries) in the CSV file:</p> <ul> <li>lass_real_evaluation.csv</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo32/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Pre-trained Weights for the Baseline System

<p><strong>== Descriptions ==</strong></p> <p>We trained the AudioSep [1] model using the <a href="../records/10887496">development set</a> (Clotho and augmented FSD50K datasets) for 200k steps with a batch size of 16 using one Nvidia A100 GPU (around 1 day). Model details can be found in the <a href="https://arxiv.org/abs/2308.05037">AudioSep paper</a>.</p> <p>Pre-trained weights for the baseline system:</p> <ul> <li>audiosep_16k,baseline,step=200000.ckpt</li> </ul> <p>Baseline codebase:</p> <ul> <li>GitHub: <a href="https://github.com/Audio-AGI/dcase2024_task9_baseline">https://github.com/Audio-AGI/dcase2024_task9_baseline</a></li> </ul> <p><strong>== Reference ==</strong></p> <p>[1] Liu X, Kong Q, Zhao Y, et al. Separate anything you describe. arXiv:2308.05037, 2023.</p> <p><strong>== Contact ==</strong></p> <p>Xubo Liu, xubo.liu@surrey.ac.uk</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Dataset for: Lessons from the Use of Natural Language Inference (NLI) in Requirements Engineering Tasks

<p>Datasets used for Requirement Engineering '24 paper, titled: Lessons from the Use of Natural Language Inference (NLI) in Requirements Engineering Tasks</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials

<p>This is the github repository hosting data and code for Task 2: Safe Biomedical Natural Language Inference for Clinical Trials at&nbsp;<a href="https://semeval.github.io/SemEval2024/" rel="nofollow">Semeval 2024</a>.</p> <p>For additional information about the task, please consult the official&nbsp;<a href="https://sites.google.com/view/nli4ct/home" rel="nofollow">website</a>.</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains: Data

<p>This repository contains the data for the following paper:</p> <p><span>Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, C&eacute;cile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norr&eacute;, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperan&ccedil;a-Rodier, Thomas Fran&ccedil;ois, Lorraine Goeuriot, J&eacute;r&ocirc;me Goulian, Mathieu Lafourcade, et al.. 2024.&nbsp;<a href="https://aclanthology.org/2024.lrec-main.827">Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains</a>. In <em>Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)</em>, pages 9463&ndash;9476, Torino, Italia. ELRA and ICCL.</span></p> <p>1) pretraining data for the Jargon specialized language models</p> <p>2) ECTHR_FR dataset for text classification in the French legal domain</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages

<p>This is the GitHub repository hosting data and code for SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages. For additional information about the task, please consult the official website.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Data-driven brain network models differentiate variability across language tasks

<p>Data and script associated with the manuscript titled &quot;Data-driven brain network models differentiate variability<br> across language tasks&quot;.&nbsp;</p>

opencc-by-4.0Sep 2018View details →
zenodo32/100

A Large Dataset of Tweets on the 2023 Presidential Elections in Nigeria for Natural Language Processing Tasks

<p>The dataset contains tweets related to the 2023 presidential elections in Nigeria. The data was retrieved from the social media&nbsp;network, Twitter (Now X) between February 4<sup>th</sup>, 2023 and April 4<sup>th</sup>, 2023. The hashtags from the official handles and other popular hashtags endorsed and/or representing the candidates of each party were considered for retrieving election related tweets using an API from Twitter social media platform. Three major political parties in Nigeria were considered and they have been labelled as Party A, Party L and Party P in this dataset. The party or group called &quot;General&quot; contains tweets from the Independent National Electoral Commission (INEC) hashtags such as <em>@inecnigeria</em> and <em>#2023election</em> which is not directly for any political party.</p> <p>The dataset has been pre-processed lightly to make it very useful to researcher for a wide range of natural language processing tasks like sentiment analysis, topic modelling, fake news detection, emotion detection, election stance, etc.</p> <p>Details of the dataset collection such as hashtags, retrieved tweets, duplicates removed, and the remaining unique tweets is presented in Table 1.&nbsp;</p> <p>&nbsp;</p> <p>Table 1: Tweets collection and duplicates removal</p> <table> <tbody> <tr> <td> <p>S/N</p> </td> <td> <p>Party</p> </td> <td> <p>Hash tags</p> </td> <td> <p>Retrieved tweets</p> </td> <td> <p>Duplicates tweets</p> </td> <td> <p>Unique tweets</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; X</p> </td> <td> <p><em>@inecnigeria</em></p> <p><em>#2023election</em></p> </td> <td> <p>64,496</p> </td> <td> <p>47,275</p> </td> <td> <p>17,195</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; A</p> <p>&nbsp;</p> </td> <td> <p><em>#TinubuIsComing</em></p> <p><em>#emilokan </em></p> <p><em>#jagabanarmy</em></p> <p><em>#RenewedHope </em></p> <p><em>#BATKSM2023</em></p> </td> <td> <p>263,870</p> </td> <td> <p>231,036</p> </td> <td> <p>32,832</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; L</p> </td> <td> <p><em>#VoteLP</em></p> <p><em>#NigeriaMustBeBright</em></p> <p><em>#PeterObiForPresident2023</em></p> <p><em>#ObiDatti2023 </em></p> <p><em>#PeterObi</em></p> </td> <td> <p>664,083</p> </td> <td> <p>310,857</p> </td> <td> <p>353,226</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>&nbsp;</p> <p>P</p> </td> <td> <p><em>#NigeriaDecides</em></p> <p><em>#VotePDP</em></p> <p><em>#AtikuOkowa2023</em></p> <p><em>#FinalPushToVictory</em></p> <p><em>#RecoverNigeria</em></p> </td> <td> <p>387,450</p> </td> <td> <p>318,425</p> </td> <td> <p>66,227</p> </td> </tr> <tr> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>1,379,899</p> </td> <td> <p>907,593</p> </td> <td> <p>468,480</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>To encourage NLP tasks, we uploaded in this Version One the following files:</p> <ol> <li>The combined dataset with pre-processed tweets and their meta data but with&nbsp;removed duplicates are in the file labelled &ldquo;Combined Dataset Pre-processed without duplicates.csv&rdquo;</li> <li>General statistics on each corpus is in the file labelled &ldquo;Dataset Statistics.xlsx&rdquo;&nbsp;&nbsp;</li> <li>The preprocessed corpus from the general group with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_GENERAL.xlsx&rdquo;</li> <li>The preprocessed corpus from Party A with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party A.xlsx&rdquo;</li> <li>The preprocessed corpus from Party L with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party L.xlsx&rdquo;</li> <li>The preprocessed corpus from Party P with the tweet contents only is in file labelled &ldquo;Preprocessed_Tweet only_Party P.xlsx&rdquo;</li> <li>The top 100 frequent tokens are in the file labelled &ldquo;Top 100 Tokens and weights.xlsx&rdquo;</li> <li>The top frequent bigrams and their weights are&nbsp;in the file labelled &ldquo;Top 100 Bigrams and weights.xlsx&rdquo;</li> <li>The top frequent trigrams and their weights are in the file labelled &ldquo;Top 100 Trigrams and weights.xlsx&rdquo;</li> </ol>

opencc-by-4.0Sep 2023View details →
zenodo28/100

REVOLUTIONIZING ENGLISH TEXT TRANSLATION: TASK-BASED ACTIVITIES FOR EFFECTIVE LANGUAGE LEARNING

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

PRIORITY TASKS OF TEACHING ENGLISH LANGUAGE

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

CodeQual: A dataset for fine-tuning Large Language Models for code quality assessment task

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo24/100

THE ROLE OF TASK-BASED LANGUAGE TEACHING (TBLT) IN ENHANCING SPEAKING SKILLS

<p><em><span>Task-Based Language Teaching (TBLT) has emerged as a significant approach in language education, focusing on the use of meaningful tasks to improve learners' language skills. This article examines the role of TBLT in enhancing speaking skills, exploring its theoretical underpinnings, implementation strategies, and effectiveness. A study involving intermediate ESL learners was conducted to evaluate the impact of TBLT on speaking proficiency. Results indicate that TBLT significantly improves learners' fluency, accuracy, and overall confidence in speaking. The discussion addresses the practical implications of these findings for language educators and provides recommendations for optimizing TBLT in classroom settings.</span></em></p>

openJul 2024View details →
zenodo20/100

WMT-SLT SRF: Training data for the WMT shared task on sign language translation (videos, subtitles)

<p>These are Standard German daily news (Tagesschau) and Swiss German weather forecast (Meteo) episodes broadcast and interpreted into Swiss German Sign Language by hearing interpreters (among them, children of Deaf adults, CODA) via Swiss National TV (Schweizerisches Radio und Fernsehen, SRF)&nbsp;(<a href="https://www.srf.ch/play/tv/sendung/tagesschau-in-gebaerdensprache?id=c40bed81-b150-0001-2b5a-1e90e100c1c0">https://www.srf.ch/play/tv/sendung/tagesschau-in-gebaerdensprache?id=c40bed81-b150-0001-2b5a-1e90e100c1c0</a>). For a more extended description of the data, visit&nbsp;<a href="https://www.wmt-slt.com/data">https://www.wmt-slt.com/data</a>.<br> <br> &nbsp;</p>

restrictedJun 2022View details →
zenodo16/100

【outputs_encode_to_representations】only two 8B LLMs, only FB task, only augmented responses language materials w/o their representations, no original prompt or response

<ul> <li>only two 8B LLMs,</li> <li>only FB task,</li> <li>only augmented responses language materials w/o their representations, no original prompt or response</li> <li>19 (valid prompts) ⨉ 2 (LLMs) = 38 (files)</li> </ul>

restrictedcc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record