Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo40/100

Sudanese dialect speech dataset

<p>This is speech dataset for the Sudanese dialect data been collected from YouTube videos represent the characteristics of the Sudanese dialect, mainly the middle of Sudan dialect -Khartoum in particular- and have some northern tendency, primarily two programs Hajj Muzakir and Dukkan Wad Elbaseer.</p> <p>Transcription is done manually by listening to the audio files repeatedly to write the captions for the collected conversations to make sure that every word is written as said by the speakers. Transcription is written without diacritics on the Arabic alphabet, in a manner that reflects the Sudanese way of speaking, therefore, any correction to the noticeable mistakes was not applied to get rid of any biases and make the data representative.</p> <p>The &#39;Dataset&#39; subdirectory contains all the audio and text files for the corpus, the files organized based on program name &#39;hm_&#39; for Hajj Muzakir program and &#39;wb_&#39; Dukkan Wad Elbaseer, each filename follows three categories first two litters for the program name &#39;hm&#39; or &#39;wb&#39;, second the number of the episode third the number of the clip, hm_01_0001.wav and wb_01_0001.wav represent first episode of each program and the first clip.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining

<p>The <em>VivesDebate-Speech</em> corpus contains the acoustic information of 29 different&nbsp;argumentative debates and the annotations of the segmentation (i.e., BIO tags) of the Argumentative Discourse Units identified in the spoken natural language discourse.&nbsp;</p>

opencc-by-nc-sa-4.0Sep 2022View details →
zenodo40/100

EEG Dataset for 'Decoding of selective attention to continuous speech from the human auditory brainstem response' and 'Neural Speech Tracking in the Theta and in the Delta Frequency Band Differentially Encode Clarity and Comprehension of Speech in Noise'.

<p>The repository contains the unprocessed EEG data recorded for the publications [1, 2]. For convenience, the onsets of the EEG data provided here are time-aligned with the onsets of the audio books in the &#39;audiobooks&#39; folder, and the EEG data are provided in HDF5 format. Please refer to the original version of this dataset for more details.</p> <p>More details, as well as the original data files, are available at the original repository&nbsp;<a href="https://doi.org/10.5281/zenodo.7086209">here</a>.</p> <p>Examples of using these data (preprocessing, fitting linear models) can be found&nbsp;<a href="https://github.com/Mike-boop/trf-examples">here</a>.</p> <p>The English conditions (clean, lb, mb, hb, fM, fW) comprised a single recording session. The Dutch conditions&nbsp;(cleanDutch, lbDutch, mbDutch, hbDutch) comprised a separate recording session. You see which participants took part in each session in session_info.json.</p> <p>Please note some details about the stimulus presentation for the various listening conditions:</p> <ul> <li>English speech-in-babble-noise (lb, mb, hb): babble noise was played by itself for one second before the audiobook track began. The babble noise was also played for one second after the audiobook track ended. Therefore, you should discard the first second and the last second from these trial during your analysis.</li> <li>Dutch speech-in-babble-noise (lbDutch, mbDutch, hbDutch): the story (narrated in Dutch) was played by itself for one second before the babble noise track began. Then, the babble noise was increased linearly in amplitude for one second. Therefore, you should discard the first two seconds from these trials during your analysis.</li> <li>Dutch in quiet, and Dutch-in-babble-noise&nbsp;(cleanDutch, lbDutch, mbDutch, hbDutch): some English sentences were embedded in the Dutch narratives in order to encourage attention. You should crop these from your analysis. The onsets and offsets of the English sentences (in samples, at 44100Hz) are provided in the audiobooks/*Dutch/english_onsets_info.json files.</li> <li>Competing-speakers conditions (fM, fW): sometimes the attended track is longer than the unattended track, or vice-versa. The onsets of both tracks are aligned. You should crop the trial to the length of the shortest track for your analysis.</li> </ul> <p>If you use this data, please cite the original publications, as well as this repository [1,2,3].</p> <p>[1] Etard O, Kegler M, Braiman C, Forte A E and Reichenbach T. &ldquo;Decoding of selective attention to continuous speech from the human auditory brainstem response&rdquo; 2019.&nbsp;<em>NeuroImage</em>&nbsp;<strong>200</strong>&nbsp;1&ndash;11</p> <p>[2] Etard O and Reichenbach T. &ldquo;Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise&rdquo; 2019.&nbsp;<em>J. Neurosci.</em>&nbsp;<strong>39</strong>&nbsp;5750&ndash;9</p> <p>[3] Etard O and Reichenbach T. &quot;EEG Dataset for &#39;Decoding of selective attention to continuous speech from the human auditory brainstem response&#39; and &#39;Neural Speech Tracking in the Theta and in the Delta Frequency Band Differentially Encode Clarity and Comprehension of Speech in Noise&quot;. Doi:&nbsp;10.5281/zenodo.7086208</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

EEG Dataset for 'Cortical Tracking of Surprisal during Continuous Speech Comprehension'

<p>The repository contains the&nbsp;unprocessed EEG data recorded for the publication [1]. For convenience, the onsets of the EEG data provided here are time-aligned with the onsets of the audio books in the &#39;audiobooks&#39; folder, and the EEG data are provided in HDF5 format. Please refer to the original version of this dataset for more details.</p> <p>A script&nbsp;is provided&nbsp;which shows how the original data were aligned (&#39;align_data.py&#39;) in Python. The order in which the audiobook chapters were presented was different for different participants. If this is important to you, the details are available in &#39;stimulus_orders.csv&#39;.</p> <p>More details, as well as the original data files, are available at the original version of this repository&nbsp;<a href="https://doi.org/10.5281/zenodo.7086168">here</a>.</p> <p>Examples for using this data (preprocessing, fitting linear models) can be found&nbsp;<a href="https://github.com/Mike-boop/trf-examples">here</a>.</p> <p>If you use this data, please cite the original publication, as well as this repository [1,2].</p> <p>[1] Weissbart H, Kandylaki KD, Reichenbach T. &ldquo;Cortical Tracking of Surprisal during Continuous Speech Comprehension&rdquo;. J Cogn Neurosci. 2020 Jan;32(1):155-166. doi: 10.1162/jocn_a_01467.</p> <p>[2] Weissbart H, Kandylaki KD, Reichenbach T. &ldquo;EEG Dataset for &#39;Cortical Tracking of Surprisal during Continuous Speech Comprehension&#39;&rdquo;. doi:&nbsp;&nbsp;10.5281/zenodo.7086167</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Appendix: Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition

<p>Appendix tables for the paper &quot;Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition&quot;.</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Automatic speech recognition datasets for Gronings, Nasal, and Besemah

<p>Automatic speech recognition datasets for Gronings, Nasal, and Besemah for experiments reported in Bartelds, San,&nbsp;McDonnell,&nbsp;Jurafsky and&nbsp;Wieling (2023).&nbsp;<em>Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</em>. ACL 2023.</p> <p>Model training code available at:&nbsp;https://github.com/Bartelds/asr-augmentation</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes

<p>The dataset was made for the purposes of the author&#39;s master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">&quot;Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo&rsquo;s Core Vocabulary&quot;</a>. The dataset includes two worksheets. The first is titled &quot;Main table&quot;, and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled &quot;Extended table&quot;, includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively).&nbsp;<br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">at&nbsp;this link</a>.&nbsp;&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Charades-STA Speech Caption Dataset

<ul> <li><strong>Dataset introduction: </strong>This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"</li> <li><strong>Associated Code: <a href="https://github.com/xian-sh/UniSDNet">https://github.com/xian-sh/UniSDNet</a></strong></li> <li><strong>Associated Paper: <a href="https://arxiv.org/abs/2403.14174">https://arxiv.org/abs/2403.14174</a><br></strong></li> <li><strong>Disclaimer: </strong>This dataset is for academic research only, non-commercial use, if you use this dataset please cite the <a href="https://arxiv.org/abs/2403.14174">associated paper</a></li> <li><strong>Cite:</strong></li> </ul> <p>@article{hu2024unified,<br>&nbsp; title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},<br>&nbsp; author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},<br>&nbsp; year={2024},<br>&nbsp; Journal={CoRR},<br>&nbsp; volume={abs/2403.14174},<br>}</p>

openJun 2023View details →
zenodo40/100

TACoS Speech Caption Dataset

<ul> <li><strong>Dataset Introduction: </strong>This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"</li> <li><strong>Associated Code: <a href="https://github.com/xian-sh/UniSDNet">https://github.com/xian-sh/UniSDNet</a></strong></li> <li><strong>Associated Paper: <a href="https://arxiv.org/abs/2403.14174">https://arxiv.org/abs/2403.14174</a><br></strong></li> <li><strong>Disclaimer: </strong>This dataset is for academic research only, non-commercial use, if you use this dataset please cite the <a href="https://arxiv.org/abs/2403.14174">associated paper</a></li> <li><strong>Cite:</strong></li> </ul> <p>@article{hu2024unified,<br>&nbsp; title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},<br>&nbsp; author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},<br>&nbsp; year={2024},<br>&nbsp; Journal={CoRR},<br>&nbsp; volume={abs/2403.14174},<br>}</p>

openJun 2023View details →
dryad40/100

Data for: A high-performance speech neuroprosthesis

<p>Brain-computer interfaces (BCIs) can restore communication to people who have lost the ability to move or speak. In this study, we demonstrated an intracortical BCI that decodes attempted speaking movements from neural activity in motor cortex and translates it to text in real-time, using a recurrent neural network decoding approach. With this BCI, our study participant, who can no longer speak intelligibly due to amyotrophic lateral sclerosis, achieved a 9.1% word error rate on a 50-word vocabulary and a 23.8% word error rate on a 125,000-word vocabulary. </p> <p>This dataset contains all of the neural activity recorded during these experiments, consisting of 10,850 spoken sentences as well as instructed delay experiments designed to investigate the neural representation of orofacial movement and speech production.</p> <p>The data have also been formatted for developing and evaluating machine learning decoding methods, and we intend to host a decoding competition. To this end, the data also contain files for reproducing our offline decoding results, including a language model and an example RNN decoder. </p> <p>Code associated with the data can be found here: <a href="https://github.com/fwillett/speechBCI">https://github.com/fwillett/speechBCI</a>.</p>

opencc-zeroJun 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from&nbsp;synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 2 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>second</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Simulated + Clutter)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains 2 distinct datasets:&nbsp;</p> <ol> <li>A&nbsp;dataset of speech mixtures containing 2-5 speakers simulated using PyRoomAcoustics. The dataset consists of 8000 training mixtures, 500 validation mixtures and 1000 testing mixtures.</li> <li>A dataset of speech mixtures containing 3-5 speakers created from synchronized recordings in reverberant rooms with objects cluttering the table.&nbsp;The dataset consists of 500 testing mixtures.</li> </ol> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>Please see the Readme for more infromation.&nbsp;Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Clear Speech Data for Syllable-Rate-Adjusted-Modulation (SRAM)

<p>This Dataset is associated with the paper &quot;Syllable-Rate-Adjusted-Modulation (SRAM) Predicts Clear and Conversational Speech Intelligibility&quot;. It contains 144 sentences recorded from two talkers (one female, one male) in both clear and conversational styles (72 sentences in each style). The sample rate was 16000Hz. The silence periods before and after the speech were removed. The speech scripts for each speech style and the human performance are included in each sub-folder.<br> SSN.wav is the steady-state noise used to create the noisy speeches.</p> <p>File Structure:<br> - Female<br> &nbsp;&nbsp;&nbsp; - Clear<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 1.wav<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - ...<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 72.wav<br> &nbsp;&nbsp;&nbsp; - Convo<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 1.wav<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - ...<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; - 72.wav<br> &nbsp;&nbsp;&nbsp; - human_results.csv<br> &nbsp;&nbsp;&nbsp; - key_words_clear.txt<br> &nbsp;&nbsp;&nbsp; - key_words_conv.txt<br> - Male<br> - SSN.wav</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Annotations for NP13 proceedings paper of Banzina & Niebuhr: How to pause in charismatic speeches: A case study of Barack Obama

<p>TextGrid file used in the above specified study as well as a readme document that includes a link to the corresponding/analyzed YouTube video and the time stamps from where to where the audio was analyzed. Both should enable users to (a) extract the audio from the YouTube video and (b) time align the extracted audio with the TextGrid file. Note that the PRAAT software tool (praat.org) is required to open TextGrid files properly.</p>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov40/100

Effects of PSAPs on Speech Processing

ClinicalTrials.gov study NCT05076045. IPD Sharing: YES. Countries: 1. Publications: 2.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov40/100

Perceptual Training to Improve Listeners' Ability to Understand Speech Produced by Individuals With Dysarthria

ClinicalTrials.gov study NCT04897711. IPD Sharing: YES. Countries: 1. Publications: 11.

controlledIPD-YESFeb 2026View details →
dryad40/100

Repeatedly experiencing the McGurk effect induces long-lasting changes in auditory speech perception

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Data for: An accurate and rapidly calibrating speech neuroprosthesis

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad40/100

Data for: A high-performance speech neuroprosthesis

Open the record for dataset details and reuse information.

publicSep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record