Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

35 results for “speech recognition”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

<p>We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words

<p>Hi,KIA dataset is a shared short Wakeup Word&nbsp;database focusing on perceived emotion in&nbsp;speech The dataset contains&nbsp;<strong>488 </strong>Wakeup Word&nbsp;speech.&nbsp;</p> <p>For more detailed information about the dataset, please refer to our paper:&nbsp;Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>:&nbsp;wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav`&nbsp;The first letter was used to express emotion.<br> &nbsp;</li> </ul> </li> <li><em><strong>annotation/</strong></em>:&nbsp;Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv&nbsp;</p> </li> <li> <p><em><strong>handcraft:</strong></em>&nbsp;Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em>&nbsp;wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p>&nbsp;</p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> &nbsp; title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> &nbsp; author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> &nbsp; booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> &nbsp; year={2022}<br> }<br> ```</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Changes in neuronal representations of phonemes in the ascending auditory system and their role speech recognition

<p>This dataset comprises neural responses to a set of speech sounds from several brain regions. Auditory nerve data was simulated using a computational model of the auditory nerve. Also included are multi-unit extracellular recordings or responses to the same stimuli in the inferior colliculus and auditory cortex of anaethetised guinea pigs.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

Appendix: Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition

<p>Appendix tables for the paper &quot;Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition&quot;.</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Automatic speech recognition datasets for Gronings, Nasal, and Besemah

<p>Automatic speech recognition datasets for Gronings, Nasal, and Besemah for experiments reported in Bartelds, San,&nbsp;McDonnell,&nbsp;Jurafsky and&nbsp;Wieling (2023).&nbsp;<em>Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</em>. ACL 2023.</p> <p>Model training code available at:&nbsp;https://github.com/Bartelds/asr-augmentation</p>

opencc-by-4.0May 2023View details →
zenodo36/100

voiceHome-2 corpus - automatic speech recognition baseline - acoustic model

<p>This entry contains the acoustic model used for evaluation of distant-microphone speech recognition performance in:</p> <p>Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent, Sunit Sivasankaran, Irina Illina, Fr&eacute;d&eacute;ric Bimbot<br> <a href="https://hal.inria.fr/hal-01923108">VoiceHome-2, an extended corpus for multichannel speech processing in real homes</a><br> <em>Speech Communication</em>, 2019, 106, pp.68-78.&nbsp;<a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">&lang;10.1016/j.specom.2018.11.002&rang;</a></p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

EEG data for "Conversation electrified: ERP correlates of speech act recognition in underspecified utterances"

<p>Please refer to the publication in Plos One for a description of the experiment and data analysis: &nbsp; Gisladottir RS, Chwilla DJ, Levinson SC (2015) Conversation Electrified: ERP Correlates of Speech Act Recognition in Underspecified Utterances. PLoS ONE 10(3): e0120068. doi: 10.1371/journal.pone.0120068</p>

opencc-by-4.0Mar 2015View details →
zenodo36/100

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).

<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Data and codes: Speech-recognition in landlide predictive modelling

<p>This is the data and codes for the manuscript &quot;Speech-recognition in landlide predictive modelling&quot;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Appendix - Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition

<p>Appendix tables for the paper &quot;Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition&quot;.</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Integration of speech separation, diarization, and recognition for multi-speaker meetings: Separated LibriCSS dataset

<p><strong>Dataset</strong></p> <p>This data repository contains separated audio streams for the LibriCSS dataset using the following window-based separation methods:</p> <p>1. <em>Mask-based MVDR</em>: Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, &ldquo;Multi-microphone neural speech separation for farfield multi-talker speech recognition,&rdquo; ICASSP 2018.</p> <p>2. <em>Sequential neural beamforming</em>: &nbsp;Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, and John R. Hershey, &ldquo;Sequential multi-frame neural beamforming for speech separation and enhancement,&rdquo; IEEE SLT 2021.</p> <p>These audio streams were used for evaluating the diarization and ASR models in our <a href="https://arxiv.org/pdf/2011.02014.pdf">JSALT 2020 paper</a>.</p> <p>The repository contains the following archive files:</p> <ul> <li>libricss_mvdr_2stream.tar.gz</li> <li>libricss_sequential_3stream.tar.gz</li> </ul> <p><strong>Citation</strong></p> <p>If you use these separated audio streams in your research, consider citing:</p> <pre><code>@article{Raj2020IntegrationOS,   title={Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis},   author={Desh Raj and Pavel Denisov and Z. Chen and H. Erdogan and Zili Huang and Mao-Kui He and Shinji Watanabe and Jun Du and T. Yoshioka and Yi Luo and N. Kanda and Jinyu Li and S. Wisdom and J. Hershey},   journal={2021 IEEE Spoken Language Technology (SLT) Workshop},   year={2021} }</code></pre> <p><br> &nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain

<p>Dear User,</p> <p>We are thrilled to introduce our latest release - the <strong>RescueSpeech</strong>&nbsp;audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.</p> <p>The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.</p> <p>1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.</p> <p>2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.</p> <p>By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.<br> &nbsp;</p>

opencc-by-nc-4.0Jun 2023View details →
zenodo32/100

Masked speech recognition by 6-13-year-olds with early-childhood otitis media: Effects of acoustic condition and otologic history

<p>Data and code used for the statistical analyses reported in the accompanying manuscript by Koiek S, Brandt C, M&ouml;ller S, Dillon H &amp; Neher T. For further details, contact tneher(AT)health.sdu.dk</p>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov32/100

Automated Telephone Outreach With Speech Recognition to Improve Diabetes Care: A Randomized Controlled Study

ClinicalTrials.gov study NCT00790530. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers

<p>Test Cases for&nbsp;Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers</p>

opencc-by-4.0May 2023View details →
ClinicalTrials.gov28/100

Speech Recognition Training in Children With Hearing Loss

ClinicalTrials.gov study NCT04041440. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov28/100

A Speech Recognition Application as a Communication Aid for Acute and Critical Care Patients With Tracheostomies

ClinicalTrials.gov study NCT06027866. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov28/100

Optimizing Soft Speech Recognition in Children With Hearing Loss

ClinicalTrials.gov study NCT05299892. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record