Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.9.0
Dataset results
35 results for “speech recognition”
Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic
<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>
BembaSpeech: A Speech Recognition Corpus for the Bemba Language
<p>We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language</p>
Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words
<p>Hi,KIA dataset is a shared short Wakeup Word database focusing on perceived emotion in speech The dataset contains <strong>488 </strong>Wakeup Word speech. </p> <p>For more detailed information about the dataset, please refer to our paper: Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>: wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav` The first letter was used to express emotion.<br> </li> </ul> </li> <li><em><strong>annotation/</strong></em>: Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv </p> </li> <li> <p><em><strong>handcraft:</strong></em> Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em> wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p> </p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> year={2022}<br> }<br> ```</p>
Changes in neuronal representations of phonemes in the ascending auditory system and their role speech recognition
<p>This dataset comprises neural responses to a set of speech sounds from several brain regions. Auditory nerve data was simulated using a computational model of the auditory nerve. Also included are multi-unit extracellular recordings or responses to the same stimuli in the inferior colliculus and auditory cortex of anaethetised guinea pigs.</p>
Appendix: Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition
<p>Appendix tables for the paper "Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition".</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>
Automatic speech recognition datasets for Gronings, Nasal, and Besemah
<p>Automatic speech recognition datasets for Gronings, Nasal, and Besemah for experiments reported in Bartelds, San, McDonnell, Jurafsky and Wieling (2023). <em>Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</em>. ACL 2023.</p> <p>Model training code available at: https://github.com/Bartelds/asr-augmentation</p>
voiceHome-2 corpus - automatic speech recognition baseline - acoustic model
<p>This entry contains the acoustic model used for evaluation of distant-microphone speech recognition performance in:</p> <p>Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent, Sunit Sivasankaran, Irina Illina, Frédéric Bimbot<br> <a href="https://hal.inria.fr/hal-01923108">VoiceHome-2, an extended corpus for multichannel speech processing in real homes</a><br> <em>Speech Communication</em>, 2019, 106, pp.68-78. <a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">⟨10.1016/j.specom.2018.11.002⟩</a></p>
EEG data for "Conversation electrified: ERP correlates of speech act recognition in underspecified utterances"
<p>Please refer to the publication in Plos One for a description of the experiment and data analysis: Gisladottir RS, Chwilla DJ, Levinson SC (2015) Conversation Electrified: ERP Correlates of Speech Act Recognition in Underspecified Utterances. PLoS ONE 10(3): e0120068. doi: 10.1371/journal.pone.0120068</p>
Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects
<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>
A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).
<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>
Data and codes: Speech-recognition in landlide predictive modelling
<p>This is the data and codes for the manuscript "Speech-recognition in landlide predictive modelling"</p>
Appendix - Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition
<p>Appendix tables for the paper "Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition".</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>
Integration of speech separation, diarization, and recognition for multi-speaker meetings: Separated LibriCSS dataset
<p><strong>Dataset</strong></p> <p>This data repository contains separated audio streams for the LibriCSS dataset using the following window-based separation methods:</p> <p>1. <em>Mask-based MVDR</em>: Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, “Multi-microphone neural speech separation for farfield multi-talker speech recognition,” ICASSP 2018.</p> <p>2. <em>Sequential neural beamforming</em>: Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, and John R. Hershey, “Sequential multi-frame neural beamforming for speech separation and enhancement,” IEEE SLT 2021.</p> <p>These audio streams were used for evaluating the diarization and ASR models in our <a href="https://arxiv.org/pdf/2011.02014.pdf">JSALT 2020 paper</a>.</p> <p>The repository contains the following archive files:</p> <ul> <li>libricss_mvdr_2stream.tar.gz</li> <li>libricss_sequential_3stream.tar.gz</li> </ul> <p><strong>Citation</strong></p> <p>If you use these separated audio streams in your research, consider citing:</p> <pre><code>@article{Raj2020IntegrationOS, title={Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis}, author={Desh Raj and Pavel Denisov and Z. Chen and H. Erdogan and Zili Huang and Mao-Kui He and Shinji Watanabe and Jun Du and T. Yoshioka and Yi Luo and N. Kanda and Jinyu Li and S. Wisdom and J. Hershey}, journal={2021 IEEE Spoken Language Technology (SLT) Workshop}, year={2021} }</code></pre> <p><br> </p>
RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain
<p>Dear User,</p> <p>We are thrilled to introduce our latest release - the <strong>RescueSpeech</strong> audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.</p> <p>The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.</p> <p>1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.</p> <p>2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.</p> <p>By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.<br> </p>
Masked speech recognition by 6-13-year-olds with early-childhood otitis media: Effects of acoustic condition and otologic history
<p>Data and code used for the statistical analyses reported in the accompanying manuscript by Koiek S, Brandt C, Möller S, Dillon H & Neher T. For further details, contact tneher(AT)health.sdu.dk</p>
Automated Telephone Outreach With Speech Recognition to Improve Diabetes Care: A Randomized Controlled Study
ClinicalTrials.gov study NCT00790530. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers
<p>Test Cases for Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers</p>
Speech Recognition Training in Children With Hearing Loss
ClinicalTrials.gov study NCT04041440. IPD Sharing: NO. Countries: 1. Publications: 0.
A Speech Recognition Application as a Communication Aid for Acute and Critical Care Patients With Tracheostomies
ClinicalTrials.gov study NCT06027866. IPD Sharing: NO. Countries: 1. Publications: 0.
Optimizing Soft Speech Recognition in Children With Hearing Loss
ClinicalTrials.gov study NCT05299892. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.