Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

8 results for “automatic speech recognition”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Automatic speech recognition datasets for Gronings, Nasal, and Besemah

<p>Automatic speech recognition datasets for Gronings, Nasal, and Besemah for experiments reported in Bartelds, San,&nbsp;McDonnell,&nbsp;Jurafsky and&nbsp;Wieling (2023).&nbsp;<em>Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</em>. ACL 2023.</p> <p>Model training code available at:&nbsp;https://github.com/Bartelds/asr-augmentation</p>

opencc-by-4.0May 2023View details →
zenodo36/100

voiceHome-2 corpus - automatic speech recognition baseline - acoustic model

<p>This entry contains the acoustic model used for evaluation of distant-microphone speech recognition performance in:</p> <p>Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent, Sunit Sivasankaran, Irina Illina, Fr&eacute;d&eacute;ric Bimbot<br> <a href="https://hal.inria.fr/hal-01923108">VoiceHome-2, an extended corpus for multichannel speech processing in real homes</a><br> <em>Speech Communication</em>, 2019, 106, pp.68-78.&nbsp;<a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">&lang;10.1016/j.specom.2018.11.002&rang;</a></p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).

<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>

opencc-by-4.0Jul 2024View details →
zenodo28/100

Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers

<p>Test Cases for&nbsp;Aster: Automatic Speech Recognition System Accessibility Testing for Stutterers</p>

opencc-by-4.0May 2023View details →
zenodo24/100

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

<p>The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda.</p> <p>The corpus of 155&nbsp;hours is publicly available online under the Creative Commons BY-NC-ND 4.0 license.&nbsp;The dataset release is comprised of the following:</p> <ol> <li>20 hours of human transcribed radio speech. The audio is 16kHZ, mono channel and with 16 bit rate.&nbsp;</li> <li>Two CSV files for the 20-hour human transcribed dataset - cleaned.csv contains cleaned transcripts and uncleaned.csv contains uncleaned transcripts. The uncleaned transcripts contain extra speech details included in tags like [laughter] for laughter, and [um] for filler pauses, which speaker is talking, where each speaker is assigned an identifier A or B.</li> <li>The transcription guide used to transcribe the radio dataset.</li> <li>&nbsp;A multi-speaker untranscribed dataset of 6 hours of radio data. 1.4 hours of women voices and 4.6 hours of men voices. Each audio is a ten-seconds clip with a single speaker.</li> <li>&nbsp;135&nbsp;hours of multi-speaker untranscribed radio data.</li> </ol> <p><strong>NOTE: You can read and cite our paper published in the&nbsp;</strong><a href="http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.208.pdf"><strong>Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)</strong></a> The Dataset is published under Creative Commons BY-NC-ND 4.0 license and in order for us to monitor who is using it for the right license we request that you reach out to us officially.&nbsp;</p>

restrictedcc-by-4.0Jan 2022View details →
ClinicalTrials.gov24/100

Noise-augmented Automatic Speech Recognition for Speech Treatment in Parkinson's Disease

ClinicalTrials.gov study NCT06540989. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record