Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5 results for “Speech augmentation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from&nbsp;synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 2 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>second</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

CpAug: Refining Copy-Paste Augmentation for Speech Anti-Spoofing

<p>Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof&rsquo; 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof&rsquo; 21LA.</p>

opencc-by-4.0Feb 2024View details →
ClinicalTrials.gov32/100

Speech Production Enhancement Using Augmentative Communication for Kids

ClinicalTrials.gov study NCT07173049. IPD Sharing: YES. Countries: 1. Publications: 6.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov24/100

Noise-augmented Automatic Speech Recognition for Speech Treatment in Parkinson's Disease

ClinicalTrials.gov study NCT06540989. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record