Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.7.1
Dataset results
7 results for “Speech separation”
First Speech Separation Challenge
<p>The first international Speech Separation Challenge took place in 2006, with results disseminated at Interspeech in Pittsburgh and later in a special issue of Computer Speech and Language (volume 24, 2010). The main focus of the challenge was to compare algorithms and human listeners on the task of identifying words in sentences from one talker when mixed with similar utterances from another talker, using a single channel (i.e., operating monaurally). Development and test data was also provided for a stationary noise masking condition. For technical details and a summary of the outcome of the Challenge, see the article by Cooke, Hershey and Rennie (pp 1--15) of the special issue (available as cooke_csl2010.pdf in this dataset). </p> <p>The current dataset consisted of the following zip files:</p> <p>twotalker_dev.zip two-talker development set<br> twotalker_test.zip two-talker test set<br> 1.zip, 2.zip etc training data for each of 34 talkers (500 sentences each)<br> ssn_dev.zip speech-shaped noise (SSN) development set<br> ssn_test.zip SSN test set<br> <br> The Challenge was funded by the EU Network of Excellence PASCAL (Pattern Analysis, Statistical Modeling and Computational Learning) </p>
PodcastMix - a dataset for separating music and speech in podcasts
<p><strong>Note: due to zenodo limitations here we host solely the metadata. the whole dataset can be found at: https://drive.google.com/drive/u/0/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z </strong></p> <p>We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. It contains audio files at 44.1kHz and the corresponding metadata. For further details check the following paper and the associated GitHub repository: </p> <ul> <li>N. Schmidt, J. Pons, M. Miron, "PodcastMix - a dataset for separating music and speech in podcasts", Interspeech (2022)</li> <li>N. Schmidt, "PodcastMix - a dataset for separating music and speech in podcasts", Masters thesis, MTG, UPF (2021) https://zenodo.org/record/5554790#.YXLHvNlByWA </li> <li>https://github.com/MTG/Podcastmix</li> </ul> <p>This dataset contains four parts. Due to zenodo file size limitation we host the training dataset on google drive. We highlight the content of the zenodo archives within brackets:</p> <ul> <li>[metadata] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[metadata] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[audio and metadata] PodcastMix-real with-reference : a test set with real podcasts with reference stems to compute evaluation metrics. The podcasts are recorded by one of the authors and the source of the music is the FMA dataset. </li> <li>[audio and metadata] PodcastMix-real no-reference: a test set with real podcasts with only the podcasts mixes for subjective evaluation. The podcasts are compiled from the internet. </li> </ul> <p>The training dataset, PodcastMix-synth may be found at our google drive repository: https://drive.google.com/drive/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z?usp=sharing . The archive comprises 450GB of audio and metadata with the following structure:</p> <ul> <li>[metadata and audio] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[metadata and audio] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> </ul> <p>Make sure you maintain the folder structure of the original dataset when you uncompress these files. </p> <p><br> This dataset is created by Nicolas Schmidt, Marius Miron, Music Technology Group - Universitat Pompeu Fabra (Barcelona) and Jordi Pons. This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 Unported License (CC BY-SA 4.0).</p> <p><br> Please acknowledge PodcastMix in Academic Research. When the present dataset is used for academic research, we would highly appreciate if authors quote the following publications:</p> <ul> <li>N. Schmidt, J. Pons, M. Miron, "PodcastMix - a dataset for separating music and speech in podcasts", Interspeech (2022)</li> <li>N. Schmidt, "PodcastMix - a dataset for separating music and speech in podcasts", Masters thesis, MTG, UPF (2021) https://zenodo.org/record/5554790#.YXLHvNlByWA </li> </ul> <p><br> The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the UPF is not liable for, and expressly excludes, all liability for loss or damage however and whenever caused to anyone by any use of the dataset or any part of it.</p> <p><br> PURPOSES. The data is processed for the general purpose of carrying out research development and innovation studies, works or projects. In particular, but without limitation, the data is processed for the purpose of communicating with Licensee regarding any administrative and legal / judicial purposes.<br> </p>
WHISPER SET 1: a dataset for multi-channel, multi-device speech separation and speech enhancement
<p>This dataset is <code>WHISPER SET 1,</code> a dataset for speech enhancement and source separation recorded with a Wireless Acoustic Sensor Network (WASN) called WHISPER <a href="https://ieeexplore.ieee.org/abstract/document/8110202">Kiselev2018</a>. The dataset contains samples for up to 4 concurrent speakers and speech in noise. The dataset was recorded in a room with low reverberation (T_60 = 0.2 s) and using 16 microphones. In general, each track contains first a calibration phase where each of the speakers sequentially is active alone for 15 seconds. Followed by 15 seconds of all the speakers together (plus noise in some cases). </p> <p>If you use this dataset please cite:</p> <ul> <li><strong>E. Ceolini, I. Kiselev and S. Liu, "Evaluating multi-channel multi-device speech separation algorithms in the wild: a hardware-software solution," in <em>IEEE/ACM Transactions on Audio, Speech, and Language Processing</em>.</strong></li> </ul> <p>===</p> <p>Each sample is a 16-channel wav file in which the order of the channel follows the following logic:</p> <p>0 - module 5 mic 1 1 - module 5 mic 2 2 - module 5 mic 3 3 - module 5 mic 4 4 - module 6 mic 1 5 - module 6 mic 2 6 - module 6 mic 3 7 - module 6 mic 4 8 - module 7 mic 1 9 - module 7 mic 2 10 - module 7 mic 3 11 - module 7 mic 4 12 - module 8 mic 1 13 - module 8 mic 2 14 - module 8 mic 3 15 - module 8 mic 4</p> <p>Refer to the <a href="https://github.com/SensorsAudioINI/WHISPER_SET_1/blob/master/WHISPER4_floor_annotated.png">floor plan</a> for a visual illustration of the microphone arrangement.</p> <p>The files are divided into two subfolders, one for the samples of speech enhancement and one for the samples of speech separation.</p> <ul> <li>In the folder of speech separation, the files are divided into subfolders defining the number of speakers in the mixtures (2, 3, or 4)</li> <li>In the folder of speech enhancement, the files are divided into subfolders following the SNR of the mixture (0, -5, -10 dB)</li> </ul> <p>Samples are ordered in folders. Each sample folder contains a 15 seconds 16-channels <code>mixture.wav</code> file, plus the 15 seconds 16-channels <code>calibX.wav</code> files one for each speaker alone or noise alone in the mixture. That is a sample with a mixture with 4 speakers will have 4 calibration files (calib1.wav, calib2.wav, calib3.wav, calib4.wav) and a mixture of a speaker plus noise will have 2 calibration files one for speech (calib1.wav) and one for noise (calib2.wav).</p> <p>== </p> <p>A Jupyter notebook is included to show an example of how to use the data of this dataset for speech separation and speech enhancement using beamforming. The notebook is dependent on <a href="https://github.com/Enny1991/beamformers">this beamforming library</a> and <a href="https://github.com/Enny1991/sep_eval">this tool</a> to evaluate the quality of the separation.</p> <p>==</p> <p>Refer to the README.md in the dataset for more information.</p> <p>For any question please contact enea.ceolini@gmail.com</p>
Relative Transfer Matrix for Low SNR Speech Separation from Noisy Sources in Reverberant Rooms
<p>This folder contains the supplementary audio files for the paper "Relative Transfer Matrix for Low SNR Speech Separation from Noisy Sources in Reverberant Rooms" submitted to <em>The Journal of the Acoustical Society of America</em>.</p>
The data and code for "Original Speech and Its Echo are Segregated and Separately Processed in the Human Brain"
<p>This dataset is associated with the manuscript "Original Speech and Its Echo are Segregated and Seperately Processed in the Human Brain", and provides the preprocessed MEG response (resampling to 100 Hz), auditory stimulus, individual quantitative observations underlying the data summarized in figures, and the analysis codes.</p>
Test dataset for separation of speech, traffic sounds, wind noise, and general sounds
<p>The dataset was generated as part of the paper:<br> Deep Complex U-Net Ensemble for Outdoor Urban Sound Source Separation,<br> K. Arendt, A. Szumaczuk, B. Jasik, P. Masztalski, K. Piaskowski, M. Matuszewski, K. Nowicki, P. Zborowski.</p> <p>It contains various sounds from the Audio Set [1] and spoken utterances from VCTK [2] and DNS [3] datasets.</p> <p>Contents:<br> sr_8k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_16k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_48k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/</p> <p>Each directory contains 512 audio samples in different sampling rate (sr_8k - 8 kHz, sr_16k - 16 kHz, sr_48k - 48 kHz).<br> The audio samples for each sampling rate are different as they were generated randomly and separately.<br> Each directory contains 5 subdirectories:<br> - mix_clean - mixed sources,<br> - s1 - source #1 (general sounds),<br> - s2 - source #2 (speech),<br> - s3 - source #3 (traffic sounds),<br> - s4 - source #4 (wind noise).</p> <p>The sound mixtures were generated by adding s2, s3, s4 to s1 with SNR ranging from -10 to 10 dB w.r.t. s1.</p> <p><br> REFERENCES:</p> <p>[1] Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman,<br> Aren Jansen, Wade Lawrence, R. Channing Moore,<br> Manoj Plakal, and Marvin Ritter, “Audio set: An ontology<br> and human-labeled dataset for audio events,” in<br> Proc. IEEE ICASSP 2017, New Orleans, LA, 2017.</p> <p>[2] Christophe Veaux, Junichi Yamagishi, and Kirsten Mac-<br> Donald, “CSTR VCTK corpus: English multi-speaker<br> corpus for CSTR voice cloning toolkit, [sound],”<br> https://doi.org/10.7488/ds/1994, University of Edinburgh.<br> The Centre for Speech Technology Research<br> (CSTR). 2017.</p> <p>[3] Chandan K. A. Reddy, Ebrahim Beyrami, Harishchandra<br> Dubey, Vishak Gopal, Roger Cheng, Ross Cutler,<br> Sergiy Matusevych, Robert Aichner, Ashkan Aazami,<br> Sebastian Braun, Puneet Rana, Sriram Srinivasan, and<br> Johannes Gehrke, “The interspeech 2020 deep noise<br> suppression challenge: Datasets, subjective speech<br> quality and testing framework,” 2020.</p>
Integration of speech separation, diarization, and recognition for multi-speaker meetings: Separated LibriCSS dataset
<p><strong>Dataset</strong></p> <p>This data repository contains separated audio streams for the LibriCSS dataset using the following window-based separation methods:</p> <p>1. <em>Mask-based MVDR</em>: Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, “Multi-microphone neural speech separation for farfield multi-talker speech recognition,” ICASSP 2018.</p> <p>2. <em>Sequential neural beamforming</em>: Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, and John R. Hershey, “Sequential multi-frame neural beamforming for speech separation and enhancement,” IEEE SLT 2021.</p> <p>These audio streams were used for evaluating the diarization and ASR models in our <a href="https://arxiv.org/pdf/2011.02014.pdf">JSALT 2020 paper</a>.</p> <p>The repository contains the following archive files:</p> <ul> <li>libricss_mvdr_2stream.tar.gz</li> <li>libricss_sequential_3stream.tar.gz</li> </ul> <p><strong>Citation</strong></p> <p>If you use these separated audio streams in your research, consider citing:</p> <pre><code>@article{Raj2020IntegrationOS, title={Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis}, author={Desh Raj and Pavel Denisov and Z. Chen and H. Erdogan and Zili Huang and Mao-Kui He and Shinji Watanabe and Jun Du and T. Yoshioka and Yi Luo and N. Kanda and Jinyu Li and S. Wisdom and J. Hershey}, journal={2021 IEEE Spoken Language Technology (SLT) Workshop}, year={2021} }</code></pre> <p><br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.