Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
94
datasets available to search
ShareScore release 0.7.1
Dataset results
94 results for “speech dataset”
TIMIT-TTS: a Text-to-Speech Dataset for Synthetic Speech Detection
<p>With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple that malicious users can create unpleasant situations with minimal effort. Also, forged media are getting more and more complex, with manipulated videos (e.g., deepfakes where both the visual and audio contents can be counterfeited) that are taking the scene over still images.<br> The multimedia forensic community has addressed the possible threats that this situation could imply by developing detectors that verify the authenticity of multimedia objects. However, the vast majority of these tools only analyze one modality at a time.<br> This was not a problem as long as still images were considered the most widely edited media, but now, since manipulated videos are becoming customary, performing monomodal analyses could be reductive. Nonetheless, there is a lack in the literature regarding multimodal detectors (systems that consider both audio and video components). This is due to the difficulty of developing them but also to the scarsity of datasets containing forged multimodal data to train and test the designed algorithms.</p> <p>In this paper we focus on the generation of an audio-visual deepfake dataset.<br> First, we present a general pipeline for synthesizing speech deepfake content from a given real or fake video, facilitating the creation of counterfeit multimodal material. The proposed method uses Text-to-Speech (TTS) and Dynamic Time Warping (DTW) techniques to achieve realistic speech tracks. Then, we use the pipeline to generate and release TIMIT-TTS, a synthetic speech dataset containing the most cutting-edge methods in the TTS field. This can be used as a standalone audio dataset, or combined with DeepfakeTIMIT and VidTIMIT video datasets to perform multimodal research. Finally, we present numerous experiments to benchmark the proposed dataset in both monomodal (i.e., audio) and multimodal (i.e., audio and video) conditions.<br> This highlights the need for multimodal forensic detectors and more multimodal deepfake data.</p> <ul> <li>For the initial version of TIMIT-TTS <strong>v1.0</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2209.08000</li> <li>TIMIT-TTS Database v1.0: https://zenodo.org/record/6560159</li> </ul> </li> </ul>
Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words
<p>Hi,KIA dataset is a shared short Wakeup Word database focusing on perceived emotion in speech The dataset contains <strong>488 </strong>Wakeup Word speech. </p> <p>For more detailed information about the dataset, please refer to our paper: Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>: wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav` The first letter was used to express emotion.<br> </li> </ul> </li> <li><em><strong>annotation/</strong></em>: Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv </p> </li> <li> <p><em><strong>handcraft:</strong></em> Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em> wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p> </p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> year={2022}<br> }<br> ```</p>
Non-acoustic Speech Dataset
<p><strong>Non-acoustic speech sensing system based on flexible piezoelectric</strong></p> <p><strong>Version 1.0.0 </strong> <br> </p> <p><br> This Read_Me.txt file briefly describes the non-acoustic speech dataset and instructions to access it. </p> <p>The non-acoustic speech sensing system based on flexible piezoelectric is designed to satisfy specific needs around testing device models (in high-noise, complex environments). The system collected vibration signals from the jaws of six males and five females containing ten different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and is able to achieve 93.7% recognition accuracy by calculation. In general, this paper provides a non-acoustic speech dataset for Mandarin, including the parts collected, the number of people collected, and the environment.</p> <p><br> The dataset is available at:</p> <p>https://10.5281/zenodo.7090120</p> <p>The data descriptor paper with details of data collection and cleaning process is under submission. For proper citation of the manuscript, please refer to the latest version of this dataset which includes the details.</p> <p>This dataset and its descriptor paper were created by:</p> <p>Shiji Yuan, Ying Sun, Dezhi Zheng, Xinlei Chen, Ying Ding,Shuai Wang, Shangchun Fan</p> <p>For questions or suggestions, please e-mail Dezhi Zheng <zhengdezhi@buaa.edu.cn></p> <p><br> <strong>Description:</strong><br> <br> Ten common words were chosen as the core of the vocabulary in this dataset. These ten command words can be used for commands in IoT or robotics applications: "forward", "backward", "right", "left", "stop", "up", "down", "draw", "drop", and "reset".</p> <p>The recording software is Adobe Audition2022,which adopts monophonic recording, 16-bit storage format, 16 kHz sampling frequency, and the recorded voice is saved in wav format. The dataset is provided with two storage rules, which are stored by subject number and corpus number as classification. In the first rule, the speech data of 11 subjects were stored in different folders with the subject serial number as the folder name. Each folder contains subfolders categorized by corpus. In the second rule, the speech data of ten corpus are stored in different folders, and the names of the folders are the corpus contents. The subject number, corpus number and record order are given for each data entry. For example, the data obtained when subject one recorded corpus 10 for the first time was labeled as 1-10_1.</p> <p>After the data collection process, a filtering algorithm for automatic detection of low non-acoustic speech data is designed to remove problematic data that are very short or very quiet. The script of the data filtering algorithm is provided in this repository. </p> <p>For specific detail of the data filtering process, please refer to the script (speech data filtering algorithm in MATLAB) in this repository and the data descriptor paper.</p> <p>The dataset in this repository is the processed version. The raw dataset and removed audio files are not included in this repository.</p> <p><br> <br> <strong>File list:</strong><br> <br> Non-acoustic Speech Dataset.zip</p> <p>speech data filtering algorithm.zip</p> <p>Readme.txt </p> <p><br> </p>
AnglistikVoices: L2 English speech dataset
<h1><strong>AnglistikVoices: an L2 English speech dataset</strong> </h1> <p>This repository contains an L2 (second language) English speech corpus consisting of <strong>74 minutes</strong> of recorded audio from <strong>15 non-native English speaking participants</strong>. The dataset was created as part of a university course, with all participants being students who are also the authors of this dataset.</p> <h2>Dataset Specifications</h2> <ul> <li><strong>Total participants</strong>: 15 non-native English speakers</li> <li><strong>Total audio duration</strong>: 74 minutes</li> <li><strong>Recordings per participant</strong>: 60 audio samples each</li> <li><strong>Sentence alignment</strong>: Available for 8 out of 15 participant</li> <li><strong>Recording equipment</strong>: Audio-Technica ATM75 microphone</li> <li><strong>Stimuli: </strong>All sentences are from the Artie Bias Corpus (https://github.com/artie-inc/artie-bias-corpus)</li> <li><strong>Recording environment</strong>: Recording booth</li> </ul> <p>The dataset contains individual recordings of non-native English speakers <strong>organized by participant ID</strong>. For 8 participants, sentence-level alignments are provided. All recordings were captured in a controlled acoustic environment using Audio-Technica ATM75 microphone to ensure high audio quality. </p> <p>The recordings consist of spoken English utterances from each participant. <strong>Detailed linguistic profiles for each participant are available in the metadata.xlsx file</strong>, which is indexed by participant ID and contains information on native language, proficiency level, language learning history, and other relevant linguistic background data.</p> <p>The audio files are organized by participant ID, matching the identifiers used in the metadata file for easy cross-referencing between the audio recordings and participant linguistic profiles.</p> <h2>Authors and Contributors</h2> <p>This dataset was created by the student participants themselves as part of their coursework.</p> <p><strong>Course Instructor</strong>: Akhilesh Kakolu Ramarao</p> <p><strong>Teaching Assistant</strong>: Anna Sophia Stein</p> <p>If you have any questions, you can contact: kakolura@hhu.de</p> <p>If you use this dataset in your research, please cite:</p> <pre><code>@dataset{kakolu_ramarao_2024_anglistikvoices, author = {Kakolu Ramarao, A. and Stein, A. S. and Tahiri, A. and Rodrigues, D. C. and Antonia Weismann, C. and Schäfer, O. S. and Kaczor, J. and Tran, N. H. and Elena Telaar, C. and Bauer, L. and Jütten, M. and Mafuta, C. and Agelopoulou, V. V. and Grabowski, Q. A. G.}, title = {AnglistikVoices: L2 English speech dataset}, publisher = {Zenodo}, version = {v1.0.0}, year = {2024}, month = jun, doi = {10.5281/zenodo.12525952}, url = {https://doi.org/10.5281/zenodo.12525952}, note = {LabPhon 19, Hanyang Institute for Phonetics and Cognitive Sciences of Language (HIPCS), Hanyang University in Seoul, Korea} }</code></pre>
The speed-curvature power law in tongue movements of repetitive speech [dataset]
<p>Files in this record contain data used to produce results presented in the<br> paper:</p> <p>Title: The speed-curvature power law in tongue movements of repetitive speech<br> Authors: Stephan R. Kuberski and Adamantios I. Gafos<br> DOI: <a href="https://doi.org/10.1371/journal.pone.0213851">https://doi.org/10.1371/journal.pone.0213851</a></p> <p>For further details refer to the included file README.txt.</p>
Fitts' law in tongue movements of repetitive speech [dataset]
<p>Files in this record contain data used to produce results presented in the paper:</p> <p> </p> <p>Title: Fitts' law in tongue movements of repetitive speech</p> <p>Authors: Stephan R. Kuberski and Adamantios I. Gafos</p> <p>DOI: TBA</p> <p> </p> <p>For further details refer to the included file README.txt.</p>
Hachidaishu Part-of-Speech Dataset
<p><strong>Full Changelog</strong>: https://github.com/yamagen/hachidaishu-pos/commits/1.0.1</p>
Amharic Hate Speech Detection Dataset
<p>Amharic Hate Speech Detection Dataset V1</p> <p>To contribute for the research and development of hate speech detection in Amharic language from social media, we are glad to release our hate speech dataset we prepared from the Ethiopian Broadcasting Corporation (EBC) Facebook page (<a href="https://www.facebook.com/EBCzena">https://www.facebook.com/EBCzena</a>), and some chosen Facebook page (<a href="https://www.facebook.com/604407519910492">https://www.facebook.com/604407519910492</a>) that we found potential hateful comments.</p> <p>We extracted comments/posts pertaining to race, religion, and ethnicity using the Facepager API, resulting in a set of 30,000 comments between April 15, 2019 and December 15, 2019. A total of 5,000 comments/posts were chosen at random for annotation. Three annotators (two candidate PhD. in Linguistics and one MSc. in Law) manually annotated the selected samples as “<strong>Hate</strong>” or “<strong>not</strong>-<strong>Hate</strong>” resulting 2,000 (1000 hate and 1000 non-hate) labeled comments because of majority vote among the annotators.</p> <p>For the labeling procedure, the annotators used Ethiopian government’s hate speech and misinformation prevention and suppression proclamation <a href="https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf">https://www.accessnow.org/cms/assets/uploads/2020/05/Hate-Speech-and-Disinformation-Prevention-and-Suppression-Proclamation.pdf</a>, as well as our definition of hate speech and the hate speech characterization lists proposed in Fino (2020) <a href="https://doi.org/10.1093/jicj/mqaa023">https://doi.org/10.1093/jicj/mqaa023</a>) were provided to the annotators.</p> <p>Accordingly, a speech is labeled as “<strong>Hate</strong>” when:</p> <ul> <li>“the speech targets a group or individual as a member of a group (ethnicity, race, religion)”</li> <li>“the speech content in the message expresses hatred”</li> <li>“the speech causes a harm”</li> <li>“the speaker intends harm or bad activity”</li> <li>“the speech incites bad actions”</li> <li>“the speech is either public and directed at a member of the group”</li> <li>“the context makes violent response possible”</li> </ul>
Lada: Ukrainian High-Quality Female Text-to-Speech Dataset
<p>The dataset has high-quality data recorded in a professional studio. </p> <p>Archives with a <strong>trimmed</strong> tag are having removed silence (aligned) using <a href="https://github.com/proger/uk">https://github.com/proger/uk</a> </p> <p><strong>Features</strong></p> <ul> <li>Quality: high</li> <li>Duration: 10h37m</li> <li>Audio formats: OPUS/WAV</li> <li>Text format: JSONL (a <code>metadata.jsonl</code> file)</li> <li>Frequency: 16000/22050/48000 Hz</li> </ul>
Voice of America: Ukrainian ASR Dataset of Broadcast Speech
<p>The dataset is based on public recordings of Voice of America (<a href="https://ukrainian.voanews.com">https://ukrainian.voanews.com</a>) extracted from their videos.</p> <p>The dataset contains 398 hours of speech.</p> <p>The dataset is created by the ASR Corpus Creator (<a href="https://zenodo.org/record/7396705">https://zenodo.org/record/7396705</a>).</p> <p>The format of files: WAV with 16 kHz.</p> <p> </p>
Sudanese dialect speech dataset
<p>This is speech dataset for the Sudanese dialect data been collected from YouTube videos represent the characteristics of the Sudanese dialect, mainly the middle of Sudan dialect -Khartoum in particular- and have some northern tendency, primarily two programs Hajj Muzakir and Dukkan Wad Elbaseer.</p> <p>Transcription is done manually by listening to the audio files repeatedly to write the captions for the collected conversations to make sure that every word is written as said by the speakers. Transcription is written without diacritics on the Arabic alphabet, in a manner that reflects the Sudanese way of speaking, therefore, any correction to the noticeable mistakes was not applied to get rid of any biases and make the data representative.</p> <p>The 'Dataset' subdirectory contains all the audio and text files for the corpus, the files organized based on program name 'hm_' for Hajj Muzakir program and 'wb_' Dukkan Wad Elbaseer, each filename follows three categories first two litters for the program name 'hm' or 'wb', second the number of the episode third the number of the clip, hm_01_0001.wav and wb_01_0001.wav represent first episode of each program and the first clip.</p>
EEG Dataset for 'Decoding of selective attention to continuous speech from the human auditory brainstem response' and 'Neural Speech Tracking in the Theta and in the Delta Frequency Band Differentially Encode Clarity and Comprehension of Speech in Noise'.
<p>The repository contains the unprocessed EEG data recorded for the publications [1, 2]. For convenience, the onsets of the EEG data provided here are time-aligned with the onsets of the audio books in the 'audiobooks' folder, and the EEG data are provided in HDF5 format. Please refer to the original version of this dataset for more details.</p> <p>More details, as well as the original data files, are available at the original repository <a href="https://doi.org/10.5281/zenodo.7086209">here</a>.</p> <p>Examples of using these data (preprocessing, fitting linear models) can be found <a href="https://github.com/Mike-boop/trf-examples">here</a>.</p> <p>The English conditions (clean, lb, mb, hb, fM, fW) comprised a single recording session. The Dutch conditions (cleanDutch, lbDutch, mbDutch, hbDutch) comprised a separate recording session. You see which participants took part in each session in session_info.json.</p> <p>Please note some details about the stimulus presentation for the various listening conditions:</p> <ul> <li>English speech-in-babble-noise (lb, mb, hb): babble noise was played by itself for one second before the audiobook track began. The babble noise was also played for one second after the audiobook track ended. Therefore, you should discard the first second and the last second from these trial during your analysis.</li> <li>Dutch speech-in-babble-noise (lbDutch, mbDutch, hbDutch): the story (narrated in Dutch) was played by itself for one second before the babble noise track began. Then, the babble noise was increased linearly in amplitude for one second. Therefore, you should discard the first two seconds from these trials during your analysis.</li> <li>Dutch in quiet, and Dutch-in-babble-noise (cleanDutch, lbDutch, mbDutch, hbDutch): some English sentences were embedded in the Dutch narratives in order to encourage attention. You should crop these from your analysis. The onsets and offsets of the English sentences (in samples, at 44100Hz) are provided in the audiobooks/*Dutch/english_onsets_info.json files.</li> <li>Competing-speakers conditions (fM, fW): sometimes the attended track is longer than the unattended track, or vice-versa. The onsets of both tracks are aligned. You should crop the trial to the length of the shortest track for your analysis.</li> </ul> <p>If you use this data, please cite the original publications, as well as this repository [1,2,3].</p> <p>[1] Etard O, Kegler M, Braiman C, Forte A E and Reichenbach T. “Decoding of selective attention to continuous speech from the human auditory brainstem response” 2019. <em>NeuroImage</em> <strong>200</strong> 1–11</p> <p>[2] Etard O and Reichenbach T. “Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise” 2019. <em>J. Neurosci.</em> <strong>39</strong> 5750–9</p> <p>[3] Etard O and Reichenbach T. "EEG Dataset for 'Decoding of selective attention to continuous speech from the human auditory brainstem response' and 'Neural Speech Tracking in the Theta and in the Delta Frequency Band Differentially Encode Clarity and Comprehension of Speech in Noise". Doi: 10.5281/zenodo.7086208</p>
EEG Dataset for 'Cortical Tracking of Surprisal during Continuous Speech Comprehension'
<p>The repository contains the unprocessed EEG data recorded for the publication [1]. For convenience, the onsets of the EEG data provided here are time-aligned with the onsets of the audio books in the 'audiobooks' folder, and the EEG data are provided in HDF5 format. Please refer to the original version of this dataset for more details.</p> <p>A script is provided which shows how the original data were aligned ('align_data.py') in Python. The order in which the audiobook chapters were presented was different for different participants. If this is important to you, the details are available in 'stimulus_orders.csv'.</p> <p>More details, as well as the original data files, are available at the original version of this repository <a href="https://doi.org/10.5281/zenodo.7086168">here</a>.</p> <p>Examples for using this data (preprocessing, fitting linear models) can be found <a href="https://github.com/Mike-boop/trf-examples">here</a>.</p> <p>If you use this data, please cite the original publication, as well as this repository [1,2].</p> <p>[1] Weissbart H, Kandylaki KD, Reichenbach T. “Cortical Tracking of Surprisal during Continuous Speech Comprehension”. J Cogn Neurosci. 2020 Jan;32(1):155-166. doi: 10.1162/jocn_a_01467.</p> <p>[2] Weissbart H, Kandylaki KD, Reichenbach T. “EEG Dataset for 'Cortical Tracking of Surprisal during Continuous Speech Comprehension'”. doi: 10.5281/zenodo.7086167</p>
Automatic speech recognition datasets for Gronings, Nasal, and Besemah
<p>Automatic speech recognition datasets for Gronings, Nasal, and Besemah for experiments reported in Bartelds, San, McDonnell, Jurafsky and Wieling (2023). <em>Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation</em>. ACL 2023.</p> <p>Model training code available at: https://github.com/Bartelds/asr-augmentation</p>
Charades-STA Speech Caption Dataset
<ul> <li><strong>Dataset introduction: </strong>This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"</li> <li><strong>Associated Code: <a href="https://github.com/xian-sh/UniSDNet">https://github.com/xian-sh/UniSDNet</a></strong></li> <li><strong>Associated Paper: <a href="https://arxiv.org/abs/2403.14174">https://arxiv.org/abs/2403.14174</a><br></strong></li> <li><strong>Disclaimer: </strong>This dataset is for academic research only, non-commercial use, if you use this dataset please cite the <a href="https://arxiv.org/abs/2403.14174">associated paper</a></li> <li><strong>Cite:</strong></li> </ul> <p>@article{hu2024unified,<br> title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},<br> author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},<br> year={2024},<br> Journal={CoRR},<br> volume={abs/2403.14174},<br>}</p>
TACoS Speech Caption Dataset
<ul> <li><strong>Dataset Introduction: </strong>This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"</li> <li><strong>Associated Code: <a href="https://github.com/xian-sh/UniSDNet">https://github.com/xian-sh/UniSDNet</a></strong></li> <li><strong>Associated Paper: <a href="https://arxiv.org/abs/2403.14174">https://arxiv.org/abs/2403.14174</a><br></strong></li> <li><strong>Disclaimer: </strong>This dataset is for academic research only, non-commercial use, if you use this dataset please cite the <a href="https://arxiv.org/abs/2403.14174">associated paper</a></li> <li><strong>Cite:</strong></li> </ul> <p>@article{hu2024unified,<br> title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},<br> author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},<br> year={2024},<br> Journal={CoRR},<br> volume={abs/2403.14174},<br>}</p>
Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)
<p>Datasets used in the paper: "Creating speech zones with self-distributing acoustic swarms"</p> <p>This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker. The recordings are captured from an array of 7 microphones, as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong> the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>
Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 2 of 2)
<p>Datasets used in the paper: "Creating speech zones with self-distributing acoustic swarms"</p> <p>This deposit contains the <strong>second</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker. The recordings are captured from an array of 7 microphones, as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong> the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>
Dataset and documented R code for "Nouns and verbs in the speech signal"
<p>The files available constitute supplementary material to the following article:</p> <p>Lohmann, Arne. Nouns and verbs in the speech signal: Are there phonetic correlates of grammatical category? <em>Linguistics</em> - <em>An Interdisciplinary Journal of the Language Sciences</em>.</p> <p>The article is to be published online in 2020, and in 2021 in the print version of the journal.</p>
WHISPER SET 1: a dataset for multi-channel, multi-device speech separation and speech enhancement
<p>This dataset is <code>WHISPER SET 1,</code> a dataset for speech enhancement and source separation recorded with a Wireless Acoustic Sensor Network (WASN) called WHISPER <a href="https://ieeexplore.ieee.org/abstract/document/8110202">Kiselev2018</a>. The dataset contains samples for up to 4 concurrent speakers and speech in noise. The dataset was recorded in a room with low reverberation (T_60 = 0.2 s) and using 16 microphones. In general, each track contains first a calibration phase where each of the speakers sequentially is active alone for 15 seconds. Followed by 15 seconds of all the speakers together (plus noise in some cases). </p> <p>If you use this dataset please cite:</p> <ul> <li><strong>E. Ceolini, I. Kiselev and S. Liu, "Evaluating multi-channel multi-device speech separation algorithms in the wild: a hardware-software solution," in <em>IEEE/ACM Transactions on Audio, Speech, and Language Processing</em>.</strong></li> </ul> <p>===</p> <p>Each sample is a 16-channel wav file in which the order of the channel follows the following logic:</p> <p>0 - module 5 mic 1 1 - module 5 mic 2 2 - module 5 mic 3 3 - module 5 mic 4 4 - module 6 mic 1 5 - module 6 mic 2 6 - module 6 mic 3 7 - module 6 mic 4 8 - module 7 mic 1 9 - module 7 mic 2 10 - module 7 mic 3 11 - module 7 mic 4 12 - module 8 mic 1 13 - module 8 mic 2 14 - module 8 mic 3 15 - module 8 mic 4</p> <p>Refer to the <a href="https://github.com/SensorsAudioINI/WHISPER_SET_1/blob/master/WHISPER4_floor_annotated.png">floor plan</a> for a visual illustration of the microphone arrangement.</p> <p>The files are divided into two subfolders, one for the samples of speech enhancement and one for the samples of speech separation.</p> <ul> <li>In the folder of speech separation, the files are divided into subfolders defining the number of speakers in the mixtures (2, 3, or 4)</li> <li>In the folder of speech enhancement, the files are divided into subfolders following the SNR of the mixture (0, -5, -10 dB)</li> </ul> <p>Samples are ordered in folders. Each sample folder contains a 15 seconds 16-channels <code>mixture.wav</code> file, plus the 15 seconds 16-channels <code>calibX.wav</code> files one for each speaker alone or noise alone in the mixture. That is a sample with a mixture with 4 speakers will have 4 calibration files (calib1.wav, calib2.wav, calib3.wav, calib4.wav) and a mixture of a speaker plus noise will have 2 calibration files one for speech (calib1.wav) and one for noise (calib2.wav).</p> <p>== </p> <p>A Jupyter notebook is included to show an example of how to use the data of this dataset for speech separation and speech enhancement using beamforming. The notebook is dependent on <a href="https://github.com/Enny1991/beamformers">this beamforming library</a> and <a href="https://github.com/Enny1991/sep_eval">this tool</a> to evaluate the quality of the separation.</p> <p>==</p> <p>Refer to the README.md in the dataset for more information.</p> <p>For any question please contact enea.ceolini@gmail.com</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.