Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
94
datasets available to search
ShareScore release 0.7.1
Dataset results
94 results for “speech dataset”
The role of isochrony in speech perception in noise - Dataset
<p>This dataset contains speech stimuli and listener data reported on in Aubanel & Schwartz (2020), DOI: <a href="http://dx.doi.org/10.1038/s41598-020-76594-1">10.1038/s41598-020-76594-1</a>. </p> <p><strong>French data</strong></p> <ul> <li>French sentences are taken from the Fharvard corpus (Aubanel et al., 2020, DOI: <a href="https://dx.doi.org/10.1016/j.specom.2020.07.004">10.1016/j.specom.2020.07.004</a>)</li> <li>Speech material and sentence recordings are available at: <a href="https://dx.doi.org/10.5281/zenodo.1462854">10.5281/zenodo.1462854</a></li> <li><strong>fr_stimuli.zip</strong> contains the stimuli presented to the listeners</li> <li><strong>fr_responses.csv</strong> contains the responses typed by listeners</li> </ul> <p><strong>English data</strong></p> <ul> <li>English sentences are taken from the Harvard corpus (Rothauser et al. 1969)</li> <li>Speech material and sentence recordings are taken from the MAVA corpus, available at: <a href="https://dx.doi.org/10.4227/139/59a4c21a896a3">10.4227/139/59a4c21a896a3</a></li> <li><strong>en_stimuli.zip</strong> contains the stimuli presented to the listeners</li> <li><strong>en_responses.csv</strong> contains the responses typed by listeners</li> </ul> <p> </p>
A Kannada Emotional Speech Dataset
<p>There was no emotional speech dataset available in Kannada. This was a limiting factor for research in the Kannada-speaking world. I introduce a Kannada emotional speech dataset and give details about its design and content. This dataset contains six different sentences, pronounced by thirteen people (four male and nine female), in five basic emotions plus one neutral emotion. They are all Kannada speakers. The dataset has been contributed by volunteers and the recordings were not made in a controlled environment. The dataset contains a total of 468 audio samples, each one in a separate audio file. The file naming convention is as follows: AA-EE-SS.wav where AA is a two-character field that gives the actor number (01 to 13), EE is a two-character field that indicates the emotion number (01 to 06), and SS is a two-character field that gives the sentence number (01 to 06). This dataset is freely available under a Creative Commons license.</p> <p>Gender and age of each of the 13 people who contributed to the dataset</p> <ul> <li>01, F, 45</li> <li>02, F, 20</li> <li>03, F, 21</li> <li>04, M, 47</li> <li>05, F, 48</li> <li>06, M, 20</li> <li>07, F, 20</li> <li>08, F, 45</li> <li>09, F, 21</li> <li>10, F, 12</li> <li>11, F, 12</li> <li>12, M, 17</li> <li>13, M, 26</li> </ul> <p>Identification characters for the emotions in the dataset</p> <ul> <li>01, Anger</li> <li>02, Sadness</li> <li>03, Surprise</li> <li>04, Happiness</li> <li>05, Fear</li> <li>06, Neutral</li> </ul> <p>Identification characters for the sentences in the dataset</p> <ul> <li>01, ರೋಗಿಗಳಿಗೆ ಚಿಕಿತ್ಸೆ ನೀಡಿ ಉಪಚರಿಸುವುದು</li> <li>02, ಈ ಕಾದಂಬರಿಯು ಎರಡು ಪಾತ್ರಗಳನ್ನು ಒಳಗೊಂಡಿದೆ</li> <li>03, ಖಾಸಗಿ ವಿಮಾನಗಳೆಂದೂ ಸಾರ್ವಜನಿಕ ವಿಮಾನಗಳೆಂದೂ ವಿಂಗಡಿಸಿದ್ದಾರೆ</li> <li>04, ನಿಮ್ಮನ್ನು ಬೀಟಿಯಾಗಿ ಬಹಳ ಸಂತೋಶ ಆಯಿತು</li> <li>05, ರಾಮನ ಎಡಬಲ ದಲ್ಲಿ ಸೀತಾ ಲಕ್ಷ್ಮಣ ರಿದ್ದಾರೆ</li> <li>06, ಕನ್ನಡವನ್ನು ಕಲಿಯಬೆಕು</li> </ul>
Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects
<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>
Chichewa speech dataset
<p>This is a speech dataset for the language "Chichewa" spoken in Malawi, Zambia, Zimbabwe and Mozambique. The zipped folder contains the following:</p> <ul> <li>chich_speech_audio_files-this is a folder with all the audio files and transcripts </li> <li>chich_speech_dataset_documentation.pdf: this is the documentation for the dataset</li> <li>chich_speech_dataset_metadata.csv: this file has basic metadata about each audio file.</li> </ul>
Dataset, test programs and analysis scripts for the related paper "Effects of reverberation on speech intelligibility in noise for hearing-impaired listeners"
<p>This dataset contains the data, test programs and analyses scripts used for a study submitted as a stage 2 registered report for Royal Society Open Science.</p>
AISHELL-Stammertalk 中文口吃数据库 A Mandarin stuttered speech dataset
<p>Dataset official website: <a href="https://aishelltech.com/aishell_6A" target="_blank" rel="noopener">https://aishelltech.com/aishell_6A</a><br><br>This Zenodo page contains dataset samples. To access and download the full dataset, please send an application here <a href="https://opendata.aishelltech.com/stammertalk" target="_blank" rel="noopener">https://opendata.aishelltech.com/stammertalk</a></p> <p>The AISHELL-Stammertalk datasets consists of recordings from 70 native mardarin AWS (Adults who stutter), including 46 males and 24 females. The total duration is 48.8 hours. Each participant engaged in a recording session lasting up to one hour, comprising two parts: conversation and voice command reading. Conversations were conducted through online interviews using platforms like Zoom or Tencent Meet, aiming to capture spontaneous speech on diverse topics. The interviewer, one of the two authors, posed questions based on a prepared list, with the flexibility to introduce impromptu questions as needed.</p> <p>In the voice command reading part, participants were tasked with reading a set of 200 commands, categorized into car navigation and smart home device interaction. To ensure variety, a new set of 200 commands was introduced for every 25 participants, resulting in a dataset featuring a total of 600 unique commands. Participants were encouraged to employ the Voluntary Stuttering technique, deliberately introducing stuttering.</p> <p>Five types of stuttering were specified by the annotation guidelines, including:<br><strong>[]</strong>: Word/phrase repetition. Designated for marking entire repeated character or phrase.<br><strong>/b</strong>: block. Gasps for air or stuttered pauses.<br><strong>/p</strong>: prolongation. Elongated phoneme.<br><strong>/r</strong>: sound repetition. Repeated phoneme that do not constitute an entire character.<br><strong>/i</strong>: interjections. Filler characters due to stuttering e.g., ‘嗯’, ‘啊’, or ‘呃’. Notably, naturally occurring interjections that don't disrupt the speech flow are excluded.</p>
A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).
<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>
A Canadian French Emotional Speech Dataset
<p>The Canadian French Emotional (CaFE) speech dataset contains six different sentences, pronounced by six male and six female actors, in six basic emotions plus one neutral emotion. The six basic emotions are acted in two different intensities: mild ("Faible") and strong ("Fort").</p> <p>This dataset is freely available under a Creative Commons license (CC BY-NC-SA 4.0).</p> <p>The dataset was digitally recorded at a high-resolution (192 kHz sampling rate, 24 bits per sample).</p> <p>A filtered and downsampled version at a lower resolution (48 kHz sampling rate, 16 bits per sample) is also available.</p> <p>This work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p> <p>To view a copy of this license, visit <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons</a></p>
Dataset for "What is Gab? A Bastion of Free Speech or an Alt-Right Echo Chamber?"
<p>This dataset was used for this project: "What is Gab? A Bastion of Free Speech or an Alt-Right Echo Chamber?". Savvas Zannettou, Barry Bradlyn, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringhini, Haewoon Kwak, Jeremy Blackburn. Workshop on Computational Methods in CyberSafety, Online Harassment and Misinformation, 2018. DOI: <a href="https://arxiv.org/ct?url=http%3A%2F%2Fdx.doi.org%2F10%252E1145%2F3184558%252E3191531&v=345a781d">10.1145/3184558.3191531</a>.</p> <p>In addition, this project has received funding from the European Union’s Horizon 2020 Research and Innovation program under the Marie Skłodowska-Curie ENCASE project (Grant Agreement No. 691025). The work reflects only the authors’ views; the Agency and the Commission are not responsible for any use that may be made of the information it contains.</p> <p>Using Gab’s API, we crawl the social network using a snowball methodology. Specifically, we obtain data for the most popular users as returned by Gab’s API and iteratively collect data from all their followers as well as their followings. Subsequently, for all users in our dataset we collect all of their the posts. Overall, we collect 22,112,812 posts from 336,752 users, between August 2016 and January 2018. This dataset is a .json file and each line has one .json object.</p>
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,” arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong> <strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>, </li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications. </p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset </strong><strong>are </strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian. </p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>
Helsinki Speech Challenge 2024 open audio dataset
<h1>Dataset</h1> <p>This training dataset is originally designed for the Helsinki Speech Challenge 2024 (HSC2024). While it was created with this challenge in mind, its applications extend far beyond, making it a valuable resource for developing and testing audio algorithms across diverse uses.</p> <p>Our dataset features clean speech samples generated by OpenAI's text-to-speech model, paired with corresponding recorded signals. These recorded signals are purposefully distorted by real-world effects such as filtering and reverb, offering a realistic testing ground for your audio processing algorithms.</p> <p>To ensure ease of use, the audio samples are organized into 10 separate zip files, each dedicated to a specific task and level of complexity. Most zip files include two folders clean and recorded: one containing clean audio data and the other housing the corresponding recorded (distorted) data. However, note that folders Task_3_Level_1 and Task_3_Level_2 only include recorded data, as their clean counterparts are identical to those in Task_2_Level_2 and Task_2_Level_3, respectively. Additionally, each zip file includes a .txt file with the original text samples associated with the audio clips.<br><br><strong>Update 29th October: </strong>We have now also added the test data that was used for evaluation in the challenge with similar structure to the rest of the data, contained in the zip files starting with "Test".</p> <p>Additionally, there is folder Impulse_Responses, which contains a clean sine sweep signal and short and a long white noise signal with recorded counterparts. <br>To help you get started, we've also provided an example folder containing samples from each task and level, giving you a comprehensive overview of the dataset's scope and variety.<br><br>The dataset also contains a python script evaluate.py. This can be used to evaluate the quality of audio files using the Mozilla Deepspeech speech recognition model. For more details on this script, see the more detailed description of the data challenge either on the website below or on arXiv <a href="https://arxiv.org/abs/2406.04123">https://arxiv.org/abs/2406.04123</a>.<br><br>Important dates:</p> <ul> <li>Data Challenge Launch: 10. June 2024.</li> <li>Sign-up deadline: 1. September 2024 (if you missed this deadline and wish to participate in the challenge, please send us an email).</li> <li>Submission deadline: 6. October 2024. We realize this deadline is a bit optimistic, but we humbly ask participants to try to make this deadline.</li> <li>Results are published: 4. November.</li> <li>Inverse days: 10.-13. December in Oulu, Finland.</li> </ul> <p><strong>Here is a link to the official webpage of the HSC2024: </strong><a href="https://blogs.helsinki.fi/helsinki-speech-challenge/" target="_blank" rel="noopener">https://blogs.helsinki.fi/helsinki-speech-challenge/</a><br><br><strong>Contact Email: </strong>hsc2024@helsinki.fi</p>
DDS (Device-Degraded Speech) Dataset - DAPS portion
<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data. </p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels. </p> <p><strong>Arxiv: </strong>https://arxiv.org/abs/2109.07931</p> <p> </p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains DAPS portion of DDS.</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong> https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong> https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong> https://zenodo.org/record/5501697</li> </ul>
Segmented DAPS (Device and Produced Speech) Dataset
<p>This is a modified version of a subset of the Device and Produced Speech (DAPS) dataset. The original dataset can be found <a href="https://zenodo.org/record/4660670#.YKuxgKhKhPZ">here</a>. This dataset contains text-aligned audio of the first script of the "clean" partition of the DAPS dataset for all 20 speakers. Phoneme and word alignments are provided as JSON files. We segment the audio and alignments into single sentences. For each sentence, we additionally provide the raw text in a txt file. Audio is provided as 44.1 kHz WAV files.</p> <p>If you use this work as part of an academic publication, please cite the paper corresponding to the original dataset:</p> <blockquote> <p>Gautham J. Mysore, <a href="https://ieeexplore.ieee.org/document/6981922">“Can We Automatically Transform Speech Recorded on Common Consumer Devices in Real-World Environments into Professional Production Quality Speech? - A Dataset, Insights, and Challenges”</a>, in the IEEE Signal Processing Letters, Vol. 22, No. 8, August 2015</p> </blockquote>
Enhanced RAVDESS Speech Dataset
<p>This is a modified version of the speech audio contained within the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) dataset. The original dataset can be found <a href="https://zenodo.org/record/1188976#.YKvCMKhKhPY">here</a>. The unmodified version of just the speech audio used as source material for this dataset can be found <a href="https://www.kaggle.com/uwrfkaggler/ravdess-emotional-speech-audio">here</a>. This dataset performs speech enhancement and bandwidth extension on the original speech using HiFi-GAN. HiFi-GAN produces high-quality speech at 48 kHz that contains significantly less noise and reverb relative to the original recordings.</p> <p>If you use this work as part of an academic publication, please cite the papers corresponding to both the original dataset as well as HiFi-GAN:</p> <blockquote> <p>Livingstone SR, Russo FA (2018) The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English. PLoS ONE 13(5): e0196391. <a href="https://doi.org/10.1371/journal.pone.0196391">https://doi.org/10.1371/journal.pone.0196391</a>.</p> <p>Su, Jiaqi, Zeyu Jin, and Adam Finkelstein. "HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks." <em>Proc. Interspeech</em>. October 2020.</p> </blockquote> <p>Note that there are two recent papers with the name "HiFi-GAN". Please be sure to cite the correct paper as listed here.</p>
FestCat speech synthesis dataset in Catalan, raw at 48kHz, 16bits/sample
<p>This dataset contains audio recordings from the 10 speakers of the FestCat project and an additional speaker "uri".</p> <p>The recordings are available in the "-raw" compressed archives and are provided in the following format:</p> <pre><code class="language-bash">SAMPLERATE=48000 # Hz BITSPERSAMPLE=16 ENCODING="SIGNED_INTEGER" AUDIOHEADER="RAW" # (no header) # For the prompts and utts: TEXT_ENCODING="ISO-8859-15"</code></pre> <p>These sampling conditions were downsampled from the original recordings captured at 96kHz and using 24 bits/sample for all the FestCat speakers. The additional speaker "uri" had recordings 16kHz and were here upsampled to 48kHz.</p> <p>The recordings have been automatically segmented. The segmentation results are available at the "-utts" archives.</p> <p>The text prompts are available in the "-prompts" archive.</p> <p>The text information is encoded using ISO-8859-15.</p> <p> </p>
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis
<p>This is the public release of the Samsung Open Mean Opinion Scores (SOMOS) dataset for the evaluation of neural text-to-speech (TTS) synthesis, which consists of audio files generated with a public domain voice from trained TTS models based on bibliography, and numbers assigned to each audio as quality (naturalness) evaluations by several crowdsourced listeners.<br><br><strong>Description</strong><br><br>The SOMOS dataset contains 20,000 synthetic utterances (wavs), 100 natural utterances and 374,955 naturalness evaluations (human-assigned scores in the range 1-5). The synthetic utterances are single-speaker, generated by training several Tacotron-like acoustic models and an LPCNet vocoder on the LJ Speech voice public dataset. 2,000 text sentences were synthesized, selected from Blizzard Challenge texts of years 2007-2016, the LJ Speech corpus as well as Wikipedia and general domain data from the Internet.<br>Naturalness evaluations were collected via crowdsourcing a listening test on Amazon Mechanical Turk in the US, GB and CA locales. The records of listening test participants (workers) are fully anonymized. Statistics on the reliability of the scores assigned by the workers are also included, generated through processing the scores and validation controls per submission page.</p> <p>To listen to audio samples of the dataset, please see <a href="https://innoetics.github.io/publications/somos-dataset/index.html">our Github page</a>.</p> <p>The dataset release comes with a carefully designed train-validation-test split (70%-15%-15%) with unseen systems, listeners and texts, which can be used for experimentation on MOS prediction.</p> <p><em>This version also contains the necessary resources to obtain the transcripts corresponding to all dataset audios.</em></p> <p><strong>Terms of use</strong></p> <ul> <li>The dataset may be used for <strong>research</strong> purposes only, for <strong>non-commercial</strong> purposes only, and may be distributed with the same terms.</li> <li>Every time you produce research that has used this dataset, please <strong>cite </strong>the dataset appropriately.</li> </ul> <p>Cite as:</p> <pre><code>@inproceedings{maniati22_interspeech, author={Georgia Maniati and Alexandra Vioni and Nikolaos Ellinas and Karolos Nikitaras and Konstantinos Klapsas and June Sig Sung and Gunu Jho and Aimilios Chalamandaris and Pirros Tsiakoulis}, title={{SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis}}, year=2022, booktitle={Proc. Interspeech 2022}, pages={2388--2392}, doi={10.21437/Interspeech.2022-10922} } </code></pre> <p><br><strong>References of resources & models used</strong></p> <p>Voice & synthesized texts:<br>K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017.</p> <p>Vocoder:<br>J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP, 2019.<br>R. Vipperla, S. Park, K. Choo, S. Ishtiaq, K. Min, S. Bhattacharya, A. Mehrotra, A. G. C. P. Ramos, and N. D. Lane, “Bunched lpcnet: Vocoder for low-cost neural text-to-speech systems,” in Proc. Interspeech, 2020.</p> <p>Acoustic models:<br>N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High quality streaming speech synthesis with low, sentence-length-independent latency,” in Proc. Interspeech, 2020.<br>Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al., “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech, 2017.<br>J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al., “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in Proc. ICASSP, 2018.<br>J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, “Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling,” arXiv preprint arXiv:2010.04301, 2020.<br>M. Honnibal and M. Johnson, “An Improved Non-monotonic Transition System for Dependency Parsing,” in Proc. EMNLP, 2015.<br>M. Dominguez, P. L. Rohrer, and J. Soler-Company, “PyToBI: A Toolkit for ToBI Labeling Under Python,” in Proc. Interspeech, 2019.<br>Y. Zou, S. Liu, X. Yin, H. Lin, C. Wang, H. Zhang, and Z. Ma, “Fine-grained prosody modeling in neural speech synthesis using ToBI representation,” in Proc. Interspeech, 2021.<br>K. Klapsas, N. Ellinas, J. S. Sung, H. Park, and S. Raptis, “WordLevel Style Control for Expressive, Non-attentive Speech Synthesis,” in Proc. SPECOM, 2021.<br>T. Raitio, R. Rasipuram, and D. Castellani, “Controllable neural text-to-speech synthesis using intuitive prosodic features,” in Proc. Interspeech, 2020.</p> <p>Synthesized texts from the Blizzard Challenges 2007, 2008, 2009, 2010, 2011, 2012, 2013, 2016:<br>M. Fraser and S. King, "The Blizzard Challenge 2007," in Proc. SSW6, 2007.<br>V. Karaiskos, S. King, R. A. Clark, and C. Mayo, "The Blizzard Challenge 2008," in Proc. Blizzard Challenge Workshop, 2008.<br>A. W. Black, S. King, and K. Tokuda, "The Blizzard Challenge 2009," in Proc. Blizzard Challenge, 2009.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2010," 2010.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2011," 2011.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2012," 2012.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2013," 2013.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2016," 2016.</p> <p><strong>Contact</strong></p> <p>Alexandra Vioni - a.vioni@samsung.com</p> <ul> <li>If you have any questions or comments about the dataset, please feel free to write to us.</li> <li>We are interested in knowing if you find our dataset useful! If you use our dataset, please email us and tell us about your research.</li> </ul>
USPDATRO: Underrepresented Speech Dataset from Romanian language Open Data
<p> USPDATRO<br> ==========</p> <p>Underrepresented Speech Dataset from Open Data: Case Study on the Romanian Language (USPDATRO) is a manually created Romanian language speech corpus.<br> It was created specifically using speech types that are underrepresented in other speech datasets.<br> Sources for this dataset are represented by open data available on multimedia platforms under a Creative Commons license.<br> The data was manually transcribed and aligned at segment level.<br> In addition to the text and audio files, we offer text annotations (lemmatization, part of speech tags, dependency parsing) in CoNLL-U Plus format.</p> <p>Each datasource is mentioned by URL in the metadata.csv file with associated license (a Creative Commons variant).</p> <p>Dataset structure:<br> - audio: Folder with audio segments in WAV format<br> - text: Folder with corresponding transcriptions<br> - conllup: Folder with corresponding token-based annotations<br> - metadata.csv: Contains information about each segment</p> <p>LICENSING</p> <p>This work (transcriptions, alignment, metadata, annotations) is provided under the license CC BY-NC-SA 4.0 (Attribution-NonCommercial-ShareAlike 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-sa/4.0/<br> and the full text here: https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode .<br> The original works considered for audio sources are available under their respective licenses (Creative Commons variants) as described in the metadata.csv file.</p> <p><br> CONTACT</p> <p>Research Institute for Artificial Intelligence "Mihai Drăgănescu", Romanian Academy<br> Web: http://www.racai.ro<br> Contact emails: vasile@racai.ro</p>
Dataset for Phonological acquisition depends on the timing of speech sound
<p>Dataset for the analysis to the manuscript: Phonological acquisition depends on the timing of<br> speech sounds: Deconvolution EEG modelling across the first five years</p>
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>We developed a tool for collecting Tunisian dialect data, prompting users to record themselves reading provided phrases. We sourced sentences from Tunisiya. These sentences are consequently removed from the LM training corpus. 89 persons have participated leading to the collection of 2631 distinct phrases. This set will be called TunSwitch TO, ``TO" standing for Tunisian Only, as these sentences do not have non-Tunisian words. </p> <p>In response to the limited availability of paired Text-Speech Tunisian datasets with code-switching, we have built a corpus through meticulous manual annotation. Whenever encountered, French and English words are enclosed within "<>" tags, and left Tunisian words without any enclosing tags. While these tags have not been used in the proposed models, they allow to have language-usage statistics and may be useful for further approaches handling code-switching. The resulting set is released as TunSwitch CS, ``CS" standing for Code-Switched.</p> <p>The TunSwitch CS dataset samples come from a set of radio shows and podcasts, representing diverse topics and a large number of unique speakers. The audio are first segmented into chunks, prioritizing word integrity using the WebRTC-VAD algorithm for silence detection. Afterward, we used a Pyannote overlap detection model to remove overlapping speech sections. Then, a music detection model is employed to eliminate music-containing chunks that could disrupt ASR model accuracy. <br> </p> <p> </p>
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>This folder contains the data used to develop and test the Tunisian Arabic Automatic Speech Recognition model developed in the following paper :</p> <p>A. A. Ben Abdallah*, A. Kabboudi, A. Kanoun, and S. Zaiem*, “Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition”, Submitted to ICASSP 2024, vol. * : These two authors have contributed equally. 2023.</p> <p><br> It contains 4 zipped folders containing audio data :<br> - TunSwitchCS.zip : containing annotated code-switched data.<br> - TunSwitchTO.zip : containing annotated Tunisian-Only data.<br> - weakly_labeled_tn.zip : containing weakly-labeled (or unlabeled) audio data. Audios may contain code-switching, but the current weak labels do not.<br> - test_wavs.zip : contains annotated testing data, divided between a code-switched part and a tunisian-only part.</p> <p><br> It also contains textual data, used for language modelling, contained in TextData.zip. Finally it also contains a language-detailed annotation of TunSwitchCS in the language_annotation.zip file .</p> <p>More details about the data are available in the paper. The current table are in a SpeechBrain-friendly format, the column path is irrelevant and has to be changed according to your local setting. Please use the provided train-dev-test splits if you work with this dataset.</p> <p>Please cite the aforementioned paper if you use or refer to this dataset. You can find models trained and tested on this dataset <a href="https://huggingface.co/SalahZa">Here</a>. Space demos are also available. </p> <p>If you use or refer to this dataset, please cite : </p> <p>```</p> <p>@misc{abdallah2023leveraging,<br> title={Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition}, <br> author={Ahmed Amine Ben Abdallah and Ata Kabboudi and Amir Kanoun and Salah Zaiem},<br> year={2023},<br> eprint={2309.11327},<br> archivePrefix={arXiv},<br> primaryClass={eess.AS}<br> }</p> <p>```</p> <p><br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.