Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Shared Acoustic Codes Underlie Emotional Communication in Music and Speech - Evidence from Deep Transfer Learning (Datasets)
<p>This repository contains the datasets used in the article "Shared Acoustic Codes Underlie Emotional Communication in Music and Speech - Evidence from Deep Transfer Learning" (Coutinho & Schuller, 2017). </p> <p>In that article four different data sets were used: SEMAINE, RECOLA, ME14 and MP (acronyms and datasets described below). The SEMAINE (speech) and ME14 (music) corpora were used for the unsupervised training of the Denoising Auto-encoders (domain adaptation stage) - only the audio features extracted from the audio files in these corpora were used and are provided in this repository. The RECOLA (speech) and MP (music) corpora were used for the supervised training phase - both the audio features extracted from the audio files and the Arousal and Valence annotations were used. In this repository, we provide the audio features extracted from the audio files for both corpora, and Arousal and Valence annotations for some of the music datasets (those that the author of this repository is the data curator).</p> <p>Below, you can find description of the various corpora, the details about the data stored in this repository and information on how to obtain the rest of the data used by Coutinho and Schuller (2017).</p> <p><strong>SEMAINE (speech)</strong></p> <p>The SEMAINE corpus (McKeown, Valstar, Cowie, Pantic & Schroder, 2012) was developed specifically to address the task of achieving emotion-rich interactions, and it is adequate for this task as it comprises a wide range of emotional speech. It includes video and speech recordings of spontaneous interactions between human and emotionally stereotyped `characters'. Coutinho & Schuller (2017) used a subset of this database (called <em>Solid-SAL</em>). The <em>Solid-SAL</em> dataset is freely available for scientific research purposes (see http://semaine-db.eu). This repository includes the audio features used in Coutinho & Schuller (2017) (under features/SEMAINE).</p> <p><strong>RECOLA (speech)</strong></p> <p>The RECOLA database (Ringeval, Sonderegger, Sauer & Lalanne, 2013) consists of multimodal recordings (audio, video, and peripheral physiological activity) of spontaneous dyadic interactions between French adults. Coutinho & Schuller (2017) used the RECOLA-Audio module which consists of the audio recordings of each participant in the dyadic phase of the task. In particular, they used the non-segmented high-quality audio signals (WAV format, 44.1kHz, 16bits), obtained through unidirectional headset microphones, of the first five minutes of each interaction. Annotations consist of time-continuous ratings of the level of Arousal and Valence dimensions of emotion perceived by each rater while seeing and listening the audio-visual recordings of each participant task. The publicly available annotated dataset includes only part of the data which amounts to a total number of 23 instances. The time frame length used by Coutinho & Schuller (2017) is 1s (the original annotations were downsampled). This repository includes the audio features used in Coutinho & Schuller (2017) (under features/RECOLA). To obtain the annotations you should contact the author of the original study (see https://diuf.unifr.ch/diva/recola/download.html for further details).</p> <p><strong>ME14 (music)</strong></p> <p>The MediaEval ``Emotion in Music'' task is dedicated to the estimation of Arousal and Valence scores continuously in time and value for song excerpts from the Free Music Archive. Coutinho and Schuller (2017) used the whole corpus (development and test sets for the 2014 challenge) which includes 1,744 songs belonging to 11 musical styles -- Soul, Blues, Electronic, Rock, Classical, Hip-Hop, International, Folk, Jazz, Country, and Pop (maximum of five songs per artist). This repository includes the audio features used in Coutinho & Schuller (2017) (under features/ME14). The full dataset (including annotations) can be obtained from http://www.multimediaeval.org/mediaeval2014/emotion2014/.</p> <p><strong>MP (music)</strong></p> <p>This is a corpus compiled specifically for this work described in Coutinho & Schuller (2017) using data collected in four previous studies. It consists of emotionally diverse full music pieces from a variety of musical styles (Classical and contemporary Western Art, Baroque, Bossa Nova, Rock, Pop, Heavy Metal, and Film Music). Annotations were obtained in controlled laboratory experiments whereby the emotional character of each piece was evaluated time-continuously in terms of levels of Arousal and Valence perceived by listeners (ranging between 35 to 52 in the four studies). In what follows, some details about the various studies are described.</p> <ul> <li>MP<sub>DB1</sub>: This subset of the MP corpus consists of the data reported by Korhonen (2004), and gently made available by the author. This dataset includes six full (or long excerpts) music pieces ranging from 151s to 315s in length (only classical music). Each piece was annotated by 35 participants (14 females). The time series correspondents to each music piece were collected at 1Hz. The golden standard for each piece was computed by averaging the individual time series across all raters. This repository includes the audio features used in Coutinho & Schuller (2017) (under features/MP/DB1). To obtain the labels please contact the author of the original study.</li> <li>MP<sub>DB2</sub>: The dataset by Coutinho & Cangelosi (2011) includes 9 full pieces (43s to 240s long) of classical music (romantic repertoire) annotated by 39 subjects (19 females). Values were recorded every time the mouse was moved with a precision of 1 ms. The resultant timeseries were then resampled (moving average) to a synchronous rate of 1 Hz. The golden standard for each piece was computed by averaging the individual time series across all raters. This repository includes the audio features (under features/MP/DB2) and labels (under annotations/MP/DB2) used in Coutinho & Schuller (2017).</li> <li>MP<sub>DB3</sub>: This dataset was collected by Coutinho & Dibben (2012) and it consists of 8 pieces of film music (84s to 130s long) taken from the late 20th century Hollywood film repertoire. Emotion ratings were given by 52 participants (26 females). The annotation procedure, data processing, and golden standard calculations were identical to MP<sub>DB2</sub>. This repository includes the audio features (under features/MP/DB3) and labels (under annotations/MP/DB3) used in Coutinho & Schuller (2017).</li> <li>MP<sub>DB4</sub>: This dataset was collected by Grewe, Nagel, Kopiez and Altenmüller (2007), and gently made available by the authors. It includes seven music pieces (127s to 502s in length) of heterogeneous styles (e.g., Rock, Pop, Heavy Metal, Classical). Each music piece was annotated by 38 participants (29 females) using an identical methodology to MP<sub>DB2</sub> and MP<sub>DB3</sub>. Data processing and golden standard calculations were also identical. This repository includes the audio features (under features/MP/DB4) used in Coutinho & Schuller (2017). To obtain the labels contact the authors of the original study</li> </ul> <p> </p> <p><strong>Bibliography</strong></p> <p>Coutinho, E., & Cangelosi, A. (2011). Musical emotions: predicting second-by-second subjective feelings of emotion from low-level psychoacoustic features and physiological measurements. <em>Emotion</em>, <em>11</em>(4), 921.</p> <p>Coutinho, E., & Dibben, N. (2013). Psychoacoustic cues to emotion in speech prosody and music. <em>Cognition & Emotion</em>, <em>27</em>(4), 658-684.</p> <p>Coutinho E, Schuller B (2017) Shared acoustic codes underlie emotional communication in music and speech—Evidence from deep transfer learning. PLoS ONE 12(6): e0179289. https://doi. org/10.1371/journal.pone.0179289.</p> <p>Grewe, O., Nagel, F., Kopiez, R., Altenmüller, E. (2007). Emotions over time: synchronicity and development of subjective, physiological, and facial affective reactions to music. <em>Emotion, 7</em>(4), pp. 774-788. DOI: 10.1037/1528-3542.7.4.774.</p> <p>Korhonen, M. (2004). Modeling Continuous Emotional Appraisals of Music Using System Identification. Available from: http://hdl.handle.net/10012/879.</p> <p>McKeown, G., Valstar, M., Cowie, R., Pantic, M., Schroder, M. (2012). The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent. <em>IEEE Transactions on Affective Computing</em>, 3, pp. 5-17. DOI: http://doi.ieeecomputersociety.org/10.1109/T-AFFC.2011.20.</p> <p>Ringeval, F., Sonderegger, A., Sauer, J. & Lalanne, D. (2013). Introducing the RECOLA Multimodal Corpus of Remote Collaborative and Affective Interactions. In <em>Proceedings of the 2nd International Workshop on Emotion Representation, Analysis and Synthesis in Continuous Time and Space (EmoSPACE 2013)</em>, Shanghai, China. IEEE</p>
Repeatedly experiencing the McGurk effect induces long-lasting changes in auditory speech perception
<p>In the McGurk effect, presentation of incongruent auditory and visual speech evokes a fusion percept different than either component modality. We show that repeatedly experiencing the McGurk effect for 14 days induces a change in auditory-only speech perception: the auditory component of the McGurk stimulus begins to evoke the fusion percept, even when presented on its own without accompanying visual speech. This perceptual change, termed fusion-induced recalibration (FIR), was talker-specific and syllable-specific and persisted for a year or more in some participants without any additional McGurk exposure. Participants who did not experience the McGurk effect did not experience FIR, showing that recalibration was driven by multisensory prediction error. A causal inference model of speech perception incorporating multisensory cue conflict accurately predicted individual differences in FIR. Just as the McGurk effect demonstrates that visual speech can alter the perception of auditory speech, FIR shows that these alterations can persist for months or years. The ability to induce seemingly permanent changes in auditory speech perception will be useful for studying plasticity in brain networks for language and may provide new strategies for improving language learning.</p>
Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics
<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims </h1> <div> <h2> </h2> <h2>Description </h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias. </p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias". </p> </div> <div> <p>The label was determined by a 75% majority vote. We classified “probably biased” and “confident biased” as biased, and “confident not biased,” “probably not biased,” and “don't know” as not biased. </p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India. </p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022. The percentage of biased tweets with the keyword 'Blacks' remained nearly the same. </p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews". </p> </div> <div> <h2> </h2> <h2>Content</h2> <h3>Cohort 2022 </h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias. </p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians. </p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks. </p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and ––as mentioned above––146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews. </p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines. </p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims. </p> </div> <div> <h3>Cohort 2023 </h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords “Asians, Blacks, Jews, Latinos and Muslims” from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias. </p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians. </p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks. </p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews. </p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines. </p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims. </p> </div> <div> <h2> </h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns: </p> <p>'TweetID': Represents the tweet ID. </p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet. </p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed). </p> </div> <div> <p>'CreateDate': Represents the date the tweet was created. </p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0). </p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0). </p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username. </p> </div> <div> <p> ‘Cohort’: Represents the year the data was annotated (class of 2022 or class of 2023) </p> </div> <div> <h2> </h2> <h2>Acknowledgements </h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy. </p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. </p> </div> <div> <p> </p> </div>
Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying
<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p> </p> <p>* To cite this dataset refer to <a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., & Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p> </p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>
Hearing Aid Noisy Speech Dataset
<p>Speech dataset, based on the Google Speech Commands Dataset, containing (simulated) noisy own voice as if it were captured by a hearing aid with 2 (front and rear) behind-the-ear microphones and 1 in-ear-canal microphone.</p>
On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection
<p>Additional resources for the paper:</p> <h3><strong><a href="https://ieeexplore.ieee.org/abstract/document/10826088">On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection.</a></strong></h3> <p>Lewis, N., Cavalcante, C. C., Boukouvalas, Z., & Corizzo, R.</p> <p><em>2024 IEEE International Conference on Big Data (BigData)</em> (pp. 3277-3281). IEEE.</p> <pre> </pre> <p> </p> <p>MMHS150K [1] is a manually labeled multimodal dataset that contains $150000$ tweets with two modalities: text, and corresponding image. Tweets are collected from September 2018 until February 2019 and are labeled according to different types of hate speech: no attacks to any community, racist, sexist, homophobic, religion-based attacks, or attacks to other communities. </p> <p>We extract vector embeddings leveraging different text (BERT, OpenAI) and image (ResNet, PVT, ViT) modele backbones and assess their effectiveness in the hate speech detection task.</p> <p> </p> <h2>Citation:</h2> <pre>@inproceedings{lewis2024effectiveness, title={On the Effectiveness of Text and Image Embeddings in Multimodal Hate Speech Detection}, author={Lewis, Nora and Cavalcante, Charles C and Boukouvalas, Zois and Corizzo, Roberto}, booktitle={2024 IEEE International Conference on Big Data (BigData)}, pages={3277--3281}, year={2024}, organization={IEEE} }</pre>
PodcastMix - a dataset for separating music and speech in podcasts
<p><strong>Note: due to zenodo limitations here we host solely the metadata. the whole dataset can be found at: https://drive.google.com/drive/u/0/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z </strong></p> <p>We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. It contains audio files at 44.1kHz and the corresponding metadata. For further details check the following paper and the associated GitHub repository: </p> <ul> <li>N. Schmidt, J. Pons, M. Miron, "PodcastMix - a dataset for separating music and speech in podcasts", Interspeech (2022)</li> <li>N. Schmidt, "PodcastMix - a dataset for separating music and speech in podcasts", Masters thesis, MTG, UPF (2021) https://zenodo.org/record/5554790#.YXLHvNlByWA </li> <li>https://github.com/MTG/Podcastmix</li> </ul> <p>This dataset contains four parts. Due to zenodo file size limitation we host the training dataset on google drive. We highlight the content of the zenodo archives within brackets:</p> <ul> <li>[metadata] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[metadata] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[audio and metadata] PodcastMix-real with-reference : a test set with real podcasts with reference stems to compute evaluation metrics. The podcasts are recorded by one of the authors and the source of the music is the FMA dataset. </li> <li>[audio and metadata] PodcastMix-real no-reference: a test set with real podcasts with only the podcasts mixes for subjective evaluation. The podcasts are compiled from the internet. </li> </ul> <p>The training dataset, PodcastMix-synth may be found at our google drive repository: https://drive.google.com/drive/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z?usp=sharing . The archive comprises 450GB of audio and metadata with the following structure:</p> <ul> <li>[metadata and audio] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> <li>[metadata and audio] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset. </li> </ul> <p>Make sure you maintain the folder structure of the original dataset when you uncompress these files. </p> <p><br> This dataset is created by Nicolas Schmidt, Marius Miron, Music Technology Group - Universitat Pompeu Fabra (Barcelona) and Jordi Pons. This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 Unported License (CC BY-SA 4.0).</p> <p><br> Please acknowledge PodcastMix in Academic Research. When the present dataset is used for academic research, we would highly appreciate if authors quote the following publications:</p> <ul> <li>N. Schmidt, J. Pons, M. Miron, "PodcastMix - a dataset for separating music and speech in podcasts", Interspeech (2022)</li> <li>N. Schmidt, "PodcastMix - a dataset for separating music and speech in podcasts", Masters thesis, MTG, UPF (2021) https://zenodo.org/record/5554790#.YXLHvNlByWA </li> </ul> <p><br> The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the UPF is not liable for, and expressly excludes, all liability for loss or damage however and whenever caused to anyone by any use of the dataset or any part of it.</p> <p><br> PURPOSES. The data is processed for the general purpose of carrying out research development and innovation studies, works or projects. In particular, but without limitation, the data is processed for the purpose of communicating with Licensee regarding any administrative and legal / judicial purposes.<br> </p>
Audio samples from generative models trained on the TIMIT speech data.
<p>This is a posting of audio snippets to accompany the paper "Benchmarking Generative Latent Variable Models for Speech".</p> <p>The snippets include samples and reconstructions. All samples are completely unconditional and utilise only the prior internal representations learned by the model. Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>
Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic
<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>
Evaluation of a Novel 8-Channel RX Coil for Speech Production at 0.55 T DATASET.
<p>This dataset contains raw imaging MRI data used in SNR evaluation for 4 subjects. For all scans, volumetric data of the upper airway was obtained using a 3D spoiled gradient echo sequence with either a speech coil, a head coil, or an integrated body coil. Imaging parameters were: flip angle = 10, TE = 5 ms, TR = 10 ms, FOV = 32x32x16 cm^3, resolution = 2.5 x 2.5 x 5 mm^3. Pre-scan noise information for each scan is also included.</p> <p> </p> <p> </p>
Database for HF Transmitted Speech with Carrier Frequency Difference
<p>This public database is designed for evaluation of carrier frequency difference estimation systems and is published as part of our paper "Open Range Pitch tracking for Carrier Frequency Difference<br> Estimation from HF Transmitted Speech". It consists of over 23 ours of real transmissions over HF links with a known carrier frequency shift during demodulation.</p> <p>To record the data we have set up a transmission system between our base station in Paderborn and several other distant base stations across Europe (see Fig. 4), transmitting utterances from the LibriSpeech corpus.</p> <p>Kiwi-software defined radio (SDR) devices at distant base stations were utilized to demodulate the received SSB HF signals and send the recorded audio signals back to our servers via a websocket connection. Audio markers had been added to the transmitted signal to allow for an automated time alignment between the transmitted<br> and received signals, easing the annotation and segmentation of the data.<br> For the transmissions a beacon, callsign DB0UPB, was used, which was supervised by a human to avoid interference with other ham radio stations. The HF signals are SSB modulated using the Lower Side Band (LSB) with a bandwidth of 2.7 kHz at carrier frequencies of 7.06 MHz − 7.063 MHz and 3.6 MHz − 3.62 MHz. To simulate a carrier frequency difference the demodulation frequency of the transmitter and the receiver were selected to differ by values from the set [0, 100, 300, 500, 1000]. Although the original speech samples have a sampling rate of 16 kHz, and the Kiwi-SDR samples the data at 12.001 Hz, the finally emitted data is band-limited to 2.7 kHz (International Telecommunication Union (ITU) regulations) which introduces a loss of the upper frequencies in case of LSB<br> transmission depending on the carrier frequency difference. The data set has a total size of 23:31 hours of which 3:28 hours contain speech activity.</p>
Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"
<p>Dataset accompanying the paper "<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>"</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>
Dendi of Parakou multi-speaker speech dataset
<p>This dataset was created for speech research purposes and contains about 676 recordings of participants reading a script in Dendi as spoken in Parakou, one sentence at a time. Each example includes the audio files and the associated text. The audio is high-quality and recorded in a quiet environment. The dataset is multi-speaker, containing recordings from 2 volunteers (male and female), where each volunteer contributed up to 338 recordings. The recordings were done in 2022 and took place in Cotonou in Benin. The recommended use cases are the building of educational material for pupils. Keys applications are Machine learning and speech technology. </p>
A Database for Reasearch on Detection and Enhancement of Speech transmitted over HF links
<p>We present an open database for the development of detection and enhancement algorithms of speech transmitted over HF radio channels.<br> It consists of audio samples recorded by various receivers at different locations across Europe, all monitoring the same single-sideband modulated transmission from a base station in Paderborn, Germany. Transmitted and received speech signals are precisely time aligned to offer parallel data for supervised training of deep learning based detection and enhancement algorithms.</p>
Fongbe speech dataset
<p>This is a speech dataset for Fongbe language spoken mostly in Benin. The folder contains the following:</p> <ul> <li>fongbe_speech_audio_files: this folder contains all the audio files in wav folder and their transcripts in lab folder</li> <li>Dataset_documentation.pdf: this is the documentation for the dataset</li> <li>fongbe_speech_dataset_metadata.csv: it provides metadata information on each audio file</li> </ul>
Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Refinement dataset SPIRA
<p>This dataset was collected over the internet and in hospital wards with the goal of detecting respiratory insufficiency (typically caused by COVID-19). This data collection is part of the SPIRA Project, whose goal is developing a system for recognizing respiratory insufficiency through speech analysis. The datasets presented here were used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger.</p> <p>The spira_trimmed_data file contains the original ~1 hour dataset collected over the internet (control) and in hospital wards (patients) by the SPIRA Project. This is as described in the paper: Deep learning against COVID-19: Respiratory insufficiency detection in Brazilian Portuguese Speech. We include it here for completeness.</p> <p>The spira_control_full_mp3 file contains the complete ~18 hours control data collected over the internet by the SPIRA project. While not useful for respiratory insufficiency detection, the dataset may be used for identifying age and gender as we mention in our paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers.</p>
Baule speech dataset
<p>The dataset was created to enable research on automatic speech recognition in Boulé (Baule) language. It contains about 565 recordings of participants reading a transcription in Baule as spoken in Côte d’Ivoire, one sentence at a time. Each example contains the audio files and the associated text. The audio is recorded in a less noisy environment by the speakers using their android phone. The dataset is multi-speaker, containing recordings from 4 volunteers (2 males and 2 females), where each volunteer contributed up to 141 recordings. The recordings took place in Abidjan, Côte d’Ivoire in April 2022.</p>
Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Pretraining Datasets raw audios from CORAA
<p>This repository contains all the pretraining datasets used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger. These datasets are part of a collection of datasets from the TaRSila project (see https://sites.google.com/view/tarsila-c4ai). The audios published here were in part also published with annotations and transcriptions as the CORAA dataset (see https://github.com/nilc-nlp/CORAA). Here we publish the original raw audios from the following datasets (without transcriptions) - ALIP, C-Oral, SP2010, NURC-Recife, NURC-São Paulo and Programa Certas Palavras. In total, the datasets contain about 800 hours of Brazilian Portuguese Speech.</p> <p>The audios have been converted to mp3 to facilitate the upload. ALIP, C-Oral and SP2010 are integrally contained in one file each. Programa Certas Palavras and NURC-Recife are split in 3 parts each, while NURC-SP is split in 7 parts of roughly equal size. More information on the datasets can be found in the paper Acoustic models of Brazilian Portuguese Speech based on Neural Transformers as well as on the original references which created these datasets.</p>
Mandarin matrix sentence test recordings: Lombard and plain speech with different speakers
<p>This dataset was recorded within the Deutsche Forschungsgemeinschaft (DFG) project: Experiments and models of speech recognition across tonal and non-tonal language systems (EMSATON, Projektnummer 415895050).</p> <p>The Lombard effect or Lombard reflex is the involuntary tendency of speakers to increase their vocal effort when speaking in loud noise to enhance the audibility of their voice. Up to date, the phenomena of Lombard effects were observed in different languages. The present database aimed at providing recordings for studying the Lombard effect with Mandarin speech.</p> <p>Eleven native-Mandarin talkers (6 female and 5 male) were recruited, both Lombard/plain speech were recorded from the same talker in the same day. </p> <p>All speakers produced fluent standard Mandarin speech (North China). All listeners were normal-hearing with pure tone thresholds of 20 dB hearing level or better at audiometric octave frequencies between 125 and 8000 Hz. All listeners provided written informed consent, approved by the Ethics Committee of Carl von Ossietzky University of Oldenburg. Listeners received an hourly compensation for their participation.</p> <p>The recording sentences were same as the official Mandarin Chinese matrix sentence test (<a href="#_ENREF_2">CMNmatrix, Hu et al. 2018</a>). </p> <p> One hundred sentences (ten base lists of ten sentences) of the CMNmatrix were recorded from each speaker in both plain and Lombard speaking styles (each base list containing all 50 words). The 100 sentences were divided into 10 blocks of 10 sentences each, and the plain and Lombard blocks were presented in an alternating order. The recording took place in a double-walled, sound-attenuated booth fulfilling ISO 8253-3 (ISO 8253-3, 2012), using a Neumann 184 microphone with a cardioid characteristic (Georg Neumann GmbH, Berlin, Germany) and a Fireface UC soundcard (with a sampling rate of 44100 Hz and resolution of 16 bits). The recording procedure generally followed the procedures of Alghamdi et al. (2018). A Mandarin-native speaker and a phonetician participated in the recording session and listened to the sentences to control the pronunciations, intonation, and speaking rate. During the recording, the speaker was instructed to read the sentence presented on a frontal screen. In case of any mispronunciation or change in the intonation, the speaker was asked via the screen to repeat the sentence again, and on average, each sentence was recorded twice. In Lombard conditions the speaker was regularly asked via a prompt to repeat a sentence, to keep the speaker in the Lombard communication situation. For the plain-speech recording blocks, the speakers were asked to pronounce the sentences with natural intonation and accentuation, and at an intermediate speaking rate, which was facilitated by a progress bar on the screen. Furthermore, the speakers were asked to keep the speaking effort constant and to avoid any exaggerated pronunciations that could lead to unnatural speech cues. For the Lombard speech recording blocks, speakers were instructed to imagine a conversation to another person in a pub-like situation. During the whole recording session, speakers wore headphones (Sennheiser HDA200) that provided the audio signal of the speaker.. In the Lombard condition, the stationary speech-shaped noise ICRA1 (Dreschler et al., 2001) was mixed with the speaker’s audio signal at a level of 80 dB SPL (calibrated with a Brüel & Kjær (B&K) 4153 artificial ear, a B&K 4134 0.5-inch inch microphone, a B&K 2669 preamplifier, and a B&K 2610). Previous studies showed that this level induced a robust Lombard speech without the danger of inducing hearing damage (Alghamdi et al., 2018).</p> <p>The sentences were cut from the recording, high-pass filtered (60 Hz cut-off frequency) and set to the average root-mean-square level of the original speech material of the Mandarin Matrix test (Hu et al., 2018). Then the best version of each sentence was chosen by native-Mandarin speakers regarding pronunciation, tempo, and intonation.</p> <p>For more detailed information, please contact hongmei.hu@uni-oldenburg, sabine.hochmuth@uni-oldenburg.de. </p> <p><em>Hu H, Xi X, Wong LLN, Hochmuth S, Warzybok A, Kollmeier B (2018) Construction and evaluation of the mandarin chinese matrix (cmnmatrix) sentence test for the assessment of speech recognition in noise. International Journal of Audiology 57:838-850. https://doi.org/10.1080/14992027.2018.1483083</em></p> <p> </p>
BembaSpeech: A Speech Recognition Corpus for the Bemba Language
<p>We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.