Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

94

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

94 results for “speech dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal

<p>This data is transcribed speech data, in Wolof, Pulaar and Sereer.</p> <p>The recordings are about agriculture. The recorded consist of farmers, agricultural advisers, and agri-food business managers.&nbsp;Type of recordings comprise interactive radio programmes, focus groups, voice messages, push messages and interviews. Therefore, spontaneous speech is prevailing. Quality of audio may vary depending on the type of programme.</p> <p>Content description :</p> <ul> <li><strong>speech_dataset_wol.tar.gz:</strong> Wolof (ISO Code 639-2: wol) speech dataset contains 55 hours of transcribed speech, including almost 13 hours of validated content check by an expert. It also contains a XSAMPA lexicon (49,132 phonetised entries) and a text corpus (1,140,508 words).</li> <li><strong>speech_dataset_fuc.tar.gz:</strong> Pulaar (ISO Code 639-2: fuc) speech dataset contains nearly 32 hours of transcribed speech, including around 11 hours of validated content check by an expert. It also contains a text corpus (742,024 words).</li> <li><strong>speech_dataset_srr.tar.gz:</strong> Sereer (ISO Code 639-2: srr) speech dataset contains 38 hours of transcribed speech, including nearly 11 hours of validated content check by an expert.<br>In total, these resources provide 125 hours of transcribed speech in the 3 most widely spoken languages in Senegal, including 35 hours of checked transcriptions.</li> </ul> <p>This work is a result of the Kallaama project, funded by Lacuna Fund for 1 year, in 2023.&nbsp;</p> <p>See the <a title="Kallaama speech dataset" href="https://github.com/gauthelo/kallaama-speech-dataset" target="_blank" rel="noopener">GitHub repository</a> for more details about the dataset.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

SaGA++ Speech-Gesture Dataset Extension

<p>This is the dataset extension release for <a href="https://pub.uni-bielefeld.de/record/2001935">the SaGA dataset</a>.</p> <p>We have extracted anonymous features from all the 25 recordings and release them all now instead of only 6 recordings that were previously released. The modalities included are text annotations, gesture annotations, prosodic audio features, and movement trajectories. More details are provided in the README document included in the dataset.</p> <p><strong>Please reference the following papers in any publication making use of the Dataset:</strong></p> <p>Kucherenko, T., Nagy, R., Neff, M., Kjellstr&ouml;m, H., &amp; Henter, G. E. (2022).&nbsp;Multimodal analysis of the predictability of hand-gesture properties. 21st International Conference on Autonomous Agents and Multiagent Systems (AAMAS).&nbsp;</p> <p>L&uuml;cking, A., Bergman, K., Hahn, F., Kopp, S., &amp; Rieser, H. (2013). Data-based analysis of speech and gesture: The Bielefeld Speech and Gesture Alignment Corpus (SaGA) and its applications.&nbsp;<em>Journal on Multimodal User Interfaces</em>,&nbsp;<em>7</em>(1), 5-18.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Introducing the COVID-19 YouTube (COVYT) speech dataset featuring the same speakers with and without infection

<p>The COVYT dataset contains speech samples from individuals who self-reported their COVID-19 infection on public social media platforms (YouTube, Xiaohongshu). These videos, as well as accompanying videos of the same people prior to infection, were mined in an attempt to gather publicly-available data for COVID-19 research. This release includes the links to the original videos along with the accompanying&nbsp;manual segmentation and diarisation that identifies the utterances of the target individuals. We are additionally releasing features derived from the segmented utterances. Finally, the dataset includes partitioning information&nbsp;according to 4 different cross-validation schemes. See the arxiv pre-print for more details:&nbsp;https://arxiv.org/abs/2206.11045</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

CLDF dataset derived from Bremer's "Sociolinguistic Survey of Six Berta Speech Varieties in Ethiopia" from 2016

<p>Cite the source of the dataset as:</p> <blockquote> <p>Bremer, Nate D. (2016): A Sociolinguistic Survey of Six Berta Speech Varieties in Ethiopia. SIL Electronic Survey Reports 2016-007. Dallas: SIL International.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

BinauRec: A dataset to test the influence of the use of room impulse responses on binaural speech enhancement

<p>BinauRec is a dataset for binaural speech enhancement. It is composed of real recordings, measured and simulated room impulse responses for the same audio scenes. Measurements are realized using behind-the-ears hearing aid shells, with and without a dummy head.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Persian Speech to Test dataset

<p>The Persian Speech to Text dataset is a collection of audio files and their corresponding transcripts, provided in CSV file format. The dataset is intended for use in training machine learning models for the task of transcribing audio files in the Persian language into text. The dataset includes 60GB of data, consisting of audio files in the WAV format and their transcripts. Each CSV file corresponds to a single ZIP or RAR file, and the name of each CSV file is the same as the corresponding ZIP or RAR file. The CSV files contain the following columns:</p> <ul> <li>wav_filename: The name of the WAV file within the ZIP or RAR file.</li> <li>wav_filesize: The size of each audio file.</li> <li>transcript: The text transcription of the audio file.</li> <li>confidence_level: A measure of the accuracy of the transcription.</li> </ul> <p>This dataset is the largest open source dataset of its kind, and it is a valuable resource for researchers and developers working on natural language processing tasks involving the Persian language. The open source nature of the dataset means that it is freely available to be used and modified by anyone, making it an important resource for advancing research and development in the field.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

BreathBase: Intra-Speech Breathing Dataset

<p>BreathBase contains 5070 breath instances detected on the recordings of 20 participants reading pre-prepared random pseudo texts in 5 different postures with 4 different microphones, simultaneously.</p> <p>It is recorded in a studio with a maximum background noise of 40 dB SPL and with professional recording equipment. It also provides tagging for 5 different postures and 4 different channels as different recording conditions for data variety.</p> <p>More than 90% of the recordings is shorter than 600 milliseconds. The minimum number of breath instances per participant is 89, the maximum number of instances is 710 and the average for all participants is 253.5 breath instances.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

TweetBLM: A Hate Speech Dataset and Analysis of BlackLivesMatter-related Microblogs on Twitter

<p>Collection of BLM related tweets and their corresponding labels of hate speech.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Shared Acoustic Codes Underlie Emotional Communication in Music and Speech - Evidence from Deep Transfer Learning (Datasets)

<p>This repository contains the datasets used in the article "Shared Acoustic Codes Underlie Emotional Communication in Music and Speech - Evidence from Deep Transfer Learning" (Coutinho &amp; Schuller, 2017). </p> <p>In that article four different data sets were used: SEMAINE, RECOLA, ME14 and MP (acronyms and datasets described below). The SEMAINE (speech) and ME14 (music) corpora were used for the unsupervised training of the Denoising Auto-encoders (domain adaptation stage) - only the audio features extracted from the audio files in these corpora were used and are provided in this repository. The RECOLA (speech) and MP (music) corpora were used for the supervised training phase -  both the audio features extracted from the audio files and the Arousal and Valence annotations were used. In this repository, we provide the audio features extracted from the audio files for both corpora, and Arousal and Valence annotations for some of the music datasets (those that the author of this repository is the data curator).</p> <p>Below, you can find description of the various corpora, the details about the data stored in this repository and information on how to obtain the rest of the data used by Coutinho and Schuller (2017).</p> <p><strong>SEMAINE (speech)</strong></p> <p>The SEMAINE corpus (McKeown, Valstar, Cowie, Pantic &amp; Schroder, 2012) was developed specifically to address the task of achieving emotion-rich interactions, and it is adequate for this task as it comprises a wide range of emotional speech. It includes video and speech recordings of spontaneous interactions between human and emotionally stereotyped `characters'. Coutinho &amp; Schuller (2017) used a subset of this database (called <em>Solid-SAL</em>). The <em>Solid-SAL</em> dataset is freely available for scientific research purposes (see http://semaine-db.eu). This repository includes the audio features used in Coutinho &amp; Schuller (2017) (under features/SEMAINE).</p> <p><strong>RECOLA (speech)</strong></p> <p>The RECOLA database (Ringeval, Sonderegger, Sauer &amp; Lalanne, 2013) consists of multimodal recordings (audio, video, and peripheral physiological activity) of spontaneous dyadic interactions between French adults. Coutinho &amp; Schuller (2017) used the RECOLA-Audio module which consists of the audio recordings of each participant in the dyadic phase of the task. In particular, they used the non-segmented high-quality audio signals (WAV format, 44.1kHz, 16bits), obtained through unidirectional headset microphones, of the first five minutes of each interaction. Annotations consist of time-continuous ratings of the level of Arousal and Valence dimensions of emotion perceived by each rater while seeing and listening the audio-visual recordings of each participant task. The publicly available annotated dataset includes only part of the data which amounts to a total number of 23 instances. The time frame length used by Coutinho &amp; Schuller (2017) is 1s (the original annotations were downsampled). This repository includes the audio features used in Coutinho &amp; Schuller (2017) (under features/RECOLA). To obtain the annotations you should contact the author of the original study (see https://diuf.unifr.ch/diva/recola/download.html for further details).</p> <p><strong>ME14 (music)</strong></p> <p>The MediaEval ``Emotion in Music'' task is dedicated to the estimation of Arousal and Valence scores continuously in time and value for song excerpts from the Free Music Archive. Coutinho and Schuller (2017) used the whole corpus (development and test sets for the 2014 challenge) which includes 1,744 songs belonging to 11 musical styles -- Soul, Blues, Electronic, Rock, Classical, Hip-Hop, International, Folk, Jazz, Country, and Pop (maximum of five songs per artist). This repository includes the audio features used in Coutinho &amp; Schuller (2017) (under features/ME14). The full dataset (including annotations) can be obtained from http://www.multimediaeval.org/mediaeval2014/emotion2014/.</p> <p><strong>MP (music)</strong></p> <p>This is a corpus compiled specifically for this work described in Coutinho &amp; Schuller (2017) using data collected in four previous studies. It consists of emotionally diverse full music pieces from a variety of musical styles (Classical and contemporary Western Art, Baroque, Bossa Nova, Rock, Pop, Heavy Metal, and Film Music). Annotations were obtained in controlled laboratory experiments whereby the emotional character of each piece was evaluated time-continuously in terms of levels of Arousal and Valence perceived by listeners (ranging between 35 to 52 in the four studies). In what follows, some details about the various studies are described.</p> <ul> <li>MP<sub>DB1</sub>: This subset of the MP corpus consists of the data reported by Korhonen (2004), and gently made available by the author. This dataset includes six full (or long excerpts) music pieces ranging from 151s to 315s in length (only classical music). Each piece was annotated by 35 participants (14 females). The time series correspondents to each music piece were collected at 1Hz. The golden standard for each piece was computed by averaging the individual time series across all raters. This repository includes the audio features used in Coutinho &amp; Schuller (2017) (under features/MP/DB1). To obtain the labels please contact the author of the original study.</li> <li>MP<sub>DB2</sub>: The dataset by Coutinho &amp; Cangelosi (2011) includes 9 full pieces (43s to 240s long) of classical music (romantic repertoire) annotated by 39 subjects (19 females). Values were recorded every time the mouse was moved with a precision of 1 ms. The resultant timeseries were then resampled (moving average) to a synchronous rate of 1 Hz. The golden standard for each piece was computed by averaging the individual time series across all raters. This repository includes the audio features (under features/MP/DB2) and labels (under annotations/MP/DB2) used in Coutinho &amp; Schuller (2017).</li> <li>MP<sub>DB3</sub>: This dataset was collected by Coutinho &amp; Dibben (2012) and it consists of 8 pieces of film music (84s to 130s long) taken from the late 20th century Hollywood film repertoire. Emotion ratings were given by 52 participants (26 females). The annotation procedure, data processing, and golden standard calculations were identical to MP<sub>DB2</sub>. This repository includes the audio features (under features/MP/DB3) and labels (under annotations/MP/DB3) used in Coutinho &amp; Schuller (2017).</li> <li>MP<sub>DB4</sub>: This dataset was collected by Grewe, Nagel, Kopiez and Altenmüller (2007), and gently made available by the authors. It includes seven music pieces (127s to 502s in length) of heterogeneous styles (e.g., Rock, Pop, Heavy Metal, Classical). Each music piece was annotated by 38 participants (29 females) using an identical methodology to MP<sub>DB2</sub> and MP<sub>DB3</sub>. Data processing and golden standard calculations were also identical. This repository includes the audio features (under features/MP/DB4) used in Coutinho &amp; Schuller (2017). To obtain the labels contact the authors of the original study</li> </ul> <p> </p> <p><strong>Bibliography</strong></p> <p>Coutinho, E., &amp; Cangelosi, A. (2011). Musical emotions: predicting second-by-second subjective feelings of emotion from low-level psychoacoustic features and physiological measurements. <em>Emotion</em>, <em>11</em>(4), 921.</p> <p>Coutinho, E., &amp; Dibben, N. (2013). Psychoacoustic cues to emotion in speech prosody and music. <em>Cognition &amp; Emotion</em>, <em>27</em>(4), 658-684.</p> <p>Coutinho E, Schuller B (2017) Shared acoustic codes underlie emotional communication in music and speech—Evidence from deep transfer learning. PLoS ONE 12(6): e0179289. https://doi. org/10.1371/journal.pone.0179289.</p> <p>Grewe, O., Nagel, F., Kopiez, R., Altenmüller, E. (2007). Emotions over time: synchronicity and development of subjective, physiological, and facial affective reactions to music. <em>Emotion, 7</em>(4), pp. 774-788. DOI: 10.1037/1528-3542.7.4.774.</p> <p>Korhonen, M. (2004). Modeling Continuous Emotional Appraisals of Music Using System Identification. Available from: http://hdl.handle.net/10012/879.</p> <p>McKeown, G., Valstar, M., Cowie, R., Pantic, M., Schroder, M. (2012). The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent. <em>IEEE Transactions on Affective Computing</em>, 3, pp. 5-17. DOI: http://doi.ieeecomputersociety.org/10.1109/T-AFFC.2011.20.</p> <p>Ringeval, F.,  Sonderegger, A., Sauer, J. &amp; Lalanne, D. (2013). Introducing the RECOLA Multimodal Corpus of Remote Collaborative and Affective Interactions. In <em>Proceedings of the 2nd International Workshop on Emotion Representation, Analysis and Synthesis in Continuous Time and Space (EmoSPACE 2013)</em>, Shanghai, China. IEEE</p>

opencc-by-4.0Mar 2017View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying

<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p>&nbsp;</p> <p>* To cite this dataset refer to&nbsp;<a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., &amp; Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p>&nbsp;</p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Hearing Aid Noisy Speech Dataset

<p>Speech dataset, based on the Google Speech Commands Dataset, containing (simulated) noisy own voice as if it were captured by a hearing aid with 2 (front and rear) behind-the-ear microphones and 1 in-ear-canal microphone.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

PodcastMix - a dataset for separating music and speech in podcasts

<p><strong>Note: due to zenodo limitations here we host solely the metadata. the whole dataset can be found at: https://drive.google.com/drive/u/0/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z </strong></p> <p>We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. It contains audio files at 44.1kHz and the corresponding metadata. For further details check the following paper and the associated GitHub repository:&nbsp;</p> <ul> <li>N. Schmidt, J. Pons, M. Miron, &quot;PodcastMix - a dataset for separating music and speech in podcasts&quot;, Interspeech&nbsp;(2022)</li> <li>N. Schmidt, &quot;PodcastMix - a dataset for separating music and speech in podcasts&quot;, Masters thesis, MTG, UPF (2021)&nbsp;https://zenodo.org/record/5554790#.YXLHvNlByWA&nbsp;</li> <li>https://github.com/MTG/Podcastmix</li> </ul> <p>This dataset contains four parts. Due to zenodo file size limitation we host the training dataset on google drive. We highlight the content of the zenodo archives within brackets:</p> <ul> <li>[metadata] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset.&nbsp;</li> <li>[metadata] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset.&nbsp;</li> <li>[audio and metadata] PodcastMix-real with-reference : a test set with real podcasts with reference stems to compute evaluation metrics. The podcasts are recorded by one of the authors and the source of the music is the FMA dataset.&nbsp;</li> <li>[audio and metadata] PodcastMix-real no-reference: a test set with real podcasts with only the podcasts mixes for subjective evaluation. The podcasts are compiled from the internet.&nbsp;</li> </ul> <p>The training dataset, PodcastMix-synth may be found at our google drive repository:&nbsp;https://drive.google.com/drive/folders/1tpg9WXkl4L0zU84AwLQjrFqnP-jw1t7z?usp=sharing . The archive comprises 450GB of audio and metadata with the following structure:</p> <ul> <li>[metadata and audio] PodcastMix-synth train: large and diverse training set that is programatically generated (with a validation partition). The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset.&nbsp;</li> <li>[metadata and audio] PodcastMix-synth test a programatically generated test set with reference stems to compute evaluation metrics. The mixtures are created programatically with music from Jamendo and speech from the VCTK dataset.&nbsp;</li> </ul> <p>Make sure you maintain the folder structure of the original dataset when you uncompress these files.&nbsp;</p> <p><br> This dataset is created by Nicolas Schmidt, Marius Miron, Music Technology Group - Universitat Pompeu Fabra (Barcelona) and Jordi Pons. This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 Unported License (CC BY-SA 4.0).</p> <p><br> Please acknowledge PodcastMix in Academic Research. When the present dataset is used for academic research, we would highly appreciate if authors quote the following publications:</p> <ul> <li>N. Schmidt, J. Pons, M. Miron, &quot;PodcastMix - a dataset for separating music and speech in podcasts&quot;, Interspeech (2022)</li> <li>N. Schmidt, &quot;PodcastMix - a dataset for separating music and speech in podcasts&quot;, Masters thesis, MTG, UPF (2021)&nbsp;https://zenodo.org/record/5554790#.YXLHvNlByWA&nbsp;</li> </ul> <p><br> The dataset and its contents are made available on an &ldquo;as is&rdquo; basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the UPF is not liable for, and expressly excludes, all liability for loss or damage however and whenever caused to anyone by any use of the dataset or any part of it.</p> <p><br> PURPOSES. The data is processed for the general purpose of carrying out research development and innovation studies, works or projects. In particular, but without limitation, the data is processed for the purpose of communicating with Licensee regarding any administrative and legal / judicial purposes.<br> &nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Evaluation of a Novel 8-Channel RX Coil for Speech Production at 0.55 T DATASET.

<p>This dataset contains raw imaging MRI&nbsp;data used in SNR evaluation for 4 subjects. For all scans, volumetric data of the upper airway was obtained using a 3D spoiled gradient echo sequence with either a speech coil, a head coil, or an integrated body coil.&nbsp;Imaging parameters were: flip angle = 10, TE = 5 ms, TR = 10 ms, FOV = 32x32x16 cm^3, resolution = 2.5 x 2.5 x 5 mm^3.&nbsp; Pre-scan noise information for each scan is also included.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Dendi of Parakou multi-speaker speech dataset

<p>This dataset was created for speech research purposes and contains about 676 recordings of participants reading a script in Dendi as spoken in Parakou, one sentence at a time. Each example includes the audio files and the associated text. The audio is high-quality and recorded in a quiet environment. The dataset is multi-speaker, containing recordings from 2 volunteers (male and female), where each volunteer contributed up to 338 recordings. The recordings were done in 2022 and took place in Cotonou in Benin.&nbsp;The recommended use cases are the building of educational material for pupils. Keys applications are Machine learning and speech technology.&nbsp;&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Fongbe speech dataset

<p>This is a speech dataset for&nbsp; Fongbe language spoken mostly in Benin. The folder contains the following:</p> <ul> <li>fongbe_speech_audio_files: this folder contains all the audio files in wav folder and their transcripts in lab folder</li> <li>Dataset_documentation.pdf: this is the documentation for the dataset</li> <li>fongbe_speech_dataset_metadata.csv: it provides metadata information on&nbsp; each audio file</li> </ul>

opencc-by-4.0May 2022View details →
zenodo40/100

Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Refinement dataset SPIRA

<p>This dataset was collected over the internet and in hospital wards with the goal of detecting respiratory insufficiency (typically caused by COVID-19). This data collection is part of the SPIRA Project, whose goal is developing a system for recognizing respiratory insufficiency through speech analysis. The datasets presented here were used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger.</p> <p>The spira_trimmed_data file contains the original ~1 hour dataset collected over the internet (control) and in hospital wards (patients) by the SPIRA Project. This is as described in the paper: Deep learning against COVID-19: Respiratory insufficiency detection in Brazilian Portuguese Speech. We include it here for completeness.</p> <p>The spira_control_full_mp3 file contains the complete ~18 hours control data collected over the internet by the SPIRA project. While not useful for respiratory insufficiency detection, the dataset may be used for identifying age and gender as we mention in our paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Baule speech dataset

<p>The dataset was created to enable research on automatic speech recognition in Boul&eacute; (Baule) language. It contains about 565 recordings of participants reading a transcription in Baule as spoken in C&ocirc;te d&rsquo;Ivoire, one sentence at a time. Each example contains the audio files and the associated text. The audio is recorded in a less noisy environment by the speakers using their android phone. The&nbsp;dataset is multi-speaker, containing recordings from 4 volunteers (2 males and 2 females), where each volunteer contributed up to 141 recordings. The recordings took place in Abidjan, C&ocirc;te d&rsquo;Ivoire in April 2022.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Pretraining Datasets raw audios from CORAA

<p>This repository contains all the pretraining datasets used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger. These datasets are part of a collection of datasets from the TaRSila project (see https://sites.google.com/view/tarsila-c4ai). The audios published here were in part also published with annotations and transcriptions as the CORAA dataset (see https://github.com/nilc-nlp/CORAA). Here we publish the original raw audios from the following datasets (without transcriptions) - ALIP, C-Oral, SP2010, NURC-Recife, NURC-S&atilde;o Paulo and Programa Certas Palavras. In total, the datasets contain about 800 hours of Brazilian Portuguese Speech.</p> <p>The audios have been converted to mp3 to facilitate the upload. ALIP, C-Oral and SP2010 are integrally contained in one file each. Programa Certas Palavras and NURC-Recife are split in 3 parts each, while NURC-SP is split in 7 parts of roughly equal size. More information on the datasets can be found in the paper Acoustic models of Brazilian Portuguese Speech based on Neural Transformers as well as on the original references which created these datasets.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record