Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “speech corpus”

Learn how ShareScore rates datasets ↗
zenodo36/100

Defining and identifying Discourse Markers in spontaneous speech. A corpus-based and experimental proposal

<p>The paper has a twofold goal: (i) to identify the cathegory of Discourse Markers (DM) and their different specific functions; (ii) to validate the proposal with a perceptual experiment. First, we propose how to identify DM and their different functions in spontaneous speech. Both the identification of the categorycategory of DM and that of specific DM functions are based on prosodic criteria. In order to identify a DM, speech segmentation is crucial, since DMs necessarily appear isolated in a prosodic unit. Besides prosodic isolation, DMs do not feature pragmatic and prosodic autonomy, but depend on the illocutionary unit of the utterance.&nbsp; Than we show that a same lexical item can fulfill more functions, while the formal cues that convey the function are prosodic ones. The last part of the chapter is devoted to presenting the methodology and the results of a perceptual experiment that tests the theoretical hypothesis by asking the listeners to recognize three different functions by means of prosodic cues only.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Defining and identifying Discourse Markers in spontaneous speech: a corpus-based and experimental proposal

<p>Audio files used as examples in the paper.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

The Grid Audio-Visual Speech Corpus

<p>The Grid Corpus is a large multitalker audiovisual sentence corpus designed to support joint computational-behavioral studies in speech perception. In brief, the corpus consists of high-quality audio and video (facial) recordings of 1000 sentences spoken by each of 34 talkers (18 male, 16 female), for a total of 34000 sentences. Sentences are of the form &quot;put red at G9 now&quot;.</p> <p>audio_25k.zip &nbsp;contains the wav format utterances at a 25 kHz sampling rate in a separate directory per talker<br> alignments.zip provides word-level time alignments, again separated by talker<br> s1.zip, s2.zip etc contain .jpg videos for each talker [note that due to an oversight, no video for talker t21 is available]</p> <p>The Grid Corpus is described in detail in the paper jasagrid.pdf included in the dataset.</p>

opencc-by-4.0Dec 2005View details →
zenodo32/100

League of Legends and hate speech: a corpus for comments in Twitch.tv

<p>League of Legends (LOL) is the most popular game on PC, drawing 8 million concurrent players. A common activity of gamers, besides playing games, is to watch other players presenting tips and tricks. Streaming platforms allow some players to show gameplays and live games. <a href="https://www.twitch.tv/">Twitch.tv</a> is the world&acute;s leading live streaming platform.&nbsp;</p> <p>Considering that hate speech is a ubiquitous problem in online gaming, we collected &nbsp;985,766 comments from five videos of the top 10 &nbsp;LOL streamers in Twitch.tv platform.&nbsp;</p> <p>The dataset is freely available in a single file, ensembling all videos/players; and divided by players as well.&nbsp;</p> <p>These comments are a rich data source for opinion mining, sentiment analysis, topic modeling, and hate speech detection (including sexism and racism).</p> <ul> </ul>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Speech Quality Apollo Corpus

<p>32 audio clips extracted from the Apollo Space Program archive&nbsp;recordings annotated with subjective intelligibility and non-intrusive objective&nbsp;metrics.</p> <p><strong>Please cite if you use this dataset:</strong></p> <p>A. Ragano, E. Benetos, and A. Hines, &quot;Development of a speech quality database under uncontrolled conditions&quot;, in Proc. 21st Annual Conference of the International Speech Communication Association (INTERSPEECH), pp. 4616-4620, Oct. 2020</p> <p>Paper available here https://www.isca-speech.org/archive/Interspeech_2020/pdfs/1899.pdf</p> <p><strong>Annotations:</strong></p> <ul> <li><strong>google_wer:&nbsp;</strong>word-error-rate of Google speech-to-text API</li> <li><strong>mosnet:</strong>&nbsp;objective quality&nbsp;score of MOSNet metric</li> <li><strong>srmr:&nbsp;</strong>objective quality score of SRMR metric</li> <li><strong>p563: &nbsp;</strong>objective quality score of ITU-T P.563</li> <li><strong>subjwer</strong>: mean&nbsp;word-error-rate of the study participants&nbsp;</li> </ul> <p>See INTERSPEECH manuscript for more information</p>

opencc-byJul 2020View details →
zenodo32/100

The Maaloula Aramaic Speech Corpus (MASC)

<p>This dataset contains the first electronic speech corpus of Maaloula Aramaic, an endangered Western Neo-Aramaic variety spoken in Syria. This 64,845-word corpus is available in four formats: (1) transcription, (2) lemmatized transcription, (3) audio files and time-aligned phonetic transcriptions, and (4) an SQLite database. The transcription files are a digitized and corrected version of authentic transcriptions of tape-recorded narratives coming from a fieldwork trip conducted in the 1980s and published in the early 1990s (Arnold, 1991a, 1991b). They contain no annotation, except for some informative tagging (e.g. to mark loanwords and misspoken words). In the lemmatized version of the files, each word form is followed by its lemma in angled brackets. The time-aligned TextGrid annotations consist of four tiers: the sentence level (Tier 1), the word level (Tiers 2 and 3), and the segment level (Tier 4). These TextGrid files are downloadable together with their audio files (for the original source of the audio data see Arnold, 2003). The SQLite database enables users to access the data on the level of tokens, types, lemmas, sentences, stories, or speakers.</p> <p>For more information, please see our paper:&nbsp;Ghattas Eid, Esther Seyffarth, Ingo Plag. 2022. The Maaloula Aramaic Speech Corpus (MASC): From Printed Material to a Lemmatized and Time-Aligned Corpus. In&nbsp;<em>Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)</em>, Marseille, France. European Language Resources Association (ELRA).</p>

openother-ncJan 2022View details →
zenodo32/100

HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers

<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss&rsquo; Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Human vocalization corpus: recordings of infant-directed and adult-directed speech and song in 21 societies

<p>This repository&nbsp;contains a corpus of 1615 audio recordings of speech and song collected in 21 societies,&nbsp;first reported&nbsp;in Moser et al. (2020;&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2020.04.09.032995v5">bioRxiv</a>) and later published in Hilton &amp; Moser et al. (2022; <a href="https://doi.org/10.1038/s41562-022-01410-x">Nature Human Behaviour</a>).&nbsp;For assistance using any of this, contact Cody Moser (<a href="mailto:cmoser2@ucmerced.edu">cmoser2@ucmerced.edu</a>), Courtney Hilton (<a href="mailto:courtney.hilton@auckland.ac.nz">courtney.hilton@auckland.ac.nz</a>), and Samuel Mehr (<a href="mailto:mehr@hey.com">mehr@hey.com</a>).</p> <p>Two versions of the audio are included: raw audio (`IDS-corpus-raw.zip`) and audio that was edited to prepare the recordings for automatic acoustic feature extraction (`IDS-corpus-edited.zip`). `IDS-textGrids.zip` contains annotation files from Praat&#39;s silence detection method, which were manually reviewed for accuracy. These files are used with the audio extraction scripts associated with the project (see code linked in paper) to build the edited audio files.</p> <p>`IDS-fieldsites.csv` contains some fieldsite-level metadata; additional metadata is in the Supplementary Information of the paper.</p> <p>In the two .zip archives, filenames have the format XXXYYZ.wav,&nbsp;where &quot;XXX&quot; is a fieldsite code, &quot;YY&quot; is a participant number, and &quot;Z&quot; is a vocalization type.</p> <p>Fieldsite codes are:</p> <blockquote> <p>MBE: Mbendjele BaYaka<br> HAD: Hadza<br> NYA: Nyangatom<br> TOP: Toposa<br> BEJ: Beijing<br> JEN: Jenu Kurubas<br> MEN: Mentawai Islanders<br> KRA: Krakow<br> LIM: Rural Poland<br> TUR: Turku<br> USD: San Diego<br> TOR: Toronto<br> VAN: Tannese Vanuatuans<br> PNG: Enga<br> WEL: Wellington<br> ARA: Arawak<br> TSI: Tsimane<br> SPA: S&aacute;para &amp; Achuar<br> QUE: Quechua<br> ACO: Afrocolombians<br> MES: Colombian Mestizos</p> </blockquote> <p>Participant numbers are padded integers, starting with 01, and are unique within fieldsites.</p> <p>Vocalization types are:</p> <blockquote> <p>A: infant-directed song<br> B: infant-directed speech<br> C: adult-directed song<br> D: adult-directed speech&nbsp;</p> </blockquote> <p>In a few cases, participants vocalized in a different language than was expected, given the primary language of their fieldsite (e.g., when the participant was multilingual, or if they sang a song that contains multiple languages, as in The Beatles&#39; &quot;Michelle&quot;). The file `IDS-unexpectedLanguages.csv` at&nbsp;<a href="https://github.com/themusiclab/infant-speech-song/blob/main/data/IDS-unexpectedLanguages.csv">https://github.com/themusiclab/infant-speech-song/blob/main/data/IDS-unexpectedLanguages.csv</a> contains an inventory of these examples from the English-speaking fieldsites. This issue only affects&nbsp;a small minority of the recordings, as it&nbsp;was typically avoided by the researchers collecting the recordings.&nbsp;</p>

openApr 2020View details →
zenodo32/100

RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain

<p>Dear User,</p> <p>We are thrilled to introduce our latest release - the <strong>RescueSpeech</strong>&nbsp;audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.</p> <p>The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.</p> <p>1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.</p> <p>2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.</p> <p>By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.<br> &nbsp;</p>

opencc-by-nc-4.0Jun 2023View details →
zenodo32/100

List of European Parliament plenary speeches selected for the corpus together with speakers' names (Nov-2014 to Apr-2018); examples of collocations of "refugee(s)", "refugié(s)", "Flüchtling(e)" and "menekült(ek)"

<p>This data relates to the article &quot;Hidden Patterns in interpreted xenophobic discourse in the European Parliament&quot; [in print].</p> <p>The data contains a chronological list of the plenary debates from which the speeches were taken as well as the names of each speaker. It also also contains examples of verbs collocating with the term <em>refugee(s</em>), <em>refugi&eacute;(s)</em>, <em>Fl&uuml;chtling(e)</em> and <em>menek&uuml;lt(ek)</em> in the four language versions (English, French, German, Hungarian). These collocations were identified by the author of the paper.</p> <p>The speeches were downloaded from the Multimedia Center on the European Parliament&#39;s pubilc website: <a href="https://multimedia.europarl.europa.eu/en/home">https://multimedia.europarl.europa.eu/en/home</a>.</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

Urdu-Sindhi Speech Emotion Corpus

<p>The <strong>Urdu-Sindhi Speech Emotion Corpus</strong> is a dataset collected at Mehran University of Engineering &amp; Technology, Pakistan by a research team led by Dr. Zafi Sherhan Syed and Dr. Sajjad Ali Memon. The dataset consists on&nbsp;1,435 audio recordings in total for seven types of emotions which include&nbsp;anger, disgust, happiness, neutral, sarcasm, sadness, and surprise in two low-resource languages of South Asia,&nbsp;that is Urdu and Sindhi.</p> <p>Due to ethical restrictions we cannot release audio recordings at the moment and instead release five feature sets from the OpenSmile toolkit.</p> <p>Please note that we&nbsp;will endeavour&nbsp;to compute any features for you locally and send those features back to you. If this interests you, please contact Dr. Zafi Syed at zafisherhan.shah@faculty.muet.edu.pk</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

Code-Switching Speech Corpus

<p><strong>German-English Code-Switching speech dataset</strong></p> <p>We provide means to resegment a subset of the German **Spoken Wikipedia Corpus** (SWC) enabling a particular focus on code-switching.&nbsp; This results in the German-English code-switching corpus, a 34h transcribed speech corpus of read Wikipedia articles which can be used as a benchmark for research on code-switching.&nbsp; The articles are read by a large and diverse group of people. The SWC is perhaps the largest corpus of freely-available aligned speech for German.&nbsp; It contains 1014 spoken articles read by more than 350 identified speakers comprising 386h of speech. This corpus is available at http://nats.gitlab.io/swc.</p> <p>In SWC, since most of the articles are long, the recordings submitted by the volunteers are also long (&sim;54min) on average.&nbsp; These audio files are manually annotated at word-level and also segment level in XML format.&nbsp; We use a language identification tool to detect code-switching in the transcription of the audio files with consecutive indices. To extract intra-sentential code-switching segments, we ensure that the detected code-switching is preceded and followed by German words or sentences. The final set consists of 34h of speech data and 12,437 code-switching segments (in Kaldi ASR toolkit data format).</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>@article{baumann2019spoken,<br> &nbsp; title={The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening},<br> &nbsp; author={Baumann, Timo and K{\&quot;o}hn, Arne and Hennig, Felix},<br> &nbsp; journal={Language Resources and Evaluation},<br> &nbsp; volume={53},<br> &nbsp; number={2},<br> &nbsp; pages={303--329},<br> &nbsp; year={2019},<br> &nbsp; publisher={Springer}<br> }</p> <p>@article{grave2018learning,<br> &nbsp; title={Learning word vectors for 157 languages},<br> &nbsp; author={Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas},<br> &nbsp; journal={arXiv preprint arXiv:1802.06893},<br> &nbsp; year={2018}<br> }</p> <p>&nbsp;</p>

opencc-by-sa-3.0Jan 2021View details →
zenodo28/100

VIVOS: Vietnamese Speech Corpus for ASR

<p><strong>VIVOS Corpus</strong></p> <p>VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.</p> <p>The corpus was published by AILAB, a computer science lab of VNUHCM - University of Science, with&nbsp;<strong>Prof. Vu Hai Quan</strong>&nbsp;is the head of.</p> <p>We publish this corpus in hope to attract more scientists to solve Vietnamese speech recognition problems. The corpus should only be used for academic purposes.</p> <p><strong>License</strong></p> <p>Creative Commons Attribution NonCommercial ShareAlike v4.0 (CC BY-NC-SA 4.0) (<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">details</a>)</p> <p><strong>Associated paper</strong></p> <p>Please cite this paper when using VIVOS corpus for research</p> <p>&quot;A non-expert Kaldi recipe for Vietnamese Speech Recognition System&quot;, Hieu-Thi Luong and Hai-Quan Vu, in Proc. <em>WLSI/OIAF4HLT2016</em> (<a href="https://aclanthology.org/W16-5207/">paper</a>)</p> <p><strong>Contact</strong></p> <p><a href="mailto:ailab@hcmus.edu.vn">ailab@hcmus.edu.vn</a></p> <p>&nbsp;</p> <p><strong>Corpus properties</strong></p> <p>Speech was recorded in a quiet environment with high quality microphone, speakers were asked to read one sentence at a time.</p> <table> <thead> <tr> <th scope="col">&nbsp;</th> <th scope="col">Training</th> <th scope="col">Testing</th> </tr> </thead> <tbody> <tr> <td>Speakers</td> <td>46</td> <td>19</td> </tr> <tr> <td>Male</td> <td>22</td> <td>12</td> </tr> <tr> <td>Female</td> <td>24</td> <td>7</td> </tr> <tr> <td>Utterances</td> <td>11660</td> <td>760</td> </tr> <tr> <td>Duration</td> <td>14:55</td> <td>00:45</td> </tr> <tr> <td>Unique Syllables</td> <td>4617</td> <td>1692</td> </tr> </tbody> </table> <p><br> <strong>Evaluations</strong></p> <p>The corpus was evaluated using our non-expert recipe for Vietnamese Speech Recognition system which is described&nbsp;in the associated paper.</p> <table> <thead> <tr> <th scope="col">&nbsp;</th> <th scope="col">baseline</th> <th scope="col">+pitch</th> <th scope="col">+tone</th> </tr> </thead> <tbody> <tr> <td>mGMM</td> <td>19.66</td> <td>15.14</td> <td>14.91</td> </tr> <tr> <td>mGMM+MMI</td> <td>18.08</td> <td>14.96</td> <td>13.91</td> </tr> <tr> <td>mGMM+SAT</td> <td>15.79</td> <td>12.07</td> <td>12.13</td> </tr> <tr> <td>mDNN+SAT</td> <td>13.34</td> <td>9.54</td> <td>9.48</td> </tr> </tbody> </table> <p><strong>Notice</strong></p> <p>This is the official replacement for http://ailab.hcmus.edu.vn/vivos/</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2016View details →
zenodo24/100

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

<p>The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda.</p> <p>The corpus of 155&nbsp;hours is publicly available online under the Creative Commons BY-NC-ND 4.0 license.&nbsp;The dataset release is comprised of the following:</p> <ol> <li>20 hours of human transcribed radio speech. The audio is 16kHZ, mono channel and with 16 bit rate.&nbsp;</li> <li>Two CSV files for the 20-hour human transcribed dataset - cleaned.csv contains cleaned transcripts and uncleaned.csv contains uncleaned transcripts. The uncleaned transcripts contain extra speech details included in tags like [laughter] for laughter, and [um] for filler pauses, which speaker is talking, where each speaker is assigned an identifier A or B.</li> <li>The transcription guide used to transcribe the radio dataset.</li> <li>&nbsp;A multi-speaker untranscribed dataset of 6 hours of radio data. 1.4 hours of women voices and 4.6 hours of men voices. Each audio is a ten-seconds clip with a single speaker.</li> <li>&nbsp;135&nbsp;hours of multi-speaker untranscribed radio data.</li> </ol> <p><strong>NOTE: You can read and cite our paper published in the&nbsp;</strong><a href="http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.208.pdf"><strong>Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)</strong></a> The Dataset is published under Creative Commons BY-NC-ND 4.0 license and in order for us to monitor who is using it for the right license we request that you reach out to us officially.&nbsp;</p>

restrictedcc-by-4.0Jan 2022View details →
zenodo20/100

STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions

<p>We present STT4SG-350 (Speech-to-Text for Swiss German), a corpus of Swiss German speech, annotated with Standard German text at the sentence level. The data is collected using a web app in which the speakers are shown Standard German sentences, which they translate to Swiss German and record. We make the corpus publicly available. It contains 343 hours of speech from all dialect regions and is the largest public speech corpus for Swiss German to date. Application areas include automatic speech recognition (ASR), text-to-speech, dialect identification, and speaker recognition. Dialect information, age group, and gender of the 316 speakers are provided. Genders are equally represented and the corpus includes speakers of all ages. Roughly the same amount of speech is provided per dialect region, which makes the corpus ideally suited for experiments with speech technology for different dialects. We provide training, validation, and test splits of the data. The test set consists of the same spoken sentences for each dialect region and allows a fair evaluation of the quality of speech technologies in different dialects. We train an ASR model on the training set and achieve an average BLEU score of 74.7 on the test set. The model beats the best published BLEU scores on 2 other Swiss German ASR test sets, demonstrating the quality of the corpus.</p>

restrictedOct 2023View details →
zenodo20/100

NeuroVoz: a Castillian Spanish corpus of parkinsonian speech

<p>The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis. This dataset is distinguished as the first of its kind to be made publicly available in Castilian Spanish, addressing a critical gap in the availability of linguistic and dialectical diversity within PD research.</p> <p>Compiled from a cohort of 112 participants, including 54 individuals diagnosed with PD and 58 healthy controls, the NeuroVoz dataset offers a rich compilation of speech recordings. All PD participants were recorded under medication (ON state), ensuring consistency and reliability in the speech samples collected. The dataset is meticulously curated to include a variety of speech tasks&mdash;ranging from sustained vowel phonations and diadochokinetic (DDK) tests to 16 structured listen-and-repeat utterances and spontaneous monologues. The inclusion of both manually transcribed listen-and-repeat tasks and Whisper-automated transcriptions for monologues underscores our commitment to data accuracy and usability.</p> <p>Encompassing 2,977 audio files, the NeuroVoz dataset provides an extensive repository, averaging 26.88 +-&nbsp;3.35 recordings per participant, making it an invaluable asset for researchers seeking to explore the nuances of PD-affected speech. The dataset's structure and composition facilitate a multifaceted analysis of speech impairments associated with PD, offering insights into phonatory, articulatory, and prosodic changes.</p> <p>In contributing to the body of knowledge with the NeuroVoz dataset, we invite the scientific community to engage with this dataset, explore the specific speech characteristics of PD in Castilian Spanish speakers, and advance the field of PD diagnosis through innovative speech analysis techniques.</p> <p>&nbsp;</p> <p>If you use this dataset, please cite both this Zenodo and the article describing the corpus:</p> <ul> <li>Mendes-Laureano, J., G&oacute;mez-Garc&iacute;a, J.A., Guerrero-L&oacute;pez, A.&nbsp;<em>et al.</em>&nbsp;NeuroVoz: a Castillian Spanish corpus of parkinsonian speech.&nbsp;<em>Sci Data</em>&nbsp;<strong>11</strong>, 1367 (2024). https://doi.org/10.1038/s41597-024-04186-z</li> <li>Zenodo dataset: Mendes-Laureano, J., G&oacute;mez-Garc&iacute;a, J. A., Guerrero-L&oacute;pez, A., Luque-Buzo, E., Arias-Londo&ntilde;o, J. D., Grandas-P&eacute;rez, F. J., &amp; Godino Llorente, J. I. (2024). NeuroVoz: a Castillian Spanish corpus of parkinsonian speech (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10777657</li> </ul>

restrictedMar 2024View details →
zenodo16/100

Radboud Maximum Speech Performance Corpus (RaMax)

<p>This corpus contains 78 native Dutch speakers&#39; maximum performance speech material (and accompanying Praat TextGrid files) of&nbsp;two tasks, namely a tongue twister and a Diadochokinesis (DDK) task.&nbsp;</p>

restrictedNov 2021View details →
zenodo16/100

EmoFilm - A multilingual emotional speech corpus

<p><strong>EmoFilm</strong> is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages. The audio clips (with a mean length of 3.5 sec. and std 1.2 sec.) were extracted in wave format (uncompressed, mono, 48 kHz sample rate and 16-bit) from 43 films (original in English and their over-dubbed Italian and Spanish versions). Genres including comedy, drama, horror, and thriller were considered; anger, contempt, happiness, fear, and sadness emotional states were taken into account. EmoFilm has been presented at Interspeech 2018:</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, and Bj&ouml;rn Schuller (2018), <em>Categorical vs Dimensional Perception of Italian Emotional Speech</em>, in Proc. of Interspeech, Hyderabad, India, pp. 3638-3642 .</p> <p>We would like to thank Linda Ratz for her contribution in the generation of the transcriptions.</p> <p>&nbsp;</p> <p><strong>How to access EmoFilm</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1pFHfsqk7snF_EVqq0WAC0Dz8FcTD3s9_/view?usp=share_link</p>

restrictedSep 2018View details →
zenodo12/100

DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception

<p>DEMoS (Database of Elicited Mood in Speech), is a corpus of induced emotional speech in Italian. DEMoS encompasses 9,365 emotional and 332 neutral samples produced by 68 native speakers (23 females, 45 males) in seven emotional states: the 'big six' anger, sadness, happiness, fear, surprise, disgust, and the secondary emotion guilt. To get more realistic productions, instead of acted speech, DEMoS contains emotional speech elicited by combinations of Mood Induction Procedures (MIP). Three elicitation methods are presented, made up by the combination of at least three MIPs, and considering six different MIPs in total. To select samples 'typical' of each emotion, evaluation strategies based on self- and external assessment were applied. The selected part of the corpus encompasses 1,564 prototypical samples produced by 59 speakers (21 females, 38 male). DEMoS has been published in the Journal Language, Resousrces, and Evalaution.</p> <p>&nbsp;</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Maximilian Schmitt, and Bj&ouml;rn Schuller (2019), <em>DEMoS: An Italian emotional speech corpus. Elicitation methods, machine learning, and perception</em>, Language, Resources, and Evaluation, Feb 2019. <a href="http://em.rdcu.be/wf/click?upn=lMZy1lernSJ7apc5DgYM8eCoqdGxOfRWEudjYRrxU-2BI-3D_Ru5N6PJ4ngeR7K-2Fncs2CW1jGAzl4dMvrVh77-2BVH-2B9g5urNss1KItQNXvWL1jiHKvcYDtUVs2c78DX20PMDTauCGehGiQvHdgrAknGggtu7pHINBqVKjp16-2BTn63kNrm22m52e-2FPV-2FidpRe8A-2FplLxPMV-2FjTR-2FLLIK8Wqe7u0-2BLSZ9w-2BWYtrAXRYn2lvPcjGTP1La8yiTxBuJKbHJpnNeFb6LmBIiNMmGRSZPIY0leXhyj4k07rx5cETF6n34aIQHP-2FwcafanNMN4BoA9QKhXGgFxvRgZQidsQ-2BCDbbTBL0PPjM3CgitSGk66qut9E3pd">https://rdcu.be/bn7oI</a></p> <p>&nbsp;</p> <p><strong>How to access DEMoS</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1v6GaCVyNcib5v802t2uXHYOioqIkoBQ-/view?usp=share_link</p>

restrictedFeb 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record