Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Domestic dogs (Canis familiaris) recognise meaningful content in monotonous streams of read speech
Open the record for dataset details and reuse information.
Predictive coding and internal error correction in speech production
Open the record for dataset details and reuse information.
Speech endpoint annotations and artefact details for ASVspoof 2017 version 2.0 dataset
<p>This repository contains speech endpoint annotations and filelists for different artefacts we found during our study on the ASVspoof 2017 v2.0 dataset as part of our work in the paper "Dataset biases in speaker verification systems: a case study on the ASVspoof 2017 benchmark" which is to be submitted to the IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM).</p> <p> </p>
The Grid Audio-Visual Speech Corpus
<p>The Grid Corpus is a large multitalker audiovisual sentence corpus designed to support joint computational-behavioral studies in speech perception. In brief, the corpus consists of high-quality audio and video (facial) recordings of 1000 sentences spoken by each of 34 talkers (18 male, 16 female), for a total of 34000 sentences. Sentences are of the form "put red at G9 now".</p> <p>audio_25k.zip contains the wav format utterances at a 25 kHz sampling rate in a separate directory per talker<br> alignments.zip provides word-level time alignments, again separated by talker<br> s1.zip, s2.zip etc contain .jpg videos for each talker [note that due to an oversight, no video for talker t21 is available]</p> <p>The Grid Corpus is described in detail in the paper jasagrid.pdf included in the dataset.</p>
Convergence in voice fundamental frequency in a joint speech production task - Dataset
<p>This dataset contains fundamental frequency values for 30 pairs of participants performing an alternate reading task.</p> <p>Fundamental frequency in each speaker's speech was artificially modified in real-time during the task. We provide both the untransformed and transformed fundamental frequency values.</p> <p>Full description of the experimental setup is found in < insert paper DOI here ></p> <p>The data is organised as follows:</p> <ul> <li>data for each pair is stored in a separate folder with the pair ID as the folder name</li> <li>each folder contains two repetitions of the task, as produced in a zero-phase and pi-phase condition, respectively</li> <li>file names ending with '<strong>f0</strong>' contain fundamental frequency data, sampled every 10 ms</li> <li>file names ending with '<strong>turns</strong>' contain time onsets of speaking turns</li> </ul> <p>The format of '<strong>f0</strong>' files is as follows:</p> <ul> <li>'<strong>t</strong>': time in seconds</li> <li>'<strong>ch</strong>': channel of the recording, indicating the participant (<em>A</em> or <em>B</em>)</li> <li>'<strong>f0_unstransf</strong>': fundamental frequency values as produced by the participant (untransformed) in Hertz</li> <li>'<strong>f0_transf</strong>': fundamental frequency values as heard by the other participant (transformed) in Hertz</li> </ul> <p>The format of '<strong>turns</strong>' files is as follows:</p> <ul> <li>'<strong>ch</strong>': channel of the recording, indicating the participant (<em>A</em> or <em>B</em>), or both participants at once (<em>joint</em>)</li> <li>'<strong>turn</strong>': index of the reading turn</li> <li>'<strong>type</strong>': turn type. Either <em>speech</em> or <em>silence</em> for each participant, or <em>turn</em> for the joint description.</li> <li>'<strong>t</strong>': turn onset in seconds</li> </ul> <p> </p>
League of Legends and hate speech: a corpus for comments in Twitch.tv
<p>League of Legends (LOL) is the most popular game on PC, drawing 8 million concurrent players. A common activity of gamers, besides playing games, is to watch other players presenting tips and tricks. Streaming platforms allow some players to show gameplays and live games. <a href="https://www.twitch.tv/">Twitch.tv</a> is the world´s leading live streaming platform. </p> <p>Considering that hate speech is a ubiquitous problem in online gaming, we collected 985,766 comments from five videos of the top 10 LOL streamers in Twitch.tv platform. </p> <p>The dataset is freely available in a single file, ensembling all videos/players; and divided by players as well. </p> <p>These comments are a rich data source for opinion mining, sentiment analysis, topic modeling, and hate speech detection (including sexism and racism).</p> <ul> </ul>
Sensitivity of occipito-temporal cortex, premotor and Broca's areas to visible speech gestures in a familiar language
<p>fMRI dataset related to the study "Sensitivity of occipito-temporal cortex, premotor and Broca’s areas to visible speech gestures in a familiar language". The dataset includes SPM analyses for the contrasts of interest described in the paper.</p>
Speech Quality Apollo Corpus
<p>32 audio clips extracted from the Apollo Space Program archive recordings annotated with subjective intelligibility and non-intrusive objective metrics.</p> <p><strong>Please cite if you use this dataset:</strong></p> <p>A. Ragano, E. Benetos, and A. Hines, "Development of a speech quality database under uncontrolled conditions", in Proc. 21st Annual Conference of the International Speech Communication Association (INTERSPEECH), pp. 4616-4620, Oct. 2020</p> <p>Paper available here https://www.isca-speech.org/archive/Interspeech_2020/pdfs/1899.pdf</p> <p><strong>Annotations:</strong></p> <ul> <li><strong>google_wer: </strong>word-error-rate of Google speech-to-text API</li> <li><strong>mosnet:</strong> objective quality score of MOSNet metric</li> <li><strong>srmr: </strong>objective quality score of SRMR metric</li> <li><strong>p563: </strong>objective quality score of ITU-T P.563</li> <li><strong>subjwer</strong>: mean word-error-rate of the study participants </li> </ul> <p>See INTERSPEECH manuscript for more information</p>
CUB-200 Speech captions (Part-I)
<p>Speech captions of CUB-200 database. These speech captions are synthesized by Tacotron-v2 according to the original textual descriptions.</p>
Data from: Severe childhood speech disorder: Gene discovery highlights transcriptional dysregulation
<p><b><span>Objective:</span></b><span> Determining the genetic basis of speech disorders provides insight into the neurobiology of human communication. Despite intensive investigation over the past two decades, the etiology of most children with speech disorder remains unexplained. Here we </span>searched for a genetic etiology in children with severe speech disorder, specifically childhood apraxia of speech (CAS).</p> <p><b>Methods:</b> Precise phenotyping together with research genome or exome analysis were performed on children referred with a primary diagnosis of CAS, as well as other medical or neurodevelopmental co-morbidities. Gene co-expression and gene set enrichment analyses analyses were conducted on high confidence gene candidates.</p> <p><b>Results:</b> 34 probands ascertained for CAS were studied. In 11/34 (32%) probands, we identified highly plausible pathogenic single nucleotide (n=10, <i>CDK13, EBF3, GNAO1, GNB1, DDX3X, MEIS2, POGZ, SETBP1, UPF2, ZNF142</i>) or copy number (n = 1, 5q14.3q21.1 locus) variants in novel genes or loci for CAS. Testing of parental DNA was available for nine probands and confirmed that the variants had arisen <i>de novo</i>. Eight genes encode proteins critical for regulation of gene transcription, and analyses of transcriptomic data found CAS-implicated genes were highly co-expressed in the developing human brain.</p> <p><b>Conclusion:</b><i> </i>We identify the likely genetic aetiology in 11 patients with CAS and implicate 9 genes for the first time. We find that CAS is often a sporadic monogenic disorder, and highly genetically heterogeneous. Highly penetrant variants implicate shared pathways in broad transcriptional regulation, highlighting the key role of transcriptional regulation in normal speech development. CAS is a distinctive, socially debilitating clinical disorder, and understanding its molecular basis is the first step towards identifying precision medicine approaches.</p>
WinoST: Evaluating Gender Bias in Speech Translation
<p>WinoST is a challenge set for evaluating gender bias in speech translation. WinoST is the speech version of WinoMT (Stanovsky et al., 2019) which is an MT challenge set and both follow an evaluation protocol to measure gender accuracy. WinoST consists of 3888 speech audios in English plus a text file with the the text of these audios. For further details, please refer to the publication entitled "Evaluating Gender Bias in Speech Translation" (Costa-jussà et al., 2020). Also cite this publication if using this corpus.</p>
Test dataset for separation of speech, traffic sounds, wind noise, and general sounds
<p>The dataset was generated as part of the paper:<br> Deep Complex U-Net Ensemble for Outdoor Urban Sound Source Separation,<br> K. Arendt, A. Szumaczuk, B. Jasik, P. Masztalski, K. Piaskowski, M. Matuszewski, K. Nowicki, P. Zborowski.</p> <p>It contains various sounds from the Audio Set [1] and spoken utterances from VCTK [2] and DNS [3] datasets.</p> <p>Contents:<br> sr_8k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_16k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_48k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/</p> <p>Each directory contains 512 audio samples in different sampling rate (sr_8k - 8 kHz, sr_16k - 16 kHz, sr_48k - 48 kHz).<br> The audio samples for each sampling rate are different as they were generated randomly and separately.<br> Each directory contains 5 subdirectories:<br> - mix_clean - mixed sources,<br> - s1 - source #1 (general sounds),<br> - s2 - source #2 (speech),<br> - s3 - source #3 (traffic sounds),<br> - s4 - source #4 (wind noise).</p> <p>The sound mixtures were generated by adding s2, s3, s4 to s1 with SNR ranging from -10 to 10 dB w.r.t. s1.</p> <p><br> REFERENCES:</p> <p>[1] Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman,<br> Aren Jansen, Wade Lawrence, R. Channing Moore,<br> Manoj Plakal, and Marvin Ritter, “Audio set: An ontology<br> and human-labeled dataset for audio events,” in<br> Proc. IEEE ICASSP 2017, New Orleans, LA, 2017.</p> <p>[2] Christophe Veaux, Junichi Yamagishi, and Kirsten Mac-<br> Donald, “CSTR VCTK corpus: English multi-speaker<br> corpus for CSTR voice cloning toolkit, [sound],”<br> https://doi.org/10.7488/ds/1994, University of Edinburgh.<br> The Centre for Speech Technology Research<br> (CSTR). 2017.</p> <p>[3] Chandan K. A. Reddy, Ebrahim Beyrami, Harishchandra<br> Dubey, Vishak Gopal, Roger Cheng, Ross Cutler,<br> Sergiy Matusevych, Robert Aichner, Ashkan Aazami,<br> Sebastian Braun, Puneet Rana, Sriram Srinivasan, and<br> Johannes Gehrke, “The interspeech 2020 deep noise<br> suppression challenge: Datasets, subjective speech<br> quality and testing framework,” 2020.</p>
Integration of speech separation, diarization, and recognition for multi-speaker meetings: Separated LibriCSS dataset
<p><strong>Dataset</strong></p> <p>This data repository contains separated audio streams for the LibriCSS dataset using the following window-based separation methods:</p> <p>1. <em>Mask-based MVDR</em>: Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, “Multi-microphone neural speech separation for farfield multi-talker speech recognition,” ICASSP 2018.</p> <p>2. <em>Sequential neural beamforming</em>: Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, and John R. Hershey, “Sequential multi-frame neural beamforming for speech separation and enhancement,” IEEE SLT 2021.</p> <p>These audio streams were used for evaluating the diarization and ASR models in our <a href="https://arxiv.org/pdf/2011.02014.pdf">JSALT 2020 paper</a>.</p> <p>The repository contains the following archive files:</p> <ul> <li>libricss_mvdr_2stream.tar.gz</li> <li>libricss_sequential_3stream.tar.gz</li> </ul> <p><strong>Citation</strong></p> <p>If you use these separated audio streams in your research, consider citing:</p> <pre><code>@article{Raj2020IntegrationOS, title={Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis}, author={Desh Raj and Pavel Denisov and Z. Chen and H. Erdogan and Zili Huang and Mao-Kui He and Shinji Watanabe and Jun Du and T. Yoshioka and Yi Luo and N. Kanda and Jinyu Li and S. Wisdom and J. Hershey}, journal={2021 IEEE Spoken Language Technology (SLT) Workshop}, year={2021} }</code></pre> <p><br> </p>
Developing an English course for Beginners with the topic Part of Speech using Google Classroom
<p>Self-paced learning means that you can study on your own time and schedule. You don't have to complete the same assignment or study at the same time as others. You can proceed from one topic or segment to the next at your own pace. Why we chose Google Classroom Media is because it's so easy to use. For people who are using Google classrooms for the first time, they will definitely have no trouble operating them.</p>
Data from: Quantification of motor speech impairment and its anatomic basis in primary progressive aphasia
Objective: To evaluate whether a quantitative speech measure is effective in identifying and monitoring motor speech impairment (MSI) in patients with primary progressive aphasia (PPA), and to investigate the neuroanatomical basis of MSI in PPA. Methods: Sixty-four patients with PPA were evaluated at baseline, with a subset (N=39) evaluated longitudinally. Articulation rate (AR), a quantitative measure derived from spontaneous speech, was measured at each timepoint. MRI was collected at baseline. Differences in baseline AR were assessed across PPA subtypes, separated by severity level. Linear mixed-effects models were conducted to assess groups differences across PPA subtypes in rate of decline in AR over a one-year period. Cortical thickness measured from baseline MRIs was used to test hypotheses about the relationship between cortical atrophy and MSI. Results: Baseline AR was reduced for patients with non-fluent variant PPA (nfvPPA), as compared to other PPA subtypes and controls, even in mild stages of disease. Longitudinal results showed a greater rate of decline in AR for the nfvPPA group over one year, as compared to logopenic and semantic variant subgroups. Reduced baseline AR was associated with cortical atrophy in left-hemisphere premotor and supplementary motor cortices. Conclusions: The AR measure is an effective quantitative index of MSI that detects MSI in mild disease stages and tracks decline in MSI longitudinally. The AR measure additionally demonstrates anatomic localization to motor-speech specific cortical regions. Our findings suggest that this quantitative measure of MSI might have utility in diagnostic evaluation and monitoring of motor speech impairments in PPA.
The EEG and fMRI signatures of neural integration: An investigation of meaningful gestures and corresponding speech
<p>One of the key features of human interpersonal communication is our ability to integrate information communicated by speech and accompanying gestures. However, it is still not fully understood how this essential combinatory process is represented in the human brain. Functional magnetic resonance imaging (fMRI) studies have unanimously attested the relevance of activation in the posterior superior temporal sulcus/middle temporal gyrus (pSTS/MTG), while electroencephalography (EEG) studies have shown oscillatory activity in specific frequency bands to be associated with multisensory integration. In the current study, we used fMRI and EEG to separately investigate the anatomical and oscillatory neural signature of integrating intrinsically meaningful gestures (IMG; e.g. “Thumbs-up gesture”) and corresponding speech (e.g., “The actor did a good job”). In both the fMRI (<em>n</em>=20) and EEG (<em>n</em>=20) study, participants were presented with videos of an actor either: performing IMG in the context of a German sentence (GG), IMG in the context of a Russian (as a foreign language) sentence (GR), or speaking an isolated German sentence without gesture (SG). The results of the fMRI experiment confirmed that gesture–speech processing of IMG activates the posterior MTG (GG>GR∩GG>SG). In the EEG experiment we found that the identical integration process (GG>GR∩GG>SG) is related to a centrally-distributed alpha (7–13 Hz) power decrease within 700–1400 ms post-onset of the critical word. These new findings suggest that BOLD response increase in the pMTG and alpha power decrease represent the neural correlates of integrating intrinsically meaningful gestures with their corresponding speech.</p>
Raw data of 'He, Y., Gebhardt, H., Steines, M., Sammer, G., Kircher, T., Nagels, A., & Straube, B. (2015). The EEG and fMRI signatures of neural integration: An investigation of meaningful gestures and corresponding speech. Neuropsychologia, 72, 27-42.'
<p>This is the raw data for</p> <p>He, Y., Gebhardt, H., Steines, M., Sammer, G., Kircher, T., Nagels, A., & Straube, B. (2015). The EEG and fMRI signatures of neural integration: An investigation of meaningful gestures and corresponding speech. Neuropsychologia, 72, 27-42.</p> <p>Both second level fMRI data and EEG data for single subjects are provided.</p> <p></p>
The influence of infant-directed speech on 12-month-olds' intersensory perception of fluent speech
<p>The present study examined whether infant-directed (ID) speech facilitates intersensory matching of audio–visual fluent speech in 12-month-old infants. German-learning infants’ audio–visual matching ability of German and French fluent speech was assessed by using a variant of the intermodal matching procedure, with auditory and visual speech information presented sequentially. In Experiment 1, the sentences were spoken in an adult-directed (AD) manner. Results showed that 12-month-old infants did not exhibit a matching performance for the native, nor for the non-native language. However, Experiment 2 revealed that when ID speech stimuli were used, infants did perceive the relation between auditory and visual speech attributes, but only in response to their native language. Thus, the findings suggest that ID speech might have an influence on the intersensory perception of fluent speech and shed further light on multisensory perceptual narrowing.</p>
Cross-modal matching of audio-visual German and French fluent speech in infancy
<p>The present study examined when and how the ability to cross-modally match audio-visual fluent speech develops in 4.5-, 6- and 12-month-old German-learning infants. In Experiment 1, 4.5- and 6-month-old infants’ audio-visual matching ability of native (German) and non-native (French) fluent speech was assessed by presenting auditory and visual speech information sequentially, that is, in the absence of temporal synchrony cues. The results showed that 4.5-month-old infants were capable of matching native as well as non-native audio and visual speech stimuli, whereas 6-month-olds perceived the audio-visual correspondence of native language stimuli only. This suggests that intersensory matching narrows for fluent speech between 4.5 and 6 months of age. In Experiment 2, auditory and visual speech information was presented simultaneously, therefore, providing temporal synchrony cues. Here, 6-month-olds were found to match native as well as non-native speech indicating facilitation of temporal synchrony cues on the intersensory perception of non-native fluent speech. Intriguingly, despite the fact that audio and visual stimuli cohered temporally, 12-month-olds matched the non-native language only. Results were discussed with regard to multisensory perceptual narrowing during the first year of life.</p>
Multilingual bottle-neck feature learning from untranscribed speech for track 1 in zerospeech2017 (system 2 -- with VTLN)
<p>We investigate the extraction of bottle-neck features (BNFs) for multiple languages without access to manual transcription. Multilingual BNFs are derived from a multi-task learning deep neural network which is trained with unsupervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models separately trained on untranscribed speech of multiple languages.</p> <blockquote> <p>In this version, the input MFCC for DPGMM is processed with VTLN.</p> </blockquote> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.