Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
94
datasets available to search
ShareScore release 0.7.1
Dataset results
94 results for “speech dataset”
ODSS: An Open Dataset of Synthetic Speech
<p>ODSS is a multilingual, multispeaker dataset of synthetic and natural speech, designed to foster research and benchmarking of novel studies on synthetic speech detection. </p> <p>ODSS comprises audio utterances generated from text by state-of-the-art synthesis methods, paired with their corresponding natural counterparts. The synthetic audio data includes several languages, with an equal representation of genders.</p> <p>Natural and synthetic speech audio files within ODSS are released under the CC-BY-SA 4.0 license: Usage, extension and redistribution by the research community are strongly encouraged.</p>
An fNIRS dataset for multimodal speech comprehension in normal hearing individuals and cochlear implant users
Open the record for dataset details and reuse information.
Speech endpoint annotations and artefact details for ASVspoof 2017 version 2.0 dataset
<p>This repository contains speech endpoint annotations and filelists for different artefacts we found during our study on the ASVspoof 2017 v2.0 dataset as part of our work in the paper "Dataset biases in speaker verification systems: a case study on the ASVspoof 2017 benchmark" which is to be submitted to the IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM).</p> <p> </p>
Convergence in voice fundamental frequency in a joint speech production task - Dataset
<p>This dataset contains fundamental frequency values for 30 pairs of participants performing an alternate reading task.</p> <p>Fundamental frequency in each speaker's speech was artificially modified in real-time during the task. We provide both the untransformed and transformed fundamental frequency values.</p> <p>Full description of the experimental setup is found in < insert paper DOI here ></p> <p>The data is organised as follows:</p> <ul> <li>data for each pair is stored in a separate folder with the pair ID as the folder name</li> <li>each folder contains two repetitions of the task, as produced in a zero-phase and pi-phase condition, respectively</li> <li>file names ending with '<strong>f0</strong>' contain fundamental frequency data, sampled every 10 ms</li> <li>file names ending with '<strong>turns</strong>' contain time onsets of speaking turns</li> </ul> <p>The format of '<strong>f0</strong>' files is as follows:</p> <ul> <li>'<strong>t</strong>': time in seconds</li> <li>'<strong>ch</strong>': channel of the recording, indicating the participant (<em>A</em> or <em>B</em>)</li> <li>'<strong>f0_unstransf</strong>': fundamental frequency values as produced by the participant (untransformed) in Hertz</li> <li>'<strong>f0_transf</strong>': fundamental frequency values as heard by the other participant (transformed) in Hertz</li> </ul> <p>The format of '<strong>turns</strong>' files is as follows:</p> <ul> <li>'<strong>ch</strong>': channel of the recording, indicating the participant (<em>A</em> or <em>B</em>), or both participants at once (<em>joint</em>)</li> <li>'<strong>turn</strong>': index of the reading turn</li> <li>'<strong>type</strong>': turn type. Either <em>speech</em> or <em>silence</em> for each participant, or <em>turn</em> for the joint description.</li> <li>'<strong>t</strong>': turn onset in seconds</li> </ul> <p> </p>
Test dataset for separation of speech, traffic sounds, wind noise, and general sounds
<p>The dataset was generated as part of the paper:<br> Deep Complex U-Net Ensemble for Outdoor Urban Sound Source Separation,<br> K. Arendt, A. Szumaczuk, B. Jasik, P. Masztalski, K. Piaskowski, M. Matuszewski, K. Nowicki, P. Zborowski.</p> <p>It contains various sounds from the Audio Set [1] and spoken utterances from VCTK [2] and DNS [3] datasets.</p> <p>Contents:<br> sr_8k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_16k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/<br> sr_48k/<br> mix_clean/<br> s1/<br> s2/<br> s3/<br> s4/</p> <p>Each directory contains 512 audio samples in different sampling rate (sr_8k - 8 kHz, sr_16k - 16 kHz, sr_48k - 48 kHz).<br> The audio samples for each sampling rate are different as they were generated randomly and separately.<br> Each directory contains 5 subdirectories:<br> - mix_clean - mixed sources,<br> - s1 - source #1 (general sounds),<br> - s2 - source #2 (speech),<br> - s3 - source #3 (traffic sounds),<br> - s4 - source #4 (wind noise).</p> <p>The sound mixtures were generated by adding s2, s3, s4 to s1 with SNR ranging from -10 to 10 dB w.r.t. s1.</p> <p><br> REFERENCES:</p> <p>[1] Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman,<br> Aren Jansen, Wade Lawrence, R. Channing Moore,<br> Manoj Plakal, and Marvin Ritter, “Audio set: An ontology<br> and human-labeled dataset for audio events,” in<br> Proc. IEEE ICASSP 2017, New Orleans, LA, 2017.</p> <p>[2] Christophe Veaux, Junichi Yamagishi, and Kirsten Mac-<br> Donald, “CSTR VCTK corpus: English multi-speaker<br> corpus for CSTR voice cloning toolkit, [sound],”<br> https://doi.org/10.7488/ds/1994, University of Edinburgh.<br> The Centre for Speech Technology Research<br> (CSTR). 2017.</p> <p>[3] Chandan K. A. Reddy, Ebrahim Beyrami, Harishchandra<br> Dubey, Vishak Gopal, Roger Cheng, Ross Cutler,<br> Sergiy Matusevych, Robert Aichner, Ashkan Aazami,<br> Sebastian Braun, Puneet Rana, Sriram Srinivasan, and<br> Johannes Gehrke, “The interspeech 2020 deep noise<br> suppression challenge: Datasets, subjective speech<br> quality and testing framework,” 2020.</p>
Integration of speech separation, diarization, and recognition for multi-speaker meetings: Separated LibriCSS dataset
<p><strong>Dataset</strong></p> <p>This data repository contains separated audio streams for the LibriCSS dataset using the following window-based separation methods:</p> <p>1. <em>Mask-based MVDR</em>: Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, “Multi-microphone neural speech separation for farfield multi-talker speech recognition,” ICASSP 2018.</p> <p>2. <em>Sequential neural beamforming</em>: Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, and John R. Hershey, “Sequential multi-frame neural beamforming for speech separation and enhancement,” IEEE SLT 2021.</p> <p>These audio streams were used for evaluating the diarization and ASR models in our <a href="https://arxiv.org/pdf/2011.02014.pdf">JSALT 2020 paper</a>.</p> <p>The repository contains the following archive files:</p> <ul> <li>libricss_mvdr_2stream.tar.gz</li> <li>libricss_sequential_3stream.tar.gz</li> </ul> <p><strong>Citation</strong></p> <p>If you use these separated audio streams in your research, consider citing:</p> <pre><code>@article{Raj2020IntegrationOS, title={Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis}, author={Desh Raj and Pavel Denisov and Z. Chen and H. Erdogan and Zili Huang and Mao-Kui He and Shinji Watanabe and Jun Du and T. Yoshioka and Yi Luo and N. Kanda and Jinyu Li and S. Wisdom and J. Hershey}, journal={2021 IEEE Spoken Language Technology (SLT) Workshop}, year={2021} }</code></pre> <p><br> </p>
THLS - An open source dataset for Brazilian Portuguese speech processing
<p>THLS Open Source Brazilian Portuguese Speech Dataset with 1000 sentences balanced phonetically.</p> <p> </p> <p>Authors:</p> <ul> <li>Luiz Felipe Vecchietti</li> <li>Thalles Melo Batista Pieroni (Voice)</li> </ul> <p> </p> <p>https://gitlab.com/lfelipesv/1000-sentences-thls-dataset</p>
Dataset Labeled For Inciting Speech
<p>This dataset is related to the paper "Understanding Inciting Speech As New Malice." The paper recently got accepted at IEEE Transactions on Computational Social Systems. Please cite the paper while using the dataset. </p>
ESCorpus-PE: A speech emotional dataset in Spanish with Peruvian accent
<p>ESCorpus-PE dataset contains emotional utterances of Spanish peruvian speech gathered from Spanish interviews, TV reports, political debate and testimonials. It contains 3749 utterances of three emotional dimensions: Valence, Arousal and Dominance. There are 80 speakers (44 male and 36 female). This data was created from Youtube audios. These audios were selected following a specific criteria specified in the paper: ESCorpus-PE: A speech emotional database in Spanish with Peruvian accent, in the section Methods/Audio/Video Selection. Anyone can use this data only for research purposes.</p> <p>More details on <br> https://github.com/Alessandra-UNSA/Peruvian_Spanish_Corpus</p>
Dataset and tools for the PSST Challenge on Post-Stroke Speech Transcription
<p>Initial Release, for archival/DOI purposes.</p>
A Large TV Dataset for Speech and Music Activity Detection
<p>Automatic speech and music activity detection (SMAD) is an enabling task that can help segment, index, and pre-process audio content in radio broadcast and TV programs. However, due to copyright concerns and the cost of manual annotation, the limited availability of diverse and sizeable datasets hinders the progress of state-of-the-art (SOTA) data-driven approaches. We address this challenge by presenting a large-scale dataset containing Mel spectrogram, VGGish, and MFCCs features extracted from around 1600 hours of professionally produced audio tracks and their corresponding noisy labels indicating the approximate location of speech and music segments. The labels are derived from several sources such as subtitles. A test set curated by human annotators is also included as a subset for evaluation. To the best of our knowledge, this dataset is the first large-scale, open-sourced dataset that contains features extracted from professionally produced audio tracks and their corresponding frame-level speech and music annotations. </p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 1)
<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 2)
<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 2
<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data. </p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels. </p> <p><strong>Arxiv:</strong> https://arxiv.org/abs/2109.07931</p> <p> </p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 2).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong> https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong> https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong> https://zenodo.org/record/5501697</li> </ul>
DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 1
<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data. </p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels. </p> <p><strong>Arxiv:</strong> https://arxiv.org/abs/2109.07931</p> <p> </p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 1).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong> https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong> https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong> https://zenodo.org/record/5501697</li> </ul>
EUParlspeech: A Dataset of Over 1 Million References to European Integration in Parliamentary Speeches
<p>As part of my PhD dissertation I developed EUParlspeech, a dataset of over 1 million references to European integration made in the plenary debates of ten national parliaments between 1989 and 2019. It is built from existing datasets of parliamentary speeches, most notably Parlspeech (Rauh and Schwalbach 2020). The dataset has applications for scholars of EU integration, party competition, political communication, and international relations. This chapter in my dissertation explains the construction of the dataset, describes its features, and demonstrates its face, convergent, and predictive validity. Automated analysis of parties' EU statements in parliament yield meaningful and well-known cross party differences, with challenger parties more likely to send clearer, more sceptical cues on integration than mainstream parties. Moreover, these automated measures correlate highly with expert assessments (CHES) and - in the case of the UK's Conservative Party - individual MPs' ideal point estimates based on EU statements in plenary debates can predict their subsequent vote and position at the 2016 referendum. I conclude that EUParlspeech data provide a promising new approach to studying party contestation over European integration.</p>
Youtube-Dataset for Language Identification in Speech Signals
<p><strong>Youtube-Dataset for Language Identification in Speech Signals</strong></p> <p>- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de</p> <p><strong>Reference</strong></p> <p>In case you use this dataset for your research, please cite</p> <p>Alexandra Draghici, Jakob Abeßer & Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020</p> <p><strong>Dataset</strong></p> <p>The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo.</p> <p>- 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders "0" - "5" encode the language id:<br> 0 - English<br> 1 - French<br> 2 - German<br> 3 - Greek<br> 4 - Italian<br> 5 - Spanish</p> <p><strong>Audio Processing</strong></p> <p>- mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br> </p>
Fluent speech commands dataset
Open the record for dataset details and reuse information.
Fourteen-channel EEG with Imagined Speech (FEIS) dataset
<pre>><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> Welcome to the FEIS (Fourteen-channel EEG with Imagined Speech) dataset. <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< The FEIS dataset comprises Emotiv EPOC+ [1] EEG recordings of: * 21 participants listening to, imagining speaking, and then actually speaking 16 English phonemes (see supplementary, below) * 2 participants listening to, imagining speaking, and then actually speaking 16 Chinese syllables (see supplementary, below) For replicability and for the benefit of further research, this dataset includes the complete experiment set-up, including participants' recorded audio and 'flashcard' screens for audio-visual prompts, Lua script and .mxs scenario for the OpenVibe [2] environment, as well as all Python scripts for the preparation and processing of data as used in the supporting studies (submitted in support of completion of the MSc Speech and Language Processing with the University of Edinburgh): * J. Clayton, "Towards phone classification from imagined speech using a lightweight EEG brain-computer interface," M.Sc. dissertation, University of Edinburgh, Edinburgh, UK, 2019. * S. Wellington, "An investigation into the possibilities and limitations of decoding heard, imagined and spoken phonemes using a low-density, mobile EEG headset," M.Sc. dissertation, University of Edinburgh, Edinburgh, UK, 2019. Each participant's data comprise 5 .csv files -- these are the 'raw' (unprocessed) EEG recordings for the 'stimuli', 'articulators' (see supplementary, below) 'thinking', 'speaking' and 'resting' phases per epoch for each trial -- alongside a 'full' .csv file with the end-to-end experiment recording (for the benefit of calculating deltas). To guard against software deprecation or inaccessability, the full repository of open-source software used in the above studies is also included. We hope for the FEIS dataset to be of some utility for future researchers, due to the sparsity of similar open-access databases. As such, this dataset is made freely available for all academic and research purposes (non-profit). ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> REFERENCING <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< If you use the FEIS dataset, please reference: * S. Wellington, J. Clayton, "Fourteen-channel EEG with Imagined Speech (FEIS) dataset," v1.0, University of Edinburgh, Edinburgh, UK, 2019. doi:10.5281/zenodo.3369178 ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> LEGAL <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< The research supporting the distribution of this dataset has been approved by the PPLS Research Ethics Committee, School of Philosophy, Psychology and Language Sciences, University of Edinburgh (reference number: 435-1819/2). This dataset is made available under the Open Data Commons Attribution License (ODC-BY): <a href="http://opendatacommons.org/licenses/by/1.0">http://opendatacommons.org/licenses/by/1.0</a> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ACKNOWLEDGEMENTS <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< The FEIS database was compiled by: Scott Wellington (MSc Speech and Language Processing, University of Edinburgh) Jonathan Clayton (MSc Speech and Language Processing, University of Edinburgh) Principal Investigators: Oliver Watts (Senior Researcher, CSTR, University of Edinburgh) Cassia Valentini-Botinhao (Senior Researcher, CSTR, University of Edinburgh) <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< METADATA ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> For participants, dataset refs 01 to 21: 01 - NNS 02 - NNS 03 - NNS, Left-handed 04 - E 05 - E, Voice heard as part of 'stimuli' portions of trials belongs to particpant 04, due to microphone becoming damaged and unusable prior to recording 06 - E 07 - E 08 - E, Ambidextrous 09 - NNS, Left-handed 10 - E 11 - NNS 12 - NNS, Only sessions one and two recorded (out of three total), as particpant had to leave the recording session early 13 - E 14 - NNS 15 - NNS 16 - NNS 17 - E 18 - NNS 19 - E 20 - E 21 - E E = native speaker of English NNS = non-native speaker of English (>= C1 level) For participants, dataset refs chinese-1 and chinese-2: chinese-1 - C chinese-2 - C, Voice heard as part of 'stimuli' portions of trials belongs to participant chinese-1 C = native speaker of Chinese <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< SUPPLEMENTARY ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> Under the international 10-20 system, the Emotiv EPOC+ headset 14 channels: F3 FC5 AF3 F7 T7 P7 O1 O2 P8 T8 F8 AF4 FC6 F4 The 16 English phonemes investigated in dataset refs 01 to 21: /i/ /u:/ /æ/ /ɔ:/ /m/ /n/ /ŋ/ /f/ /s/ /ʃ/ /v/ /z/ /ʒ/ /p /t/ /k/ The 16 Chinese syllables investigated in dataset refs chinese-1 and chinese-2: mā má mǎ mà mēng méng měng mèng duō duó duǒ duò tuī tuí tuǐ tuì All references to 'articulators' (e.g. as part of filenames) refer to the 1-second 'fixation point' portion of trials. The name is a layover from preliminary trials which were modelled on the KARA ONE database (<a href="http://www.cs.toronto.edu/~complingweb/data/karaOne/karaOne.html">http://www.cs.toronto.edu/~complingweb/data/karaOne/karaOne.html</a>) [3]. <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< <>< ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> ><> [1] Emotiv EPOC+. <a href="https://emotiv.com/epoc">https://emotiv.com/epoc</a>. Accessed online 14/08/2019. [2] Y. Renard, F. Lotte, G. Gibert, M. Congedo, E. Maby, V. Delannoy, O. Bertrand, A. Lécuyer. “OpenViBE: An Open-Source Software Platform to Design, Test and Use Brain-Computer Interfaces in Real and Virtual Environments”, Presence: teleoperators and virtual environments, vol. 19, no 1, 2010. [3] S. Zhao, F. Rudzicz. "Classifying phonological categories in imagined and articulated speech." In Proceedings of ICASSP 2015, Brisbane Australia, 2015.</pre>
Dataset for "Capturing Formality in Speech Across Domains and Languages"
<p>We share the data used in our paper "Capturing Formality in Speech Across Domains and Languages", previously hosted on Google Drive. Corpora available in this dataset release include:</p> <ol> <li>All-India Radio (Hindi) [<a href="../api/records/13298510/draft/files/all_india_radio-20240812T152206Z-001.zip/content" target="_blank" rel="noopener noreferrer">all_india_radio-20240812T152206Z-001.zip</a>]</li> <li>Bangor Miami (Spanish-English) [<a href="../api/records/13298510/draft/files/bangor_miami_clean-20240812T152204Z-001.zip/content" target="_blank" rel="noopener noreferrer">bangor_miami_clean-20240812T152204Z-001.zip</a>]</li> <li>CallHome (English; Spanish) [<a href="../api/records/13298510/draft/files/callhome-20240812T152201Z-001.zip/content" target="_blank" rel="noopener noreferrer">callhome-20240812T152201Z-001.zip</a>; <a href="../api/records/13298510/draft/files/callhome-20240812T152201Z-002.zip/content" target="_blank" rel="noopener noreferrer">callhome-20240812T152201Z-002.zip</a>]</li> <li>CALLFriend (Hindi) [<a href="../api/records/13298510/draft/files/cf_hindi-20240812T152159Z-001.zip/content" target="_blank" rel="noopener noreferrer">cf_hindi-20240812T152159Z-001.zip</a>]</li> <li>HUB4-SE (Spanish) [<a href="../api/records/13298510/draft/files/hub4_se-20240812T152156Z-001.zip/content" target="_blank" rel="noopener noreferrer">hub4_se-20240812T152156Z-001.zip</a>]</li> <li>HUB5 (Mandarin) [<a href="../api/records/13298510/draft/files/hub5_transcript-20240812T152154Z-001.zip/content" target="_blank" rel="noopener noreferrer">hub5_transcript-20240812T152154Z-001.zip</a>]</li> <li>Multilingual TEDx (English; Spanish) [<a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-001.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-001.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-002.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-002.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-003.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-003.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-004.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-004.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-005.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-005.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-006.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-006.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-007.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-007.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-008.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-008.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-009.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-009.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-010.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-010.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-011.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-011.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-012.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-012.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-013.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-013.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-014.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-014.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-015.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-015.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-016.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-016.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-017.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-017.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-018.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-018.zip</a>, <a href="../api/records/13298510/draft/files/mtedx_es-en-20240812T152152Z-019.zip/content" target="_blank" rel="noopener noreferrer">mtedx_es-en-20240812T152152Z-019.zip</a>]</li> <li>Multitarget TED (English; Mandarin) [<a href="../api/records/13298510/draft/files/multitarget-ted-20240812T152149Z-001.zip/content" target="_blank" rel="noopener noreferrer">multitarget-ted-20240812T152149Z-001.zip</a>]</li> <li>IIT-B (Hindi) [<a href="../api/records/13298510/draft/files/parallel-n-20240812T042826Z-001.zip/content" target="_blank" rel="noopener noreferrer">parallel-n-20240812T042826Z-001.zip</a>]</li> <li>TDT4 (English; Mandarin) [<a href="../api/records/13298510/draft/files/tdt4_multilingual_news-20240812T152131Z-001.zip/content" target="_blank" rel="noopener noreferrer">tdt4_multilingual_news-20240812T152131Z-001.zip</a>]</li> <li>TED Talks India (Hindi) [<a href="../api/records/13298510/draft/files/ted_talks_hindi-20240812T042346Z-001.zip/content" target="_blank" rel="noopener noreferrer">ted_talks_hindi-20240812T042346Z-001.zip</a>]</li> <li>UN (Mandarin) [<a href="../api/records/13298510/draft/files/UNv1.0.en-zh-002.en.zip/content" target="_blank" rel="noopener noreferrer">UNv1.0.en-zh-002.en.zip</a>; <a href="../api/records/13298510/draft/files/UN-20240812T040956Z-002.zip/content" target="_blank" rel="noopener noreferrer">UN-20240812T040956Z-002.zip</a>; <a href="../api/records/13298510/draft/files/UN-20240812T040956Z-003.zip/content" target="_blank" rel="noopener noreferrer">UN-20240812T040956Z-003.zip</a>]</li> <li>YouTube (English; Spanish; Hindi; Mandarin) [<a href="../api/records/13298510/draft/files/youtube-20240812T040730Z-001.zip/content" target="_blank" rel="noopener noreferrer">youtube-20240812T040730Z-001.zip</a>; <a href="../api/records/13298510/draft/files/youtube-20240812T040730Z-002.zip/content" target="_blank" rel="noopener noreferrer">youtube-20240812T040730Z-002.zip</a>; <a href="../api/records/13298510/draft/files/youtube-20240812T040730Z-003.zip/content" target="_blank" rel="noopener noreferrer">youtube-20240812T040730Z-003.zip</a>]</li> <li>All-CS (Hindi-English) [<a href="../api/records/13298510/draft/files/All-CS.json/content" target="_blank" rel="noopener noreferrer">All-CS.json</a>]</li> <li>Europarl v7 (Spanish) [<a href="../api/records/13298510/draft/files/europarl-v7.es-en.es/content" target="_blank" rel="noopener noreferrer">europarl-v7.es-en.es</a>]</li> </ol> <p>If using our YouTube and/or TED Talks India corpora, please cite our paper:</p> <p>Bhattacharya, D., Chi, J., Hirschberg, J., Bell, P. (2023) Capturing Formality in Speech Across Domains and Languages. Proc. INTERSPEECH 2023, 1030-1034, doi: 10.21437/Interspeech.2023-1852</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.