Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo32/100

TAU Sound Events and Speech Privacy Preservation

<div>The TAU Sound Events and Speech Privacy Preservation Dataset is a collection of audio data used in the work "Adversarial Representation Learning for Robust Privacy Preservation in Audio" by S. Gharib, M. Tran, D. Luong, K. Drossos, and T. Virtanen. The dataset is created by merging subsets of the <a href="../records/4060432">Freesound 50k Dataset (FSD50K)</a> and the <a href="https://www.openslr.org/12">LibriSpeech corpus</a>. Both FSD50K and LibriSpeech are licensed under the Creative Commons license.</div> <div>The dataset contains of ~5000 one-second sound event samples with or without speech content provided in WAV and NumPy array format (approximately half of the samples contains speech). The creation of the dataset ensures an equal number of samples for male and female speakers across each sound event class. The sound event classes included in this dataset are:</div> <ul> <li>dog barking</li> <li>glass breaking</li> <li>gun shot</li> <li>cough</li> <li>slam</li> <li>applause</li> <li>dished pot pan</li> <li>toilet flush</li> <li>cat meowing</li> <li>doorbell</li> <li>crying</li> <li>drill</li> </ul> <p>Please check the README for better understanding of the dataset.</p>

opencc-by-nc-4.0Dec 2023View details →
zenodo32/100

Crowdsourced LibriTTS Speech Prominence Annotations

<p>Dataset corresponding to the ICASSP 2024 paper "Crowdsourced and Automatic Speech Prominence Estimation" <a href="https://arxiv.org/abs/2310.08464">[link]</a></p> <p>This dataset is useful for training machine learning models to perform automatic emphasis annotaiton, as well as downstream tasks such as&nbsp;emphasis-controlled TTS, emotion recognition, and text summarization. The dataset is described in Section 3 (Emphasis Annotation Dataset). The contents of this section are copied below for convenience.</p> <p>We used our crowdsourced annotation system to perform human annotation on one eighth of the train-clean-100 partition of the LibriTTS [1] dataset. Specifically, participants annotated 3,626 utterances with a total length of 6.42 hours and 69,809 words from 18 speakers (9 male and 9 female). We collected at least one annotation of all 3,626 utterances, at least two annotations of 2,259 of those utterances, at least four annotations of 974 utterances, and at least eight annotations of 453 utterances. We did this in order to explore (in Section 6) whether it is more cost-effective to train a system on multiple annotations of fewer utterances or fewer annotations of more utterances. We paid 298 annotators to annotate batches of 20 utterances, where each batch takes approximately 15 minutes. We paid $3.34 for each completed batch (estimated $13.35 per hour). Annotators each annotated between one and six batches. We recruited on MTurk US residents with an approval rating of at least 99 and at least 1000 approved tasks. Today, microlabor platforms like MTurk are plagued by automated task-completion software agents (bots) that randomly fill out surveys. We filtered out bots by excluding annotations from an additional 107 annotators that marked more than 2/3 of words as emphasized in eight or more utterances of the 20 utterances in a batch. Annotators who fail the bot filter are blocked from performing further annotation. We also recorded participants' native country and language, but note these may be unreliable as many MTurk workers use VPNs to subvert IP region filters on MTurk [2].</p> <p>The average Cohen Kappa score for annotators with at least one overlapping utterance is 0.226 (i.e., ``Fair'' agreement)---but not all annotators annotate the same utterances, and this overemphasizes pairs of annotators with low overlap. Therefore, we use a one-parameter logistic model (i.e., a Rasch model) computed via py-irt [3], which predicts heldout annotations from scores of overlapping annotators with 77.7% accuracy (50% is random).</p> <p>The structure of this dataset is a single JSON file of word-aligned emphasis annotations. The JSON references file stems of the LibriTTS dataset, which can be found <a href="https://www.openslr.org/60/">here</a>. All code used in the creation of the dataset can be found <a href="https://github.com/interactiveaudiolab/emphases">here</a>. The format of the JSON file is as follows.</p> <p>&nbsp;</p> <pre><code>{ &lt;anonymized_participant_id_0&gt;: { "annotations": [ { "score": [ &lt;word_0_prominence&gt;,<br> &lt;word_1_prominence&gt;, &nbsp; &nbsp; &nbsp;... ], "stem": &lt;libritts_file_stem&gt;, "words": [ [ &lt;word_0&gt;, &lt;word_0_start_time&gt;, &lt;word_0_end_time&gt; ], [ &lt;word_1&gt;, &lt;word_1_start_time&gt;, &lt;word_1_end_time&gt; ],<br> ... &nbsp; &nbsp; ] },<br> ... ], "country": &lt;participant_0_country&gt;, "language": &lt;participant_0_language&gt; }, ... }</code></pre> <p><br>[1] Zen et al., &ldquo;LibriTTS: A corpus derived from LibriSpeech for text-to-speech,&rdquo; in Interspeech, 2019.<br>[2] Moss et al., &ldquo;Bots or inattentive humans? Identifying sources of low-quality data in online platforms,&rdquo; PsyArXiv preprint PsyArXiv:wr8ds, 2021.<br>[3] John Patrick Lalor and Pedro Rodriguez, &ldquo;py-irt: A scalable item response theory library for Python,&rdquo; INFORMS Journal on Computing, 2023.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

THLS - An open source dataset for Brazilian Portuguese speech processing

<p>THLS Open Source Brazilian Portuguese Speech Dataset with 1000 sentences balanced phonetically.</p> <p>&nbsp;</p> <p>Authors:</p> <ul> <li>Luiz Felipe Vecchietti</li> <li>Thalles Melo Batista Pieroni (Voice)</li> </ul> <p>&nbsp;</p> <p>https://gitlab.com/lfelipesv/1000-sentences-thls-dataset</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

LittlEARS questionnaire on early speech production for the Serbian language in children with normal hearing

<p>This dataset includes data for validation of LittlEARS questionnaire on early speech production (LEESPQ) for the Serbian language in children with normal hearing.</p> <p>Dataset includes data about 206 children, age 0 to 18 months. Information was provided by their parents.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Forms and functions of pre-speech gestures in 12- to 15-months old infants

<p>Supplementary material 2 for manuscript titled "<strong><span>Forms and functions of pre-speech gestures in 12- to 15-months old infants"</span></strong></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Resources for the paper "Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo32/100

Dataset Labeled For Inciting Speech

<p>This dataset is related to the paper "Understanding Inciting Speech As New Malice." The paper recently got accepted at IEEE Transactions on Computational Social Systems. Please cite the paper while using the dataset.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Splicing Detection and Localization for Speech Deepfakes using Audio Novelty

<p>With the rapid progress of artificial intelligence over the last few years, the possibility of generating highly realistic multimedia content is within everyone's reach. This has paved the way for the creation of deepfakes, synthetic data produced using deep learning techniques that realistically represent people in behaviors not belonging to them.&nbsp;<br>In the audio case, speech deepfakes can be used alone or combined with splicing techniques.<br>This consists of replacing portions of authentic speech with synthetic elements, thereby altering the conveyed message and leading to significant threats.<br>In this work, we address the problem of splicing detection and localization in speech deepfakes.<br>We consider spliced audio tracks created by substituting parts of a pristine speech with synthetically generated segments. We design a system able to detect whether a manipulation takes place and localize it in time.<br>The proposed method, trained exclusively on splicing-free audio tracks, extracts embeddings from the input signal through a sliding window. Then, it employs audio novelty to measure the similarity among consecutive signal sections and uses it to detect and localize splicing points.<br>We evaluate our method on a state-of-the-art dataset as well as a novel, bias-free corpus specifically developed and released in this paper.&nbsp;<br>The proposed approach is benchmarked against multiple baselines for both splicing detection and localization tasks.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

ESCorpus-PE: A speech emotional dataset in Spanish with Peruvian accent

<p>ESCorpus-PE dataset contains emotional utterances of Spanish peruvian speech gathered from Spanish interviews, TV reports, political debate and testimonials. It contains 3749 utterances of three emotional dimensions: Valence, Arousal and Dominance. There are 80 speakers (44 male and 36 female). This data was created from Youtube audios. These audios were selected following a specific criteria specified in the paper: ESCorpus-PE: A speech emotional database in Spanish with Peruvian accent, in the section Methods/Audio/Video Selection. Anyone can use this data only for research purposes.</p> <p>More details on&nbsp;<br> https://github.com/Alessandra-UNSA/Peruvian_Spanish_Corpus</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Quantitative comparison of respiratory aerosol implicates speech as a principal driver in asymptomatic transmission of airborne disease

<p>Video recordings of breath and speech droplets.&nbsp;</p> <p>The first set of recordings shows breath droplets, enlarged in size due to nucleated condensation, while crossing a 0.7-mm thickness light-sheet at the original recording speed (120 f/s) and in slow motion (24 f/s).</p> <p>The second set of recordings shows droplets generated by speaking a single word, &lsquo;popeye&rsquo;, while traversing a low intensity (~0.1W/cm<sup>2</sup>) light-sheet of 13 mm thickness. The recordings are shown at the original recording speed (120 f/s), in slow motion (24 f/s) and after background subtraction of Raleigh scattering and other time-invariant sources of light.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Toward General Speech Restoration With VoiceFixer - Training Data

<p>This data set contains speech data and noise data for the paper:&nbsp;VoiceFixer: Toward General Speech Restoration with Neural Vocoder.</p> <p>If you found this dataset helpful, please consider citing:&nbsp;</p> <blockquote> <pre>@article{liu2021voicefixer, title={VoiceFixer: Toward General Speech Restoration with Neural Vocoder}, author={Liu, Haohe and Kong, Qiuqiang and Tian, Qiao and Zhao, Yan and Wang, DeLiang and Huang, Chuanzeng and Wang, Yuxuan}, journal={arXiv preprint arXiv:2109.13731}, year={2021} }</pre> </blockquote>

opencc-by-4.0Sep 2021View details →
zenodo32/100

Dataset and tools for the PSST Challenge on Post-Stroke Speech Transcription

<p>Initial Release, for archival/DOI purposes.</p>

openother-openMar 2022View details →
zenodo32/100

The Maaloula Aramaic Speech Corpus (MASC)

<p>This dataset contains the first electronic speech corpus of Maaloula Aramaic, an endangered Western Neo-Aramaic variety spoken in Syria. This 64,845-word corpus is available in four formats: (1) transcription, (2) lemmatized transcription, (3) audio files and time-aligned phonetic transcriptions, and (4) an SQLite database. The transcription files are a digitized and corrected version of authentic transcriptions of tape-recorded narratives coming from a fieldwork trip conducted in the 1980s and published in the early 1990s (Arnold, 1991a, 1991b). They contain no annotation, except for some informative tagging (e.g. to mark loanwords and misspoken words). In the lemmatized version of the files, each word form is followed by its lemma in angled brackets. The time-aligned TextGrid annotations consist of four tiers: the sentence level (Tier 1), the word level (Tiers 2 and 3), and the segment level (Tier 4). These TextGrid files are downloadable together with their audio files (for the original source of the audio data see Arnold, 2003). The SQLite database enables users to access the data on the level of tokens, types, lemmas, sentences, stories, or speakers.</p> <p>For more information, please see our paper:&nbsp;Ghattas Eid, Esther Seyffarth, Ingo Plag. 2022. The Maaloula Aramaic Speech Corpus (MASC): From Printed Material to a Lemmatized and Time-Aligned Corpus. In&nbsp;<em>Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)</em>, Marseille, France. European Language Resources Association (ELRA).</p>

openother-ncJan 2022View details →
dryad32/100

Data from: Parallel processing in speech perception with local and global representations of linguistic context

<p>Speech processing is highly incremental. It is widely accepted that human listeners continuously use the linguistic context to anticipate upcoming concepts, words, and phonemes. However, previous evidence supports two seemingly contradictory models of how a predictive context is integrated with the bottom-up sensory input: Classic psycholinguistic paradigms suggest a two-stage process, in which acoustic input initially leads to local, context-independent representations, which are then quickly integrated with contextual constraints. This contrasts with the view that the brain constructs a single coherent, unified interpretation of the input, which fully integrates available information across representational hierarchies, and thus uses contextual constraints to modulate even the earliest sensory representations. To distinguish these hypotheses, we tested magnetoencephalography responses to continuous narrative speech for signatures of local and unified predictive models. Results provide evidence that listeners employ both types of models in parallel. Two local context models uniquely predict some part of early neural responses, one based on sublexical phoneme sequences, and one based on the phonemes in the current word alone; at the same time, even early responses to phonemes also reflect a unified model that incorporates sentence-level constraints to predict upcoming phonemes. Neural source localization places the anatomical origins of the different predictive models in nonidentical parts of the superior temporal lobes bilaterally, with the right hemisphere showing a relative preference for more local models. These results suggest that speech processing recruits both local and unified predictive models in parallel, reconciling previous disparate findings. Parallel models might make the perceptual system more robust, facilitate processing of unexpected inputs, and serve a function in language acquisition.</p>

opencc-zeroMay 2022View details →
zenodo32/100

Speeches

<p>Lengths of all speeches in words in a set of 275 plays</p>

opencc-by-4.0Jun 2022View details →
dryad32/100

Data from: Whistling shares a common tongue with speech: bioacoustics from real-time MRI of the human vocal tract

Most human communication is carried by modulations of the voice. However, a wide range of cultures has developed alternate forms of communication that make use of a whistled sound source. For example, whistling is used as a highly salient signal for capturing attention, can have iconic cultural meanings such as the cat-call, enact a formal code as in boatswain's calls, or stand as a proxy for speech in whistled languages. We used real-time magnetic resonance imaging to examine the muscular control of whistling to describe a strong association between the shape of the tongue and the whistled frequency. This bioacoustic profile parallels the use of the tongue in vowel production. This is consistent with the role of whistled languages as proxies for spoken languages, in which one of the acoustical features of speech sounds are substituted with a frequency modulated whistle. Furthermore, previous evidence that non-human apes may be capable of learning to whistle from humans suggests that these animals may have similar sensorimotor abilities to those that are used to support speech in humans.

opencc-zeroSep 2019View details →
zenodo32/100

Seeing speech: The cerebral substrate of tickertape synesthesia

<p>We report the first functional MRI study of a tickertape synesthete. These synesthetes were described by Galton as &quot;persons [who] see mentally in print every word that is uttered and they read them off usually as from a long imaginary strip of paper&quot;.</p> <p>Our synesthete and 35 other non-synesthetes controls were presented different auditory stimuli. We performed univariate and multivariate analysis, as dynamic causal modelling for the synesthete.</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

Central Bank Digital Speech Analysis for Measuring Monetary Sovereignty Concern for France, Japan, the UK, and the EMU

<p>This dataset was compiled to measure central banks&#39; shifts in monetary sovereignty concern. The scores received by each speech range from -1 to +1, with -1 being shrinking concern for monetary sovereignty and +1 being growing concern. 0 is neutral. The dataset runs from December 2018 to mid-January 2020 and was used in a study to determine correlation between changes in monetary sovereignty and retail CBDC programs amongst France, the UK, and Japan.&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

A Large TV Dataset for Speech and Music Activity Detection

<p>Automatic speech and music activity detection (SMAD) is an enabling task that can help segment, index, and pre-process audio content in radio broadcast and TV programs. However, due to copyright concerns and the cost of manual annotation, the limited availability of diverse and sizeable datasets hinders the progress of state-of-the-art (SOTA) data-driven approaches. We address this challenge by presenting a large-scale dataset containing Mel spectrogram, VGGish, and MFCCs features extracted from around 1600 hours of professionally produced audio tracks and their corresponding noisy labels indicating the approximate location of speech and music segments. The labels are derived from several sources such as subtitles. A test set curated by human annotators is also included as a subset for evaluation.&nbsp;To the best of our knowledge, this dataset is the first large-scale, open-sourced dataset that contains features extracted from professionally produced audio tracks and their corresponding frame-level speech and music annotations.&nbsp;</p>

openapache2.0Dec 2021View details →
zenodo32/100

Demo data and models for: Automated speech detection in eco-acoustic data enables privacy protection and human disturbance quantification

<p>Folder containing a <strong>demo dataset</strong> and the <strong>model weights</strong> resulting from the ecoVAD pipeline. The data contained in this folder allows for full reproducibility of the pipeline described on the <a href="https://github.com/NINAnor/ecoVAD">ecoVAD GitHub repository</a>.</p> <p>If you have any questions or issues with the dataset, please open an issue on the ecoVAD GitHub repository.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record