Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.7.1
Dataset results
859 results for “speech”
Speech disfluencies: Neurophysiological aspect in normal population
Open the record for dataset details and reuse information.
Salvaging the Internet Hate Machine: Using the discourse of extremist online subcultures to identify emergent extreme speech
<p>This dataset accompanies a paper submitted to the WebSci 20 conference. In this paper, we present a lexicon of 'extreme speech' that may be used to detect hate speech and extreme speech on online platforms. We outline a cross-disciplinary research protocol through which this lexicon is initially extracted from a corpus of 3,335,265 posts from 4chan's /pol/ sub-forum using a hybrid method comprising word2vec modeling and subsequent snowballing of nearest neighbours of a small initial expert seed list of extreme language. The choice of corpus is significant, as 4chan is a space of rapid language innovation and obscure extreme vernacular, complicating generalised approaches. Our lexicon detects significantly more extreme posts within a corpus from a more mainstream platform (Reddit) than another popular lexicon, Hatebase, with similar accuracy. Our lexicon and the method of its creation thus provide a contribution to the study of the toxicity of online subcultures similar to 4chan, as well as more mainstream platforms. As we demonstrate, the lexicon allows for more effective detecting of extreme speech in these spaces. This method and the lexicon have further been made available through an open-source web tool for the study of online social platforms, 4CAT. The computational methods and lexicon on offer here can thus be used by a wide academic audience, fostering interdisciplinary approaches to the study of online hate and extreme speech. </p> <p>The dataset comprises the following items:</p> <ul> <li>The 4chan corpus from which the extreme speech lexicon was generated (posts from /pol/, 1 October 2019 - 1 November 2019)</li> <li>The Reddit corpus used to verify and test the lexicon (posts from the_donald, theredpill, politics and chapotraphouse, 1 October 2019 - 1 November 2019)</li> <li>The word2vec model from which the extreme speech lexicon was generated</li> <li>The extreme speech lexicon that was generated</li> </ul>
EEG Data for: "Cortical oscillations and entrainment in speech processing during working memory load"
<p>This repository contains EEG and audio data used and described in:</p> <p><strong>Hjortkjær, J, Märcher-Rørsted, J, Fuglsang, SA, Dau, T (2018). Cortical oscillations and entrainment in speech processing during working memory load. European Journal of Neuroscience. </strong><strong>doi</strong><strong>:10.1111/ejn.13855</strong></p> <p>Please cite this article when using the data</p> <p> </p> <p>The MAT-files contain the aligned EEG and audio data for each subject (N=22). The envelopes of the speech audio (without noise) have been extracted as described in the paper. Each file (data_N.mat) contains a Matlab struct in the format of the Fieldtrip toolbox containing the following fields:</p> <p> </p> <p>data.trial: EEG and audio data for all 40 trials [channels x timepoints]</p> <ul> <li>channels 1-64: scalp EEG</li> <li>channel 65: left mastoid electrode</li> <li>channel 66: right mastoid electrode</li> <li>channel 67: horizontal EOG</li> <li>channel 68: vertical EOG for left eye</li> <li>channel 69: vertical EOG for right eye</li> <li>channel 70: audio envelopes</li> </ul> <p>data.trialinfo: Experimental condition in each trial</p> <ul> <li>1 = low noise, 1-back</li> <li>2 = low noise, 2-back</li> <li>3 = high noise, 1-back</li> <li>4 = high noise, 2-back</li> </ul> <p>data.time: Sample indices for each trial in seconds</p> <p>data.label: Name of each channel in data.trial</p> <p>data.fsample: EEG/audio sampling rate in Hz (128)</p>
Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)
<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from 'Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais' de J.-F. Bladé, 'Coundes biarnés, couéilhuts aüs parsàas miéytadès dou péys dé Biarn' de J.-V. Lalanne, 'Contes populaires du Languedoc' de L. Lambert and 'Contes populaires recueillis dans la Grande-Lande' de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>
Netanyahu's and Abbas' speeches at the UNGA 2010-19, coded using a populism framework
<p>Databaset with speeches of Benjamin Netanyahu and Mahmoud Abbas before the United Nations General Assembly (UNGA) (2010-2019) coded using MAXQDA following populism multidimensional comparative framework by Olivas Osuna (2021).</p> <p>In this database syntactic units —sentences—are individually coded whenever they match the criteria corresponding to any of populism/anti-populism, re-bordering/de-bordering, religion and securitisation codes defined previously. See Olivas Osuna and Rama (2021) and Olivas Osuna (2022) for reference to the methodology and previous empirical applications.</p>
Sculpted Speech
<p>This dataset contains the stimuli heard by participants in a study of sculpted speech. The zip file contains 240 Spanish sentences from the Sharvard corpus in each of eight conditions that vary in the amount of target speech information they contain. </p> <p>For illustration purposes (for those not wanting to download the entire zip), the dataset also contains as separate wav files for one example sentence in each of the eight conditions. The names for these wav files map to the subdirectories as follows:</p> <p>baseline -> 'mix'<br> speech -> 'mask_speech'<br> mix -> 'mask_mix'<br> fine -> 'fine_mix'<br> envelope -> 'fine_masker'<br> ssn -> 'mask_ssn'<br> music -> 'mask_music'<br> compspeech -> 'mask_babble'<br> </p>
WOLOF TTS(Text To Speech) Data
<p>This contains a WOLOF Text To Speech(TTS) dataset, it contains recordings from two natif Wolof actos (a male and female voice).<br> Each actor recored more than 20 000 sentences.<br> The notebook accompanying the dataset contains a brief analysis of the dataset and the code creating the appropriate train/validation and test set.<br> The file [male, female]train, [male, female]validation and [male, female]test are also present to extract the corespondig audios inside the data-commonvoice.zip</p> <p>The text dataset come from news website, Wikipedia and self curated text. We made sure with the help of our Wolof expert that the text dataset cover the different phonemes in the Wolof language.</p>
Speech stimuli (ACI experiment)
<p>Speech stimuli involved in the Auditory Classification Image experiment (Alda/Alga/Arda/Arga). Male speaker, wav format, 48 kHz</p>
Children speech recording (English, spontaneous speech + pre-defined sentences)
<p>The dataset contains audio recordings (lossless WAV) of 11 young children (age M=4.9 years old; 5 females, 6 males).</p> <p>Recordings include:</p> <ul> <li>free speech (retelling a picture book, ‘Frog, Where Are You?’ by Mercer Mayer)</li> <li>repeating 5 pre-defined short sentences (like 'the horse is in the stable')</li> <li>telling the numbers from 1 to 10</li> </ul> <p>The recordings are in English and the participants include both native and non-native speakers.</p> <p>Each sample is recorded from 3 sources:</p> <ul> <li>A studio-grade microphone (Rode NT1-A)</li> <li>A portable microphone (Zoom H1)</li> <li>The two front microphones of the Aldebaran NAO robot</li> </ul> <p>(note that, due to technical issues, a few (sample/microphone) combinations are missing).</p> <p> </p> <p>For the free-speech recording, a manual segmentation of the utterances is provided as well.</p>
Original Recording of Freiburg Words for Testing Hearing with Speech
<p>The files contain recordings of the monosyllabic words, polysyllabic numbers, and CCITT noise (ITU, 1993) of the Freiburg Speech Test (Hahlbrock, 1953). The words are given in DIN 45621-1:1995. The 400 monosyllables are organized in 20 test lists à 20 nouns. They are stored in "Einsilbige Wörter.zip". The name of each wav-file includes the number of the test list (L01, L02, …) and the position of the word within the test list (W01, W02, …). The 100 polysyllables are organized in 10 test lists à 10 numbers. They are stored in "Mehrsilbige Wörter (Zahlen).zip" with similar nomenclature. DIN 45626-1:1995 describes the recordings. Speech signals were recorded in 1969 in the studios of Norddeutscher Rundfunk, Hamburg, Germany, with the speaker Claus Wunderlich (Brinkmann, 1974). The recordings were processed by Physikalisch-Technische Bundesanstalt, Braunschweig, Germany, and Polygram International, Hannover, Germany. Later, the recordings were digitalized by Siemens AG and distributed on compact disc by "Siemens Audiologische Technik GmbH" (legal successor WS Audiology A/S) with the title "Wörter für Gehörprüfung mit Sprache" ("Words for hearing tests with speech") under item no. 7970155. The words were cut out as accurately as possible, i.e., with as little background before and after each word as possible (see Winkler and Holube, 2016a). The CCITT noise was originally included for calibration purposes, but is often used as noise masker when the monosyllables are presented in background noise.</p>
A challenging data set for evaluating part-of-speech taggers
<p>This data set contains 2,227 sentences, with a part-of-speech (POS) tag specified for a single word in the sentence. The data file is a tab-separated text file where each row (after the header row) is formatted as follows:</p><p><i>sentence <TAB> POS tag <TAB> (optional) motivation</i></p><p>Note that, in the sentence (= a string of space-separated characters), the POS-tagged word is indicated by the POS tag in brackets, placed just after the word to which it refers. Example:</p><p><i>The road bends [VERB] to the right . VERB</i></p><p>In this example, the optional motivation is not included, as the tagged word can easily be identified as being a verb.</p>
Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal
<p>This data is transcribed speech data, in Wolof, Pulaar and Sereer.</p> <p>The recordings are about agriculture. The recorded consist of farmers, agricultural advisers, and agri-food business managers. Type of recordings comprise interactive radio programmes, focus groups, voice messages, push messages and interviews. Therefore, spontaneous speech is prevailing. Quality of audio may vary depending on the type of programme.</p> <p>Content description :</p> <ul> <li><strong>speech_dataset_wol.tar.gz:</strong> Wolof (ISO Code 639-2: wol) speech dataset contains 55 hours of transcribed speech, including almost 13 hours of validated content check by an expert. It also contains a XSAMPA lexicon (49,132 phonetised entries) and a text corpus (1,140,508 words).</li> <li><strong>speech_dataset_fuc.tar.gz:</strong> Pulaar (ISO Code 639-2: fuc) speech dataset contains nearly 32 hours of transcribed speech, including around 11 hours of validated content check by an expert. It also contains a text corpus (742,024 words).</li> <li><strong>speech_dataset_srr.tar.gz:</strong> Sereer (ISO Code 639-2: srr) speech dataset contains 38 hours of transcribed speech, including nearly 11 hours of validated content check by an expert.<br>In total, these resources provide 125 hours of transcribed speech in the 3 most widely spoken languages in Senegal, including 35 hours of checked transcriptions.</li> </ul> <p>This work is a result of the Kallaama project, funded by Lacuna Fund for 1 year, in 2023. </p> <p>See the <a title="Kallaama speech dataset" href="https://github.com/gauthelo/kallaama-speech-dataset" target="_blank" rel="noopener">GitHub repository</a> for more details about the dataset.</p>
Challenges to freedom of speech and journalists in Ukraine in times of war – Non-representative online expert survey of Ukrainian journalists (January 2023)
The expert survey of journalists was conducted from 18 to 27 January 2023 using a self-completion questionnaire in Google Forms. The survey was conducted by the Ilko Kucheriv Democratic Initiatives Foundation on the request of the Human Rights Centre ZMINA with the support of Freedom House Ukraine. A total of 132 people participated in the survey. The respondents were selected using the method of voluntary selection and snowballing to the point of saturation. The sample represents only the opinion of the respondents, but it also allows us to talk about certain trends and common assessments of certain phenomena and processes in the journalistic field. The survey includes questions about freedom of speech and self-censorship in the media environment during the Russian-Ukrainian war. The data collection contains original survey data. The Excel file (.xlsx) is the original file with the respondents' answers in Ukrainian, provided by the Ilko Kucheriv Democratic Initiatives Foundation. The documentation includes the questions and answer options of the original questionnaire in Ukrainian and English. Additionally, the data collection contains the "Summary" file, which is an analytical report prepared by the Ilko Kucheriv Democratic Initiatives Foundation and the Human Rights Centre ZMINA. The report uses data from an expert survey of journalists in 2019 and 2023, and the results of focus groups in 2022.
Oral cancer speech corpus for paper "Detecting and analysing spontaneous oral cancer speech in the wild"
<p>This is the oral cancer speech corpus used in the paper <em>"Detecting and analysing spontaneous oral cancer speech in the wild".</em></p> <p><strong>Description</strong></p> <p>This dataset contains approximately 3 hours of oral cancer speech data collected from YouTube, including a file with additional metadata. We use this dataset to perform an oral cancer speech detection task in our paper.</p> <p><strong>Funding</strong></p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under Marie Sklodowska-Curie grant agreement No 766287. The Department of Head and Neck Oncology and surgery of the Netherlands Cancer Institute receives a research grant from Atos Medical (Horby, Sweden),<br> which contributes to the existing infrastructure for quality of life research.</p> <p><strong>Citation:</strong></p> <p>If you use this dataset please cite:</p> <pre><code>@misc{halpern2020detecting, title={Detecting and analysing spontaneous oral cancer speech in the wild}, author={Bence Mark Halpern and Rob van Son and Michiel van den Brekel and Odette Scharenborg}, year={2020}, eprint={2007.14205}, archivePrefix={arXiv}, primaryClass={eess.AS} }</code></pre> <p> </p>
Google Speech Commands-Musan test set
<p>This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2]. It is introduced in our ICASSP 2022 paper [3]. </p> <p>Specifically, we created this test set by mixing the speech in the Google Speech Commands v2 test set with random noise in the Musan dataset at different signal to noise ratio -12.5,-10,0,10,20,30 and 40 decibel (dB). </p> <p>The Google Speech Commands v2 dataset is under the Creative Commons BY 4.0 license. It could be downloaded at: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz</p> <p>The Musan dataset is under Attribution 4.0 International (CC BY 4.0). It could be downlowned at https://www.openslr.org/17/</p> <p>Citations:</p> <p>[1] Pete Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,” arXiv preprint arXiv:1804.03209, 2018.</p> <p>[2] David Snyder, Guoguo Chen, and Daniel Povey, “Musan: A music, speech, and noise corpus,” arXiv preprint arXiv:1510.08484, 2015.</p> <p>[3] V. A. Trinh, H. Salami Kavaki and M. I. Mandel, "Importantaug: A Data Augmentation Agent for Speech," <em>ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</em>, 2022, pp. 8592-8596, doi: 10.1109/ICASSP43922.2022.9747003.</p>
SaGA++ Speech-Gesture Dataset Extension
<p>This is the dataset extension release for <a href="https://pub.uni-bielefeld.de/record/2001935">the SaGA dataset</a>.</p> <p>We have extracted anonymous features from all the 25 recordings and release them all now instead of only 6 recordings that were previously released. The modalities included are text annotations, gesture annotations, prosodic audio features, and movement trajectories. More details are provided in the README document included in the dataset.</p> <p><strong>Please reference the following papers in any publication making use of the Dataset:</strong></p> <p>Kucherenko, T., Nagy, R., Neff, M., Kjellström, H., & Henter, G. E. (2022). Multimodal analysis of the predictability of hand-gesture properties. 21st International Conference on Autonomous Agents and Multiagent Systems (AAMAS). </p> <p>Lücking, A., Bergman, K., Hahn, F., Kopp, S., & Rieser, H. (2013). Data-based analysis of speech and gesture: The Bielefeld Speech and Gesture Alignment Corpus (SaGA) and its applications. <em>Journal on Multimodal User Interfaces</em>, <em>7</em>(1), 5-18.</p>
Introducing the COVID-19 YouTube (COVYT) speech dataset featuring the same speakers with and without infection
<p>The COVYT dataset contains speech samples from individuals who self-reported their COVID-19 infection on public social media platforms (YouTube, Xiaohongshu). These videos, as well as accompanying videos of the same people prior to infection, were mined in an attempt to gather publicly-available data for COVID-19 research. This release includes the links to the original videos along with the accompanying manual segmentation and diarisation that identifies the utterances of the target individuals. We are additionally releasing features derived from the segmented utterances. Finally, the dataset includes partitioning information according to 4 different cross-validation schemes. See the arxiv pre-print for more details: https://arxiv.org/abs/2206.11045</p>
British English Cued Speech--interspeech 2019
<p>The first British English Cued Speech (CS) dataset is recorded for this work in Cued Speech UK association without using any artificial mark, and it is also the first one especially for the continuous recognition in British English CS. A professional CS interpreter (with no hearing impairment) is asked to simultaneously utter and encode a set of 97 British English sentences (e.g., I feel it is a time to move to a new chapter in my career). 97 sentences are chosen from the uploaded text file. There are totally 907 monophthongs and 138 diphthongs in the dataset. Color video images of the interpreter’s upper body are recorded at 25 fps, with a spatial resolution of 720x1280. </p> <p>If you use this dataset, please cite the following:</p> <pre>@inproceedings{Liu2019, author={Li Liu and Jianze Li and Gang Feng and Xiao-Ping Zhang}, title={{Automatic Detection of the Temporal Segmentation of Hand Movements in British English Cued Speech}}, year=2019, booktitle={Proc. Interspeech 2019}, pages={2285--2289}, doi={10.21437/Interspeech.2019-2353}, url={http://dx.doi.org/10.21437/Interspeech.2019-2353} }</pre>
DeepPredSpeech: computational models of predictive speech coding based on deep learning
<p>This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: </p> <p>Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (<a href="https://doi.org/10.1101/471581">biorXiv preprint https://doi.org/10.1101/471581</a>). </p> <ul> <li>Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). <ul> <li>Audio recordings are available in the audio_clean/ directory</li> <li>Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)</li> <li>Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf</li> </ul> </li> <li>Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. </li> <li>Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories</li> <li>Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of <span class="math-tex">\(\tau_p\)</span> or <span class="math-tex">\(\tau_f\)</span> are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory</li> </ul> <p>Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/</p> <p>All directories have been zipped before upload.</p> <p>Feel free to contact me for more details.</p> <p>Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, thomas.hueber@gipsa-lab.fr </p>
CLDF dataset derived from Bremer's "Sociolinguistic Survey of Six Berta Speech Varieties in Ethiopia" from 2016
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bremer, Nate D. (2016): A Sociolinguistic Survey of Six Berta Speech Varieties in Ethiopia. SIL Electronic Survey Reports 2016-007. Dallas: SIL International.</p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.