Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Stimuli for "Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video"
<p>This dataset contains all stimuli used in "Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video" by Matthew Groh, Aruna Sankaranarayanan, Nikhil Singh, Dong Young Kim, Andrew Lippman, and Rosalind Picard.</p> <p>The videos are contained in the Materials folder and the Data Folder provides two .csv files that map the video filenames with their metadata. </p> <p>For video sources, see Table 19 in the arXiv version of the paper: https://arxiv.org/abs/2202.12883</p>
Speeches SJ and MZ, annotation version 02 Mar 2018
<p>First version of annotated keynote speech excerpts of SJ amd MZ (4 audio files and the corresponding PRAAT Textgrid files created by Jana Voße, Esther Novák-Tót, Jana Thumm, and Simón Gonzalez Ochoa)</p>
LANGUAGE FOUNDATIONS OF ORAL SPEECH DEVELOPMENT
<p><span>Children create<span> </span>strong language skills when they are involved in playful, language-rich situations with opening<span> </span>to learn new words. Hands-on experiences encourage learning and provide a context for new words to be explored. For example, it’s easier for children to learn vegetable names when they are touching or tasting them. </span></p>
EEG data of continuous listening of music and speech
<p>This dataset contains EEG recordings from 18 subjects listening to continuous sound, either speech or music. Continuous audio stimuli were presented to listeners in trials of 70 seconds from one loudspeaker located 150 cm in from of them. They were instructed to attentively listen to the sound during the whole trial.</p> <p>All listeners were native Danish speakers and were presented with 5 different types of audio stimuli:</p> <ul> <li>Instrumental music: Excerpt of Disney songs: Excerpts of polyphonic Disney songs with no lyrics. The melody line from the original version was replaced by a similar melody played by a synthetic cello (Referred to as MC: Music Cello in the dataset).</li> <li>Music with understood lyrics: Excerpts of polyphonic Disney songs with lyrics in Danish, understood by the listeners (Referred to as MD: Music Danish in the dataset).</li> <li>Music with non-understood lyrics: Excerpts of polyphonic Disney songs with lyrics in Finnish, not understood by the listeners (Referred to as MF: Music Finnish in the dataset).</li> <li>Understood speech: Excerpts of an audiobook in Danish read by a woman, understood by the listeners (Referred to as SD: Speech Danish in the dataset).</li> <li>Non-Understood speech: Excerpts of an audiobook in Finnish read by a woman, understood by the listeners (Referred to as SF: Speech Danish in the dataset).</li> </ul> <p> </p> <p>Data were recorded using a 64-channels g.HIamp-Research system and digitalized at a sampling rate of 2400 Hz.</p> <p>The dataset contains pre-processed EEG data (see pre-processing step applies to the data below), for each listener. Trials with large noise artefacts have been removed.</p> <p>If you use this data, please cite the original publications, as well as this repository</p> <p>Simon, A., Bech, S., Loquet, G., & Østergaard, J. (2024). Cortical linear encoding and decoding of sounds: Similarities and differences between naturalistic speech and music listening. European Journal of Neuroscience, 59(8), 2059–2074. https://doi.org/10.1111/ejn.16265</p> <p> </p> <p>Simon, A., Bech, S., Loquet, G., Østergaard, J. (<span>2022</span>). EEG data of continuous listening of music and speech. Zenodo. 10.5281/zenodo.7500806</p> <p>The processed data contains EEG data and an aligned audio envelope for each category of audio stimuli. The MATLAB script contains the processing applied to obtain it.</p> <p>The dataset was created within the <a href="https://iosr.surrey.ac.uk/projects/InHEAR/">InHear</a> project.</p> <p>For more information, amds@es.aau.dk</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><em>Preprocessing done</em></p> <p> </p> <p><em>-re-reference to average channels</em></p> <p><em>-downsampling to 512Hz</em></p> <p><em>-bandpass filter 0.5-45 Hz</em></p> <p><em>-ICA decomposition using SOBI algorithm</em></p> <p><em>-removed eyes and noise components</em></p>
EmoFilm - A multilingual emotional speech corpus
<p><strong>EmoFilm</strong> is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages. The audio clips (with a mean length of 3.5 sec. and std 1.2 sec.) were extracted in wave format (uncompressed, mono, 48 kHz sample rate and 16-bit) from 43 films (original in English and their over-dubbed Italian and Spanish versions). Genres including comedy, drama, horror, and thriller were considered; anger, contempt, happiness, fear, and sadness emotional states were taken into account. EmoFilm has been presented at Interspeech 2018:</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, and Björn Schuller (2018), <em>Categorical vs Dimensional Perception of Italian Emotional Speech</em>, in Proc. of Interspeech, Hyderabad, India, pp. 3638-3642 .</p> <p>We would like to thank Linda Ratz for her contribution in the generation of the transcriptions.</p> <p> </p> <p><strong>How to access EmoFilm</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1pFHfsqk7snF_EVqq0WAC0Dz8FcTD3s9_/view?usp=share_link</p>
Hateful Messages: A Conversational Data Set of Hate Speech produced by Adolescents on Discord
<p>With the rise of social media, a rise of hateful content can be observed. Even though the understanding and definitions of hate speech varies, platforms, communities, and legislature all acknowledge the problem. Therefore, adolescents are a new and active group of social media users. The majority of adolescents experience or witness online hate speech. Research in the field of automated hate speech classification has been on the rise and focuses on aspects such as bias, generalizability, and performance. To increase generalizability and performance, it is important to understand biases within the data. This research addresses the bias of youth language within hate speech classification and contributes by providing a modern and anonymized hate speech youth language data set consisting of 88.395 annotated chat messages. The data set consists of publicly available online messages from the chat platform Discord. ~6,42\% of the messages were classified by a self-developed annotation schema as hate speech. For 35.553 messages, the user profiles provided age annotations setting the average author age to under 20 years old.</p>
An Exploratory Comparison of Attitudes Toward Stuttering and Cluttering of Chinese Practicing Speech and Language Therapists (SLTs) and SLT Students
Open the record for dataset details and reuse information.
Data for Reported user-generated online hate speech: The 'ecosystem', frames, and ideologies
<p>This is the dataset for the article entitled Reported user-generated online hate speech: The 'ecosystem', frames, and ideologies. The same dataset is provided in two different formats: comma-separated values (.csv) and Excel format (.xlsx). A basic legend to the data is provided separately in the corresponding PDF document. An extended legend to the data is available at: <strong><a href="https://doi.org/10.5281/zenodo.6656185">https://doi.org/10.5281/zenodo.6656185</a></strong>.</p>
The Berlin Dataset of Lombard and Masked Speech (BELMASK)
<p>The Berlin Dataset of Lombard and Masked Speech (BELMASK) is a phonetically controlled audiovisual dataset of speech produced in adverse speaking conditions. The dataset contains in total 128 min of audio and video recordings of 10 German native speakers (4 female, 6 male) with a mean age of 30.2 years (SD: 6.3 years), uttering matrix sentences in cued, uninstructed speech in four conditions: (i) with a Filtering Facepiece P2 (FFP2), (ii) without an FFP2 mask in silence, (iii) with an FFP2 mask while exposed to noise, iv) without an FFP2 mask while exposed to noise. Noise consisted of mixed-gender six-talker babble played over headphones to the speakers, triggering the Lombard effect. All conditions are readily available in face-and-voice and voice-only formats. The speech material is annotated, employing a multi-layer architecture, and was originally conceptualized to be used for the administration of a working memory task. The dataset is available for academic research in the area of speech communication, acoustics, psychology and related disciplines upon request, <strong>after signing an End User License Agreement (EULA)</strong>.</p>
DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception
<p>DEMoS (Database of Elicited Mood in Speech), is a corpus of induced emotional speech in Italian. DEMoS encompasses 9,365 emotional and 332 neutral samples produced by 68 native speakers (23 females, 45 males) in seven emotional states: the 'big six' anger, sadness, happiness, fear, surprise, disgust, and the secondary emotion guilt. To get more realistic productions, instead of acted speech, DEMoS contains emotional speech elicited by combinations of Mood Induction Procedures (MIP). Three elicitation methods are presented, made up by the combination of at least three MIPs, and considering six different MIPs in total. To select samples 'typical' of each emotion, evaluation strategies based on self- and external assessment were applied. The selected part of the corpus encompasses 1,564 prototypical samples produced by 59 speakers (21 females, 38 male). DEMoS has been published in the Journal Language, Resousrces, and Evalaution.</p> <p> </p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Maximilian Schmitt, and Björn Schuller (2019), <em>DEMoS: An Italian emotional speech corpus. Elicitation methods, machine learning, and perception</em>, Language, Resources, and Evaluation, Feb 2019. <a href="http://em.rdcu.be/wf/click?upn=lMZy1lernSJ7apc5DgYM8eCoqdGxOfRWEudjYRrxU-2BI-3D_Ru5N6PJ4ngeR7K-2Fncs2CW1jGAzl4dMvrVh77-2BVH-2B9g5urNss1KItQNXvWL1jiHKvcYDtUVs2c78DX20PMDTauCGehGiQvHdgrAknGggtu7pHINBqVKjp16-2BTn63kNrm22m52e-2FPV-2FidpRe8A-2FplLxPMV-2FjTR-2FLLIK8Wqe7u0-2BLSZ9w-2BWYtrAXRYn2lvPcjGTP1La8yiTxBuJKbHJpnNeFb6LmBIiNMmGRSZPIY0leXhyj4k07rx5cETF6n34aIQHP-2FwcafanNMN4BoA9QKhXGgFxvRgZQidsQ-2BCDbbTBL0PPjM3CgitSGk66qut9E3pd">https://rdcu.be/bn7oI</a></p> <p> </p> <p><strong>How to access DEMoS</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1v6GaCVyNcib5v802t2uXHYOioqIkoBQ-/view?usp=share_link</p>
Hate speech and personal attack dataset in French social media
<p>This dataset contains 29109 French tweet ids and corresponding annotations for Hate Speech label and 39109 French tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
Hate speech and personal attack dataset in Spanish social media
<p>This dataset contains 37688 Spanish tweet ids and corresponding annotations for Hate Speech label and 37688 Spanish tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
Hate speech and personal attack dataset in German social media
<p>This dataset contains 43735 German tweet ids and corresponding annotations for Hate Speech label and 43734 German tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
Hate speech and personal attack dataset in English social media
<p>This dataset contains 92022 English tweet ids and corresponding annotations for Hate Speech label and 90892 English tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
Hate speech and personal attack dataset in Greek social media
<p>This dataset contains 61481 Greek tweet ids and corresponding annotations for Hate Speech label and 61481 Greek tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language
Open the record for dataset details and reuse information.
Rehabilitation of post-stroke aphasia by a single protocol targeting phonological, lexical, and semantic deficits with speech output tasks
<p>This study assessed the effectiveness of a novel rehabilitation protocol (PHOLEXSEM), focused on PHonological, SEmantic, and LExical deficits, aiming at improving lexical retrieval, and, generally, spoken output. The study is published here: https://doi.org/10.23736/S1973-9087.24.08576-9</p> <p><span>These data cannot be made publicly available because they include sensitive information, but can be made available to interested researchers upon reasonable request. Please, forward your request to Prof. Nadia Bolognini (n.bolognini@auxologico.it).</span></p>
data set related to article Neural Changes Induced by a Speech Motor Treatment in Childhood Apraxia of Speech A Case Series
<p>This record contains raw data related to article Neural Changes Induced by a Speech Motor Treatment in Childhood Apraxia of Speech A Case Series</p>
ProHaters - Proactive Profiling of Hate Speech Spreaders
<p>In the framework of the ProHaters - Proactive Profiling of Hate Speech Spreaders project funded by CDTI under grant IDI-20210776, the language resources for stance, polarization, hate speech, and bots have been created and annotated.</p> <p>These datasets cover English, German, and Spanish, including different dialects of English and Spanish.</p> <p>The datasets were created based on publicly available data collected from open sources and/or social media platforms, following their redistribution policies and the General Data Protection Regulation.</p> <p>The created datasets will be used for report writing and building planned tools and prototypes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.