Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo32/100

RAVDESS Speech 16K

<p>This is the 16kHz-version of RAVDESS Speech dataset. The original version is sampled at 48kHz. See link below for detail.</p> <p><a href="../records/1188976">https://zenodo.org/records/1188976&nbsp;</a></p> <p>Why 16k version?</p> <ul> <li>Most deep learning model requires 16kHz of sampling rate.</li> <li>Lower sampling rate saves disk (70 MB vs 200 MB).</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo32/100

United Nations General Assembly Meeting Speeches 1993 -- 2023

<p>This is a collection of discussions/speeches of UN General Assembly meetings 1993-2023 sourced from the UN Digital Library (https://digitallibrary.un.org/search?ln=en&amp;cc=Speeches). Each speech is labelled by speaker name, organisation/country represented by a speaker, speaker title, date of speech and speech content.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 1)

<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJun 2024View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 2)

<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJul 2024View details →
zenodo32/100

The impact of exploiting spectro-temporal context in computational speech segregation

<p>The experimental data&nbsp;from the study:</p> <p>https://asa.scitation.org/doi/10.1121/1.5020273</p> <p>Group 1 contains results, masks and audio from the models of the 16 GMM component segregation system<br> Group 2 contains results, masks and audio from the models of the 64 GMM component segregation system</p> <p>There are three folders:</p> <p>Audio:<br> The CLUE sentences that were used for the listener study</p> <p>IBM = Ideal Binary Mask, UP = UnProcessed, EBM = Estimated Binary Mask.&nbsp;</p> <p>The IBM and UP are stored in one of the configuration folders (Front-end), that is:</p> <p>Audio\Group1\Front-end\icra_01_10sec_matched\UP<br> Audio\Group1\Front-end\icra_01_10sec_matched\IBM<br> Audio\Group1\Front-end\icra_01_10sec_matched\EBM</p> <p>Results:<br> The computed metrics for group 1 &amp; 2 as well as Word Recognition Scores (WRSs) from the listener study</p> <p>BinaryMasks:</p> <p>a priori SNR masks, IBMs and EBMs from group 1 and 2.</p> <p><br> Developed with Matlab R2016a.</p>

opencc-by-4.0Dec 2017View details →
zenodo32/100

Raw Data from - Speech Auditory Brainstem Responses in Adult Hearing Aid Users: Effects of Aiding and Background Noise, and Prediction of Behavioral Measures

<p><em><strong>Folder and Data Description for dataset of:</strong></em></p> <p><strong>Speech Auditory Brainstem Responses in Adult Hearing Aid Users: Effects of Aiding and Background Noise, and Prediction of Behavioral Measures</strong></p> <p>Ghada BinKhamis, Antonio Elia Forte, Tobias Reichenbach, Martin O&rsquo;Driscoll, and Karolina Kluk</p> <p><strong>Please site the&nbsp;paper when using this dataset</strong>&nbsp;(DOI: 10.1177/2331216519848297)</p> <p>&nbsp;</p> <p><strong>Shared dataset is as follows:</strong></p> <ul> <li><strong>Behavioral data is in the excel spread sheet entitled:</strong>&nbsp;&ldquo;BinKhamis_et_al_behavioral_data .xlsx&rdquo;<br> &nbsp;</li> <li><strong>Speech-ABRs&nbsp;(raw EEG (speech-ABR) data) are contained within the five &#39;zip&#39; folders.</strong></li> </ul> <p><strong>Description of the &ldquo;Speech-ABRs&rdquo; folders, subfolders, and raw EEG files:</strong></p> <p><strong>&ldquo;Speech-ABRs&rdquo; Folder Information:</strong></p> <ul> <li><strong>Each Speech_ABR&nbsp;folder</strong>&nbsp;contains subfolders from a subset of participants (e.g. Speech_ABR_1_20.zip contains data from participant number&nbsp;1 to participant number 20)</li> <li><strong>Subfolders:</strong> <ul> <li>Each subfolder starts with the participant code: e.g. HA1, HA2, HA3, HA4, HA5, &hellip;, HA98</li> <li>Next is the background condition: noise or quiet</li> <li>Next is whether recordings were: aided or unaided</li> </ul> </li> <li><strong>Example subfolder names:</strong> <ul> <li><strong><em>HA1 noise aided:</em></strong> participant number 1, aided speech-ABRs in background noise</li> <li><strong><em>HA4 noise unaided:</em></strong> participant number 4, unaided speech-ABRs in background noise</li> <li><strong><em>HA55 quiet aided:</em></strong> participant number 55, aided speech-ABRs in quiet</li> <li><strong><em>HA97 quiet unaided</em></strong>: participant number 97, unaided speech-ABRs in quiet</li> </ul> </li> <li>Each participant has 4 subfolders for the four recording conditions (aided quiet, aided noise, unaided quiet, unaided noise) <ul> <li><strong>Each subfolder contains four &lsquo;.mat&rsquo; files, &lsquo;.mat&rsquo; file names:</strong> <ul> <li>Each &lsquo;.mat&rsquo; file starts with the participant code: e.g. HA1, HA2, HA3, HA4, HA5, &hellip;, HA98</li> <li>Next is the stimulus: 40 da</li> <li>Next is &lsquo;unaided&rsquo; only if recordings were without HA</li> <li>Next is &lsquo;noise&rsquo; only if the background condition was noise</li> <li>Next is the stimulus polarity: <ul> <li>&lsquo;Pos&rsquo; for positive/standard</li> <li>&lsquo;Neg&rsquo; for negative (reversed polarity stimulus)</li> </ul> </li> <li>And finally the test ear and recording number for that polarity <ul> <li>R1 is the first recording from the right ear, R2 is the second recording from the right ear</li> <li>L1 is the first recording from the left ear, L2 is the second recording from the left ear</li> </ul> </li> <li><strong>Example &lsquo;.mat&rsquo; file name:</strong> <ul> <li><strong><em>HA1 40 da Neg Noise R1.mat: </em></strong>participant number 1, aided speech-ABR in response to the 40 ms [da], reversed stimulus polarity, in background noise, right ear recording number 1.</li> <li><strong><em>HA4 40 da unaided Pos Noise L2.mat:</em></strong> participant number 1, unaided speech-ABR in response to the 40 ms [da], standard stimulus polarity, left ear recording number 2.</li> <li><strong><em>HA7 40 da Neg R2.mat:</em></strong> participant number 7, aided speech-ABR in response to the 40 ms [da], reversed stimulus polarity, right ear recording number 2.</li> <li><strong><em>HA10 40 da unaided Pos Noise L1.mat:</em></strong> participant number 10, unaided speech-ABR in response to the 40 ms [da], standard stimulus polarity, in background noise, left ear recording number 1.</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p><strong>File Information:</strong></p> <p><strong>Description of &lsquo;.mat&rsquo; files that can be accessed and processed using MATLAB (MathWorks):</strong></p> <p>Each &lsquo;.mat&rsquo; file is a structure that contains the following fields:</p> <ul> <li>The first nine fields are informational, for example:</li> <li><strong><em>xunits</em></strong>: &lsquo;s&rsquo; indicates that the recording time window is in seconds, conversion to milliseconds would be required to plot the data in milliseconds</li> <li><strong><em>start: </em></strong>&lsquo;0&rsquo; indicates that both stimulus and recording start at 0 seconds</li> <li><strong><em>points:</em></strong> <strong>2200</strong> is the number of sample points</li> <li><strong><em>chans:</em></strong> 2 is the number of channels <ul> <li><em>Right ear:</em> channel 2, <em>Left ear:</em> channel 1</li> </ul> </li> <li><strong><em>frames:</em></strong> 2500 is the number of epochs</li> <li>The last filed <strong>&lsquo;values&rsquo;</strong> is what contains the raw EEG data (2200x2x2500) <ul> <li><strong>2200 </strong>is the number of samples</li> <li><strong>2 </strong>is the number of channels (channel one is recorded from the left ear lobe (A1) and channel two is from the right ear lobe (A2))</li> <li><strong>2500 </strong>is the number of epochs <ul> <li>Stimulus starts at 0 seconds per epoch, pre-stimulus baseline may be extracted from the end of each epoch (i.e. before the next stimulus).</li> <li>Data are in Volts; conversion to <strong>&mu;Volts </strong>(multiply by 1000) is required.</li> </ul> </li> </ul> </li> </ul> <p><strong>Date of data collection:&nbsp;</strong>October 2017 to July 2018</p>

opencc-by-4.0Apr 2019View details →
zenodo32/100

HaterNet a system for detecting and analyzing hate speech in Twitter

<p>This dataset consists of&nbsp; two corpuses used in the paper &quot;Detecting and analyzing hate speech in Twitter: HaterNet a system in the Spanish prevention of hate crime office&quot;. A first one based on tweets collected at different random dates between February 2017 and December 2017 with a final size of 2 million tweets. A second one with&nbsp;6,000 tweets labeled as described in the paper as hate containing or not.</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

Participant Data and Code for "Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video"

<p>All code produced to analyze the participant response data and the participant response data itself are included in this repository.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers

<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss&rsquo; Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Spatial–temporal dynamics of gesture–speech integration: a simultaneous EEG-fMRI study

<p>The semantic integration between gesture and speech (GSI) is mediated by the left posterior temporal sulcus/middle temporal gyrus (pSTS/MTG) and the left inferior frontal gyrus (IFG). Evidence from electroencephalography (EEG) suggests that oscillations in the alpha and beta bands may support processes at different stages of GSI. In the present study, we investigated the relationship between electrophysiological oscillations and blood-oxygen-level-dependent (BOLD) activity during GSI. In a simultaneous EEG-fMRI study, German participants (n = 19) were presented with videos of an actor either performing meaningful gestures in the context of a comprehensible German (GG) or incomprehensible Russian sentence (GR), or just speaking a German sentence (SG). EEG results revealed reduced alpha and beta power for the GG vs. SG conditions, while fMRI analyses showed BOLD increase in the left pSTS/MTG for GG &gt; GR &cap; GG &gt; SG. In time-window-based EEG-informed fMRI analyses, we further found a positive correlation between single-trial alpha power and BOLD signal in the left pSTS/MTG, the left IFG, and several sub-cortical regions. Moreover, the alpha-pSTS/MTG correlation was observed in an earlier time window in comparison to the alpha-IFG correlation, thus supporting a two-stage processing model of GSI. Our study shows that EEG-informed fMRI implies multiple roles of alpha oscillations during GSI, and that the method is a best candidate for multidimensional investigations on complex cognitive functions such as GSI.</p>

opencc-by-4.0Jun 2021View details →
dryad32/100

Data from: Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech

People routinely hear and understand speech at rates of 120–200 words per minute [1, 2]. Thus, speech comprehension must involve rapid, online neural mechanisms that process words' meanings in an approximately time-locked fashion. However, in the context of continuous speech, electrophysiological evidence for such time-locked processing has been lacking. Whilst valuable insights into the semantic processing of speech have been provided by the "N400 component" of the event-related potential [3-6], this literature has been dominated by paradigms using incongruous words within specially constructed sentences, and may not accurately reflect natural, narrative speech comprehension. Building on the discovery that cortical activity "tracks" the dynamics of running speech [7-9], and psycholinguistic work both demonstrating [10-12] and modeling [13-15] how context rapidly impacts on word processing, we describe a new approach for deriving an electrophysiological correlate of natural speech comprehension. We used a computational model [16] to quantify the meaning carried by each word based on how semantically dissimilar it was to its preceding context and then regressed this quantity against electroencephalographic (EEG) data recorded from subjects as they listened to narrative speech. This produced a prominent negativity at a time-lag of 200–600 ms on centro-parietal EEG channels, characteristics common to the N400. Applying this approach to EEG datasets involving time-reversed speech, cocktail party attention and audiovisual speech-in-noise demonstrated that this response was very sensitive to whether or not subjects understood the speech they heard. These findings demonstrate that, when successfully comprehending natural speech, the human brain responds to the contextual semantic content of each word in a relatively time-locked fashion.

opencc-zeroDec 2017View details →
zenodo32/100

Sample Stimuli Presentation for a Remote Speech Segmentation Study

<p>A video used in a remote eye-tracking study to assess speech segmentation abilities in children with Down Syndrome.</p>

openother-ncJul 2021View details →
zenodo32/100

DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 2

<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data.&nbsp;</p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels.&nbsp;</p> <p><strong>Arxiv:</strong>&nbsp;https://arxiv.org/abs/2109.07931</p> <p>&nbsp;</p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 2).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong>&nbsp;https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong>&nbsp;https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong>&nbsp;https://zenodo.org/record/5501697</li> </ul>

openodc-bySep 2021View details →
zenodo32/100

DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 1

<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data.&nbsp;</p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels.&nbsp;</p> <p><strong>Arxiv:</strong>&nbsp;https://arxiv.org/abs/2109.07931</p> <p>&nbsp;</p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 1).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong>&nbsp;https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong>&nbsp;https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong>&nbsp;https://zenodo.org/record/5501697</li> </ul>

openodc-bySep 2021View details →
zenodo32/100

Human vocalization corpus: recordings of infant-directed and adult-directed speech and song in 21 societies

<p>This repository&nbsp;contains a corpus of 1615 audio recordings of speech and song collected in 21 societies,&nbsp;first reported&nbsp;in Moser et al. (2020;&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2020.04.09.032995v5">bioRxiv</a>) and later published in Hilton &amp; Moser et al. (2022; <a href="https://doi.org/10.1038/s41562-022-01410-x">Nature Human Behaviour</a>).&nbsp;For assistance using any of this, contact Cody Moser (<a href="mailto:cmoser2@ucmerced.edu">cmoser2@ucmerced.edu</a>), Courtney Hilton (<a href="mailto:courtney.hilton@auckland.ac.nz">courtney.hilton@auckland.ac.nz</a>), and Samuel Mehr (<a href="mailto:mehr@hey.com">mehr@hey.com</a>).</p> <p>Two versions of the audio are included: raw audio (`IDS-corpus-raw.zip`) and audio that was edited to prepare the recordings for automatic acoustic feature extraction (`IDS-corpus-edited.zip`). `IDS-textGrids.zip` contains annotation files from Praat&#39;s silence detection method, which were manually reviewed for accuracy. These files are used with the audio extraction scripts associated with the project (see code linked in paper) to build the edited audio files.</p> <p>`IDS-fieldsites.csv` contains some fieldsite-level metadata; additional metadata is in the Supplementary Information of the paper.</p> <p>In the two .zip archives, filenames have the format XXXYYZ.wav,&nbsp;where &quot;XXX&quot; is a fieldsite code, &quot;YY&quot; is a participant number, and &quot;Z&quot; is a vocalization type.</p> <p>Fieldsite codes are:</p> <blockquote> <p>MBE: Mbendjele BaYaka<br> HAD: Hadza<br> NYA: Nyangatom<br> TOP: Toposa<br> BEJ: Beijing<br> JEN: Jenu Kurubas<br> MEN: Mentawai Islanders<br> KRA: Krakow<br> LIM: Rural Poland<br> TUR: Turku<br> USD: San Diego<br> TOR: Toronto<br> VAN: Tannese Vanuatuans<br> PNG: Enga<br> WEL: Wellington<br> ARA: Arawak<br> TSI: Tsimane<br> SPA: S&aacute;para &amp; Achuar<br> QUE: Quechua<br> ACO: Afrocolombians<br> MES: Colombian Mestizos</p> </blockquote> <p>Participant numbers are padded integers, starting with 01, and are unique within fieldsites.</p> <p>Vocalization types are:</p> <blockquote> <p>A: infant-directed song<br> B: infant-directed speech<br> C: adult-directed song<br> D: adult-directed speech&nbsp;</p> </blockquote> <p>In a few cases, participants vocalized in a different language than was expected, given the primary language of their fieldsite (e.g., when the participant was multilingual, or if they sang a song that contains multiple languages, as in The Beatles&#39; &quot;Michelle&quot;). The file `IDS-unexpectedLanguages.csv` at&nbsp;<a href="https://github.com/themusiclab/infant-speech-song/blob/main/data/IDS-unexpectedLanguages.csv">https://github.com/themusiclab/infant-speech-song/blob/main/data/IDS-unexpectedLanguages.csv</a> contains an inventory of these examples from the English-speaking fieldsites. This issue only affects&nbsp;a small minority of the recordings, as it&nbsp;was typically avoided by the researchers collecting the recordings.&nbsp;</p>

openApr 2020View details →
zenodo32/100

Textless Speech-to-Music Retrieval Using Emotion Similarity (Pretrained Model)

<p>- github:&nbsp;https://github.com/SeungHeonDoh/speech-to-music&nbsp;<br> - Demo :&nbsp;https://seungheondoh.github.io/speech-to-music-demo/<br> - ArXiv : update soon</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Neural preprocessed data associated to "Decoding grasp and speech signals from the cortical grasp circuit in a tetraplegic human"

<p>This dataset is composed of electrophysiology data from a tetraplegic human participant implanted with three 96 channel Utah arrays (Blackrock) in the supramarginal gyrus (SMG), ventral premotor cortex (PMV) and&nbsp;somatosensory cortex.&nbsp;The dataset includes preprocessed (spike sorted)&nbsp;firing rate data for 96 recorded channels, as described in Wandelt et al (2022), &quot;Decoding grasp and speech signals from the cortical grasp circuit in a tetraplegic human&quot; published in Neuron (<a href="https://doi.org/10.1016/j.neuron.2022.03.009">10.1016/j.neuron.2022.03.009</a>).</p> <p>To run the code associated with the processed data, download it here https://zenodo.org/record/6330179.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain

<p>Dear User,</p> <p>We are thrilled to introduce our latest release - the <strong>RescueSpeech</strong>&nbsp;audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.</p> <p>The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.</p> <p>1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.</p> <p>2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.</p> <p>By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.<br> &nbsp;</p>

opencc-by-nc-4.0Jun 2023View details →
zenodo32/100

List of European Parliament plenary speeches selected for the corpus together with speakers' names (Nov-2014 to Apr-2018); examples of collocations of "refugee(s)", "refugié(s)", "Flüchtling(e)" and "menekült(ek)"

<p>This data relates to the article &quot;Hidden Patterns in interpreted xenophobic discourse in the European Parliament&quot; [in print].</p> <p>The data contains a chronological list of the plenary debates from which the speeches were taken as well as the names of each speaker. It also also contains examples of verbs collocating with the term <em>refugee(s</em>), <em>refugi&eacute;(s)</em>, <em>Fl&uuml;chtling(e)</em> and <em>menek&uuml;lt(ek)</em> in the four language versions (English, French, German, Hungarian). These collocations were identified by the author of the paper.</p> <p>The speeches were downloaded from the Multimedia Center on the European Parliament&#39;s pubilc website: <a href="https://multimedia.europarl.europa.eu/en/home">https://multimedia.europarl.europa.eu/en/home</a>.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Distinct Dimensional Encoding of Speech in the Dorsal and Ventral Auditory Streams

<p>Data and code for&nbsp;&quot;Distinct Dimensional Encoding of Speech in the Dorsal and Ventral Auditory Streams&quot;</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record