Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo36/100

Intonational Speech Prosody Encoding in Human Auditory Cortex

<p>This dataset contains data and results associated with the manuscript, "Intonational Speech Prosody Encoding in Human Auditory Cortex", as well as code used to analyze the data and generate the figures of the manuscript. </p> <p>intonatang-2017.7.17.tar.gz contains the entire project, including neural data, stimulus sound files, analysis code, and documentation. The code and documentation can also be viewed on Github at https://github.com/ChangLabUcsf/intonatang.</p> <p>We additionally included each block of neural data and a zipped file containing the experimental, acoustic stimuli as separate files in this dataset. The data comprise three experiment types, "Speech", "Non-speech control", and "Non-speech missing f0 control". The neural data files are named with a subject identification number and a block number.</p> <p>Speech:</p> <ol> <li>EC113_B13</li> <li>EC113_B20</li> <li>EC113_B21</li> <li>EC118_B3</li> <li>EC118_B7</li> <li>EC118_B13</li> <li>EC122_B30</li> <li>EC122_B40</li> <li>EC122_B43</li> <li>EC122_B53</li> <li>EC123_B4</li> <li>EC123_B5</li> <li>EC123_B10</li> <li>EC125_B13</li> <li>EC125_B1044</li> <li>EC129_B10</li> <li>EC129_B16</li> <li>EC129_B37</li> <li>EC131_B47</li> <li>EC131_B48</li> <li>EC137_B7</li> <li>EC137_B10</li> <li>EC142_B36</li> <li>EC142_B37</li> <li>EC143_B9</li> <li>EC143_B11</li> <li>EC143_B13</li> </ol> <p>Non-speech control:</p> <ol> <li>EC122_B33</li> <li>EC122_B45</li> <li>EC123_B11</li> <li>EC123_B16</li> <li>EC125_B30</li> <li>EC129_B40</li> <li>EC129_B42</li> <li>EC131_B54</li> <li>EC131_B59</li> </ol> <p>Non-speech missing f0 control:</p> <ol> <li>EC137_B9</li> <li>EC137_B11</li> <li>EC142_B38</li> <li>EC142_B40</li> <li>EC143_B10</li> <li>EC143_B12</li> <li>EC143_B14</li> </ol> <p>These .mat files contain the following variables: </p> <ul> <li> badTimeSegments - (n_badTimeSegments x 2) <ul> <li>This variable contains manually marked time segments containing epileptiform, electrical, or movement artifacts. Each row indicates one bad time segment, with the start time and end time in seconds.</li> </ul> </li> <li>bcs - (n_bcs) <ul> <li>This array contains manually marked bad channels. These channels from the ECoG grid either had continuous epileptiform activity or signal indistinguishable from noise. The channels are indexed from 0.</li> </ul> </li> <li>ECXXX_BXX_hg_100Hz - (n_chans x n_timepoints) <ul> <li>This variable contains the mean high-gamma analytic amplitude signal for each channel, sampled at 100Hz. The mean is taken across 8 bands between 70-150Hz. The variable name contains the subject number, ECXXX, and block number BXX. </li> </ul> </li> <li>ECXXX_BXX_log_hg_100Hz - (n_chans x n_timepoints) <ul> <li>This variable contains the mean of the natural logarithm of the high-gamma analytic amplitude signal for each channel. The log is taken for each of the 8 bands between 70-150Hz and then averaged.</li> </ul> </li> <li>experiment <ul> <li>This variable holds the experiment type and is either "Speech", "Non-speech control", or "Non-speech missing f0 control".</li> </ul> </li> <li>sentence_numbers - (n_trials) <ul> <li>The integers in this array are the sentence number condition for each trial in this block. The sentence number conditions depend on the experiment type.</li> </ul> <ol> <li>For the "Speech" experiment, the four sentences indicated by 1, 2, 3, and 4 are "Humans value genuine behavior", "Movies demand minimal energy", "Lawyers give a relevant opinion", and "Reindeer are a visual animal".</li> <li>For the "Non-speech control" experiment, the sentence number conditions indicate which sentence from the main experiment the amplitude contour for the control stimuli came from. A sentence number of 5 means that the amplitude contour was flat. </li> <li>For the "Non-speech missing f0 control", the sentence number holds information about the composition of the stimulus (which harmonics were present), whether noise was added, and how much the pitch range was stretched.  <ul> <li>0: 4h + 5h + 6h, no noise, stretch = 1</li> <li>1: f0 + 2h + 3h, no noise, stretch = 1</li> <li>2: 4h + 5h + 6h, noise, stretch = 1</li> <li>3: 4h + 5h + 6h, noise, stretch = 0.5</li> <li>4: 4h + 5h + 6h, noise, stretch = 2</li> </ul> </li> </ol> </li> <li>sentence_types - (n_trials) <ul> <li>The sentence type is the intonation contour condition. Across all experiment types, a sentence type of 1 is Neutral, 2 is Question, 3 is Emphasis 1, and 4 is Emphasis 3.</li> </ul> </li> <li>speakers - (n_trial) <ul> <li>The speaker conditions depend on the experiment. <ul> <li>The speaker condition for the "Speech" experiment is an integer between 1 and 3. 1 is the low-formant, low-pitch male speaker. 2 is the high-formant, high-pitch female speaker. 3 is the low-formant, high-pitch female speaker. The absolute pitch values of speakers 2 and 3 match, while the formant values of speaker 1 and 3 match.</li> <li>The speaker condition for both of the two non-speech experiments are either 1 or 2. 1 means low absolute pitch (male) and 2 means high absolute pitch (female).</li> </ul> </li> </ul> </li> <li>stims - (n_trials) <ul> <li>This array holds the stimulus name that was played for each trial. The names refer to the wav files in the tokens, tokens_nonspeech, and tokens_missing_f0 folders, for the "Speech", "Non-speech control", and "Non-speech missing f0 control" experiments, respectively.</li> </ul> </li> <li>times - (n_trials) <ul> <li>This array contains the onset times of each trial in seconds.</li> </ul> </li> </ul>

opencc-by-sa-4.0Jul 2017View details →
zenodo36/100

Relative Transfer Matrix for Low SNR Speech Separation from Noisy Sources in Reverberant Rooms

<p>This folder contains the supplementary audio files for the paper "Relative Transfer Matrix for Low SNR Speech Separation from Noisy Sources in Reverberant Rooms" submitted to <em>The Journal of the Acoustical Society of America</em>.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Corona speeches mini corpus

<div> <div> <div> <p>This data collection (mini corpus) collects fifteen speeches of three speakers, Emmanuel Macron, Pedro S&aacute;nchez, and Angela Merkel, with five speeches per speaker. Four of their speeches are from between March and June, 2020, and one speech per speaker is from October or November, 2020. The speeches share important parallels in content and the speakers have similar intentions.</p> </div> </div> </div>

opencc-by-nc-nd-4.0Dec 2023View details →
zenodo36/100

The data and code for "Original Speech and Its Echo are Segregated and Separately Processed in the Human Brain"

<p>This dataset is associated with the manuscript "Original Speech and Its Echo are Segregated and Seperately Processed in the Human Brain", and provides the preprocessed MEG response (resampling to 100 Hz), auditory stimulus, individual quantitative observations underlying the data summarized in figures, and the analysis codes.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

German Parliamentary Speeches

<div> <div>These datasets are part of the thesis entitled "A natural language processing analysis of parliamentary speeches in the German Bundestag from 1949 to 2023 with regards to gender equality and women's politics". The focus of the thesis lies on contrasting the thematic preferences of men and women, as well as their connection to women's political issues based on their speeches and highlighting their historical development. Furthermore, the speakers&rsquo; reception in parliament is analyzed and the gender-specific distinguishability of speech style is verified. The data is available in a pandas DataFrame format.</div> <div>&nbsp;</div> <div> <div> <div>The plenary protocol data has been extracted from the following two sources:</div> <div> <ul> <li>Blaette, Andreas (2017): GermaParl. Corpus of Plenary Protocols of the German Bundestag. TEI files, availables at:&nbsp;<a href="https://github.com/PolMine/GermaParlTEI">https://github.com/PolMine/GermaParlTEI</a></li> </ul> </div> <div> <ul> <li>Deutscher Bundestag. Open Data - Plenarprotokolle der 20. Wahlperiode und Stammdaten aller Abgeordneten seit 1949.&nbsp;<a href="https://www.bundestag.de/services/opendata">https://www.bundestag.de/services/opendata</a></li> </ul> </div> <div>&nbsp;</div> <div>The following datasets are contained:</div> <div> <ul> <li><strong>data_merged_revised.pkl:</strong> DataFrame containing the extracted and cleaned data from the sources mentioned above&nbsp;</li> <li><strong>data_topics_revised.pkl</strong>: DataFrame containing the data from the sources mentioned above and the topic distributions after topic modelling</li> <li><strong>data_undersampled_processed.pkl</strong>: Undersampled and preprocessed dataset which can be used for training classification models</li> </ul> </div> </div> </div> </div>

opencc-by-sa-4.0Mar 2024View details →
zenodo36/100

CpAug: Refining Copy-Paste Augmentation for Speech Anti-Spoofing

<p>Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof&rsquo; 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof&rsquo; 21LA.</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Data for: Speech naturalness detection and language representation in the dog brain

<p>Abstract<br> Family dogs are exposed to a continuous flow of human speech throughout their lives. However, the extent of their abilities in speech perception is unknown. Here, we used functional magnetic resonance imaging (fMRI) to test speech detection and language representation in the dog brain. Dogs (n = 18) listened to natural speech and scrambled speech in a familiar and an unfamiliar language. Speech scrambling distorts auditory regularities specific to speech and to a given language, but keeps spectral voice cues intact. We hypothesized that if dogs can extract auditory regularities of speech, and of a familiar language, then there will be distinct patterns of brain activity for natural speech vs. scrambled speech, and also for familiar vs. unfamiliar language. Using multivoxel pattern analysis (MVPA) we found that bilateral auditory cortical regions represented natural speech and scrambled speech differently; with a better classifier performance in longer-headed dogs in a right auditory region. This neural capacity for speech detection was not based on preferential processing for speech but rather on sensitivity to sound naturalness.<br> Furthermore, in case of natural speech, distinct activity patterns were found for the two languages in the secondary auditory cortex and in the precruciate gyrus; with a greater difference in responses to the familiar and unfamiliar languages in older dogs, indicating a role for the amount of language exposure. No regions represented differently the scrambled versions of the two languages, suggesting that the activity difference between languages in natural speech reflected sensitivity to language-specific regularities rather than to spectral voice cues. These findings suggest that separate cortical regions support speech naturalness detection and language representation in the dog brain.<br> <br> This dataset contains</p> <ul> <li>Raw data (four functional runs and Matlab logs n = 18)</li> <li>Dog brain&nbsp;template</li> <li>Stimuli (natural and scrambled speech in Hungarian and Spanish)&nbsp;</li> <li>MVPA main results maps (Speech detection and Language discrimination, n = 18)</li> <li>GLM All sounds &gt; Silence contrast (n = 18)&nbsp;</li> </ul>

opencc-by-3.0Jan 2022View details →
zenodo36/100

Corpus of Distorted Speech

<p>This dataset contains audio stimuli (wav format, 16 kHz) used to test the perception of speech that has undergone artificial distortion. The corpus consists of 240 sentences for each of eight types of distortion. The original (unmodified) sentences come from the public-domain Sharvard Corpus, sentence numbers 241-480.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

A Kannada Emotional Speech Dataset

<p>There was&nbsp;no emotional speech dataset available in Kannada. This was&nbsp;a limiting factor for research in the Kannada-speaking world. I introduce&nbsp;a Kannada emotional&nbsp;speech dataset and give&nbsp;details about its design and content. This dataset contains six different sentences, pronounced by thirteen people (four&nbsp;male and nine&nbsp;female), in five&nbsp;basic emotions plus one neutral emotion. They are all Kannada speakers.&nbsp;The dataset has been contributed by volunteers and the recordings were not made in a controlled environment. The dataset contains a total of 468&nbsp;audio samples, each one in a separate audio file. The file naming convention is as follows:&nbsp;AA-EE-SS.wav where AA is a two-character field that gives the actor number (01 to 13), EE is a two-character field that indicates the emotion number (01 to 06), and SS is a two-character field that gives the sentence number (01 to 06). This dataset is freely available under a Creative Commons license.</p> <p>Gender and age of each of the 13 people who contributed to the dataset</p> <ul> <li>01, F, 45</li> <li>02, F, 20</li> <li>03, F, 21</li> <li>04, M, 47</li> <li>05, F, 48</li> <li>06, M, 20</li> <li>07, F, 20</li> <li>08, F, 45</li> <li>09, F, 21</li> <li>10, F, 12</li> <li>11, F, 12</li> <li>12, M, 17</li> <li>13, M, 26</li> </ul> <p>Identification characters&nbsp;for the emotions in the dataset</p> <ul> <li>01, Anger</li> <li>02, Sadness</li> <li>03, Surprise</li> <li>04, Happiness</li> <li>05, Fear</li> <li>06, Neutral</li> </ul> <p>Identification characters&nbsp;for the sentences in the dataset</p> <ul> <li>01,&nbsp;ರೋಗಿಗಳಿಗೆ ಚಿಕಿತ್ಸೆ ನೀಡಿ ಉಪಚರಿಸುವುದು</li> <li>02,&nbsp;ಈ ಕಾದಂಬರಿಯು ಎರಡು ಪಾತ್ರಗಳನ್ನು ಒಳಗೊಂಡಿದೆ</li> <li>03,&nbsp;ಖಾಸಗಿ ವಿಮಾನಗಳೆಂದೂ ಸಾರ್ವಜನಿಕ ವಿಮಾನಗಳೆಂದೂ ವಿಂಗಡಿಸಿದ್ದಾರೆ</li> <li>04,&nbsp;ನಿಮ್ಮನ್ನು ಬೀಟಿಯಾಗಿ ಬಹಳ ಸಂತೋಶ ಆಯಿತು</li> <li>05,&nbsp;ರಾಮನ ಎಡಬಲ ದಲ್ಲಿ ಸೀತಾ ಲಕ್ಷ್ಮಣ ರಿದ್ದಾರೆ</li> <li>06,&nbsp;ಕನ್ನಡವನ್ನು ಕಲಿಯಬೆಕು</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

PB2007 French acoustic-articulatory speech database

<p><strong>PB2007 acoustic-articulatory speech dataset</strong></p> <p>Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.&nbsp;<br> Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France<br> <br> LICENSE:<br> ========<br> This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license</p> <p><br> CREDITS - ATTRIBUTION:<br> ======================<br> If using this dataset, please cite one of the following studies (all of them exploit this dataset)&nbsp;<br> - Ben Youssef, A., Badin, P., Bailly, G. &amp; Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.<br> - Ben Youssef, A., Badin, P. &amp; Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose &amp; S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.<br> - Hueber T., Bailly G., Badin P., Elisei F., &quot;Speaker Adaptation of an Acoustic-Articulatory Inversion Model<br> using Cascaded Gaussian Mixture Regressions&quot;, Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757.&nbsp;</p> <p>&nbsp;</p> <p>DATA FILES DESCRIPTION:<br> =======================<br> /_seq/:&nbsp;<br> &nbsp;&nbsp; &nbsp;Electro-magnetic Articulography data, recorded at 100Hz<br> &nbsp;&nbsp; &nbsp;Sensors :<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR01 : LT_x (lower incisor, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR02 : tip_x (tongue tip, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR03 : mid_x (tongue dorsum, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR04 : bck_x (tongue back, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR05 : LL_vis_x (lower lips, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR06 : UL_vis_x (upper lips, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR07 : LT_z (lower incisor, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR08 : tip_z (tongue tip, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR09 : mid_z (tongue dorsum, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR10 : bck_z (tongue back, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR11 : LL_vis_z (lower lips, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR12 : UL_vis_z (upper lips, z coordinate)</p> <p>/_wav16:&nbsp;<br> &nbsp;&nbsp; &nbsp;subject audio signal, synchronized with the EMA data<br> &nbsp;&nbsp; &nbsp;Format: PCA wav, 16kHz, 16bits</p> <p>/_lab: phonetic segmentation using the following set<br> __ (long pause), _ (short pause), a, e^ (as in &quot;lait&quot;), e (as in &quot;bl&eacute;&quot;), i, y (as in &quot;voiture&quot;), u (as in &quot;loup&quot;), o^ (as in &quot;pomme&quot;),x (as in &quot;pneu&quot;), x^ (as in &quot;coeur&quot;), a~ (as in &quot;flan&quot;), e~ (as in &quot;in&quot;), x~ (as in &quot;un&quot;), o~ (as in &quot;mon&quot;), p, t, k, f, s, s^ (as in &quot;CHat&quot;), b, d, g, v, z, z^ (as in &quot;les Gens&quot;), m, n, r, l, w, h, j, o, q (schwa)<br> &nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Preprocessed Data and Pretrained Models for Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings

<p>This is preprocessed data and pretrained models from two of our papers:</p> <p>&quot;Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings,&quot; by Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. (ICASSP 2020)<br> <a href="https://arxiv.org/abs/1910.10838">https://arxiv.org/abs/1910.10838</a></p> <p>&nbsp;&quot;Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis,&quot; by Erica Cooper, Xin Wang, Yi Zhao, Yusuke Yasuda, and Junichi Yamagishi. (arXiv)&nbsp;<a href="https://arxiv.org/abs/2011.04839">https://arxiv.org/abs/2011.04839</a></p> <p>This data is meant to be used with our open-source implementation, which can be found here: &nbsp;https://github.com/nii-yamagishilab/multi-speaker-tacotron</p> <p>More information about the directory structure and how to use the data can be found in the READMEs on GitHub.</p>

openother-openMar 2022View details →
zenodo36/100

Principal component decomposition of acoustic and neural representations of time-varying pitch reveals adaptive efficient coding of speech covariation patterns

<p>Data and scripts reported by Llanos and colleagues in their&nbsp;Brain and Language&nbsp;study&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Synthetic Göttingen Sentence Test material created with a text-to-speech system

<p>The speech material of the German G&ouml;ttingen Sentence Test [1] was synthesized using a commercial text-to-speech system (Acapela Cloud Service). More details will be found in [2].</p> <p>Files:</p> <p>goesa_synth_female.zip<br> contains all 200 sentences with a synthetic female voice and the corresponding speech adjusted noise, which was generated by superimposing the speech material 30 times according to [3].</p> <p>goesa_synth_male.zip<br> contains all 200 sentences with a synthetic male voice and the corresponding speech adjusted noise, which was generated by superimposing the speech material 30 times according to [3].</p> <p>&nbsp;</p>

opencc-by-nc-4.0May 2022View details →
zenodo36/100

Chichewa speech dataset

<p>This is a speech dataset for the language &quot;Chichewa&quot; spoken in Malawi, Zambia, Zimbabwe and Mozambique. The zipped folder contains the following:</p> <ul> <li>chich_speech_audio_files-this is a folder with all the audio files and transcripts&nbsp;</li> <li>chich_speech_dataset_documentation.pdf: this is the documentation for the dataset</li> <li>chich_speech_dataset_metadata.csv: this file has basic metadata about each audio file.</li> </ul>

opencc-by-4.0May 2022View details →
zenodo36/100

Dataset, test programs and analysis scripts for the related paper "Effects of reverberation on speech intelligibility in noise for hearing-impaired listeners"

<p>This dataset contains the data, test programs and analyses scripts used for a study submitted as a stage 2 registered report for Royal Society Open Science.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Materials used in "Measuring audio-visual speech intelligibility under dynamic listening conditions using virtual reality"

<p>The materials in this record are the audio and video files, together with various configuration files, used by the &quot;SEAT&quot; software in the study described in</p> <p>Moore, Green, Brookes &amp; Naylor (2022)&nbsp;&quot;Measuring audio-visual speech intelligibility under dynamic listening conditions using virtual reality&quot;</p> <p>They are shared in this form so that the experiment may be reproduced.&nbsp; For any other use please contact the authors to obtain the original database(s) from which these materials are derived.</p> <p>The materials were created to be compatible with v0.3 of SEAT, which is available from&nbsp;<a href="https://github.com/ImperialCollegeLondon/sap-elospheres-audiovisual-test/releases/tag/v0.3">GitHub</a>. Note that the materials must be placed at <code>C:\seat_experiments\cafe_AV.</code></p>

opencc-by-nc-nd-4.0Aug 2022View details →
zenodo36/100

Training a Text-to-Speech System for Dialectal Arabic with a Focus on the Iraqi Dialect

<p>This research introduces a novel approach to Text-to-Speech (TTS) synthesis, focusing on the phonetic complexities of Arabic dialects, with particular emphasis on the Iraqi dialect. While existing Arabic speech corpora provide substantial coverage of Modern Standard Arabic (MSA), they fall short in capturing the phonetic richness of regional dialects. To address this gap, we utilized Nawar Halabi's Arabic Speech Corpus as a base dataset and enriched it with custom-recorded samples of the Iraqi dialect, incorporating distinctive phonemes such as گ ,ڤ ,پ ,چ ,ۆ ,ڵ ,ێ, and using the Tatweel character (ـ) as a vowel. Our approach, powered by the FastPitch model and a customized phonetiser, successfully synthesized the Iraqi dialect while also demonstrating adaptability to other Arabic dialects, including Egyptian, Khaliji, Syrian, and more. The results of this research signify a promising advancement in Arabic TTS technology, expanding its scope to authentically represent the diverse linguistic landscape of the Arabic-speaking world.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Dimensional coding of native and non-native speech categories: data and code

<p><strong>Data</strong>: (i) fMRI nii files, (ii) categorization responses , and (iii) stimuli</p> <p><strong>Code</strong>: (i) Dimensionality analyses, (ii) statistical analyses, and (iii) MVPC analyses, and (iv) data visualization</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones: Audio Examples

<p>This upload contains audio examples for the preprint "Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones".</p> <p>The audio files correspond to subplots of the spectrogram shown in the default preview, starting from the upper left corner (subplot 0) to the upper right corner (subplot 1) and so on.</p> <h2>Abstract</h2> <p>Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user. However, due to the hearable occluding the ear canal, the in-ear microphone mostly records body-conducted speech, typically suffering from band-limitation effects and amplification at low frequencies. Since the occlusion effect is determined by the ratio between the air-conducted and body-conducted components of own voice, the own voice transfer characteristics between the outer face of the hearable and the in-ear microphone depend on the speech content and the individual talker. In this paper, we propose a speech-dependent model of the own voice transfer characteristics based on phoneme recognition, assuming a linear time-invariant relative transfer function for each phoneme. We consider both individual models as well as models averaged over several talkers. Experimental results based on recordings with a prototype hearable show that the proposed speech-dependent model enables to simulate in-ear signals more accurately than a speech-independent model in terms of technical measures, especially under utterance mismatch and talker mismatch. Additionally, simulation results show that talker-averaged models generalize better to different talkers than individual models.</p> <p>&nbsp;</p> <p>The examples are also available here: <a href="https://m-ohlenbusch.github.io/own_voice_modeling_examples/" target="_blank" rel="noopener">https://m-ohlenbusch.github.io/own_voice_modeling_examples/</a></p> <p>Arxiv preprint: <a href="https://arxiv.org/abs/2310.06554">https://arxiv.org/abs/2310.06554</a></p>

opencc-by-nc-nd-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record