Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Speech and Noise Corpora for Pitch Estimation of Human Speech
<p><em>Part of the dissertation <a href="http://localhost:8000/index.html">Pitch of Voiced Speech in the Short-Time Fourier Transform: Algorithms, Ground Truths, and Evaluation Methods</a>.<br> © 2020, Bastian Bechtold. All rights reserved.</em></p> <p> </p> <p>This dataset contains common speech and noise corpora for evaluating fundamental frequency estimation algorithms as convenient <a href="https://jbof.readthedocs.io/en/latest/">JBOF</a> dataframes. Each corpus is available freely on its own, and allows redistribution:</p> <ul> <li><a href="http://www.festvox.org/cmu_arctic/">CMU-ARCTIC</a> (<em>BSD license) [1]</em></li> <li><a href="http://www.cstr.ed.ac.uk/research/projects/fda/">FDA</a> (<em>free to download)</em> [2]</li> <li><a href="https://lost-contact.mit.edu/afs/nada.kth.se/dept/tmh/corpora/KeelePitchDB/">KEELE</a> (<em>free for noncommercial use</em>) [3]</li> <li><a href="http://www.cstr.ed.ac.uk/research/projects/artic/mocha.html">MOCHA-TIMIT</a> (<em>free for noncommercial use</em>) [4]</li> <li><a href="https://www.spsc.tugraz.at/databases-and-tools/ptdb-tug-pitch-tracking-database-from-graz-university-of-technology.html">PTDB-TUG</a> (<em>ODBL license</em>) [5]</li> <li><a href="http://www.speech.cs.cmu.edu/comp.speech/Section1/Data/noisex.html">NOISEX</a> (<em>free to download</em>) [7]</li> <li><a href="https://research.qut.edu.au/saivt/databases/qut-noise-databases-and-protocols/">QUT-NOISE</a> (<em>CC-BY-SA license</em>) [8]</li> </ul> <p>Additionally, this dataset contains <em>PDAs-0.0.1-py3-none-any.whl</em>, a Python ≥ 3.6 module for Linux, containing several well-known fundamental frequency estimation algorithms:</p> <ul> <li>AUTOC [9]</li> <li>AMDF [10]</li> <li><a href="http://www2.ece.rochester.edu/projects/wcng/code.html">BANA</a> [11]</li> <li>CEP [12]</li> <li><a href="https://github.com/marl/crepe">CREPE</a> [13]</li> <li><a href="http://www.kki.yamanashi.ac.jp/~mmorise/world/english/">DIO</a> [14]</li> <li><a href="http://web.cse.ohio-state.edu/pnl/software.html">DNN</a> [15]</li> <li><a href="https://github.com/LvHang/pitch">KALDI</a> [16]</li> <li>MAPS</li> <li><a href="http://www.seas.ucla.edu/spapl/shareware.html">MBSC</a> [17]</li> <li><a href="https://github.com/jkjaer/fastF0Nls">NLS</a> [18]</li> <li><a href="http://www.ee.ic.ac.uk/hp/staff/dmb/voicebox/voicebox.html">PEFAC</a> [19]</li> <li><a href="https://github.com/praat/praat">PRAAT</a> [20]</li> <li><a href="http://www.speech.kth.se/wavesurfer/links.html">RAPT</a> [21]</li> <li><a href="http://labrosa.ee.columbia.edu/projects/SAcC/">SACC</a> [22]</li> <li><a href="http://www.seas.ucla.edu/spapl/weichu/safe/">SAFE</a> [23]</li> <li><a href="https://mathworks.com/matlabcentral/fileexchange/1230">SHR</a> [24]</li> <li>SIFT [25]</li> <li><a href="https://github.com/covarep/covarep">SRH</a> [26]</li> <li><a href="https://github.com/HidekiKawahara/legacy_straight">STRAIGHT</a> [27]</li> <li><a href="http://www.cise.ufl.edu/~acamacho/english/curriculum.html">SWIPE</a> [28]</li> <li><a href="http://www.ws.binghamton.edu/zahorian/yaapt.htm">YAAPT</a> [29]</li> <li><a href="http://audition.ens.fr/adc/">YIN</a> [30]</li> </ul> <p>The algorithms are included in their native programming language (Matlab for BANA, DNN, MBSC, NLS, NLS2, PEFAC, RAPT, RNN, SACC, SHR, SRH, STRAIGHT, SWIPE, YAAPT, and YIN; C for KALDI, PRAAT, and SAFE; Python for AMDF, AUTOC, CEP, CREPE, MAPS, and SIFT), and adapted to a common Python interface. AMDF, AUTOC, CEP, and SIFT are our partial re-implementations as no original source code could be found.</p> <p>All algorithms have been released as open source software, and are covered by their respective licenses.</p> <p>All of these files are published as part of my dissertation, "<a href="https://bastibe.github.io/Dissertation-Website/">Pitch of Voiced Speech in the Short-Time Fourier Transform: Algorithms, Ground Truths, and Evaluation Methods</a>", and in support of the <a href="https://github.com/bastibe/Replication-Dataset-Scripts">Replication Dataset for Fundamental Frequency Estimation</a>.</p> <p>References:</p> <ol> <li>John Kominek and Alan W Black. CMU ARCTIC database for speech synthesis, 2003.</li> <li>Paul C Bagshaw, Steven Hiller, and Mervyn A Jack. Enhanced Pitch Tracking and the Processing of F0 Contours for Computer Aided Intonation Teaching. In EUROSPEECH, 1993.</li> <li>F Plante, Georg F Meyer, and William A Ainsworth. A Pitch Extraction Reference Database. In Fourth European Conference on Speech Communication and Technology, pages 837–840, Madrid, Spain, 1995.</li> <li>Alan Wrench. MOCHA MultiCHannel Articulatory database: English, November 1999.</li> <li>Gregor Pirker, Michael Wohlmayr, Stefan Petrik, and Franz Pernkopf. A Pitch Tracking Corpus with Evaluation on Multipitch Tracking Scenario. page 4, 2011.</li> <li>John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren, and Victor Zue. TIMIT Acoustic-Phonetic Continuous Speech Corpus, 1993.</li> <li>Andrew Varga and Herman J.M. Steeneken. Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recog- nition systems. Speech Communication, 12(3):247–251, July 1993.</li> <li>David B. Dean, Sridha Sridharan, Robert J. Vogt, and Michael W. Mason. The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithms. Proceedings of Interspeech 2010, 2010.</li> <li>Man Mohan Sondhi. New methods of pitch extraction. Audio and Electroacoustics, IEEE Transactions on, 16(2):262—266, 1968.</li> <li>Myron J. Ross, Harry L. Shaffer, Asaf Cohen, Richard Freudberg, and Harold J. Manley. Average magnitude difference function pitch extractor. Acoustics, Speech and Signal Processing, IEEE Transactions on, 22(5):353—362, 1974.</li> <li>Na Yang, He Ba, Weiyang Cai, Ilker Demirkol, and Wendi Heinzelman. BaNa: A Noise Resilient Fundamental Frequency Detection Algorithm for Speech and Music. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(12):1833–1848, December 2014.</li> <li>Michael Noll. Cepstrum Pitch Determination. The Journal of the Acoustical Society of America, 41(2):293–309, 1967.</li> <li>Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello. CREPE: A Convolutional Representation for Pitch Estimation. arXiv:1802.06182 [cs, eess, stat], February 2018. arXiv: 1802.06182.</li> <li>Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications. IEICE Transactions on Information and Systems, E99.D(7):1877–1884, 2016.</li> <li>Kun Han and DeLiang Wang. Neural Network Based Pitch Tracking in Very Noisy Speech. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(12):2158–2168, Decem- ber 2014.</li> <li>Pegah Ghahremani, Bagher BabaAli, Daniel Povey, Korbinian Riedhammer, Jan Trmal, and Sanjeev Khudanpur. A pitch extraction algorithm tuned for automatic speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, pages 2494–2498. IEEE, 2014.</li> <li>Lee Ngee Tan and Abeer Alwan. Multi-band summary correlogram-based pitch detection for noisy speech. Speech Communication, 55(7-8):841–856, September 2013.</li> <li>Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, and Søren Holdt Jensen. Fast fundamental frequency estimation: Making a statistically efficient estimator computationally efficient. Signal Processing, 135:188–197, June 2017.</li> <li>Sira Gonzalez and Mike Brookes. PEFAC - A Pitch Estimation Algorithm Robust to High Levels of Noise. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(2):518—530, February 2014.</li> <li>Paul Boersma. Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound. In Proceedings of the institute of phonetic sciences, volume 17, page 97—110. Amsterdam, 1993.</li> <li>David Talkin. A robust algorithm for pitch tracking (RAPT). Speech coding and synthesis, 495:518, 1995.</li> <li>Byung Suk Lee and Daniel PW Ellis. Noise robust pitch tracking by subband autocorrelation classification. In Interspeech, pages 707–710, 2012.</li> <li>Wei Chu and Abeer Alwan. SAFE: a statistical algorithm for F0 estimation for both clean and noisy speech. In INTERSPEECH, pages 2590–2593, 2010.</li> <li>Xuejing Sun. Pitch determination and voice quality analysis using subharmonic-to-harmonic ratio. In Acoustics, Speech, and Signal Processing (ICASSP), 2002 IEEE International Conference on, volume 1, page I—333. IEEE, 2002.</li> <li>Markel. The SIFT algorithm for fundamental frequency estimation. IEEE Transactions on Audio and Electroacoustics, 20(5):367—377, December 1972.</li> <li>Thomas Drugman and Abeer Alwan. Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics. In Interspeech, page 1973—1976, 2011.</li> <li>Hideki Kawahara, Masanori Morise, Toru Takahashi, Ryuichi Nisimura, Toshio Irino, and Hideki Banno. TANDEM-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation. In Acous- tics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on, pages 3933–3936. IEEE, 2008.</li> <li>Arturo Camacho. SWIPE: A sawtooth waveform inspired pitch estimator for speech and music. PhD thesis, University of Florida, 2007.</li> <li>Kavita Kasi and Stephen A. Zahorian. Yet Another Algorithm for Pitch Tracking. In IEEE International Conference on Acoustics Speech and Signal Processing, pages I–361–I–364, Orlando, FL, USA, May 2002. IEEE.</li> <li>Alain de Cheveigné and Hideki Kawahara. YIN, a fundamental frequency estimator for speech and music. The Journal of the Acoustical Society of America, 111(4):1917, 2002.</li> </ol>
CUB-200 Speech captions (Part-II)
<p>Speech captions of CUB-200 database. These speech captions are synthesized by Tacotron-v2 according to the original textual descriptions.</p>
Nganasan and Kamas Speech Recognition Models
<p>These are the models trained in our paper</p> <p>Partanen, N., Hämäläinen, M. and Klooster, T. (2020) Speech Recognition for Endangered and Extinct Samoyedic languages. In <em>Proceedings of the 34th Pacific Asia Conference on Language, Information and Computation</em>.</p> <p>See the readme for more</p> <p><sub>Based on corpora from</sub></p> <p><sub>Gusev, Valentin; Klooster, Tiina; Wagner-Nagy, Beáta. 2019. "INEL Kamas Corpus." Version 1.0. Publication date 2019-12-15. http://hdl.handle.net/11022/0000-0007-DA6E-9. Archived in Hamburger Zentrum für Sprachkorpora. In: Wagner-Nagy, Beáta; Arkhipov, Alexandre; Ferger, Anne; Jettka, Daniel; Lehmberg, Timm (eds.). The INEL corpora of indigenous Northern Eurasian languages.</sub></p> <p><sub>Maria Brykina, Valentin Gusev, Sandor Szeverényi, and Beáta Wagner-Nagy. 2018. Nganasan spoken language corpus (nslc). Archived in Hamburger Zentrumfür Sprachkorpora. Version 0.2. Publication date, 12.</sub></p>
Cross-modal matching of audio-visual German and French fluent speech in infancy
<p>The present study examined when and how the ability to cross-modally match audio-visual fluent speech develops in 4.5-, 6- and 12-month-old German-learning infants. In Experiment 1, 4.5- and 6-month-old infants’ audio-visual matching ability of native (German) and non-native (French) fluent speech was assessed by presenting auditory and visual speech information sequentially, that is, in the absence of temporal synchrony cues. The results showed that 4.5-month-old infants were capable of matching native as well as non-native audio and visual speech stimuli, whereas 6-month-olds perceived the audio-visual correspondence of native language stimuli only. This suggests that intersensory matching narrows for fluent speech between 4.5 and 6 months of age. In Experiment 2, auditory and visual speech information was presented simultaneously, therefore, providing temporal synchrony cues. Here, 6-month-olds were found to match native as well as non-native speech indicating facilitation of temporal synchrony cues on the intersensory perception of non-native fluent speech. Intriguingly, despite the fact that audio and visual stimuli cohered temporally, 12-month-olds matched the non-native language only. Results were discussed with regard to multisensory perceptual narrowing during the first year of life.</p>
FORMATION AND DEVELOPMENT OF GRAMMATICALLY CORRECT SPEECH IN CHILDREN WITH SPEECH DEFECTS
Open the record for dataset details and reuse information.
Supplementary material to 'Automatic Identification of Hate Speech – A Case-Study of Alt-Right YouTube Videos'
<p>The associated files have been created for and is analysed in a fortcoming article entitled <em>Automatic Identification of Hate Speech – A Case-Study of Alt-Right YouTube Videos'. </em>The material is divided into six tables as follows:</p> <table> <tbody> <tr> <td>Sentence top 5%</td> <td>The 19th 20-quantile predicted most hateful sentences</td> </tr> <tr> <td>Sentence bottom 5%</td> <td>The bottom 20-quantile predicted moste hatefull sentences (the least likely to contain hatespeech)</td> </tr> <tr> <td>Paragraphs</td> <td>Prediction and annotation of paragraphs</td> </tr> <tr> <td>Video top 10%</td> <td>Titles of the top decile predicted hateful videos</td> </tr> <tr> <td>Video bottom 10%</td> <td>Titles of the bottom decile predicted hateful videos</td> </tr> <tr> <td>Video bottom 10% - Alt right</td> <td>Titles of the bottom decile predicted hateful videos without History</td> </tr> </tbody> </table> <p>The data is uploaded in two formats:</p> <p><strong>Excel file: </strong>Automatic_Detection_of_Hate_Speech_a_Case-Study_of_Alt-Right_Videos.xlsx contains all six tables in one file, with a supplementary <em>codebook. </em></p> <p><strong>Tab Separated Values (TSV):</strong> Each file correspond to a single sheet from the excel file, and are named accordingly. UTF-8 Encoded.<strong><br></strong></p>
THE ANTHROPOLOGICAL STUDIES ON THE IMPACT OF GENDER DIFFERENCE ON SPEECH
<p><span>It should be noted that the issues of gender difference on speech also carried out in the fields of two directions, anthropology and dialectology, in contrast to sociolinguistic models. Anthropologists have mainly studied the influence of gender difference on language through social behavior at a certain level. However, dialectologists have studied the difference of speech not only in terms of gender, but also in the group of speakers of urban and rural areas. </span></p>
VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language
<p> Speech BCIs based on implanted electrodes hold significant promise for enhancing spoken communication through high temporal resolution and invasive neural sensing. Despite the potential, acquiring such data is challenging due to its invasive nature, and publicly available datasets, particularly for tonal languages, are limited. In this study, we introduce <em>VocalMind</em>, a stereotactic electroencephalography (sEEG) dataset focused on Mandarin Chinese, a tonal language. This dataset includes sEEG-speech parallel recordings from three distinct speech modes, namely vocalized speech, mimed speech, and imagined speech, at both word and sentence levels, totaling over one hour of intracranial neural recordings related to speech production. This paper also presents a baseline model as the reference model for future studies, at the same time, ensuring the integrity of the dataset. The diversity of tasks and the substantial data volume provide a valuable resource for developing advanced algorithms for speech decoding, thereby advancing BCI research for spoken communication.</p>
The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition
<p>The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda.</p> <p>The corpus of 155 hours is publicly available online under the Creative Commons BY-NC-ND 4.0 license. The dataset release is comprised of the following:</p> <ol> <li>20 hours of human transcribed radio speech. The audio is 16kHZ, mono channel and with 16 bit rate. </li> <li>Two CSV files for the 20-hour human transcribed dataset - cleaned.csv contains cleaned transcripts and uncleaned.csv contains uncleaned transcripts. The uncleaned transcripts contain extra speech details included in tags like [laughter] for laughter, and [um] for filler pauses, which speaker is talking, where each speaker is assigned an identifier A or B.</li> <li>The transcription guide used to transcribe the radio dataset.</li> <li> A multi-speaker untranscribed dataset of 6 hours of radio data. 1.4 hours of women voices and 4.6 hours of men voices. Each audio is a ten-seconds clip with a single speaker.</li> <li> 135 hours of multi-speaker untranscribed radio data.</li> </ol> <p><strong>NOTE: You can read and cite our paper published in the </strong><a href="http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.208.pdf"><strong>Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)</strong></a> The Dataset is published under Creative Commons BY-NC-ND 4.0 license and in order for us to monitor who is using it for the right license we request that you reach out to us officially. </p>
ISSUES OF SPEECH ETIQUETTE IN UZBEK FOLKLORE PROVERBS
Open the record for dataset details and reuse information.
Multilingual test set for language identification and speech recognition from European Parliament recordings
<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings. </p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes. </p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>
Synthetic Speech Dataset
<p>This contains the speech data synthesized for the following paper:</p> <p>"Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models" by Loren Lugosch, Brett Meyer, Derek Nowrouzezahrai, and Mirco Ravanelli</p>
The effect of spatial energy spread on sound image size and speech intelligibility [dataset]
<p>This dataset contains the data and supplementary material for the publication titled "The effect of spatial energy spread on sound image size and speech intelligibility" published in the Journal of the Acoustical Society of America.</p> <p>This repository contains the results from the three experiments as well as a figure showing the adaptive tracks of the adaptive procedure and the estimated psychometric functions in experiment 3. The subject identifiers are common across the three experiments.</p> <p><strong>Experiment 1: </strong></p> <p>Factors: <em>Subject/Listener, ambisonics order, audio/stimulus type, spatial location of stimulus, room condition, repetition</em></p> <p>Measured variables: <em>Source image size, source direction (azimuth), source distance</em></p> <p><strong>Experiment 2:</strong></p> <p>Factors: <em>Subject/Listener, ambisonics order, interferer location, room condition</em></p> <p>Measured variable: <em>Speech reception threshold (SRT)</em></p> <p><strong>Experiment 3:</strong></p> <p>Factors: <em>Subject/Listener, ambisonics order, repetition</em></p> <p>Measured variable: <em>Speech reception angle (SRA)</em></p> <p>The figure (exp3_adaptiveTracks.png) shows the adaptive tracks in experiment 3. Each panel shows one of the ambisonics orders. The x-axis is the trial number and the y-axis the separation angle between target and interfering talkers. Each color indicates one subject.</p> <p>The figure (exp3_psychometricFcns.png) shows the estimated psychometric functions in experiment 3. Each panel shows one of the ambisonics orders. The x-axis is the separation angle between target and interfering talkers and the y-axis is the percent correct words. Each color indicates one subject. The black line indicates the median over the subjects and repetitions and the black cross the SRA predicted with this method.</p>
Data from: The visual speech head start improves perception and reduces superior temporal cortex responses to auditory speech
Visual information about speech content from the talker's mouth is often available before auditory information from the talker's voice. Previously we demonstrated that audiovisual speech selectively enhances activity in regions of the early visual cortex representing the mouth of the talker (Ozker et al., 2018b). Here we examined perceptual and neural responses to words with and without this visual head start. For both types of words, perception was enhanced by viewing the talker's face, but the enhancement was significantly greater for words with a head start. Neural responses were measured from electrodes implanted over auditory association cortex in the posterior superior temporal gyrus (pSTG) of epileptic patients. The presence of visual speech suppressed responses to auditory speech, more so for words with a visual head start. We suggest that the head start inhibits representations of incompatible auditory phonemes, increasing perceptual accuracy and decreasing total neural responses. Together with previous work showing visual cortex modulation (Ozker et al., 2018b) these results from pSTG demonstrate that multisensory interactions are a powerful modulator of activity throughout the speech perception network.
Audio-tactile speech EEG Dataset
<p>Dataset setting out to investigate behavioural and EEG responses to continuous speech coupled with tactile pulses. Support functions to process data and derive neural responses to both sensory features can be found on <a href="https://github.com/phg17/Pieeg">this repository</a>. The raw behavioural and EEG data processed to this end is provided here, as well as the stimuli used in the behavioural and electrophysiological experiments.</p> <p>Note : For a library that features encoding and decoding from EEG data but less specific about the data format, please see <a href="https://github.com/phg17/sPyEEG">this repository</a>.</p> <p><strong>#Introduction</strong></p> <p>This dataset contains the behavioural response to one speech comprehension task in which subjects received both audio and tactile stimuli. It also contains the EEG response to 2 sessions of audio-book listening accompanied by those same tactile stimuli. Finally, the stimuli for both experiments are also available in the dataset. There were 18 subjects.</p> <p>Naming Conventions:</p> <ul> <li>The subjects IDs are : 'al', 'yr', 'alio', 'chap', 'sep', 'phil', 'lad', 'calco', 'hudi', 'nima', 'ogre', 'raqu', 'nikf', 'zartan', 'naga', 'miya', 'elios', 'olio'</li> <li>The labels of the condition generally have three fields : type, delay and correlation. The type indicates the nature of stimuli (audio, tactile or audio-tactile). The delay indicates the delay between the two stimuli in the case there are two, it is fixed at 0 otherwise. The correlation indicates whether the tactile stimuli is correlated to the audio or not, it is at True by default, except for the sham stimuli described in the paper.</li> </ul> <p><strong>#Content</strong></p> <p>The dataset contains a .zip named after each subject : these are the eeg recordings for each of them. It also contains files called Stimuli_EEG and Stimuli_Behavioural which contains the stimuli for each experiment. Finally Behavioural_Experiment.zip contains the data from the behavioural experiment. Their respective content is described in more details down below. </p> <ul> <li><strong>Behavioural Stimuli: </strong>This folder contains files named 'sent_4k__xxxyyy.zzz'. sent_4k refers to sent stimuli at 44.1kHz. xxx refers to the number of the sentence. yyy.zzz refers to the nature and format of the file: TextGrid, wav, audio.npy, phone.npy. TextGrid are text files indicating the timing of each phoneme and words, they can be vizualized alongside the audio file using <a href="https://www.fon.hum.uva.nl/praat/">Praat</a>. .wav are simple audio files at 44.1kHz. audio.npy is just the Audio file put in an .npy format and resampled at the internal rate of the TDT RX8 system used to send the stimuli (39062kHz). phone.npy are the tactile pulses at the onset of vowels and resampled accordingly.</li> </ul> <p> </p> <ul> <li><strong>EEG Stimuli:</strong> This folder contains files named 'Odin_x_y_zzz.vvv'. x and y refers to the chapter in the story and the part of that chapter. zzz.vvv refers to the nature and format. Most notations are the same as for the <strong>Behavioural Stimuli</strong> file. In addition, there are sparse matrices containing the timing of all phonetic features and phonemes in the .npz format. Finally, this folder also contains comprehension questions and their answers for each part of each chapter.</li> </ul> <p> </p> <ul> <li><strong>Behavioural Experiment: </strong>This folder contains two types of files : 'xxx_subjective.pkl' and 'xxx_total_behavioural.csv' where xxx refers to the subject ID. 'xxx_subjective.pkl' is a dictionary containing the average comfort score for each condition. 'xxx_total_behavioural.csv' is a table containing for each trial: the sentence number (same notation as in <strong>Behavioural Stimuli</strong>), the type, delay and correlation nature of the stimuli, the real sentence and the ratio of understood words of the subjects.</li> </ul> <p> </p> <ul> <li><strong>EEG Experiment</strong>: All of the 'xxx.zip' files where xxx refers to the subject ID. Each folder contains one folder for each of the two sessions (xxx_EEG_1 and xxx_EEG_2). Inside of those, .eeg, .vhdr and .vmrk files can be found. These correspond to the recorded data. More information on these file formats can be found on the <a href="https://www.fieldtriptoolbox.org/getting_started/brainvision/">Brian vision webpage</a>. It also contains a .csv file containing for each segment of ~2mn30: the chapter and part sent to the subject, the type, delay and correlation nature of the stimuli, the onset of the target tactile stimuli (for the behavioural task in the tactile-only condition) and answers of the subjects to the comprehension questions (for the behavioural task in the audio-only condition). Finally, the order of chapter and parameters in a .npy format is also provided although redundant with the .csv file.</li> </ul> <p> </p> <p><strong>#EEG Data Format</strong></p> <p>We recorded continuous EEG from the participants over an hour at 1kHz. </p> <p>In addition to the 63 channels, there is also Stimtrack channel labelled as 'Sound' which was tracking an addition of the auditory and tactile stimuli and allowed to track the sent stimuli as well as the alignment. The corresponding files can be found using the provided .csv file.</p> <p> The .vmrk file provided contains the timing of the triggers sent at the start of each 2mn30 segment and can help with tracking the progress in the experiment. Overall .vmrk and stimtrack are used together to align the EEG response with the corresponding stimuli and checking the state of the experiment.</p> <p>The subjects could also press a button to signal that they detected the tactile pattern they were ask to find, this response can be found as an additional channel called 'Button'.</p> <p><strong>#Note </strong></p> <p>Dysfunctions of the equipment were detected and should be taken into account:</p> <ul> <li>The button experienced malfunction when 'chap' was using it during the tactile pattern detection task</li> <li>The channels 'CPz', 'FCz', 'FC6', 'AF3', 'AF7' needed to be interpolated.</li> <li>For some subjects, the recording was interrupted and therefore the recording was separated in half. The missing sessions were as follow: 'yr' : 12, 'naga' : 0, 13, 'elios' : 14. Note that in those cases, you might need to modify the trigger timing extracted from the .vmrk file.</li> </ul>
Toward General Speech Restoration With VoiceFixer - Testsets
<p>This is the test set for the paper VoiceFixer: Toward General Speech Restoration with Neural Vocoder<em>.</em></p> <p>If you found this dataset helpful, please consider citing: </p> <blockquote> <pre>@article{liu2021voicefixer, title={VoiceFixer: Toward General Speech Restoration with Neural Vocoder}, author={Liu, Haohe and Kong, Qiuqiang and Tian, Qiao and Zhao, Yan and Wang, DeLiang and Huang, Chuanzeng and Wang, Yuxuan}, journal={arXiv preprint arXiv:2109.13731}, year={2021} }</pre> </blockquote>
Supplementary Materials for article entitles 'Nominal and verbal syntax in translation and interpreting. Evidence from English speeches made in the European Parliament and their German translations and interpretations', submitted to Languages
<p>The Supplementary Materials contain the transcriptions (raw and tagged, 'sample_df.tsv'), the data frames with POS-frequencies, with ('pos_freqs_PART_split.tsv') and without ('pos_freqs.tsv') the PART-split in the German data, the data frames for the identification of interpreters ('voice_embeddings.csv'), an R-script for the statistical analysis and generation of plots ('stats.R') as well as the plots (folder 'plots').</p>
Synthetic phrase-based test material created with a text-to-speech system
<p>New speech material consisting of phrases was synthesized using a commercial text-to-speech system (Acapela Cloud Service). More details will be found in [1].</p> <p>Files:</p> <p>syntheticPhrases.zip<br>contains all 772 phrases with a synthetic female voice and the corresponding speech adjusted noise, which was generated by superimposing the speech material 30 times according to [2].</p>
Chief Logan Elm-Logan's Speech
The Logan Elm that stood near Circleville in Pickaway County, Ohio, was one of the largest American elm trees (Ulmus americana) recorded. The 65-foot-tall (20 m) tree had a trunk circumference of 24 feet (7.3 m) and a crown spread of 180 feet (55 m).[1] Weakened by Dutch Elm Disease, the tree died from storm damage in 1964.[1] The Logan Elm State Memorial commemorates the site and preserves various associated markers and monuments.[1] According to tradition, Chief Logan of the Mingo tribe delivered a passionate speech at a peace-treaty meeting under this elm in 1774,[1] said to be the most famous speech ever given by a Native American,[citation needed] now known as "Logan's Lament": Source: Objaverse 1.0 / Sketchfab
Sensory-Motor Integration for Speech Rehabilitation in Patients with Post-stroke Aphasia
ClinicalTrials.gov study NCT04433351. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.