Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo40/100

TIMIT-TTS: a Text-to-Speech Dataset for Synthetic Speech Detection

<p>With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple that malicious users can create unpleasant situations with minimal effort. Also, forged media are getting more and more complex, with manipulated videos (e.g., deepfakes where both the visual and audio contents can be counterfeited) that are taking the scene over still images.<br> The multimedia forensic community has addressed the possible threats that this situation could imply by developing detectors that verify the authenticity of multimedia objects. However, the vast majority of these tools only analyze one modality at a time.<br> This was not a problem as long as still images were considered the most widely edited media, but now, since manipulated videos are becoming customary, performing monomodal analyses could be reductive. Nonetheless, there is a lack in the literature regarding multimodal detectors (systems that consider both audio and video components). This is due to the difficulty of developing them but also to the scarsity of datasets containing forged multimodal data to train and test the designed algorithms.</p> <p>In this paper we focus on the generation of an audio-visual deepfake dataset.<br> First, we present a general pipeline for synthesizing speech deepfake content from a given real or fake video, facilitating the creation of counterfeit multimodal material. The proposed method uses Text-to-Speech (TTS) and Dynamic Time Warping (DTW) techniques to achieve realistic speech tracks. Then, we use the pipeline to generate and release TIMIT-TTS, a synthetic speech dataset containing the most cutting-edge methods in the TTS field. This can be used as a standalone audio dataset, or combined with DeepfakeTIMIT and VidTIMIT video datasets to perform multimodal research. Finally, we present numerous experiments to benchmark the proposed dataset in both monomodal (i.e., audio) and multimodal (i.e., audio and video) conditions.<br> This highlights the need for multimodal forensic detectors and more multimodal deepfake data.</p> <ul> <li>For the initial version of TIMIT-TTS&nbsp;<strong>v1.0</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2209.08000</li> <li>TIMIT-TTS Database v1.0: https://zenodo.org/record/6560159</li> </ul> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words

<p>Hi,KIA dataset is a shared short Wakeup Word&nbsp;database focusing on perceived emotion in&nbsp;speech The dataset contains&nbsp;<strong>488 </strong>Wakeup Word&nbsp;speech.&nbsp;</p> <p>For more detailed information about the dataset, please refer to our paper:&nbsp;Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>:&nbsp;wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav`&nbsp;The first letter was used to express emotion.<br> &nbsp;</li> </ul> </li> <li><em><strong>annotation/</strong></em>:&nbsp;Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv&nbsp;</p> </li> <li> <p><em><strong>handcraft:</strong></em>&nbsp;Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em>&nbsp;wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p>&nbsp;</p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> &nbsp; title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> &nbsp; author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> &nbsp; booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> &nbsp; year={2022}<br> }<br> ```</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Non-acoustic speech sensing system based on flexible piezoelectric

<p>The non-acoustic speech sensing system based on flexible piezoelectric is designed to satisfy specific needs around testing device models&nbsp;(in high-noise, complex environments). The system collected vibration signals from the jaws of six males and five females containing ten&nbsp;different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and is able to achieve 93.7%&nbsp;recognition accuracy by calculation. In general, this paper provides a non-acoustic speech dataset for Mandarin, including the parts&nbsp;collected, the number of people collected, and the environment.</p> <p><br> The dataset is available at:</p> <p>https://doi.org/10.5281/zenodo.7095762</p> <p><br> The data descriptor paper with details of data collection and cleaning process is under submission. For proper citation of the manuscript,&nbsp;please refer to the latest version of this dataset which includes the details.</p> <p>This dataset and its descriptor paper were created by:</p> <p>Shiji Yuan, Ying Sun, Dezhi Zheng, Xinlei Chen,Ying Ding, Shuai Wang, Shangchun Fan</p> <p>For questions or suggestions, please e-mail Dezhi Zheng &lt;zhengdezhi@buaa.edu.cn&gt;</p> <p><br> <strong>Description:</strong><br> <br> Ten common words were chosen as the core of the vocabulary in this dataset. These ten command words can be used for commands in&nbsp;IoT or robotics applications: &quot;forward&quot;, &quot;backward&quot;, &quot;right&quot;, &quot;left&quot;, &quot;stop&quot;, &quot;up&quot;, &quot;down&quot;, &quot;draw&quot;, &quot;drop&quot;, and &quot;reset&quot;.</p> <p>The recording was carried on by software named Adobe Audition 2022. We set monophonic recording, 16-bit storage format, and 16 kHz&nbsp;sampling frequency before recording and saved the recorded voice in wav format. &nbsp;The dataset is provided with two storage rules, which&nbsp;are stored by subject number and command number as classification. In the first rule, the speech data of 11 subjects were stored in&nbsp;different folders with the subject serial number as the folder name. Each folder contains subfolders categorized by command. In the&nbsp;second rule, the speech data of ten commands are stored in different folders, and the names of the folders are the command contents.&nbsp;</p> <p><br> The subject number, command number and record order are given for each data entry. For example, the data obtained when subject 1&nbsp;recorded command 10 for the first time was labeled as &quot;1-10_1&quot;.</p> <p>After the data collection process, a filtering algorithm for automatic detection of low non-acoustic speech data was designed to remove problematic data that were very short or very quiet.The script of the data filtering algorithm is provided in this repository. &nbsp;</p> <p>For specific detail of the data filtering process, please refer to the script (speech data filtering algorithm in MATLAB) in this repository and the data descriptor paper.</p> <p>The dataset in this repository is the processed version. The raw dataset and removed audio files are not included in this repository.</p> <p><br> <br> <strong>File list:</strong></p> <p><br> Non-acoustic Speech Dataset.zip</p> <p>speech data filtering algorithm.zip</p> <p>Readme.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Non-acoustic Speech Dataset

<p><strong>Non-acoustic speech sensing system based on flexible piezoelectric</strong></p> <p><strong>Version 1.0.0&nbsp;&nbsp; &nbsp;</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;<br> &nbsp;</p> <p><br> This Read_Me.txt file briefly describes the non-acoustic speech dataset and instructions to access it.&nbsp;</p> <p>The non-acoustic speech sensing system based on flexible piezoelectric is designed to satisfy specific needs around testing device models (in high-noise, complex environments). The system collected vibration signals from the jaws of six males and five females containing ten different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and is able to achieve 93.7% recognition accuracy by calculation. In general, this paper provides a non-acoustic speech dataset for Mandarin, including the parts collected, the number of people collected, and the environment.</p> <p><br> The dataset is available at:</p> <p>https://10.5281/zenodo.7090120</p> <p>The data descriptor paper with details of data collection and cleaning process is under submission. For proper citation of the manuscript,&nbsp;please refer to the latest version of this dataset which includes the details.</p> <p>This dataset and its descriptor paper were created by:</p> <p>Shiji Yuan, Ying Sun, Dezhi Zheng, Xinlei Chen, Ying Ding,Shuai Wang, Shangchun Fan</p> <p>For questions or suggestions, please e-mail Dezhi Zheng &lt;zhengdezhi@buaa.edu.cn&gt;</p> <p><br> <strong>Description:</strong><br> <br> Ten common words were chosen as the core of the vocabulary in this dataset. These ten command words can be used for commands in&nbsp;IoT or robotics applications: &quot;forward&quot;, &quot;backward&quot;, &quot;right&quot;, &quot;left&quot;, &quot;stop&quot;, &quot;up&quot;, &quot;down&quot;, &quot;draw&quot;, &quot;drop&quot;, and &quot;reset&quot;.</p> <p>The recording software is Adobe Audition2022,which adopts monophonic recording, 16-bit storage format, 16 kHz sampling frequency, and the recorded voice is saved in wav&nbsp;format. The dataset is provided with two storage rules, which are stored by subject&nbsp;number and corpus number as classification. In the first rule, the speech data of 11 subjects were stored in different folders with the&nbsp;subject serial number as the folder name. Each folder contains subfolders categorized by corpus. In the second rule, the speech data&nbsp;of ten corpus are stored in different folders, and the names of the folders are the corpus contents. The subject number, corpus number&nbsp;and record order are given for each data entry. For example, the data obtained when subject one recorded corpus 10 for the first time&nbsp;was labeled as 1-10_1.</p> <p>After the data collection process, a filtering algorithm for automatic detection of low non-acoustic speech data is designed to remove&nbsp;problematic data that are very short or very quiet. The script of the data filtering algorithm is provided in this repository. &nbsp;</p> <p>For specific detail of the data filtering process, please refer to the script (speech data filtering algorithm in MATLAB) in this repository and the data descriptor paper.</p> <p>The dataset in this repository is the processed version. The raw dataset and removed audio files are not included in this repository.</p> <p><br> <br> <strong>File list:</strong><br> <br> Non-acoustic Speech Dataset.zip</p> <p>speech data filtering algorithm.zip</p> <p>Readme.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;</p> <p><br> &nbsp;&nbsp; &nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Non-acoustic speech sensing system based on flexible piezoelectric

<p>The non-acoustic speech sensing system based on flexible piezoelectric is designed to satisfy specific needs around testing device models&nbsp;(in high-noise, complex environments). The system collected vibration signals from the jaws of six males and five females containing ten&nbsp;different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and is able to achieve 93.7%&nbsp;recognition accuracy by calculation. In general, this paper provides a non-acoustic speech dataset for Mandarin, including the parts&nbsp;collected, the number of people collected, and the environment.</p> <p><br> The dataset is available at:</p> <p>https://doi.org/10.5281/zenodo.7185663</p> <p><br> The data descriptor paper with details of data collection and cleaning process is under submission. For proper citation of the manuscript,&nbsp;please refer to the latest version of this dataset which includes the details.</p> <p>This dataset and its descriptor paper were created by:</p> <p>Shiji Yuan, Ying Sun, Shuai Wang, Xinlei Chen,Ying Ding,Dezhi Zheng , Shangchun Fan</p> <p>For questions or suggestions, please e-mail Shuai Wang &lt;wangshuai@buaa.edu.cn&gt;</p> <p><br> <strong>Description:</strong><br> Ten common words were chosen as the core of the vocabulary in this dataset. These ten command words can be used for commands in&nbsp;IoT or robotics applications: &quot;forward&quot;, &quot;backward&quot;, &quot;right&quot;, &quot;left&quot;, &quot;stop&quot;, &quot;up&quot;, &quot;down&quot;, &quot;draw&quot;, &quot;drop&quot;, and &quot;reset&quot;.</p> <p>The recording was carried on by software named Adobe Audition2022. We set monophonic recording, 16-bit storage format, and 16 kHz&nbsp;sampling frequency before recording and saved the recorded voice in wav format. &nbsp;The dataset is provided with two storage rules, which&nbsp;are stored by subject number and command number as classification. In the first rule, the speech data of 11 subjects were stored in&nbsp;different folders with the subject serial number as the folder name. Each folder contains subfolders categorized by command. In the&nbsp;second rule, the speech data of ten commands are stored in different folders, and the names of the folders are the command contents.&nbsp;The subject number, command number and record order are given for each data entry. For example, the data obtained when subject 1&nbsp;recorded command 10 for the first time was labeled as &quot;1-10_1&quot;.</p> <p>After the data collection process, a filtering algorithm for automatic detection of low non-acoustic speech data was designed to remove problematic data that were very short or very quiet.The script of the data filtering algorithm is provided in this repository. &nbsp;</p> <p>For specific detail of the data filtering process, please refer to the script (speech data filtering algorithm in MATLAB) in this repository and the data descriptor paper.</p> <p>The dataset in this repository is the processed version. The raw dataset and removed audio files are not included in this repository.</p> <p><br> <br> <strong>File list:</strong><br> <br> Non-acoustic Speech Dataset.zip</p> <p>speech data filtering algorithm.zip</p> <p>Readme.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;</p> <p><br> &nbsp;&nbsp; &nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Raw Data from - Speech-ABRs in Cochlear Implant Recipients: Feasibility Study

<p><strong>Speech-ABRs in Cochlear Implant Recipients: Feasibility Study</strong></p> <p>Ghada BinKhamis, Emanuele Perugia, Martin O&rsquo;Driscoll, Karolina Kluk</p> <p><strong>Please site the&nbsp;paper when using this dataset</strong>&nbsp;(DOI:&nbsp;10.1080/14992027.2019.1619100)</p> <p>&nbsp;</p> <p><strong>Participant Recordings (BinKhamis et al Speech_ABR in CI raw data.zip)</strong></p> <ul> <li><strong>Description of raw EEG (speech-ABR) data main folder, subfolders, and raw EEG files</strong> <ul> <li><strong>Folder Information</strong> <ul> <li><strong>Main folder:</strong> Contains 12 subfolders with the raw data from 12 participants</li> <li><strong>Subfolders:</strong> Each subfolder contains raw EEG recordings from one participant, subfolders names are:</li> <li>CI1, CI2, CI3, CI4, CI5, CI6, CI7, CI8, CI9, CI10, CI11, CI12</li> </ul> </li> <li><strong>Each participant subfolder contains four &lsquo;.mat&rsquo; files:</strong> <ul> <li>Each &lsquo;.mat&rsquo; file starts with the participant ID <ul> <li>CI1, CI2, CI3, CI4, CI5, CI6, CI7, CI8, CI9, CI10, CI11, CI12</li> </ul> </li> <li>Followed by the stimulus: 40 da</li> <li>Followed by the stimulus polarity <ul> <li>Pos for positive/standard</li> <li>Neg for negative (reversed polarity stimulus)</li> </ul> </li> <li>Followed by the ipsilateral/implanted ear <ul> <li>R for right ear</li> <li>L for left ear</li> </ul> </li> <li>Finally the recording number for that polarity <ul> <li>1 is the first recording</li> <li>2 is the second recording</li> </ul> </li> </ul> </li> <li><strong>Example &lsquo;.mat&rsquo; file name:</strong> <ul> <li><em>CI1 40 da Neg R1.mat:</em> Participant number 1, reversed polarity stimulus, implanted ear &ndash; right ear, recording number one</li> <li><em>CI9 40 da Pos L2.mat:</em> Participant number 9, standard/positive polarity stimulus, implanted ear &ndash; left ear, recording number two</li> </ul> </li> <li><strong>Description of &lsquo;.mat&rsquo; files that can be accessed and processed using MATLAB (MathWorks): </strong>Each &lsquo;.mat&rsquo; file is a structure that contains the following fields: <ul> <li>The first nine fields are informational, for example: <ul> <li>xunits: &lsquo;s&rsquo; indicates that the recording time window is in seconds, conversion to milliseconds would be required to plot the data in milliseconds</li> <li>start: &lsquo;0&rsquo; indicates that both stimulus and recording start at 0 seconds</li> <li>points: <strong>13750</strong> is the number of sample points</li> <li>chans: 2 is the number of channels (channel 2 is the ipsilateral channel for participants with a Right CI, and channel 1 is the ipsilateral channel for participants with a Left CI)</li> <li>frames: 2500 is the number of epochs (repetitions)</li> </ul> </li> <li>The last filed <strong>&lsquo;values&rsquo;</strong> is what contains the raw EEG data (13750x2x2500) <ul> <li><strong>13750 </strong>is the number of samples</li> <li><strong>2 </strong>is the number of channels (channel one is recorded from the left ear lobe (A1) and channel two is from the right ear lobe (A2))</li> <li><strong>2500 </strong>is the number of epochs</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p><strong>Date of data collection:&nbsp;</strong>September 2017</p> <p>&nbsp;</p> <p><strong>Artificial-head artefact Recordings (BinKhamis et al Speech_ABR in CI artefact data.zip)</strong></p> <ul> <li><strong>This folder contains six artificial-head artefact recordings</strong> <ul> <li><strong>CI07_Artefact_40da_neg.mat</strong> <ul> <li>Artefact based on closest approximation to CI07&#39;s MAP</li> <li>Recorded in response to a&nbsp;reversed&nbsp;polarity (neg) 40 ms [da]</li> </ul> </li> <li><strong>CI07_Artefact_40da_pos.mat</strong> <ul> <li>Artefact based on closest approximation to CI07&#39;s MAP</li> <li>Recorded in response to a&nbsp;standard&nbsp;polarity (pos) 40 ms [da]</li> </ul> </li> <li><strong>CI08_Artefact_40da_neg.mat</strong> <ul> <li>Artefact based on closest approximation to CI08&#39;s MAP</li> <li>Recorded in response to a&nbsp;reversed&nbsp;polarity (neg) 40 ms [da]</li> </ul> </li> <li><strong>CI08_Artefact_40da_pos.mat</strong> <ul> <li>Artefact based on closest approximation to CI08&#39;s MAP</li> <li>Recorded in response to a&nbsp;standard&nbsp;polarity (pos) 40 ms [da]</li> </ul> </li> <li><strong>CI09_Artefact_40da_neg.mat</strong> <ul> <li>Artefact based on closest approximation to CI09&#39;s MAP</li> <li>Recorded in response to a&nbsp;reversed&nbsp;polarity (neg) 40 ms [da]</li> </ul> </li> <li><strong>CI09_Artefact_40da_pos.mat</strong> <ul> <li>Artefact based on closest approximation to CI09&#39;s MAP</li> <li>Recorded in response to a&nbsp;standard&nbsp;polarity (pos) 40 ms [da]<br> &nbsp;</li> </ul> </li> </ul> </li> <li><strong>Artefact recordings are in &#39;.mat&#39; format (MATLAB, MathWorks)</strong> <ul> <li>Each &#39;.mat&#39; file is a 10988x2500 matrix <ul> <li><strong>10988</strong>&nbsp;is the number of sample points</li> <li><strong>2500</strong> is the number of epochs&nbsp;(repetitions)</li> </ul> </li> </ul> </li> </ul> <p><strong>Artefacts recorded in: </strong>June&nbsp;2018</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Figure 10. Conceptual diagram illustrating vector quantization codeword formation-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>After the conversion process converts the analogue speech signal into a series words,<br> then converting recognized word to facial animation based on VRML.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 9. Mel Cepstral Coefficients in time domain-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>In this section, Feature matching Algorithm has been discussed. The goal of feature<br> matching is to classify objects into one of a number of categories or classes. In this project, Vector<br> Quantization approach will be used and the best matching result will be the desired voice.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 6. Change the frequency at Hertz scale to Mel Scale.-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Following computing the spectrum power and applying the above equation on frequency<br> axis, some mediating filters equal to identical overlapping are applied on the scaled spectrum and<br> each filters energy is computed as particularity. This is because conception of a particular frequency<br> by the auditory system is affected by a critical band of frequencies surrounding it. Number of filters<br> is usually between 20 and 30. Logarithmic non linear change operations on obtained particulates for<br> adjusting amplified of particularities and their important though coordinating them with the<br> structure of the auditory system following computing energy of each filter is done as follows: in the<br> following equation F1 is filter in I th, and e (i) is logarithm of energy at ith band.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 4. Hamming Window applied to each frame-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Use of speech spectrum for modifying work domain on signals from time to frequency is<br> made possible using Fourier coefficients. At such applications the rapid and practical way of<br> estimating the spectrum is use of rapid Fourier changes.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 3 Frame blocking of the speech signal-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>The next frame will begin M samples (i.e. 156 samples) after the first frame, and it will<br> overlap the first frame by N-M samples (256 &ndash; 156 = 100 samples). Then the third frame will start<br> at 2M samples after the first frame and it will overlap first frame by N-2M. The fourth frame will<br> start at 3M samples after the first, and it will overlap it by N-3M. The process will continue until all<br> input signal is accounted for. The result of this step plotted using MATLAB plot command and<br> displayed in Figure 3.<br> Figure 3 Frame blocking of the speech signal<br> The next step in the processing is to window each individual frame so as to minimize the<br> signal discontinuities at the beginning and end of each frame. The concept here is to minimize the<br> spectral distortion by using the window to taper the signal to zero at the beginning and end of each<br> frame. If we define the window as w(n), 0 &le; n &le; N &minus;1, where N is the number of samples in each<br> frame, then the result of windowing is the signal<br> y (n) = x (n)w(n), 0 &le; n &le; N &minus;1 l l<br> &nbsp;</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 1. The Structure of the Automatic Translate Voice to Sign Language Animation System-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>All technologies of voice recognition, speaker identification and verification, each has its<br> own advantages and disadvantages and may requires different treatments and techniques. The<br> choice of which technology to use is application-specific. At the highest level, all voice recognition<br> systems contain two main modules: feature extraction and feature matching. Feature extraction is<br> the process that extracts a small amount of data from the voice signal that can later be used to<br> represent each word. Feature matching involves the actual procedure to identify the unknown word<br> by comparing extracted features from his/her voice input with the ones from a set of known words.<br> A wide range of possibilities exist for parametrically representing the speech signal for the<br> voice recognition task, such as Linear Prediction Coding (LPC), RASTA-PLP and Mel-Frequency<br> Cepstrum Coefficients (MFCC).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 8. Speech signal after frequency wrapping-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Use of speech spectrum for changing the work domain on the signal from time to frequency<br> is used via Fourier conventions. The rapid and practical way of estimating Spectrum in such<br> applications is employment of rapid Fourier transform. The final stage is extracting particularities is<br> use of discrete cosine to return particularities to the time domain and converse FFT approximation.<br> Major advantage of this method is decrease of number of particularities of number of filter from f N<br> to c N in which c f N &le; N . In addition, doing so, includes making independent of the obtained<br> particularity and rendering them non dependant which leads matrix covariance features to become<br> axial. The following equation shows this point.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 7. Speech signal after FFT-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Use of speech spectrum for changing the work domain on the signal from time to frequency<br> is used via Fourier conventions. The rapid and practical way of estimating Spectrum in such<br> applications is employment of rapid Fourier transform. The final stage is extracting particularities is<br> use of discrete cosine to return particularities to the time domain and converse FFT approximation.<br> Major advantage of this method is decrease of number of particularities of number of filter from f N<br> to c N in which c f N &le; N . In addition, doing so, includes making independent of the obtained<br> particularity and rendering them non dependant which leads matrix covariance features to become<br> axial. The following equation shows this point.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 2. MFCC Block Diagram Step 1-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Mel Frequency Cepstral Coefficients (MFCC) are coefficients that represent audio, based on<br> perception. It is derived from the Fourier Transform (FFT) or the Discrete Cosine Transform (DCT)<br> of the audio clip. The basic difference between the FFT/DCT and the MFCC is that in the MFCC,<br> the frequency bands are positioned logarithmically (on the Mel scale) which approximates the<br> human auditory system&#39;s response more closely than the linearly spaced frequency bands of FFT or<br> DCT. This allows for better processing of data. The main purpose of the MFCC processor is to<br> mimic the behavior of the human ears. Overall the MFCC process has 5 steps that show in figure 2.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 11. A sample of speech to facial animation system-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>To increase the autonomy of deaf and hard of hearing people in their day-to-day<br> professional and social lives, in this paper design and initial implementation of a new approach<br> based on MFCC and Vector Quantization Method is described. This approach includes<br> analyses of speech to animate the talking head. Our future work will include the conception of<br> new test types and performance patterns. We are particularly interested in extending this<br> approach to testing to include implementation of applications under real-time constraints</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Figure 5. Mel-spaced filter bank-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives

<p>Physiologic changes show that human comprehension of frequency content of sound does<br> not obey a linear space. Therefore for each individual it is computed and measure with a real<br> frequency of sound peak at Mel seal. Using the below equation one can change the frequency at<br> Hertz scale to Mel Scale.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Attempted Speech on Single Words [intracortical array][T16][BG2]

<h2>Summary of the Data and Experiment</h2> <h3>General Overview</h3> <p>The data provided come from an attempted speech experiment involving a participant (T16) in the BrainGate2 clinical trial. The dataset includes neural recordings from four microelectrode arrays implanted in the left precentral gyrus. The primary focus is on the array located in the ventral precentral gyrus (6v), which is associated with speech-related activity.</p> <h3>Data Description</h3> <p>The neural data consists of two main streams:</p> <ol> <li>Log Local Field Potential (LFP) Power: <ul> <li>Description: Extracellular neural data from the lower frequency bands.</li> <li>Sampling Rate: Preprocessed to 2 kHz.</li> <li>Filtering: Bandpass filtered with a 150-450 Hz band.</li> <li>Scale: Log-scale</li> <li>Bin size: 20ms Summed</li> <li>Shape: 128457x64 (time by channels).</li> </ul> </li> <li>Binned Threshold Crossings: <ul> <li>Description: Extracellularly collected neural spiking activity related to attempted speech.</li> <li>Threshold: -4.5 RMS</li> <li>Bin size: 20ms Summed</li> <li>Shape: 128457x64 (time by channels)</li> </ul> </li> </ol> <p>The included data are within a single mat file with the keys `trial_info_df`, `lfp_matrix`, `spikes_matrix`, and `timevector`.</p> <h3>Trial Information</h3> <p>The `trial_info_df` key contains information about each trial in the form of a pandas DataFrame with the following columns:</p> <ul> <li>block_num</li> <li>cue</li> <li>start_time</li> <li>go_cue_time</li> <li>end_time</li> </ul> <h3>Participant and Experiment Information</h3> <ul> <li>Participant T16: Right-handed, 52-year-old woman with tetraplegia and dysarthria due to a pontine stroke.</li> <li>Implantation: Four 64-channel intracortical microelectrode arrays in the left precentral gyrus.</li> <li>Recording Day: Data collected 69 days post-implantation.</li> <li>Recording Platform: Backend for Realtime Asynchronous Neural Decoding (BRAND) platform.</li> <li>Spiking Data Extraction: Linear regression referencing, bandpass filtering (250-5000 Hz), and threshold crossings identification (-4.5 RMS threshold).</li> <li>LFP power Extraction: Band pass filtering (150 - 450 Hz), and notch filtering at harmonics of 60 Hz.</li> </ul> <h3>Experimental Task</h3> <ul> <li>Task: Cued speech task where Participant T16 vocalized words presented on a screen.</li> <li>Word Bank: Compiled from a 50-word vocabulary by Moses et al.</li> <li>Trial Structure:<br> <ul> <li>A red square appeared below a word.</li> <li>After 1500 ms, the square turned green, cueing the participant to vocalize the word.</li> <li>The trial ended after the participant finished speaking, with a 1000 ms interval before the next trial.</li> </ul> </li> </ul> <h2>Loading and Handling the Data</h2> <p>The data can be loaded into Python using `scipy.io.loadmat`:</p> <h3>Loading the Data</h3> <blockquote> <p>import scipy.io<br>import pandas as pd</p> <p># Load the .mat file<br>data = scipy.io.loadmat('path_to_mat_file.mat')</p> <p># Extracting the trial information<br>trial_info_df = pd.DataFrame(data['trial_info_df'])<br>trial_info_df.columns = ['block_num', 'cue', 'start_time', 'go_cue_time', 'end_time']</p> <p># Extracting neural data matrices<br>lfp_matrix = data['lfp_matrix']<br>spikes_matrix = data['spikes_matrix']</p> <p># Extracting time vector<br>timevector = data['timevector']</p> </blockquote> <h3>Wrapping Trial Information in a DataFrame</h3> <blockquote> <p>trial_info_df = pd.DataFrame(data['trial_info_df'], columns=['block_num', 'cue', 'start_time', 'go_cue_time', 'end_time'])</p> </blockquote> <h3>Exploring the Data</h3> <ul> <li>Log Local Field Potential (LFP) Power: `lfp_matrix`</li> <li>Threshold Crossings (Spikes): `spikes_matrix`</li> <li>Time Vector: `timevector`</li> </ul> <h2>Support</h2> <p><span>This work was supported by NIH-NINDS/OD DP2NS127291 (CP), NIH-NICHD F32HD112173 (SRN), NIH K08NS060223 (MS), R01NS112942 (MS), and RF1NS125026 (MS). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, or the Department of Veterans Affairs, or the United States Government.</span></p> <p><span>*Caution: Investigational Device. Limited by federal law to investigational use. </span></p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

AnglistikVoices: L2 English speech dataset

<h1><strong>AnglistikVoices: an L2 English speech dataset</strong>&nbsp;</h1> <p>This repository contains an L2 (second language) English speech corpus consisting of&nbsp;<strong>74 minutes</strong>&nbsp;of recorded audio from&nbsp;<strong>15 non-native English speaking participants</strong>.&nbsp;The dataset was created as part of a university course, with all participants being students who are also the authors of this dataset.</p> <h2>Dataset Specifications</h2> <ul> <li><strong>Total participants</strong>: 15 non-native English speakers</li> <li><strong>Total audio duration</strong>: 74 minutes</li> <li><strong>Recordings per participant</strong>: 60 audio samples each</li> <li><strong>Sentence alignment</strong>: Available for 8 out of 15 participant</li> <li><strong>Recording equipment</strong>: Audio-Technica ATM75 microphone</li> <li><strong>Stimuli:&nbsp;</strong>All sentences are from the Artie Bias Corpus (https://github.com/artie-inc/artie-bias-corpus)</li> <li><strong>Recording environment</strong>: Recording booth</li> </ul> <p>The dataset contains individual recordings of non-native English speakers&nbsp;<strong>organized by participant ID</strong>. For 8 participants, sentence-level alignments are provided. All recordings were captured in a controlled acoustic environment using Audio-Technica ATM75 microphone to ensure high audio quality.&nbsp;</p> <p>The recordings consist of spoken English utterances from each participant. <strong>Detailed linguistic profiles for each participant are available in the metadata.xlsx file</strong>, which is indexed by participant ID and contains information on native language, proficiency level, language learning history, and other relevant linguistic background data.</p> <p>The audio files are organized by participant ID, matching the identifiers used in the metadata file for easy cross-referencing between the audio recordings and participant linguistic profiles.</p> <h2>Authors and Contributors</h2> <p>This dataset was created by the student participants themselves as part of their coursework.</p> <p><strong>Course Instructor</strong>: Akhilesh Kakolu Ramarao</p> <p><strong>Teaching Assistant</strong>: Anna Sophia Stein</p> <p>If you have any questions, you can contact: kakolura@hhu.de</p> <p>If you use this dataset in your research, please cite:</p> <pre><code>@dataset{kakolu_ramarao_2024_anglistikvoices, author = {Kakolu Ramarao, A. and Stein, A. S. and Tahiri, A. and Rodrigues, D. C. and Antonia Weismann, C. and Sch&auml;fer, O. S. and Kaczor, J. and Tran, N. H. and Elena Telaar, C. and Bauer, L. and J&uuml;tten, M. and Mafuta, C. and Agelopoulou, V. V. and Grabowski, Q. A. G.}, title = {AnglistikVoices: L2 English speech dataset}, publisher = {Zenodo}, version = {v1.0.0}, year = {2024}, month = jun, doi = {10.5281/zenodo.12525952}, url = {https://doi.org/10.5281/zenodo.12525952}, note = {LabPhon 19, Hanyang Institute for Phonetics and Cognitive Sciences of Language (HIPCS), Hanyang University in Seoul, Korea} }</code></pre>

opencc-by-sa-4.0Jun 2024View details →
zenodo40/100

Emozionalmente: a crowdsourced Italian speech emotional corpus

<p>This repository contains Emozionalmente: an extensive simulated speech emotional corpus in Italian. The dataset comprises 6,902 labeled samples acted out by 431 amateur actors, each verbalizing 18 different sentences to express the Big Six emotions (anger, disgust, fear, joy, sadness, surprise) plus neutrality. The labels represent the emotional communicative intention of the actors (i.e., the seven emotional states).</p> <p>Key details about the dataset:<br>- **Recording specifications**: The recordings were generally obtained with non-professional equipment. They are .wav files, mono-channel, with a sample size of 16 bits and a sample rate of 16,000 Hz. Each audio recording lasts 3.81 seconds on average (SD = 0.99 seconds).<br>- **Validation**: To validate the emotional content of the clips, 829 humans evaluated each audio recording, providing five evaluations per audio. The general Unweighted Average Recall (UAR) achieved by the evaluators was 66%, which is comparable to previous literature in the field.</p> <p>The repository includes the following additional resources:<br>1. **Demographic information**: Three .csv files describing the demographics of the actors and evaluators, as well as the emotions they expressed and recognized for each audio sample.<br>2. **Data splits**: A speaker-independent train-dev-test split, stratified by emotion, gender, and age.</p> <p>If you use this dataset, please cite the following paper:</p> <blockquote> <p>F. Catania, J. W. Wilke and F. Garzotto,<br><em>"Emozionalmente: A Crowdsourced Corpus of Simulated Emotional Speech in Italian,"</em><br>IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1142&ndash;1155, 2025.<br>doi: <a target="_new" rel="noopener">10.1109/TASLPRO.2025.3540662</a></p> </blockquote>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record