Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
503
datasets available to search
ShareScore release 0.9.0
Dataset results
503 results for “Voice”
Examples: Bridging Communication Gaps: The Role of Voice-Enabled AI in Medicine
<p><strong>Illustrative examples of potential application cases of advanced voice mode in Clinical Practice. </strong></p>
PADE Fine details voice quality
<p>This corpus gathers sentences uttered by two female speakers (S3 and S6) performing six speech acts in USA English using the sentence “Mary was dancing”. The wav files are provided with their syllabic and phonemic alignment (in TextGrids). The dataset was developed as part of the French ANR project PADE.<br>The recording protocol was described in Rilliard et al. (2013); analysis can be found in Rilliard et al. 2017; this subset was used for the publication Erickson et al. (submitted). The recordings of six pairs of opposed voice qualities, on the word “dancing”, by a female USA English speaker are also included, and were used for the perceptual evaluation presented in Erickson et al. (2024).<br>References:</p> <p>Erickson, D., Rilliard, A., Thurgood, E., de Moraes, J. A., & Shochi, T. (2024). Acoustic and perceptual profiles of American English social affective expressions: a case study. <em>Journal of Speech Sciences</em>, <em>13</em>(00), e024004. https://doi.org/10.20396/joss.v13i00.20015<br>Rilliard, A., Erickson, D., Shochi, T., & de Moraes, J. A. de. (2013). Social face to face communication—American English attitudinal prosody. Interspeech 2013, 1648–1652. https://doi.org/10.21437/Interspeech.2013-427<br>Rilliard, A., Erickson, D., de Moraes, J. A. de, & Shochi, T. (2017). Perception of expressive prosodic speech acts performed in USA English by L1 and L2 speakers. Journal of Speech Sciences, 6(1), 27–45. https://doi.org/10.20396/joss.v6i1.14981</p>
The human Voice Areas: spatial organisation and inter-individual variability in temporal and extra-temporal cortices
Open the record for dataset details and reuse information.
GRB assessment of the Saarbrücken Voice Database
<p>GRB evaluations of the Saarbruecken database. Annex to the paper: <em>Multimodal and multi-output deep learning architectures for the automatic assessment of voice quality using the GRB scale. </em></p> <p>Published in IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. 14, NO. 2, FEBRUARY 2020: <strong>DOI: </strong><a href="https://doi.org/10.1109/JSTSP.2019.2956410">10.1109/JSTSP.2019.2956410</a> <br> </p> <p>If you use these labels please cite as:</p> <p>J. D. Arias-Londoño, J. A. Gómez-García and J. I. Godino-Llorente, "Multimodal and Multi-Output Deep Learning Architectures for the Automatic Assessment of Voice Quality Using the GRB Scale," in <em>IEEE Journal of Selected Topics in Signal Processing</em>, vol. 14, no. 2, pp. 413-422, Feb. 2020.</p>
Voice Conversion Challenge 2020 database v1.0
<pre>Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform. In 2016, we have launched the Voice Conversion Challenge (VCC) 2016 [1][2] at Interspeech 2016. The objective of the 2016 challenge was to better understand different VC techniques built on a freely-available common dataset to look at a common goal, and to share views about unsolved problems and challenges faced by the current VC techniques. The VCC 2016 focused on the most basic VC task, that is, the construction of VC models that automatically transform the voice identity of a source speaker into that of a target speaker using a parallel clean training database where source and target speakers read out the same set of utterances in a professional recording studio. 17 research groups had participated in the 2016 challenge. The challenge was successful and it established new standard evaluation methodology and protocols for bench-marking the performance of VC systems. In 2018, we have launched the second edition of VCC, the VCC 2018 [3]. In the second edition, we revised three aspects of the challenge. First, we educed the amount of speech data used for the construction of participant's VC systems to half. This is based on feedback from participants in the previous challenge and this is also essential for practical applications. Second, we introduced a more challenging task refereed to a Spoke task in addition to a similar task to the 1st edition, which we call a Hub task. In the Spoke task, participants need to build their VC systems using a non-parallel database in which source and target speakers read out different sets of utterances. We then evaluate both parallel and non-parallel voice conversion systems via the same large-scale crowdsourcing listening test. Third, we also attempted to bridge the gap between the ASV and VC communities. Since new VC systems developed for the VCC 2018 may be strong candidates for enhancing the ASVspoof 2015 database, we also asses spoofing performance of the VC systems based on anti-spoofing scores. In 2020, we launched the third edition of VCC, the VCC 2020 [4][5]. In this third edition, we constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. The dataset for intra-lingual VC consists of a smaller parallel corpus and a larger nonparallel corpus, where both of them are of the same language. The dataset for cross-lingual VC consists of a corpus of the source speakers speaking in the source language and another corpus of the target speakers speaking in the target language. As a more challenging task than the previous ones, we focused on cross-lingual VC, in which the speaker identity is transformed between two speakers uttering different languages, which requires handling completely nonparallel training over different languages. This repository contains the training and evaluation data released to participants, target speaker’s speech data in English for reference purpose, and the transcriptions for evaluation data. For more details about the challenge and the listening test results please refer to [4] and README file. </pre> <pre>[1] Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "The Voice Conversion Challenge 2016" in Proc. of Interspeech, San Francisco. [2] Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "Analysis of the Voice Conversion Challenge 2016 Evaluation Results" in Proc. of Interspeech 2016. [3] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, Zhenhua Ling, "The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods", Proc Speaker Odyssey 2018, June 2018. [4] Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda. "Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion" Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 80-98, DOI: 10.21437/VCC_BC.2020-14.</pre>
Voice Conversion Challenge 2020 Listening Test Data
<pre>Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform. In 2016, we have launched the Voice Conversion Challenge (VCC) 2016 [1][2] at Interspeech 2016. The objective of the 2016 challenge was to better understand different VC techniques built on a freely-available common dataset to look at a common goal, and to share views about unsolved problems and challenges faced by the current VC techniques. The VCC 2016 focused on the most basic VC task, that is, the construction of VC models that automatically transform the voice identity of a source speaker into that of a target speaker using a parallel clean training database where source and target speakers read out the same set of utterances in a professional recording studio. 17 research groups had participated in the 2016 challenge. The challenge was successful and it established new standard evaluation methodology and protocols for bench-marking the performance of VC systems. In 2018, we have launched the second edition of VCC, the VCC 2018 [3]. In the second edition, we revised three aspects of the challenge. First, we educed the amount of speech data used for the construction of participant's VC systems to half. This is based on feedback from participants in the previous challenge and this is also essential for practical applications. Second, we introduced a more challenging task refereed to a Spoke task in addition to a similar task to the 1st edition, which we call a Hub task. In the Spoke task, participants need to build their VC systems using a non-parallel database in which source and target speakers read out different sets of utterances. We then evaluate both parallel and non-parallel voice conversion systems via the same large-scale crowdsourcing listening test. Third, we also attempted to bridge the gap between the ASV and VC communities. Since new VC systems developed for the VCC 2018 may be strong candidates for enhancing the ASVspoof 2015 database, we also asses spoofing performance of the VC systems based on anti-spoofing scores. In 2020, we launched the third edition of VCC, the VCC 2020 [4][5]. In this third edition, we constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. The dataset for intra-lingual VC consists of a smaller parallel corpus and a larger nonparallel corpus, where both of them are of the same language. The dataset for cross-lingual VC consists of a corpus of the source speakers speaking in the source language and another corpus of the target speakers speaking in the target language. As a more challenging task than the previous ones, we focused on cross-lingual VC, in which the speaker identity is transformed between two speakers uttering different languages, which requires handling completely nonparallel training over different languages. As for listening test, we subcontracted the crowd-sourced perceptual evaluation with English and Japanese listeners to Lionbridge TechnologiesInc. and Koto Ltd., respectively. Given the extremely large costs required for the perceptual evaluation, we selected 5 utterances (E30001, E30002, E30003,E30004, E30005) only from each speaker of each team. To evaluate the speaker similarity of the cross-lingual task, we used audio in both the English language and in the target speaker’s L2language as reference. For each source-target speaker pair, we selected three English recordings and two L2 language recordings as the natural reference for the converted five utterances. </pre> <p>This data repository includes the audio files used for the crowd-sourced perceptual evaluation and raw listening test scores. </p> <pre>[1] Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "The Voice Conversion Challenge 2016" in Proc. of Interspeech, San Francisco. [2] Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "Analysis of the Voice Conversion Challenge 2016 Evaluation Results" in Proc. of Interspeech 2016. [3] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, Zhenhua Ling, "The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods", Proc Speaker Odyssey 2018, June 2018. [4] Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda. "Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion" Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 80-98, DOI: 10.21437/VCC_BC.2020-14. [5] Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang, Zhenhua Ling, Junichi Yamagishi, Yi Zhao, Xiaohai Tian, and Tomoki Toda. "Predictions of subjective ratings and spoofing assessments of voice conversion challenge 2020 submissions." Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 99-120, DOI: 10.21437/VCC_BC.2020-15. </pre>
German own voice recordings with hearable microphones
<p>This dataset is supplementary material to the article "Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with an In-ear Microphone" published in Acta Acustica, vol. 8 (2024).</p> <p>The dataset consists of recordings of own voice speech of 18 talkers (5 female, 13 male) wearing hearable devices in both ears. All talkers were native German speakers. The dataset was recorded in a sound-proof listening booth using the Hearpiece prototype device (closed vent variant) [1].</p> <p>The speech uttered by the talkers is pre-determined text read from a screen. The talkers press a button on the screen to start the recording, read the sentence out loud in a normal voice, and press another button to stop the recording. It was possible for the talkers to re-record a sentence if desired. The sentences read by the talkers originate from the following sources:</p> <ul> <li>The north wind and the sun (German), 6 sentences</li> <li>Berlin and Marburg Sentences [2] (German), 2x100 sentences</li> <li>100 sentences for language learners [3] (German), 100 sentences</li> <li>Some held-out vowels and consonants</li> <li>[some seconds of silence]</li> </ul> <p>The full text read by each talker is written in <code>full_text.txt</code>.</p> <p>The recordings are contained in the folder <code>speech</code>. Each subfolder contains recordings from a different talker (e.g., <code>VP_01</code>). The sentences uttered by each talker are numbered following this scheme: <code>VP_01_0.wav</code> to <code>VP_01_313.wav</code>. Talkers where the device could not be inserted, or where the fit did not provide sufficient attenuation of external sounds to the in-ear microphone, were excluded.</p> <p>A DPA 6060 lavalier clip microphone and a Tbone SC140 cardiod microphone were recorded as reference signals. From two Hearpiece devices (closed vent), the concha and in-ear microphones were recorded. Audio was recorded at a sampling frequency of 44100 Hz.</p> <p>The channels of the recordings, counting from 0, recorded the following microphones:</p> <ul> <li>0: Lavalier-microphone clipped to the shirt neck, shirt collar etc. of the talker</li> <li>1: Reference microphone about 50 cm in front of the talker</li> <li>2: Left in-ear microphone Hearpiece</li> <li>3: Left concha microphone Hearpiece</li> <li>4: Right in-ear microphone Hearpiece</li> <li>5: Right concha microphone Hearpiece</li> </ul> <p>[1] F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier: "A One-Size-Fits-All Earpiece with Multiple Microphones and Drivers for Hearing Device Research". In: Proc. AES International Conference on Headphone Technology. San Francisco, USA, Aug. 2019.</p> <p>[2] A. P. Simpson, K. J. Kohler, and T. Rettstadt. "The Kiel Corpus of Read/Spontaneous Speech: Acoustic Data Base, Processing Tools, and Analysis Results". In: Arbeitsberichte Institut für Phonetik Und Digitale Sprachverarbeitung Universität Kiel. Vol. 32. IPDS, Nov. 1997, pp. 243-247.</p> <p>[3] A. Neustein. "100 Sätze Reichen Für Ein Ganzes Leben" (Blog-post). https://deutschlernerblog.de/100-saetze-reichen-fuer-ein-ganzes-leben/. Aug. 2019.</p> <p> </p> <p> </p>
ASVspoof2019LA-Sim: Augmented Dataset for An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
<p>This is the dataset we augmented to study the channel effects for anti-spoofing. For more details, please refer to our Interspeech 2021 paper: "An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems".</p> <p>Proceeding: <a href="https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html">https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html</a></p> <p>Arxiv: <a href="https://arxiv.org/pdf/2104.01320.pdf">https://arxiv.org/pdf/2104.01320.pdf</a></p> <p>Code: <a href="https://github.com/yzyouzhang/Empirical-Channel-CM">https://github.com/yzyouzhang/Empirical-Channel-CM</a></p> <p>Contact: you.zhang@rochester.edu</p> <p><strong>Version 1.0</strong> contains the <strong>training</strong> and the <strong>development</strong> set. We have added the <strong>evaluation</strong> set in <strong>version 1.1 </strong>but deleted the training set due to the size limitation, but you can still access the training set in version 1.0.</p> <p>Please check it out.</p> <p>To extract the files, please use the following commands:</p> <pre><code class="language-bash">cat eval.tar.gz-part* > eval.tar.gz tar -xvzf *.tar.gz</code></pre> <p>After concatenation, to make sure the download is complete, you can check with the following:</p> <pre><code>md5sum *.tar.gz 15dea7d28b126994bb6b159778f706af dev.tar.gz 0615052b34ca6c7f58505eaa8647844f eval.tar.gz 3058dd9d407f3c9ae697acca8c34a6c3 train.tar.gz</code></pre> <p>Thanks.</p>
How do native and non-native speakers recognize emotions in the instructor's voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners [dataset]
<p>Dataset for the journal article <em>How do native and non-native speakers recognize emotions in the instructor’s voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners.</em></p>
Jamendo Corpus for Singing Voice Detection
<p>This is a public corpus of 93 creative-commons licensed music pieces annotated<br> by voice (sung or spoken) and no-voice.</p>
Thorsten-Voice Dataset 2022.10
<p>The goal of project "Thorsten-Voice" is to provide voice datasets and TTS models for free and high quality german artificial voice. This dataset "Thorsten-Voice dataset 2022.10" is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.</p> <blockquote> <p><strong>"I contribute my personal voice as a person believing in a world where all people are equal. No matter of gender, sexual orientation, religion, skin color and geocoordinates of birth location. A global world where everybody is warmly welcome on any place on this planet and open and free knowledge and education is available to everyone." (Thorsten Müller)</strong></p> </blockquote> <p> </p> <p><strong>Dataset details</strong>:</p> <ul> <li>ljspeech file and directory structure</li> <li>12.450 recorded phrases (wav files)</li> <li>more than 11 hours of pure audio</li> <li>samplerate 22.050Hz</li> <li>mono</li> <li>normalized to -24dB</li> <li>no silence at beginning/ending</li> <li>avg spoken chars per second: 17,5</li> </ul> <p>See more details on my <a href="https://github.com/thorstenMueller/Thorsten-Voice">Github page</a> or <a href="https://www.Thorsten-Voice.de">Thorsten-Voice</a> project website.</p>
Impact of Passive Voice in Requirements Engineering
<p>This repository contains instrumentation material and results for the experiment described in the paper "On The Impact of Passive Voice Requirements on Domain Modelling" by Henning Femmer, Jan Kucera, and Antonio Vetrò from Technische Universität München.</p>
COALA voice data and transcripts Dutch
<p>This dataset contains audio files and transcripts in Dutch and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut für Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: "Have you ever worked in a factory, manufacturing setting?" (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: "If yes, for how many years?" (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>
COALA voice data and transcripts Italian
<p>This dataset contains audio files and transcripts in Italian and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut für Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: "Have you ever worked in a factory, manufacturing setting?" (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: "If yes, for how many years?" (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>
InterTVA. A multimodal MRI dataset for the study of inter-individual differences in voice perception and identification.
Open the record for dataset details and reuse information.
RealVAD: A Real-world Dataset for Voice Activity Detection
<p><strong>RealVAD: A Real-world Dataset for Voice Activity Detection</strong></p> <p>The task of automatically detecting “Who is Speaking and When” is broadly named as Voice Activity Detection (VAD). Automatic VAD is a very important task and also the foundation of several domains, e.g., human-human, human-computer/ robot/ virtual-agent interaction analyses, and industrial applications.</p> <p>RealVAD dataset is constructed from a YouTube video composed of a panel discussion lasting approx. 83 minutes. The audio is available from a single channel. There is one static camera capturing all panelists, the moderator and audiences.</p> <p>Particular aspects of RealVAD dataset are:</p> <ul> <li> It is composed of panelists with different nationalities (British, Dutch, French, German, Italian, American, Mexican, Columbian, Thai). This aspect allows studying the effect of ethnic origin variety to the automatic VAD.</li> <li> There is a gender balance such that there are four female and five male panelists.</li> <li> The panelists are sitting in two rows and they can be gazing audience, other panelists, their laptop, the moderator or anywhere in the room while speaking or not-speaking. Therefore, they were captured not only from frontal-view but also from side-view varying based on their instant posture and head orientation.</li> <li> The panelists are moving freely and are doing various spontaneous actions (e.g., drinking water, checking their cell phone, using their laptop, etc.), resulting in different postures.</li> <li> The panelists’ body parts are sometimes partially occluded by their/other's body part or belongings (e.g., laptop).</li> <li> There are also natural changes of illumination and shadow rising on the wall behind the panelists in the back row.</li> <li> Especially, for the panelists sitting in the front row, there is sometimes background motion occurring when the person(s) behind them moves.</li> </ul> <p>The annotations includes:</p> <ul> <li> The upper body detection of nine panelists in bounding box form.</li> <li> Associated VAD ground-truth (speaking, not-speaking) for nine panelists.</li> <li> Acoustic features extracted from the video: MFCC and raw filterbank energies.</li> </ul> <p><em>All info regarding the annotations are given in the ReadMe.txt and Acoustic Features README.txt files.</em></p> <p><strong>When using this dataset for your research, please cite the following paper in your publication:</strong></p> <ol> <li>C. Beyan, M. Shahid and V. Murino, "RealVAD: A Real-world Dataset and A Method for Voice Activity Detection by Body Motion Analysis", in IEEE Transactions on Multimedia, 2020.</li> </ol>
Papuan Voices Media Files
<p>Papuan Voices Media Files (wav) - Supplement to dataset <a href="https://doi.org/10.5281/zenodo.4350691">Papuan Voices</a><br> <br> Papuan Voices presents phonetically-transcribed primary recordings, from numerous places throughout Papua island<br> <br> Available online at: <a href="https://papuanvoices.clld.org">https://papuanvoices.clld.org</a></p>
Mixe-Zoquean Voices Media Files
<p>Mixe-Zoquean Voices Media Files (wav) - Supplement to dataset <a href="https://doi.org/10.5281/zenodo.4350652">Mixe-Zoquean Voices</a><br> <br> Mixe-Zoquean Voices presents primary recordings of languages from the Mixe-Zoque language family.<br> <br> Available online at: <a href="https://mixezoqueanvoices.clld.org">https://mixezoqueanvoices.clld.org</a></p>
Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns"
<p>The .pdf file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> </ul> <p> </p> <p>The .zip files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p> </p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots. </li> </ul>
Emotional Voice Messages (EMOVOME) database
<p>The Emotional Voice Messages (EMOVOME) database is a speech dataset collected for emotion recognition in real-world conditions. It contains 999 spontaneous voice messages from 100 Spanish speakers, collected from real conversations on a messaging app. EMOVOME includes both expert and non-expert emotional annotations, covering valence and arousal dimensions, along with emotion categories for the expert annotations. Detailed participant information is provided, including sociodemographic data and personality trait assessments using the NEO-FFI questionnaire. Moreover, EMOVOME provides audio recordings of participants reading a given text, as well as transcriptions of all 999 voice messages. Additionally, baseline models for valence and arousal recognition are provided, utilizing both speech and audio transcriptions.</p> <h2>Description</h2> <p>For details on the EMOVOME database, please refer to the article:</p> <blockquote> <p><em>"EMOVOME Database: Advancing Emotion Recognition in Speech Beyond Staged Scenarios". Lucía Gómez-Zaragozá, Rocío del Amor, María José Castro-Bleda, Valery Naranjo, Mariano Alcañiz Raya, Javier Marín-Morales. (pre-print available in <a href="https://doi.org/10.48550/arXiv.2403.02167" target="_blank" rel="noopener">https://doi.org/10.48550/arXiv.2403.02167</a>)<br></em></p> </blockquote> <h2>Content</h2> <p>The Zenodo repository contains four files:</p> <ul> <li><strong>EMOVOME_agreement.pdf</strong>: agreement file required to access the original audio files, detailed in section Usage Notes. </li> <li><strong>labels.csv</strong>: ratings of the three non-experts and the expert annotator, independently and combined.</li> <li><strong>participants_ids.csv</strong>: table mapping each numerical file ID to its corresponding alphanumeric participant ID.</li> <li><strong>transcriptions.csv</strong>:<strong> </strong>transcriptions of each audio.</li> </ul> <p>The repository also includes three folders:</p> <ul> <li><strong>Audios</strong>: it contains the file <strong>features_eGeMAPSv02.csv</strong> corresponding to the standard acoustic feature set used in the baseline model, and two folders: <ul> <li><strong>Lecture</strong>: contains the audio files corresponding to the text readings, with each file named according to the participant's ID.</li> <li><strong>Emotions</strong>: contains the voice recordings from the messaging app provided by the user, named with a file ID.</li> </ul> </li> <li><strong>Questionnaires</strong>: it contains two files: 1) <strong>sociodemographic_spanish.csv</strong> and <strong>sociodemographic_english.csv</strong> are the sociodemographic data of participants in Spanish and English, respectively, including the demographic information; and 2) <strong>NEO-FFI_spanish.csv</strong> includes the participants’ answers to the Spanish version of the NEO-FFI questionnaire. The three files include a column indicating the participant's ID to link the information.</li> <li><strong>Baseline_emotion_recognition</strong>: it includes three files and two folders. The file <strong>partitions.csv</strong> specifies the proposed data partition. Particularly, the dataset is divided into 80% for development and 20% for testing using a speaker-independent approach, i.e., samples from the same speaker are not included in both development and test. The development set includes 80 participants (40 female, 40 male) containing the following distribution of labels: 241 negative, 305 neutral and 261 positive valence; and 148 low, 328 neutral and 331 high arousal. The test set includes 20 participants (10 female, 10 male) with the distribution of labels that follows: 57 negative, 62 neutral and 73 positive valence; and 13 low, 70 neutral and 109 high arousal. Files <strong>baseline_speech.ipynb</strong> and <strong>baseline_text.ipynb</strong> contain the code used to create the baseline emotion recognition models based on speech and text, respectively. The actual trained models for valence and arousal prediction are provided in folders <strong>models_speech</strong> and <strong>models_text</strong>. </li> </ul> <p><em>Audio files in “Lecture” and “Emotions” are only provided to the users that complete the agreement file in section Usage Notes. Audio files are in Ogg Vorbis format at 16-bit and 44.1 kHz or 48 kHz. The total size of the “Audios” folder is about 213 MB. </em></p> <h2>Usage Notes</h2> <p>All the data included in the EMOVOME database is publicly available under the Creative Commons Attribution 4.0 International license. The only exception is the original raw audio files, for which an additional step is required as a security measure to safeguard the speakers' privacy. To request access, interested authors should first complete and sign the agreement file <strong>EMOVOME_agreement.pdf</strong> and send it to the corresponding author (<em><a href="mailto:jamarmo@htech.upv.es" target="_blank" rel="noopener">jamarmo@htech.upv.es</a></em>). The data included in the EMOVOME database is expected to be used for research purposes only. Therefore, the agreement file states that the authors are not allowed to share the data with profit-making companies or organisations. They are also not expected to distribute the data to other research institutions; instead, they are suggested to kindly refer interested colleagues to the corresponding author of this article. By agreeing to the terms of the agreement, the authors also commit to refraining from publishing the audio content on the media (such as television and radio), in scientific journals (or any other publications), as well as on other platforms on the internet. The agreement must bear the signature of the legally authorised representative of the research institution (e.g., head of laboratory/department). Once the signed agreement is received and validated, the corresponding author will deliver the "Audios" folder containing the audio files through a download procedure. A direct connection between the EMOVOME authors and the applicants guarantees that updates regarding additional materials included in the database can be received by all EMOVOME users.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.