Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

503

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

503 results for “Voice”

Learn how ShareScore rates datasets ↗
zenodo48/100

Examples: Bridging Communication Gaps: The Role of Voice-Enabled AI in Medicine

<p><strong>Illustrative examples of potential application cases of advanced voice mode in Clinical Practice.&nbsp;</strong></p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

PADE Fine details voice quality

<p>This corpus gathers sentences uttered by two female speakers (S3 and S6) performing six speech acts in USA English using the sentence &ldquo;Mary was dancing&rdquo;. The wav files are provided with their syllabic and phonemic alignment (in TextGrids). The dataset was developed as part of the French ANR project PADE.<br>The recording protocol was described in Rilliard et al. (2013); analysis can be found in Rilliard et al.&nbsp; 2017; this subset was used for the publication Erickson et al. (submitted). The recordings of six pairs of opposed voice qualities, on the word &ldquo;dancing&rdquo;, by a female USA English speaker are also included, and were used for the perceptual evaluation presented in Erickson et al. (2024).<br>References:</p> <p>Erickson, D., Rilliard, A., Thurgood, E., de Moraes, J. A., &amp; Shochi, T. (2024). Acoustic and perceptual profiles of American English social affective expressions: a case study.&nbsp;<em>Journal of Speech Sciences</em>, <em>13</em>(00), e024004. https://doi.org/10.20396/joss.v13i00.20015<br>Rilliard, A., Erickson, D., Shochi, T., &amp; de Moraes, J. A. de. (2013). Social face to face communication&mdash;American English attitudinal prosody. Interspeech 2013, 1648&ndash;1652. https://doi.org/10.21437/Interspeech.2013-427<br>Rilliard, A., Erickson, D., de Moraes, J. A. de, &amp; Shochi, T. (2017). Perception of expressive prosodic speech acts performed in USA English by L1 and L2 speakers. Journal of Speech Sciences, 6(1), 27&ndash;45. https://doi.org/10.20396/joss.v6i1.14981</p>

opencc-by-4.0Mar 2022View details →
OpenNeuro44/100

The human Voice Areas: spatial organisation and inter-individual variability in temporal and extra-temporal cortices

Open the record for dataset details and reuse information.

openPDDLJan 2019View details →
zenodo44/100

GRB assessment of the Saarbrücken Voice Database

<p>GRB&nbsp;evaluations of the Saarbruecken database. Annex to the paper:&nbsp;<em>Multimodal and multi-output deep learning architectures for the automatic assessment of voice quality using the GRB scale.&nbsp;</em></p> <p>Published in&nbsp;IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. 14, NO. 2, FEBRUARY 2020:&nbsp;<strong>DOI:&nbsp;</strong><a href="https://doi.org/10.1109/JSTSP.2019.2956410">10.1109/JSTSP.2019.2956410</a>&nbsp;<br> &nbsp;</p> <p>If you use these labels&nbsp;please cite as:</p> <p>J. D. Arias-Londo&ntilde;o, J. A. G&oacute;mez-Garc&iacute;a and J. I. Godino-Llorente, &quot;Multimodal and Multi-Output Deep Learning Architectures for the Automatic Assessment of Voice Quality Using the GRB Scale,&quot; in&nbsp;<em>IEEE Journal of Selected Topics in Signal Processing</em>, vol. 14, no. 2, pp. 413-422, Feb. 2020.</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

Voice Conversion Challenge 2020 database v1.0

<pre>Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform. In 2016, we have launched the Voice Conversion Challenge (VCC) 2016 [1][2] at Interspeech 2016. The objective of the 2016 challenge was to better understand different VC techniques built on a freely-available common dataset to look at a common goal, and to share views about unsolved problems and challenges faced by the current VC techniques. The VCC 2016 focused on the most basic VC task, that is, the construction of VC models that automatically transform the voice identity of a source speaker into that of a target speaker using a parallel clean training database where source and target speakers read out the same set of utterances in a professional recording studio. 17 research groups had participated in the 2016 challenge. The challenge was successful and it established new standard evaluation methodology and protocols for bench-marking the performance of VC systems. In 2018, we have launched the second edition of VCC, the VCC 2018 [3]. In the second edition, we revised three aspects of the challenge. First, we educed the amount of speech data used for the construction of participant&#39;s VC systems to half. This is based on feedback from participants in the previous challenge and this is also essential for practical applications. Second, we introduced a more challenging task refereed to a Spoke task in addition to a similar task to the 1st edition, which we call a Hub task. In the Spoke task, participants need to build their VC systems using a non-parallel database in which source and target speakers read out different sets of utterances. We then evaluate both parallel and non-parallel voice conversion systems via the same large-scale crowdsourcing listening test. Third, we also attempted to bridge the gap between the ASV and VC communities. Since new VC systems developed for the VCC 2018 may be strong candidates for enhancing the ASVspoof 2015 database, we also asses spoofing performance of the VC systems based on anti-spoofing scores. In 2020, we launched the third edition of VCC, the VCC 2020 [4][5]. In this third edition, we constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. The dataset for intra-lingual VC consists of a smaller parallel corpus and a larger nonparallel corpus, where both of them are of the same language. The dataset for cross-lingual VC consists of a corpus of the source speakers speaking in the source language and another corpus of the target speakers speaking in the target language. As a more challenging task than the previous ones, we focused on cross-lingual VC, in which the speaker identity is transformed between two speakers uttering different languages, which requires handling completely nonparallel training over different languages. This repository contains the training and evaluation data released to participants, target speaker&rsquo;s speech data in English for reference purpose, and the transcriptions for evaluation data. For more details about the challenge and the listening test results please refer to [4] and README file. </pre> <pre>[1] Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi &quot;The Voice Conversion Challenge 2016&quot; in Proc. of Interspeech, San Francisco. [2] Mirjam Wester, Zhizheng Wu, Junichi Yamagishi &quot;Analysis of the Voice Conversion Challenge 2016 Evaluation Results&quot; in Proc. of Interspeech 2016. [3] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, Zhenhua Ling, &quot;The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods&quot;, Proc Speaker Odyssey 2018, June 2018. [4] Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda. &quot;Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion&quot; Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 80-98, DOI: 10.21437/VCC_BC.2020-14.</pre>

openother-openDec 2020View details →
zenodo44/100

Voice Conversion Challenge 2020 Listening Test Data

<pre>Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform. In 2016, we have launched the Voice Conversion Challenge (VCC) 2016 [1][2] at Interspeech 2016. The objective of the 2016 challenge was to better understand different VC techniques built on a freely-available common dataset to look at a common goal, and to share views about unsolved problems and challenges faced by the current VC techniques. The VCC 2016 focused on the most basic VC task, that is, the construction of VC models that automatically transform the voice identity of a source speaker into that of a target speaker using a parallel clean training database where source and target speakers read out the same set of utterances in a professional recording studio. 17 research groups had participated in the 2016 challenge. The challenge was successful and it established new standard evaluation methodology and protocols for bench-marking the performance of VC systems. In 2018, we have launched the second edition of VCC, the VCC 2018 [3]. In the second edition, we revised three aspects of the challenge. First, we educed the amount of speech data used for the construction of participant&#39;s VC systems to half. This is based on feedback from participants in the previous challenge and this is also essential for practical applications. Second, we introduced a more challenging task refereed to a Spoke task in addition to a similar task to the 1st edition, which we call a Hub task. In the Spoke task, participants need to build their VC systems using a non-parallel database in which source and target speakers read out different sets of utterances. We then evaluate both parallel and non-parallel voice conversion systems via the same large-scale crowdsourcing listening test. Third, we also attempted to bridge the gap between the ASV and VC communities. Since new VC systems developed for the VCC 2018 may be strong candidates for enhancing the ASVspoof 2015 database, we also asses spoofing performance of the VC systems based on anti-spoofing scores. In 2020, we launched the third edition of VCC, the VCC 2020 [4][5]. In this third edition, we constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. The dataset for intra-lingual VC consists of a smaller parallel corpus and a larger nonparallel corpus, where both of them are of the same language. The dataset for cross-lingual VC consists of a corpus of the source speakers speaking in the source language and another corpus of the target speakers speaking in the target language. As a more challenging task than the previous ones, we focused on cross-lingual VC, in which the speaker identity is transformed between two speakers uttering different languages, which requires handling completely nonparallel training over different languages. As for listening test, we subcontracted the crowd-sourced perceptual evaluation with English and Japanese listeners to Lionbridge TechnologiesInc. and Koto Ltd., respectively. Given the extremely large costs required for the perceptual evaluation, we selected 5 utterances (E30001, E30002, E30003,E30004, E30005) only from each speaker of each team. To evaluate the speaker similarity of the cross-lingual task, we used audio in both the English language and in the target speaker&rsquo;s L2language as reference. For each source-target speaker pair, we selected three English recordings and two L2 language recordings as the natural reference for the converted five utterances. </pre> <p>This data repository includes the audio files used for the&nbsp;crowd-sourced perceptual evaluation and raw&nbsp;listening test scores.&nbsp;</p> <pre>[1] Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi &quot;The Voice Conversion Challenge 2016&quot; in Proc. of Interspeech, San Francisco. [2] Mirjam Wester, Zhizheng Wu, Junichi Yamagishi &quot;Analysis of the Voice Conversion Challenge 2016 Evaluation Results&quot; in Proc. of Interspeech 2016. [3] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, Zhenhua Ling, &quot;The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods&quot;, Proc Speaker Odyssey 2018, June 2018. [4] Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda. &quot;Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion&quot; Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 80-98, DOI: 10.21437/VCC_BC.2020-14. [5] Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang, Zhenhua Ling, Junichi Yamagishi, Yi Zhao, Xiaohai Tian, and Tomoki Toda. &quot;Predictions of subjective ratings and spoofing assessments of voice conversion challenge 2020 submissions.&quot; Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 99-120, DOI: 10.21437/VCC_BC.2020-15. </pre>

openother-openDec 2020View details →
zenodo44/100

German own voice recordings with hearable microphones

<p>This dataset is supplementary material to the article "Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with an In-ear Microphone" published in Acta Acustica, vol. 8 (2024).</p> <p>The dataset consists of recordings of own voice speech of 18 talkers (5 female, 13 male) wearing hearable devices in both ears. All talkers were native German speakers. The dataset was recorded in a sound-proof listening booth using the Hearpiece prototype device (closed vent variant) [1].</p> <p>The speech uttered by the talkers is pre-determined text read from a screen. The talkers press a button on the screen to start the recording, read the sentence out loud in a normal voice, and press another button to stop the recording. It was possible for the talkers to re-record a sentence if desired. The sentences read by the talkers originate from the following sources:</p> <ul> <li>The north wind and the sun (German), 6 sentences</li> <li>Berlin and Marburg Sentences [2] (German), 2x100 sentences</li> <li>100 sentences for language learners [3] (German), 100 sentences</li> <li>Some held-out vowels and consonants</li> <li>[some seconds of silence]</li> </ul> <p>The full text read by each talker is written in <code>full_text.txt</code>.</p> <p>The recordings are contained in the folder <code>speech</code>. Each subfolder contains recordings from a different talker (e.g., <code>VP_01</code>). The sentences uttered by each talker are numbered following this scheme:&nbsp;<code>VP_01_0.wav</code> to <code>VP_01_313.wav</code>. Talkers where the device could not be inserted, or where the fit did not provide sufficient attenuation of external sounds to the in-ear microphone, were excluded.</p> <p>A DPA 6060 lavalier clip microphone and a Tbone SC140 cardiod microphone were recorded as reference signals. From two Hearpiece devices (closed vent), the concha and in-ear microphones were recorded. Audio was recorded at a sampling frequency of 44100 Hz.</p> <p>The channels of the recordings, counting from 0, recorded the following microphones:</p> <ul> <li>0: Lavalier-microphone clipped to the shirt neck, shirt collar etc. of the talker</li> <li>1: Reference microphone about 50 cm in front of the talker</li> <li>2: Left in-ear microphone Hearpiece</li> <li>3: Left concha microphone Hearpiece</li> <li>4: Right in-ear microphone Hearpiece</li> <li>5: Right concha microphone Hearpiece</li> </ul> <p>[1] F. Denk, M. Lettau, H. Schepker, S. Doclo, R. Roden, M. Blau, J.-H. Bach, J. Wellmann, and B. Kollmeier: "A One-Size-Fits-All Earpiece with Multiple Microphones and Drivers for Hearing Device Research". In: Proc. AES International Conference on Headphone Technology. San Francisco, USA, Aug. 2019.</p> <p>[2] A. P. Simpson, K. J. Kohler, and T. Rettstadt. "The Kiel Corpus of Read/Spontaneous Speech: Acoustic Data Base, Processing Tools, and Analysis Results". In: Arbeitsberichte Institut f&uuml;r Phonetik Und Digitale Sprachverarbeitung Universit&auml;t Kiel. Vol. 32. IPDS, Nov. 1997, pp. 243-247.</p> <p>[3] A. Neustein. "100 S&auml;tze Reichen F&uuml;r Ein Ganzes Leben" (Blog-post). https://deutschlernerblog.de/100-saetze-reichen-fuer-ein-ganzes-leben/. Aug. 2019.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Mar 2024View details →
zenodo44/100

ASVspoof2019LA-Sim: Augmented Dataset for An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems

<p>This is the dataset we augmented to study the channel effects for anti-spoofing. For more details, please refer to our Interspeech 2021 paper: &quot;An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems&quot;.</p> <p>Proceeding: <a href="https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html">https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html</a></p> <p>Arxiv: <a href="https://arxiv.org/pdf/2104.01320.pdf">https://arxiv.org/pdf/2104.01320.pdf</a></p> <p>Code:&nbsp;<a href="https://github.com/yzyouzhang/Empirical-Channel-CM">https://github.com/yzyouzhang/Empirical-Channel-CM</a></p> <p>Contact: you.zhang@rochester.edu</p> <p><strong>Version 1.0</strong> contains the <strong>training</strong> and the <strong>development</strong> set. We have added the <strong>evaluation</strong> set in <strong>version 1.1 </strong>but deleted the training set due to the size limitation, but you can still access the training set in version 1.0.</p> <p>Please check it out.</p> <p>To extract the files, please use the following commands:</p> <pre><code class="language-bash">cat eval.tar.gz-part* &gt; eval.tar.gz tar -xvzf *.tar.gz</code></pre> <p>After concatenation, to make sure the download is complete, you can check with the following:</p> <pre><code>md5sum *.tar.gz 15dea7d28b126994bb6b159778f706af dev.tar.gz 0615052b34ca6c7f58505eaa8647844f eval.tar.gz 3058dd9d407f3c9ae697acca8c34a6c3 train.tar.gz</code></pre> <p>Thanks.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

How do native and non-native speakers recognize emotions in the instructor's voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners [dataset]

<p>Dataset for the journal article&nbsp;<em>How do native and non-native speakers recognize emotions in the instructor&rsquo;s voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners.</em></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Jamendo Corpus for Singing Voice Detection

<p>This is a public corpus of 93 creative-commons licensed music pieces annotated<br> by voice (sung or spoken) and no-voice.</p>

opencc-by-4.0Mar 2008View details →
zenodo44/100

Thorsten-Voice Dataset 2022.10

<p>The goal of project &quot;Thorsten-Voice&quot; is to provide voice datasets and TTS models for free and high quality german artificial voice. This dataset &quot;Thorsten-Voice dataset 2022.10&quot; is a neutrally spoken voice dataset recorded by Thorsten M&uuml;ller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.</p> <blockquote> <p><strong>&quot;I contribute my personal voice as a person believing in a world where all people are equal. No matter of gender, sexual orientation, religion, skin color and geocoordinates of birth location. A global world where everybody is warmly welcome on any place on this planet and open and free knowledge and education is available to everyone.&quot; (Thorsten M&uuml;ller)</strong></p> </blockquote> <p>&nbsp;</p> <p><strong>Dataset details</strong>:</p> <ul> <li>ljspeech file and directory structure</li> <li>12.450 recorded phrases (wav files)</li> <li>more than 11 hours of pure audio</li> <li>samplerate 22.050Hz</li> <li>mono</li> <li>normalized to -24dB</li> <li>no silence at beginning/ending</li> <li>avg spoken chars per second: 17,5</li> </ul> <p>See more details on my <a href="https://github.com/thorstenMueller/Thorsten-Voice">Github page</a> or <a href="https://www.Thorsten-Voice.de">Thorsten-Voice</a> project website.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Impact of Passive Voice in Requirements Engineering

<p>This repository contains instrumentation material and results for the experiment described in the paper &quot;On The Impact of Passive Voice Requirements on Domain Modelling&quot; by Henning Femmer, Jan Kucera, and Antonio Vetr&ograve; from Technische Universit&auml;t M&uuml;nchen.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

COALA voice data and transcripts Dutch

<p>This dataset contains audio files and transcripts in Dutch and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut f&uuml;r Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years&nbsp; (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: &quot;Have you ever worked in a factory, manufacturing setting?&quot; (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: &quot;If yes, for how many years?&quot; (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>

opencc-zeroOct 2023View details →
zenodo44/100

COALA voice data and transcripts Italian

<p>This dataset contains audio files and transcripts in Italian and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut f&uuml;r Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years&nbsp; (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: &quot;Have you ever worked in a factory, manufacturing setting?&quot; (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: &quot;If yes, for how many years?&quot; (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>

opencc-zeroOct 2023View details →
OpenNeuro40/100

InterTVA. A multimodal MRI dataset for the study of inter-individual differences in voice perception and identification.

Open the record for dataset details and reuse information.

openhttps://creativecommons.org/licenses/by-nc-sa/4.0/Jan 2019View details →
zenodo40/100

RealVAD: A Real-world Dataset for Voice Activity Detection

<p><strong>RealVAD: A Real-world Dataset for Voice Activity Detection</strong></p> <p>The task of automatically detecting &ldquo;Who is Speaking and When&rdquo; is broadly named as Voice Activity Detection (VAD). Automatic VAD is a very important task and also the foundation of several domains, e.g., human-human, human-computer/ robot/ virtual-agent interaction analyses, and industrial applications.</p> <p>RealVAD dataset is constructed from a YouTube video composed of a panel discussion lasting approx. 83 minutes. The audio is available from a single channel. There is one static camera capturing all panelists, the moderator and audiences.</p> <p>Particular aspects of RealVAD dataset are:</p> <ul> <li>&nbsp;It is composed of panelists with different nationalities (British, Dutch, French, German, Italian, American, Mexican, Columbian, Thai). This aspect allows studying the effect of ethnic origin variety to the automatic VAD.</li> <li>&nbsp;There is a gender balance such that there are four female and five male panelists.</li> <li>&nbsp;The panelists are sitting in two rows and they can be gazing audience, other panelists, their laptop, the moderator or anywhere in the room while speaking or not-speaking. Therefore, they were captured not only from frontal-view but also from side-view varying based on their instant posture and head orientation.</li> <li>&nbsp;The panelists are moving freely and are doing various spontaneous actions (e.g., drinking water, checking their cell phone, using their laptop, etc.), resulting in different postures.</li> <li>&nbsp;The panelists&rsquo; body parts are sometimes partially occluded by their/other&#39;s body part or belongings (e.g., laptop).</li> <li>&nbsp;There are also natural changes of illumination and shadow rising on the wall behind the panelists in the back row.</li> <li>&nbsp;Especially, for the panelists sitting in the front row, there is sometimes background motion occurring when the person(s) behind them moves.</li> </ul> <p>The annotations includes:</p> <ul> <li>&nbsp;The upper body detection of nine panelists in bounding box form.</li> <li>&nbsp;Associated VAD ground-truth (speaking, not-speaking) for nine panelists.</li> <li>&nbsp;Acoustic features extracted from the video: MFCC and raw filterbank energies.</li> </ul> <p><em>All info regarding the annotations are given in the ReadMe.txt and Acoustic Features README.txt files.</em></p> <p><strong>When using this dataset for your research, please cite the following paper in your publication:</strong></p> <ol> <li>C. Beyan, M. Shahid and V. Murino, &quot;RealVAD: A Real-world Dataset and A Method for Voice Activity Detection by Body Motion Analysis&quot;, in IEEE Transactions on Multimedia, 2020.</li> </ol>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Papuan Voices Media Files

<p>Papuan Voices Media Files (wav) - Supplement to dataset <a href="https://doi.org/10.5281/zenodo.4350691">Papuan Voices</a><br> <br> Papuan Voices presents phonetically-transcribed primary recordings, from numerous places throughout Papua island<br> <br> Available online at: <a href="https://papuanvoices.clld.org">https://papuanvoices.clld.org</a></p>

opencc-by-nc-4.0Dec 2020View details →
zenodo40/100

Mixe-Zoquean Voices Media Files

<p>Mixe-Zoquean Voices Media Files (wav) - Supplement to dataset <a href="https://doi.org/10.5281/zenodo.4350652">Mixe-Zoquean Voices</a><br> <br> Mixe-Zoquean Voices presents primary recordings of languages from the Mixe-Zoque language family.<br> <br> Available online at: <a href="https://mixezoqueanvoices.clld.org">https://mixezoqueanvoices.clld.org</a></p>

opencc-by-nc-4.0Dec 2020View details →
zenodo40/100

Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns"

<p>The .pdf&nbsp;file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> </ul> <p>&nbsp;</p> <p>The .zip&nbsp;files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p>&nbsp;</p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots.&nbsp;</li> </ul>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Emotional Voice Messages (EMOVOME) database

<p>The Emotional Voice Messages (EMOVOME) database is a speech dataset collected for emotion recognition in real-world conditions. It contains 999 spontaneous voice messages from 100 Spanish speakers, collected from real conversations on a messaging app. EMOVOME includes both expert and non-expert emotional annotations, covering valence and arousal dimensions, along with emotion categories for the expert annotations. Detailed participant information is provided, including sociodemographic data and personality trait assessments using the NEO-FFI questionnaire. Moreover, EMOVOME provides audio recordings of participants reading a given text, as well as transcriptions of all 999 voice messages. Additionally, baseline models for valence and arousal recognition are provided, utilizing both speech and audio transcriptions.</p> <h2>Description</h2> <p>For details on the EMOVOME database, please refer to the article:</p> <blockquote> <p><em>"EMOVOME Database: Advancing Emotion Recognition in Speech Beyond Staged Scenarios". Luc&iacute;a G&oacute;mez-Zaragoz&aacute;, Roc&iacute;o del Amor, Mar&iacute;a Jos&eacute; Castro-Bleda, Valery Naranjo, Mariano Alca&ntilde;iz Raya, Javier Mar&iacute;n-Morales. (pre-print available in <a href="https://doi.org/10.48550/arXiv.2403.02167" target="_blank" rel="noopener">https://doi.org/10.48550/arXiv.2403.02167</a>)<br></em></p> </blockquote> <h2>Content</h2> <p>The Zenodo repository contains four files:</p> <ul> <li><strong>EMOVOME_agreement.pdf</strong>: agreement file required to access the original audio files, detailed in section Usage Notes.&nbsp;</li> <li><strong>labels.csv</strong>: ratings of the three non-experts and the expert annotator, independently and combined.</li> <li><strong>participants_ids.csv</strong>: table mapping each numerical file ID to its corresponding alphanumeric participant ID.</li> <li><strong>transcriptions.csv</strong>:<strong> </strong>transcriptions of each audio.</li> </ul> <p>The repository also includes three folders:</p> <ul> <li><strong>Audios</strong>: it contains the file&nbsp;<strong>features_eGeMAPSv02.csv</strong> corresponding to the standard acoustic feature set used in the baseline model, and two folders: <ul> <li><strong>Lecture</strong>: contains the audio files corresponding to the text readings, with each file named according to the participant's ID.</li> <li><strong>Emotions</strong>: contains the voice recordings from the messaging app provided by the user, named with a file ID.</li> </ul> </li> <li><strong>Questionnaires</strong>: it contains two files: 1)&nbsp;<strong>sociodemographic_spanish.csv</strong> and&nbsp;<strong>sociodemographic_english.csv</strong> are the sociodemographic data of participants in Spanish and English, respectively, including the demographic information; and &nbsp;2) <strong>NEO-FFI_spanish.csv</strong> includes the participants&rsquo; answers to the Spanish version of the NEO-FFI questionnaire. The three files include a column indicating the participant's ID to link the information.</li> <li><strong>Baseline_emotion_recognition</strong>: it includes three files and two folders. The file&nbsp;<strong>partitions.csv</strong> specifies the proposed data partition. Particularly, the dataset is divided into 80% for development and 20% for testing using a speaker-independent approach, i.e., samples from the same speaker are not included in both development and test. The development set includes 80 participants (40 female, 40 male) containing the following distribution of labels: 241 negative, 305 neutral and 261 positive valence; and 148 low, 328 neutral and 331 high arousal. The test set includes 20 participants (10 female, 10 male) with the distribution of labels that follows: 57 negative, 62 neutral and 73 positive valence; and 13 low, 70 neutral and 109 high arousal. Files&nbsp;<strong>baseline_speech.ipynb</strong> and&nbsp;<strong>baseline_text.ipynb</strong> contain the code used to create the baseline emotion recognition models based on speech and text, respectively. The actual trained models for valence and arousal prediction are provided in folders&nbsp;<strong>models_speech</strong> and&nbsp;<strong>models_text</strong>.&nbsp;</li> </ul> <p><em>Audio files in &ldquo;Lecture&rdquo; and &ldquo;Emotions&rdquo; are only provided to the users that complete the agreement file in section Usage Notes. Audio files are in Ogg Vorbis format at 16-bit and 44.1 kHz or 48 kHz. The total size of the &ldquo;Audios&rdquo; folder is about 213 MB.&nbsp;</em></p> <h2>Usage Notes</h2> <p>All the data included in the EMOVOME database is publicly available under the Creative Commons Attribution 4.0 International license. The only exception is the original raw audio files, for which an additional step is required as a security measure to safeguard the speakers' privacy. To request access, interested authors should first complete and sign the agreement file <strong>EMOVOME_agreement.pdf</strong> and send it to the corresponding author (<em><a href="mailto:jamarmo@htech.upv.es" target="_blank" rel="noopener">jamarmo@htech.upv.es</a></em>). The data included in the EMOVOME database is expected to be used for research purposes only. Therefore, the agreement file states that the authors are not allowed to share the data with profit-making companies or organisations. They are also not expected to distribute the data to other research institutions; instead, they are suggested to kindly refer interested colleagues to the corresponding author of this article. By agreeing to the terms of the agreement, the authors also commit to refraining from publishing the audio content on the media (such as television and radio), in scientific journals (or any other publications), as well as on other platforms on the internet. The agreement must bear the signature of the legally authorised representative of the research institution (e.g., head of laboratory/department). Once the signed agreement is received and validated, the corresponding author will deliver the "Audios" folder containing the audio files through a download procedure. A direct connection between the EMOVOME authors and the applicants guarantees that updates regarding additional materials included in the database can be received by all EMOVOME users.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record