Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
503
datasets available to search
ShareScore release 0.9.0
Dataset results
503 results for “Voice”
Papuan Voices
<p>Cite the source of the dataset as:</p> <blockquote> <p>Yusuf Sawaki, Jeanete Lekeneny, Apriani Arilaha, Jimmi Kirihio, Emanuel Tutorop, Boas Wabia, Infak Mayor, Simon Tabuni, Stevani Karesina, Paul Heggarty, Darja Dërmaku-Appelganz (2020). Papuan Voices.</p> </blockquote>
Legacies of Catalogue Descriptions and Curatorial Voice: Infographics
<p>The project '<a href="https://cataloguelegacies.github.io/">Legacies of Catalogue Descriptions and Curatorial Voice: Opportunities for Digital Scholarship</a>' (2020-22, AHRC, Project Reference AH/T013036/1) sought to develop a platform for a transformational impact in digital scholarship within cultural institutions by opening up new and important directions for computational, critical, and curatorial analysis of collection catalogues. Our pilot research investigated the temporal and spatial legacy of a landmark catalogue: the ‘Catalogue of Political and Personal Satires Preserved in the Department of Prints and Drawings in the British Museum’, entries in which form the basis of related catalogue data at institutions including the Lewis Walpole Library and the British Library.</p> <p>Towards the end of the project we invited together members of our community to discuss shared agendas and actions: where we should focus our collective resources, what things we need to work differently, and the role of computational technologies in both shaping and constraining change.</p> <p>These infographics respond to those conversations. 'Legacy catalogue entries: actions and agendas' presents priorities that emerged from a longlist of project findings. And 'Tracing the transmission of legacy catalogue entries' visualises the journey of legacy catalogue data in our case study from conception, through cataloging infrastructures, and into the future.</p> <p>Both graphics were designed by Lucy Havens, with editorial support from James Baker and Rossitza Atanassova. The process is described at Lucy Havens, 'Designing Infographics for the “Legacies of Catalogue Descriptions” Project' (<a href="https://traininthedistance.wordpress.com/2022/03/15/designing-infographics-for-the-legacies-of-catalogue-descriptions-project/">March 2022</a>).</p> <p>We thank our community of participants for their encouragement and suggestions during the design process.</p> <p>Larger version of each graphic are available for print purposes on request. Contact <a href="mailto:j.w.baker@soton.ac.uk">j.w.baker@soton.ac.uk</a> to request a copy.</p>
Electrobyte for Singing Voice Detection
<p>This is a public dataset of 90 copyright-free electronic songs with vocal annotations (sing/no sing).</p>
See no evil in the voice-to-voice customer service context
<p>A sample of more than 28,000 front-line employee (FLE) - customer interactions, extrapolating from foundational framing, we pit conventional service approaches against one another to propose a dual-process model, situating customer frustration/satisfaction as mediators of the indirect relationships between resolution/relational tactics and call duration – a key customer service efficiency outcome.</p>
Annotated-VocalSet: A Singing Voice Dataset
<p>This dataset provides annotations for the <a href="https://doi.org/10.5281/zenodo.1442513">VocalSet dataset</a>, which is available online at</p> <pre><a href="https://doi.org/10.5281/zenodo.1442513">https://doi.org/10.5281/zenodo.1442513</a></pre> <p>.</p> <p>The annotations generated for the VocalSet audio files include fundamental frequency contour, note onset, note offset, the transition between notes, note F0, note duration, Midi pitch, and lyrics.</p> <p><a href="https://doi.org/10.5281/zenodo.1442513">VocalSet</a> consists of more than 10 hours of monophonic recorded audio of professional singers in a variety of vocal techniques (n = 17) and several singers (m = 20) with several WAV files (p = 3560). However, although several categories, including techniques, singers, tempo, and loudness, are considered in the dataset, the sung notes were not annotated. Therefore, this dataset aims to annotate VocalSet to make it a more powerful dataset for researchers.</p> <p>Details of the dataset are provided in the following academic journal paper.</p> <p><a href="https://www.mdpi.com/2076-3417/12/18/9257">Faghih, Behnam, and Joseph Timoney. 2022. "Annotated-VocalSet: A Singing Voice Dataset" <em>Applied Sciences</em> 12, no. 18: 9257. https://doi.org/10.3390/app12189257</a></p> <p>Please use the above paper to cite this dataset.</p>
GWAS summary statistics and code for "Sequence variants affecting voice pitch in humans"
<p>Contents: GWAS summary statistics for voice pitch (median F0 in reading) and code for acoustic analysis</p> <p>Please refer to the corresponding publication:</p> <p>Gisladottir et al. Sequence variants affecting voice pitch in humans. <em>Science Advances</em></p> <p>The GWAS summary statistics is also available at: https://www.decode.com/summarydata/</p> <p>The code for acoustic analysis is also available at: https://github.com/cadia-lvl/deCODE</p> <p> </p> <p> </p> <p> </p>
Figure 1. The Structure of the Automatic Translate Voice to Sign Language Animation System-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives
<p>All technologies of voice recognition, speaker identification and verification, each has its<br> own advantages and disadvantages and may requires different treatments and techniques. The<br> choice of which technology to use is application-specific. At the highest level, all voice recognition<br> systems contain two main modules: feature extraction and feature matching. Feature extraction is<br> the process that extracts a small amount of data from the voice signal that can later be used to<br> represent each word. Feature matching involves the actual procedure to identify the unknown word<br> by comparing extracted features from his/her voice input with the ones from a set of known words.<br> A wide range of possibilities exist for parametrically representing the speech signal for the<br> voice recognition task, such as Linear Prediction Coding (LPC), RASTA-PLP and Mel-Frequency<br> Cepstrum Coefficients (MFCC).</p>
Efficacy of Transformational Breath® for anxiety management in professional voice users
<p>Raw data for publication of original research entitled "<span>Efficacy of Transformational Breath<sup>®</sup> for anxiety management in professional voice users".</span></p> <p><span>Randomised controlled trial</span></p> <p><span>Quantitative and qualitative data sets relating to responses of treatment and control groups.</span></p>
Axiom voice recognition dataset
<p>The AXIOM Voice Dataset has the main purpose of gathering audio recordings from Italian natural language speakers. This voice data collection intended to obtain audio reconding sample for the training and testing of VIMAR algorithm implemented for the Smart Home scenario for the Axiom board. The final goal was to developing an efficient voice recognition system using machine learning algorithms. A team of UX researchers of the University of Siena collected data for five months and tested the voice recognition system on the AXIOM board [1]. The data acquisition process involved natural Italian speakers who provided their written consent to participate in the research project. The participants were selected in order to maintain a cluster with different characteristics in gender, age, region of origin and background. </p>
Voice-activated applications and Multipath TCP: A good match?
<p>The attachment is the dataset of the paper "Voice-activated applications and Multipath TCP: A good match?"<br> which has been accepted in Workshop on Mobile Network Measurement 2018.</p> <p>Homepage: https://www.info.ucl.ac.be/~tranviet/</p> <p>LKL-MPTCP stack: https://github.com/hoang-tranviet/mptcp/tree/lkl_4.13-mptcp_v0.93_API<br> Simulated voice traffic program: https://github.com/hoang-tranviet/iperf-siri<br> Client Docker script: https://github.com/hoang-tranviet/lkl-docker-monroe</p>
VocalSet: A Singing Voice Dataset
<p><strong>NEW IN VocalSet 1.2:</strong> We now have 3 file organization versions:</p> <ol> <li>Files organized by singer</li> <li>Files organized by technique</li> <li>Files organized by vowel</li> </ol> <p>We hope that this will ease the process of training and testing models using these different attributes of the dataset.</p> <p> </p> <p><strong>Overview:</strong></p> <p>We present VocalSet, a singing voice dataset consisting of 10.1 hours of monophonic recorded audio of professional singers demonstrating both standard and extended vocal techniques on all 5 vowels. Existing singing voice datasets aim to capture a focused subset of singing voice characteristics, and generally consist of just a few singers. VocalSet contains recordings from 20 different singers (9 male, 11 female) and a range of voice types. VocalSet aims to improve the state of existing singing voice datasets and singing voice research by capturing not only a range of vowels, but also a diverse set of voices on many different vocal techniques, sung in contexts of scales, arpeggios, long tones, and excerpts.</p> <p>We have included two .txt files 'train_singers_technique.txt 'and 'test_singers_technique.txt' in which you will find a list of the singers we used to train and test our technique classifier on. 'DataSetVocalises.pdf' contains the sheet singers sang from in their recording sessions. 'readme-anon.txt' contains more information about the dataset, including the mapping from filename to singer voice type as well as more information on the vocalises that will help you map files to sheet music. Enjoy and please cite accordingly!</p>
InterTVA. Single-trial beta maps for event-related voice localizer.
<p>The InterTVA dataset has been acquired with two main objectives. First, from a neuroscientific perspective, it aims at studying the inter-individual differences observed in people's ability at performing voice perception and voice identification tasks. Secondly, from a methodological perspective, it should allow benchmarking multi-view machine learning methods. Indeed, it includes several MRI modalities: anatomical MRI, diffusion MRI and several sessions of functional MRI -- one resting state run, one event-related voice localizer run (passive listening of vocal and non-vocal sounds), and four runs during which the subject performed a voice identification task.</p> <p>The present dataset contains a pre-processed version of the data acquired during the event-related voice localizer run. A GLM was performed using one regressor for each of the 144 trials, during each of which the participant was passively listening to a vocal or non-vocal stimulus. We therefore provide 144 beta maps for 39 subjects out of the 40 participants (one was excluded for excessive motion). This allows performing group-level MVPA using an inter-subject pattern analysis (ISPA) approach, as described in the following paper:</p> <p>Q. Wang, B. Cagna, T. Chaminade, et S. Takerkart, « Inter-subject pattern analysis: A straightforward and powerful scheme for group-level MVPA », <em>NeuroImage</em>, vol. 204, p. 116205, janv. 2020. https://doi.org/10.1016/j.neuroimage.2019.116205</p> <p>The source code to perform ISPA is also available: http://www.github.com/SylvainTakerkart/inter_subject_pattern_analysis</p> <p>If you use this data, please cite the following paper which describes it exhaustively:</p> <p>V. Aglieri, B. Cagna, P. Belin, S. Takerkart, "Single-trial fMRI activation maps measured during the InterTVA event-related voice localizer. A data set ready for inter-subject pattern analysis", Data in Brief, vol. 29, p. 105170, April 2020. https://doi.org/10.1016/j.dib.2020.105170</p>
Mobile Device Voice Recordings at King's College London (MDVR-KCL) from both early and advanced Parkinson's disease patients and healthy controls
<p><strong>Dataset description</strong></p> <p>The dataset description will start with describing the local conditions and other metadata, then will continue with describing the recording procedure and annotation methodology. Finally, a brief description of the dataset deployment and publication will be given.</p> <p><strong>Meta Information</strong></p> <p>The dataset was recorded at King's College London (KCL) Hospital, Denmark Hill, Brixton, London SE5 9RS in the period from 26 to 29 September 2017. We used a typical examination room with about ten square meters area and a typical reverberation tome of approx. 500ms to perform the voice recordings. Due to the fact, that the voice recordings are performed in the realistic situation of doing a phone call (i.e. participant holds the phone to the preferred ear and microphone is in direct proximity to the mouth), one can assume that all recordings were performed within the reverberation radius and thus can be considered as “clean”.</p> <p><strong>Recording Procedure</strong></p> <p>We used a Motorola Moto G4 Smartphone as recording device. To perform the voice recordings on the device, we developed a “Toggle Recording App”, which uses the same functionalities as the voice recording module used within the i-PROGNOSIS Smartphone application, but deployed as a standalone android application. This means, that the voice capturing service runs as a standalone background service on the recording device and triggers voice recordings via on- and off-hook signals of the Smartphone. Due to the fact, that we directly record the microphone signal, and not the GSM (“Global System for Mobile Communications”) compressed stream, we end up with high quality recordings with a sample rate of 44.1 kHz and a bit depth of 16 Bit (audio CD quality). The raw, uncompressed data is directly written to the external storage of the Smartphone (SD-card) using the well-known WAVE file format (.wav). We used the following workflow to perform a voice recording:</p> <ul> <li>Ask the participant to relax a bit and then to make a phone call to the test executor (off-hook signal triggered).}</li> <li>Ask the participant to read out “The North Wind and the Sun”</li> <li>Depending on the constitution of the participant either ask to read out “Tech. Engin. Computer applications in geography snippet”</li> <li>Start a spontaneous dialog with the participant, the test executor starts asking random questions about places of interest, local traffic, or personal interests if acceptable.</li> <li>Test executor ends call by farewell (on-hook signal triggered).</li> </ul> <p><strong>Annotation Scheme</strong></p> <p>For each HC and PD participant, we labeled the data regarding scores on the Hoehn & Yahr (H&Y), as well as the UPDRS II part 5 and UPDRS III part 18 scale. The voice recordings are labeled in the following scheme:</p> <p>SI_ HS_ HYR_ UPDRS II-5_UPDRS III-18</p> <p>with</p> <ul> <li>SI as subject identification in the form ID<em>NN</em>, <em>N</em> in [0, 9]</li> <li>HS as the health status label (hc or pd accordingly)</li> <li>HYR as the expert assessed H&Y scale rating</li> <li>UPDRS II-5 as the according expert peer-reviewed score</li> <li>UPDRS III-18 as the according expert assessed score</li> </ul> <p>For example, an audio recording with the file name “ID02_pd_1_2_1.wav” represents a recording of the third participant (First participant was anonymized as ID00), which has PD and a H&Y rating of 1, a UPDRS II-5 score of 2 and a UPDRS III-18 score of 1. At this point, it should be noted, that also all healthy controls were evaluated with regard to the introduced scales, because Parkinson's disease and voice degradation correlate, but don't match exactly. This means, that the data set includes one HC participant (ID31) with UPDRS II-5 and III-18 rating of 1, and also includes PD patients with UPDRS II-5 and III-18 ratings of 0. It should be emphasized, that this does not mean the data set includes ambiguous information, but that an expert was not able to hear voice degradation that would end up in a UPDRS rating greater than zero. Machine learning approaches may be able to nevertheless classify correctly, or at least learn to correlate, but not match PD and voice degradation at any time.</p> <p><strong>Appendix</strong></p> <p>North Wind and the Sun (Orthographic Version):</p> <p>“The North Wind and the Sun were disputing which was the stronger, when a traveler came along wrapped in a warm cloak. They agreed that the one who first succeeded in making the traveler take his cloak off should be considered stronger than the other. Then the North Wind blew as hard as he could, but the more he blew the more closely did the traveler fold his cloak around him; and at last the North Wind gave up the attempt. Then the Sun shone out warmly, and immediately the traveler took off his cloak. And so the North Wind was obliged to confess that the Sun was the stronger of the two.”</p> <p>BNC – Tech. Engin. Computer applications in geography snippet:</p> <p>“[...] This is because there is less scattering of blue light as the atmospheric path length and consequently the degree of scattering of the incoming radiation is reduced. For the same reason, the sun appears to be whiter and less orange-coloured as the observer's altitude increases; this is because a greater proportion of the sunlight comes directly to the observer's eye. Figure 5.7 is a schematic representation of the path of electromagnetic energy in the visible spectrum as it travels from the sun to the Earth and back again towards a sensor mounted on an orbiting satellite. The paths of waves representing energy prone to scattering (that is, the shorter wavelengths) as it travels from sun to Earth are shown. To the sensor it appears that all the energy has been reflected from point P on the ground whereas, in fact, it has not, because some has been scattered within the atmosphere and has never reached the ground at all. [...]”</p>
Dataset for "Acoustic voice variation within and between speakers"
<p>This dataset accompanies the article "Acoustic voice variation within and between speakers" in the <em>Journal of Acoustical Society of America</em>.</p>
Webis-Voice-based-and-Conversational-Argument-Search-20
<p>Interface, questionnaires, and collected data for the paper "<a href="https://webis.de/publications.html#?q=stein2020b">Investigating Expectations for Voice-based and Conversational Argument Search on the Web</a>".</p> <p>Data is mostly in tab-separated values format (like csv, just with tabs). For the transcripts of the user study, only the sequence of action labels (in xml) is released for privacy reasons. The labels are described in user-study-tags-participant.txt and user-study-tags-system.txt.</p>
Santiago Laxopa Zapotec voice quality data
<p>This contains a csv file containing the acoustic measures that were generated from VoiceSauce. </p>
Linked collectors and determiners for: A new Amazonian species of Adenomera (Anura: Leptodactylidae) from the Brazilian state of Pará: a tody-tyrant voice in a frog.
Natural history specimen data linked to collectors and determiners held within, "A new Amazonian species of Adenomera (Anura: Leptodactylidae) from the Brazilian state of Pará: a tody-tyrant voice in a frog". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/ff77f600-4cf8-4437-ab29-fe9f679bbe4c">https://bionomia.net/dataset/ff77f600-4cf8-4437-ab29-fe9f679bbe4c</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/ff77f600-4cf8-4437-ab29-fe9f679bbe4c">https://gbif.org/dataset/ff77f600-4cf8-4437-ab29-fe9f679bbe4c</a>. Formatted as a Frictionless Data package.
Thorsten-Voice Dataset 2021.02
<p>Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.</p> <blockquote> <p><strong>"I contribute my personal voice as a person believing in a world where all people are equal. No matter of gender, sexual orientation, religion, skin color and geocoordinates of birth location. A global world where everybody is warmly welcome on any place on this planet and open and free knowledge and education is available to everyone." (Thorsten Müller)</strong></p> </blockquote> <p> </p> <p><strong>Dataset details</strong>:</p> <ul> <li>ljspeech file and directory structure</li> <li>22.668 recorded phrases (wav files)</li> <li>more than 23 hours of pure audio</li> <li>samplerate 22.050Hz</li> <li>mono</li> <li>normalized to -24dB</li> <li>phrase length (min/avg/max): 2 / 52 / 180 chars</li> <li>no silence at beginning/ending</li> <li>avg spoken chars per second: 14</li> <li>sentences with question mark: 2.780</li> <li>sentences with exclamation mark: 1.840</li> </ul> <p>See more details on my <a href="https://github.com/thorstenMueller/Thorsten-Voice">Github page</a> or <a href="https://www.Thorsten-Voice.de">Thorsten-Voice</a> project website.</p>
Vanuatu Voices
<p>Cite the source of the dataset as:</p> <blockquote> <p>Lana Takau, Tom Fitzpatrick, Mary Walworth, Aviva Shimelman, Sandrine Bessis, Tom Ennever, Iveth Rodriguez, Hans-Jörg Bibiko, Daria Dërmaku, Murray Garde, Marie-France Duhamel, Giovanni Abete, Laura Wägerle, Kaitip W. Kami, Tihomir Rangelov, & Russell Gray. (2025). Vanuatu Voices (v1.4.1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4309140</p> </blockquote>
Voice of America: Ukrainian ASR Dataset of Broadcast Speech
<p>The dataset is based on public recordings of Voice of America (<a href="https://ukrainian.voanews.com">https://ukrainian.voanews.com</a>) extracted from their videos.</p> <p>The dataset contains 398 hours of speech.</p> <p>The dataset is created by the ASR Corpus Creator (<a href="https://zenodo.org/record/7396705">https://zenodo.org/record/7396705</a>).</p> <p>The format of files: WAV with 16 kHz.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.