Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

16

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

16 results for “singing voice”

Learn how ShareScore rates datasets ↗
zenodo44/100

Jamendo Corpus for Singing Voice Detection

<p>This is a public corpus of 93 creative-commons licensed music pieces annotated<br> by voice (sung or spoken) and no-voice.</p>

opencc-by-4.0Mar 2008View details →
zenodo40/100

Electrobyte for Singing Voice Detection

<p>This is a public dataset of 90 copyright-free electronic songs with vocal annotations (sing/no sing).</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Annotated-VocalSet: A Singing Voice Dataset

<p>This dataset provides annotations for the <a href="https://doi.org/10.5281/zenodo.1442513">VocalSet dataset</a>, which is available online at</p> <pre><a href="https://doi.org/10.5281/zenodo.1442513">https://doi.org/10.5281/zenodo.1442513</a></pre> <p>.</p> <p>The annotations generated for the VocalSet audio files include fundamental frequency contour, note onset, note offset, the transition between notes, note F0, note duration, Midi pitch, and lyrics.</p> <p><a href="https://doi.org/10.5281/zenodo.1442513">VocalSet</a> consists of more than 10 hours of monophonic recorded audio of professional singers in a variety of vocal techniques (n = 17) and several singers (m = 20) with several WAV files (p = 3560). However, although several categories, including techniques, singers, tempo, and loudness, are considered in the dataset, the sung notes were not annotated. Therefore, this dataset aims to annotate VocalSet to make it a more powerful dataset for researchers.</p> <p>Details of the dataset are provided in the following academic journal paper.</p> <p><a href="https://www.mdpi.com/2076-3417/12/18/9257">Faghih, Behnam, and Joseph Timoney. 2022. &quot;Annotated-VocalSet: A Singing Voice Dataset&quot;&nbsp;<em>Applied Sciences</em>&nbsp;12, no. 18: 9257. https://doi.org/10.3390/app12189257</a></p> <p>Please use the above paper to cite this dataset.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

VocalSet: A Singing Voice Dataset

<p><strong>NEW IN VocalSet 1.2:</strong> We now have 3 file organization versions:</p> <ol> <li>Files organized by singer</li> <li>Files organized by technique</li> <li>Files organized by vowel</li> </ol> <p>We hope that this will ease the process of training and testing models using these different attributes of the dataset.</p> <p>&nbsp;</p> <p><strong>Overview:</strong></p> <p>We present VocalSet, a singing voice dataset consisting of 10.1 hours of monophonic recorded audio of professional singers demonstrating both standard and extended vocal techniques on all 5 vowels. Existing singing voice datasets aim to capture a focused subset of singing voice characteristics, and generally consist of just a few singers. VocalSet contains recordings from 20 different singers (9 male, 11 female) and a range of voice types. &nbsp;VocalSet aims to improve the state of existing singing voice datasets and singing voice research by capturing not only a range of vowels, but also a diverse set of voices on many different vocal techniques, sung in contexts of scales, arpeggios, long tones, and excerpts.</p> <p>We have included two .txt files &#39;train_singers_technique.txt &#39;and &#39;test_singers_technique.txt&#39;&nbsp;in which you will find a list of the singers we used to train and test our technique classifier on. &#39;DataSetVocalises.pdf&#39; contains the sheet singers sang from in their recording sessions. &#39;readme-anon.txt&#39; contains more information about the dataset, including the mapping from filename to singer voice type as well as more information on the vocalises that will help you map files to sheet music. Enjoy and please cite accordingly!</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

MedleyVox: An Evaluation Dataset for Multiple Singing Voices Separation

<p>MedleyVox is an evaluation dataset for multiple singing voices separation, which consists of&nbsp;381 segments (1.1hour),&nbsp;containing 23 songs from MedleyDB v1 and v2 (https://medleydb.weebly.com).</p> <p>&nbsp;</p> <p>MedleyVox&nbsp;contains 1) unison, 2) duet, 3) main vs. rest (folder name &#39;rest&#39;) , and 4) N-singing (&#39;unison&#39; + &#39;duet&#39; + &#39;rest&#39;, please check our dataloader code&nbsp;for N-singing in&nbsp;https://github.com/jeonchangbin49/MedleyVox/blob/main/svs/data/test_dataset.py)&nbsp;categories explained in our ICASSP 2023 paper. For more details, please check our paper (https://arxiv.org/pdf/2211.07302.pdf) and code repository (https://github.com/jeonchangbin49/MedleyVox).</p> <p>&nbsp;</p> <p>Acknowledgment</p> <p>We are grateful to Rachel Bittner, the author of the original MedleyDB data, for allowing us to publish the&nbsp;MedleyVox&nbsp;dataset.</p> <p>&nbsp;</p> <p>License</p> <p>This work is licensed under a&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</a>.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge (WildSVDD Track)

<p>For more information about SVDD Challenge 2024, please refer to https://challenge.singfake.org/.<br><br>WildSVDD track dataset is an extension of <a href="https://singfake.org/">SingFake</a> dataset.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset for Interspeech 2018 submission: Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions

<p>This dataset contains the materials for training, testing the joint and HSMM models mentioned in the paper &quot;<em>Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions&quot;</em>.</p> <p>The filename list of this dataset can be found in the function <em>get_train_test_recordings_joint()</em> of <em>./general/trainTestSeparation.py</em> file. The dataset contains the Praat TextGrids and .wavs of the variables: <em>train_primary_school, val_primary_school</em> and <em>test_primary_school</em>. For accessing other datasets such as <em>train_nacta_2017, train_nacta</em> and <em>train_sepa</em>, please download them from the links:</p> <p>jingju dataset part1:&nbsp;<a href="https://zenodo.org/record/1185154">https://zenodo.org/record/1185154</a></p> <p>jingju dataset part2:&nbsp;<a href="https://doi.org/10.5281/zenodo.842229">https://doi.org/10.5281/zenodo.842229</a></p> <p>Once you have downloaded these three datasets, you need to set the paths in <em>./general/filePathShared.py</em>.</p> <p>Set <em>path_jingju_dataset</em> to the parent path of these three datasets.</p> <p>Set <em>primarySchool_dataset_root_path</em> to the path of the interspeech2018 dataset (the current dataset).</p> <p>Set <em>nacta_dataset_root_path</em> to the path of the jingju&nbsp;dataset part1.</p> <p>Set <em>nacta2017_dataset_root_path</em> to the path the jingju&nbsp;dataset part2.</p> <p>For more information on this paper, please refer to the Github page:&nbsp;<a href="https://github.com/ronggong/interspeech2018_submission01">https://github.com/ronggong/interspeech2018_submission01</a></p> <p>&nbsp;</p>

opencc-by-nc-4.0Feb 2018View details →
zenodo36/100

Jingju a cappella singing voice test dataset for "An efficient deep learning model for musical onset detection"

<p>Jingju a cappella singing voice test dataset used in the paper &quot;An efficient deep learning model for musical onset detection&quot;.</p> <p>Arxiv paper link:&nbsp;<a href="https://arxiv.org/abs/1806.06773">https://arxiv.org/abs/1806.06773</a></p> <p>Supplementary information and code for the paper:&nbsp;<a href="https://github.com/ronggong/musical-onset-efficient">https://github.com/ronggong/musical-onset-efficient</a></p> <p><strong>Content:</strong></p> <ol> <li>ismir_2018_dataset_for_reviewing.zip: audio, syllable boundary and label annotation</li> <li>jingju dataset train test split filenames.xlsx: train and test split filename list</li> </ol> <p><strong>Citation:</strong></p> <pre>@article{gong2018towards, title={Towards an efficient deep learning model for musical onset detection}, author={Gong, Rong and Serra, Xavier}, journal={arXiv preprint arXiv:1806.06773}, year={2018} } </pre> <p><strong>Contact:</strong></p> <p>Rong Gong: rong.gong&lt;at&gt;upf.edu</p>

opencc-by-nc-4.0Aug 2018View details →
zenodo32/100

Data to: Aerosol emission of child voices during speaking, singing and shouting

<p>This dataset contains raw data of emitted aerosols measured via a laser particle counter. Further, R-code (RMarkdown) for statistical analyses is available. A pre-print of an article based on these data was deposited here:</p> <p>https://www.medrxiv.org/content/early/2020/09/18/2020.09.17.20196733</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo32/100

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge (CtrSVDD Track, Test Set)

<p>For more information about SVDD Challenge 2024, please refer to https://challenge.singfake.org/.<br><br>We have released the test set here.</p> <p>The training and development set is at https://zenodo.org/records/10467648.</p> <p>The Interspeech paper that describes the dataset details and baseline analysis is https://arxiv.org/abs/2406.02438.</p>

opencc-by-nc-nd-4.0Mar 2024View details →
ClinicalTrials.gov28/100

Effects of Finger Kazoo Exercise With and Without Oropharyngeal Enlargement in the Operatic Singing Voice

ClinicalTrials.gov study NCT07129668. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
dryad28/100

Data from: Lower vocal tract morphologic adjustments are relevant for voice timbre in singing

Open the record for dataset details and reuse information.

publicJul 2016View details →
ClinicalTrials.gov24/100

Singing-voice Disorders and Aerodynamic Profiles in Dysodic Singers

ClinicalTrials.gov study NCT04036864. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo16/100

singing voice detection dataset ALL-Pub-SVD-In-One

<p>singing voice detection datasets:</p> <p>it contains the public dataset of Jamendo Dataset, and the RWC dataset , and the iKala dataset, and the MIR-1K dataset , and teh MedlyDB dataset.</p> <p>Jamendo are same with the raw.</p> <p>RWC use the simply mode same with the lable in Jamendo just sing or nosing.</p> <p>iKala and MIR-1K are full wavfiles with the generated lables from the PitchLables. 0 for nosing and other for sing.</p> <p>MedlyDB use part wavfiles from the RAW dataset and also generated lables from the MELODY_2 lables with 0 for nosing and other for nosing.</p> <p>&nbsp;&nbsp;</p>

restrictedApr 2019View details →
zenodo16/100

Singing Voice Audio Dataset

<p>This dataset is for the purpose of the analysis of singing voice. It is our hope that the publication of this dataset will encourage further work into the area of singing voice audio analysis by removing one of the main impediments in this research area - the lack of data (unaccompanied singing).</p> <p>It contains over 70 original vocal recordings by 28 professional, semi-professional and amateur singers. Singing style is predominantly Chinese Opera but some recordings are Western Opera. All recordings are 44.1 KHz sample rate and have been amplitude normalised.</p>

restrictedNov 2019View details →
zenodo8/100

PMD-Singing: Singing Voice Dataset for Phonation Mode Detection

<p>The Sung Phonation Mode Dataset (noted as PMSing in paper) is the first multi-phonation audio dataset for PMD task.&nbsp;</p> <p>The total duration of PMSing is 1.51 h, containing 42 songs with an average duration of 2.16 min. Compared to existing PMC datasets, the PMSing dataset contains a longer duration, and the duration of the phonation modes varies from 0.01 to 6.89 s. Additionally, all the audio files in PMSing contain multiple phonation modes.&nbsp;</p> <p>A sample audio and its annotation file&nbsp;are provided.</p>

restrictedFeb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record