Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

65

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

65 results for “audio recordings”

Learn how ShareScore rates datasets ↗
zenodo16/100

ACPAS dataset: Aligned Classical Piano Audio and Score (real recording subset)

<p><strong>ACPAS</strong>&nbsp;is a dataset with aligned audio and scores for classical piano music containing 497 distinct music scores aligned with 2189 performances, in total 179.77 hours. For each performance, we provide the corresponding performance audio (real recording or synthesized recording), performance MIDI, and MIDI score, together with rhythm and key annotations.</p> <p>This is the&nbsp;<strong>Real recording subset</strong>&nbsp;of the ACPAS dataset. To download the full dataset and for dataset details, please refer to the dataset webpage at&nbsp;<a href="https://cheriell.github.io/research/ACPAS_dataset">https://cheriell.github.io/research/ACPAS_dataset</a></p> <p>For any questions, suggestions, or comments, please do not hesitate to contact&nbsp;<a href="mailto:lele.liu@qmul.ac.uk">lele.liu@qmul.ac.uk</a></p> <p><strong>How to cite:</strong></p> <p>-&nbsp;Lele Liu, Veronica Morfi, and Emmanouil Benetos, &quot;ACPAS: A Dataset of Aligned Classical Piano Audio and Scores for Audio-to-Score Transcription,&quot; in ISMIR Late-breaking Demo, 2021.</p> <p><strong>Funding:</strong></p> <p>L. Liu is a research student at the UKRI Centre for Doctoral Training in Artificial Intelligence and Music, supported jointly by the China Scholarship Council and Queen Mary University of London.</p> <p>&nbsp;</p>

restrictedOct 2021View details →
zenodo16/100

Seesjärvi Karelian Audio Recordings

<p>This data set contains the digitalized mp3- and wav-audio recordings of the fieldwork interviews conducted in 1989-1991 by Anneli Sarhimaa (https://www.sneb.uni-mainz.de/team/anneli-sarhimaa/) among speakers of the South Karelian dialects used around lake Segozero (Fin. Seesj&auml;rvi)&nbsp; in the Republic of Karelia, the Russian Federation. The interviews were made to gain authentic empirical research material on the linguistic state of the Karelian language at the time and on the ways the community-wide Karelian-Russian biligualism and the rather widespread knowledge of the closely related Finnish language were reflected in everyday language use.</p>

restrictedNov 2022View details →
zenodo16/100

Tver Karelian Audio Recordings

<p>This data set contains the digitalized mp3 recordings of the fieldwork interviews conducted in 1992 by Anneli Sarhimaa and Lea Siilin among speakers of the Tver Karelian dialects in the Central Tver Oblast&#39; in the Russian Federation. The interviews were made to gain authentic empirical research material on the linguistic state of the Karelian language at the time and on the ways the community-wide Karelian-Russian biligualism was reflected in everyday language use.</p>

restrictedNov 2022View details →
zenodo12/100

Jingju Audio Recordings Collection

<p>The <strong>Jingju Audio Recordings Collection</strong> (<strong>JARC</strong>) is part of the <a href="http://compmusic.upf.edu/corpora"><strong>Jingju Music Corpus</strong></a> created in the <a href="http://compmusic.upf.edu/"><strong>CompMusic</strong> project</a>. It is formed by 91 commercial CDs. The purpose of the <strong>JARC</strong>&nbsp;is the computational research of singing melody in jingju arias from the traditional repertoire (传统戏). Therefore, the <strong>JARC</strong> is comprised of CDs consisting of aria compilations (excluding recordings of full plays or of arias from modern plays (现代戏). In order to apply computational techniques, the CDs were selected with the required recording quality. Therefore, the <strong>JARC</strong> contains releases from the 1980s onwards.</p> <p><strong>Content of the JARC</strong></p> <p>The <strong>JARC</strong> is comprised by the ripped tracks of the 91 realeases in FLAC format. Each release is accompanied by a <code>Cover Art</code> folder including scanned copies of its cover art. All the corresponding editorial metadata are available in the MusicBrainz collection <a href="https://musicbrainz.org/collection/40d0978b-0796-4734-9fd4-2b3ebe0f664c"><strong>Dunya Jingju</strong></a> in its original Chinese script. In order to ease non Chinese speakers browsing the collection, all releases have a pseudo-release including the romanized version of the release&#39;s and each recording&#39;s titles using the <a href="https://en.wikipedia.org/wiki/Pinyin">Hanyu Pinyin system</a>. The romanized version of the artist&#39;s name can be obtained from the MusicBrainz field &#39;sort name.&#39; Romanizations of the related works are available as aliases. Consequently releases, recordings, artists and works can be searched both in Chinese and Latin scripts. The editorial metadata stored in MusicBrainz, including the front image from the cover art, are used to tag the FLAC files using <a href="https://picard.musicbrainz.org/">MusicBrainz Picard</a>.</p> <p>Since the purpose of the <strong>JARC</strong> is the research of jingju singing melody, all recordings are tagged with their corresponding <em>shengqiang</em> and <em>banshi</em>, and all artists are tagged with their corresponding role type. When the release makes it clearly explicit, artists are also tagged with their corresponding school (流派).</p> <p>In order to give an overview of the <strong>JARC</strong>&#39;s coverage, the following numbers summarize some of its content:</p> <ul> <li>It comprises <strong>91 releases</strong>, covering <strong>1687 recordings</strong>, which account for more than <strong>154 hours</strong> of music.</li> <li>It contains <strong>95 performers</strong> (considering only actors and actresses, not instrumentalists), covering the most representative role types in terms of singing: <ul> <li><strong><em>dan</em></strong>: 34 artists, 634 recordings</li> <li><strong><em>laosheng</em></strong>: 25 artists, 562 recordings</li> <li><strong><em>laodan</em></strong>: 11 artists, 262 recordings</li> <li><strong><em>xiaosheng</em></strong>: 10 artists, 71 recordings</li> <li><strong><em>jing</em></strong>: 8 artists, 138 recordings</li> </ul> </li> <li>It presents a wide coverage of the two more representative <em>shengqiang</em> in jingju, namely <em>erhuang</em> (612 recordings) and <em>xipi</em> (811 recordings), and it also covers others such as <em>fan&#39;erhuang</em> (123 recordings), <em>fanxipi</em> (29 recordings), <em>sipingdiao</em> (64 recordings), <em>nanbangzi</em> (60 recordings), <em>gaobozi</em> (9 recordings) and others.</li> <li>It comprises the recording of 792 jingju arias from 299 different plays.</li> </ul> <p><strong>Using the JARC</strong></p> <p>Since the <strong>JARC</strong> is sourced from copyrighted releases, it can only be shared for non commercial research purposes (see the form below).</p> <p>For referencing the <strong>JARC</strong>, and obtaining a more detailed description of it, please refer to</p> <blockquote> <p>Caro Repetto, Rafael (2018) <em>The musical dimension of Chinese traditional theatre: An analysis from computer aided musicology</em>. PhD thesis, Universitat Pompeu Fabra, Barcelona, Spain.</p> </blockquote> <p><strong>Acknowledgements</strong></p> <p>The creation of the <strong>JARC</strong> is funded by the European Research Council under the European Union&rsquo;s Seventh Framework Program (FP7/2007-2013), as part of the CompMusic project (ERC grant agreement 267583).</p>

restrictedOct 2018View details →
zenodo8/100

Audio files for A Dataset of EEG and EOG recordings from an Auditory EOG-based Communication System for Patients in Locked-In State

<ol> <li>Origin of the data</li> </ol> <p>These Audio files are related to the publication &quot;A dataset of EEG and EOG from an auditory EOG-based communication system for patients in locked-in state&quot; by Andres Jaramillo-Gonzalez, Shize Wu, Alessandro Tonin, Aygul Rana, Majid Khalili-Ardali, Niels Birbaumer, and Ujwal Chaudhary. To use the audio data, please contact&nbsp; Dr. Ujwal Chaudhary, as mentioned in the manuscript.</p> <ol> <li>Details of the data</li> </ol> <p>The description of the nature and details of the *.txt files are in the section &quot;Data Records&quot; of the manuscript. As mentioned in the manuscript, before the beginning of the study, at least 100 questions with known &quot;yes&quot; or &quot;no&quot; answers were formulated and recorded with a family member or caretaker&#39;s voice close to the patient.&nbsp; Each question with a &quot;yes&quot; answer is paired with a similar question with a &quot;no&quot; answer (e.g., &quot;Paris is the capital of France&quot; and &quot;Berlin is the capital of France&quot;).</p> <p>Each question is saved as an audio file with an explicit identifier, a question with a &quot;yes&quot; answer is saved with a 001_NUMBER identifier, and a question with a &quot;no&quot; answer is saved with a 002_NUMBER identifier. The NUMBER label is composed of 5 digits; the first two determine the number of the patient, and the last three digits indicate the number of the question. The value of the label NUMBER is the same for a semantically paired sentence. In the example, &quot;001_11007.wav&quot; indicates a &quot;yes&quot; type of answer, question number 07 from P11. Sentences are then placed in a specific folder of the used laptop&#39;s storage, accessed and played by the communication system along with the sessions during sentence presentation.</p>

restrictedNov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record