Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5 results for “symbolic music”

Learn how ShareScore rates datasets ↗
zenodo40/100

PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

<p>We introduce&nbsp;<strong>PDMX</strong>: a <strong>P</strong>ublic&nbsp;<strong>D</strong>omain&nbsp;<strong>M</strong>usic<strong>X</strong>ML dataset for symbolic music processing. Refer to our <a title="PDMX Paper" href="https://arxiv.org/abs/2409.10831" target="_blank" rel="noopener">paper</a> for more information, and our <a title="PDMX GitHub Repository" href="https://github.com/pnlong/PDMX/" target="_blank" rel="noopener">GitHub repository</a> for any code-related details. Please cite both our paper and <a href="https://arxiv.org/abs/2410.02084" target="_blank" rel="noopener">our collaborators' paper</a> if you use this dataset (see our GitHub for more information).</p> <p>Upon further use of the PDMX dataset, we discovered a discrepancy between the public-facing copyright metadata on the <a href="https://musescore.com/">MuseScore website</a> and the internal copyright data of the MuseScore files themselves, which affected 31,221 (12.29% of) songs. We have decided to proceed with the former given its public visibility on Musescore (i.e. this is what the MuseScore website presents its users with). We have noted files with conflicting internal licenses in the&nbsp;<em><strong>license_conflict</strong></em> column of PDMX. We recommend using the&nbsp;<em><strong>no_license_conflict</strong></em> subset of PDMX (which still includes 222,856 songs) moving forward.</p> <p>Additionally, for each song in PDMX, we not only provide the <em>MusicRender</em> and metadata JSON files, but we also try to include the associated compressed MusicXML (MXL), sheet music (PDF), and MIDI (MID) files when available. Due to the corruption of 42 of the original MuseScore files,&nbsp;these songs lack those associated files (since they could not be converted to those formats) and only include the <em>MusicRender</em> and metadata JSON files. The&nbsp;<em><strong>all_valid</strong></em> subset of PDMX describes the songs where all associated files are valid.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Emotion4MIDI: A Lyrics-Based Emotion-Labeled Symbolic Music Dataset

<p>This dataset includes emotion labels for the publicly available MIDI dataset, namely Lakh MIDI Dataset and Reddit MIDI dataset. The values represent the probability of containing a particular emotion. For a single song, more than one emotion can be present, hence the values don't add up to 1.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Audio Rendered Using Symbolic Music Data Sampled From DMelodies

<p>This is a dataset of audio files rendered using&nbsp;a subset of&nbsp;symbolic music data from&nbsp;<a href="https://github.com/ashispati/dmelodies_dataset">DMelodies</a>.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Material for the paper "MIDI2vec: Learning MIDI Embeddings for Reliable Prediction of Symbolic Music Metadata"

<p>Edgelists and trained embeddings used in the paper</p>

opencc-by-4.0Jul 2021View details →
zenodo12/100

UMD-350MB: Refined MIDI Dataset for Symbolic Music Generation

<p><strong>UMD-350MB</strong></p> <p>The Universal MIDI Dataset 350MB (UMD-350MB) is a proprietary collection of 85,618 MIDI files curated for research and development within our organization. This collection is a subset sampled from a larger dataset developed for pretraining symbolic music models.</p> <p>The field of symbolic music generation is constrained by limited data compared to language models. Publicly available datasets, such as the Lakh MIDI Dataset, offer large collections of MIDI files sourced from the web. While the sheer volume of musical data might appear beneficial, the actual amount of valuable data is less than anticipated, as many songs contain less desirable melodies with erratic and repetitive events.</p> <p>The UMD-350MB employs an attention-based approach to achieve more desirable output generations by focusing on human-reviewed training examples of single-track melodies, chord progressions, leads and arpeggios with an average duration of 8 bars. This was achieved by refining the dataset over 24 months, ensuring consistent quality and tempo alignment. Moreover, the dataset is normalized by setting the timing information to 120 BPM with a tick resolution (PPQ) of 96 and transposing the musical scales to C major and A minor (natural scales).</p> <p><strong>Melody Styles</strong></p> <p>A major portion of the dataset is composed of newly produced private data to represent modern musical styles.</p> <ul> <li>Pop: 1970s to 2020s Pop music</li> <li>EDM: Trance, House, Synthwave, Dance, Arcade</li> <li>Jazz: Bebop, Ballad, Latin-Jazz, Bossa-Jazz, Ragtime</li> <li>Soul: 80s Classic, Neo-Soul, Latin-Soul</li> <li>Urban: Pop, Hip-Hop, Trap, R&amp;B, Afrobeat</li> <li>World: Latin, Bossa Nova, European</li> <li>Other: Film, Cinematic, Game music and piano references</li> </ul> <p><em>Actual MIDI files are unlabeled for unsupervised training.</em></p> <p><strong>Dataset Access</strong></p> <p>Please note that this is a closed-source dataset with very limited access. Considerations for access include proposals for data augmentation, chord extraction and other enhancement methods, whether through scripts, algorithmic techniques, manual editing in a DAW or additional processing methods.</p> <p>For inquiries about this dataset, please email us.</p>

restrictedJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record