Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6 results for “midi dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Expanded Groove MIDI Dataset

<p>The Expanded Groove MIDI Dataset (E-GMD), a large dataset of human drum performances, with audio recordings annotated in MIDI. E-GMD contains 444 hours of audio from 43 drum kits and is an order of magnitude larger than similar datasets. It is also the first human-performed drum dataset with annotations of velocity.</p> <p>Additional information is available on the Magenta website: <a href="https://g.co/magenta/e-gmd">The Expanded Groove MIDI Dataset</a></p> <p>If you use the E-GMD dataset in your work, please cite the paper where it was introduced:</p> <p>Lee Callender, Curtis Hawthorne, and Jesse Engel. &quot;<a href="https://goo.gl/magenta/e-gmd-paper">Improving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Dataset.</a>&quot; 2020. arXiv:2004.00188.</p> <p>You can also use the following BibTeX entry:</p> <pre>@misc{callender2020improving, title={Improving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Dataset}, author={Lee Callender and Curtis Hawthorne and Jesse Engel}, year={2020}, eprint={2004.00188}, archivePrefix={arXiv}, primaryClass={cs.SD} } </pre> <p>Please also make sure to specify which version of the dataset you are using.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Violin MIDI Dataset

<h3>Violin MIDI Transcription Dataset from <strong><em><a title="High-Resolution Violin Transcription using Weak Labels" href="https://repositori.upf.edu/handle/10230/58121">High-Resolution Violin Transcription using Weak Labels</a></em></strong></h3> <h3><strong>Description</strong></h3> <p><strong>Overview</strong></p> <p>A descriptive transcription of a violin performance requires detecting not only the notes but also the fine-grained pitch variations, such as vibrato. Most existing deep learning methods for music transcription do not capture these variations and often need frame-level annotations, which are scarce for the violin. To enable the development of transcription methods tailored for the analysis of violin performances, this dataset includes MIDI files aligned to violin performances with 5.8 ms frame resolution and 10-cent frequency resolution, and includes pitch bends that represent fine-grained deviations such as vibrato and intonation choice.</p> <p><strong>Dataset Creation</strong></p> <p>The dataset consists of performances of three violin etude books by 22 violinists.&nbsp;The etude books included are:</p> <ul> <li>Paganini, Op. 1</li> <li>Wohlfahrt, Op. 35</li> <li>Kayser, Op. 20</li> </ul> <p>The original MIDI files pre-alignment are sourced from IMSLP and MuseScore (published under Creative Commons license). In the dataset, we provide these open-source MIDI files after they are aligned with the performance links from YouTube. Notably, each MIDI file includes pitch bends to offer a more realistic representation of the expressive performances, and to facilitate intonation analysis for pedagogical applications. The filenames are structured to include reconstructable links to the performances:</p> <p><code>{composer}_{catalog_number}_{performer}_{YouTube_ID}-{YouTube_start_sec}-{YouTube_end_sec}.mid</code></p> <p><strong>Contributions and Findings</strong></p> <p>We target the task of descriptive violin transcription with pitch bends, and demonstrate that our method (1) outperforms generic systems in the proxy tasks of violin transcription and pitch estimation, and (2) can automatically generate new training labels by aligning its feature representations with unseen scores. Additionally, we share our model along with 34 hours of score-aligned solo violin performance dataset, notably including the 24 Paganini Caprices.</p> <h3><strong>Citation</strong></h3> <p>If you find this dataset useful, please cite the following paper:</p> <blockquote> <p>Nazif Can Tamer, Yigitcan &Ouml;zer,&nbsp;Meinard M&uuml;ller, Xavier Serra, &ldquo;High-Resolution Violin Transcription using Weak Labels&rdquo;, in Proc. of the 24th Int. Society for Music Information&nbsp;Retrieval Conf., Milan, Italy, 2023.</p> </blockquote> <p><code>@inproceedings{tamer2023high,</code><br><code>&nbsp; title={High-Resolution Violin Transcription Using Weak Labels},</code><br><code>&nbsp; author={Tamer, Nazif Can and {\"O}zer, Yigitcan and M{\"u}ller, Meinard and Serra, Xavier},</code><br><code>&nbsp; booktitle={International Society for Music Information Retrieval Conference (ISMIR)},</code><br><code>&nbsp; year={2023}</code><br><code>}</code></p> <p>&nbsp;</p>

opencc-by-sa-4.0Oct 2023View details →
zenodo40/100

The MIDI Linked Data Cloud dataset

<p>The study of music is highly interdisciplinary, and thus requires the combination of datasets from multiple musical domains, such as catalog metadata (authors, song titles, dates), industrial records (labels, producers, sales), and music notation (scores). While today an abundance of music metadata exists on the Linked Open Data cloud, datasets containing interoperable symbolic descriptions of music itself, i.e. music notation with note and instrument level information, are scarce. This is the MIDI Linked Data Cloud, a dataset that represents multiple collections of digital music in the MIDI standard format as Linked Data. At the time of writing, the dataset comprises 10,215,557,355 triples of 308,443 interconnected MIDI scores, and provides Web-compatible descriptions of their MIDI events.</p>

opencc-by-4.0May 2017View details →
zenodo36/100

Groove2Groove MIDI Dataset: synthetic accompaniments in 3k styles

<p>The&nbsp;<em>Groove2Groove MIDI Dataset</em>&nbsp;is a parallel corpus of synthetic MIDI accompaniments in almost 3000 different styles,&nbsp;created as described in the paper&nbsp;<em><a href="https://doi.org/10.1109/TASLP.2020.3019642">Groove2Groove: One-Shot Accompaniment Style Transfer with Supervision from Synthetic Data</a></em>&nbsp;[<a href="https://groove2groove.telecom-paris.fr/data/paper.pdf">pdf</a>]. See the <code>README.md</code> file or the&nbsp;<em><a href="https://groove2groove.telecom-paris.fr/#Dataset">Groove2Groove website</a></em> for more information.</p> <p>The dataset is split into the following sections:</p> <ul> <li><code>train</code>&nbsp;contains 5744 MIDI files in 2872 styles (exactly 2 files per style). Each file contains 252 measures&nbsp;following a 2 measure count-in.</li> <li><code>val</code>&nbsp;and&nbsp;<code>test</code>&nbsp;each contain 1200 files in 40 styles (exactly 30 files per style, 16 bars per file after the count-in). The sets of styles are disjoint from each other and from those in&nbsp;<code>train</code>.</li> <li><code>itest</code>&nbsp;is generated from the same chord charts as&nbsp;<code>test</code>, but in 40 styles from the training set.</li> </ul> <p>Chord charts for all MIDI files are provided in the ABC format&nbsp;and the Band-in-a-Box (MGU) format. Each chord chart corresponds to at least 2 MIDI files in different styles.</p> <p>The code used to automate Band-in-a-Box is available in the <a href="https://github.com/cifkao/pybiab">pybiab</a> package.</p> <p>If you use the data in your research, please reference the paper (not just&nbsp;the Zenodo record):</p> <pre><code>@article{groove2groove, author={Ond\v{r}ej C\'{i}fka and Umut \c{S}im\c{s}ekli and Ga\"{e}l Richard}, title={{Groove2Groove}: One-Shot Music Style Transfer with Supervision from Synthetic Data}, journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}, publisher={IEEE}, year={2020}, volume={28}, pages={2638--2650}, doi={10.1109/TASLP.2020.3019642}, url={https://doi.org/10.1109/TASLP.2020.3019642} }</code></pre>

opencc-by-nc-4.0Aug 2020View details →
zenodo16/100

Los Angeles MIDI Dataset

<p>SOTA kilo-scale MIDI dataset for MIR and Music AI purposes</p> <p>&nbsp;</p> <p>Main Features:</p> <p>1) ~600000 MIDIs to explore :)</p> <p>2) Each MIDI file was read-checked and 100% de-duped</p> <p>3) Extensive meta-data for each MIDI file</p> <p>4) Detailed features matrix for each MIDI file</p> <p>5) Helper Python code</p> <p>&nbsp;</p> <p>Project Los Angeles</p> <p>Tegridy Code 2023</p>

restrictedDec 2022View details →
zenodo12/100

UMD-350MB: Refined MIDI Dataset for Symbolic Music Generation

<p><strong>UMD-350MB</strong></p> <p>The Universal MIDI Dataset 350MB (UMD-350MB) is a proprietary collection of 85,618 MIDI files curated for research and development within our organization. This collection is a subset sampled from a larger dataset developed for pretraining symbolic music models.</p> <p>The field of symbolic music generation is constrained by limited data compared to language models. Publicly available datasets, such as the Lakh MIDI Dataset, offer large collections of MIDI files sourced from the web. While the sheer volume of musical data might appear beneficial, the actual amount of valuable data is less than anticipated, as many songs contain less desirable melodies with erratic and repetitive events.</p> <p>The UMD-350MB employs an attention-based approach to achieve more desirable output generations by focusing on human-reviewed training examples of single-track melodies, chord progressions, leads and arpeggios with an average duration of 8 bars. This was achieved by refining the dataset over 24 months, ensuring consistent quality and tempo alignment. Moreover, the dataset is normalized by setting the timing information to 120 BPM with a tick resolution (PPQ) of 96 and transposing the musical scales to C major and A minor (natural scales).</p> <p><strong>Melody Styles</strong></p> <p>A major portion of the dataset is composed of newly produced private data to represent modern musical styles.</p> <ul> <li>Pop: 1970s to 2020s Pop music</li> <li>EDM: Trance, House, Synthwave, Dance, Arcade</li> <li>Jazz: Bebop, Ballad, Latin-Jazz, Bossa-Jazz, Ragtime</li> <li>Soul: 80s Classic, Neo-Soul, Latin-Soul</li> <li>Urban: Pop, Hip-Hop, Trap, R&amp;B, Afrobeat</li> <li>World: Latin, Bossa Nova, European</li> <li>Other: Film, Cinematic, Game music and piano references</li> </ul> <p><em>Actual MIDI files are unlabeled for unsupervised training.</em></p> <p><strong>Dataset Access</strong></p> <p>Please note that this is a closed-source dataset with very limited access. Considerations for access include proposals for data augmentation, chord extraction and other enhancement methods, whether through scripts, algorithmic techniques, manual editing in a DAW or additional processing methods.</p> <p>For inquiries about this dataset, please email us.</p>

restrictedJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record