Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “midi dataset”
Expanded Groove MIDI Dataset
<p>The Expanded Groove MIDI Dataset (E-GMD), a large dataset of human drum performances, with audio recordings annotated in MIDI. E-GMD contains 444 hours of audio from 43 drum kits and is an order of magnitude larger than similar datasets. It is also the first human-performed drum dataset with annotations of velocity.</p> <p>Additional information is available on the Magenta website: <a href="https://g.co/magenta/e-gmd">The Expanded Groove MIDI Dataset</a></p> <p>If you use the E-GMD dataset in your work, please cite the paper where it was introduced:</p> <p>Lee Callender, Curtis Hawthorne, and Jesse Engel. "<a href="https://goo.gl/magenta/e-gmd-paper">Improving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Dataset.</a>" 2020. arXiv:2004.00188.</p> <p>You can also use the following BibTeX entry:</p> <pre>@misc{callender2020improving, title={Improving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Dataset}, author={Lee Callender and Curtis Hawthorne and Jesse Engel}, year={2020}, eprint={2004.00188}, archivePrefix={arXiv}, primaryClass={cs.SD} } </pre> <p>Please also make sure to specify which version of the dataset you are using.</p>
Violin MIDI Dataset
<h3>Violin MIDI Transcription Dataset from <strong><em><a title="High-Resolution Violin Transcription using Weak Labels" href="https://repositori.upf.edu/handle/10230/58121">High-Resolution Violin Transcription using Weak Labels</a></em></strong></h3> <h3><strong>Description</strong></h3> <p><strong>Overview</strong></p> <p>A descriptive transcription of a violin performance requires detecting not only the notes but also the fine-grained pitch variations, such as vibrato. Most existing deep learning methods for music transcription do not capture these variations and often need frame-level annotations, which are scarce for the violin. To enable the development of transcription methods tailored for the analysis of violin performances, this dataset includes MIDI files aligned to violin performances with 5.8 ms frame resolution and 10-cent frequency resolution, and includes pitch bends that represent fine-grained deviations such as vibrato and intonation choice.</p> <p><strong>Dataset Creation</strong></p> <p>The dataset consists of performances of three violin etude books by 22 violinists. The etude books included are:</p> <ul> <li>Paganini, Op. 1</li> <li>Wohlfahrt, Op. 35</li> <li>Kayser, Op. 20</li> </ul> <p>The original MIDI files pre-alignment are sourced from IMSLP and MuseScore (published under Creative Commons license). In the dataset, we provide these open-source MIDI files after they are aligned with the performance links from YouTube. Notably, each MIDI file includes pitch bends to offer a more realistic representation of the expressive performances, and to facilitate intonation analysis for pedagogical applications. The filenames are structured to include reconstructable links to the performances:</p> <p><code>{composer}_{catalog_number}_{performer}_{YouTube_ID}-{YouTube_start_sec}-{YouTube_end_sec}.mid</code></p> <p><strong>Contributions and Findings</strong></p> <p>We target the task of descriptive violin transcription with pitch bends, and demonstrate that our method (1) outperforms generic systems in the proxy tasks of violin transcription and pitch estimation, and (2) can automatically generate new training labels by aligning its feature representations with unseen scores. Additionally, we share our model along with 34 hours of score-aligned solo violin performance dataset, notably including the 24 Paganini Caprices.</p> <h3><strong>Citation</strong></h3> <p>If you find this dataset useful, please cite the following paper:</p> <blockquote> <p>Nazif Can Tamer, Yigitcan Özer, Meinard Müller, Xavier Serra, “High-Resolution Violin Transcription using Weak Labels”, in Proc. of the 24th Int. Society for Music Information Retrieval Conf., Milan, Italy, 2023.</p> </blockquote> <p><code>@inproceedings{tamer2023high,</code><br><code> title={High-Resolution Violin Transcription Using Weak Labels},</code><br><code> author={Tamer, Nazif Can and {\"O}zer, Yigitcan and M{\"u}ller, Meinard and Serra, Xavier},</code><br><code> booktitle={International Society for Music Information Retrieval Conference (ISMIR)},</code><br><code> year={2023}</code><br><code>}</code></p> <p> </p>
The MIDI Linked Data Cloud dataset
<p>The study of music is highly interdisciplinary, and thus requires the combination of datasets from multiple musical domains, such as catalog metadata (authors, song titles, dates), industrial records (labels, producers, sales), and music notation (scores). While today an abundance of music metadata exists on the Linked Open Data cloud, datasets containing interoperable symbolic descriptions of music itself, i.e. music notation with note and instrument level information, are scarce. This is the MIDI Linked Data Cloud, a dataset that represents multiple collections of digital music in the MIDI standard format as Linked Data. At the time of writing, the dataset comprises 10,215,557,355 triples of 308,443 interconnected MIDI scores, and provides Web-compatible descriptions of their MIDI events.</p>
Groove2Groove MIDI Dataset: synthetic accompaniments in 3k styles
<p>The <em>Groove2Groove MIDI Dataset</em> is a parallel corpus of synthetic MIDI accompaniments in almost 3000 different styles, created as described in the paper <em><a href="https://doi.org/10.1109/TASLP.2020.3019642">Groove2Groove: One-Shot Accompaniment Style Transfer with Supervision from Synthetic Data</a></em> [<a href="https://groove2groove.telecom-paris.fr/data/paper.pdf">pdf</a>]. See the <code>README.md</code> file or the <em><a href="https://groove2groove.telecom-paris.fr/#Dataset">Groove2Groove website</a></em> for more information.</p> <p>The dataset is split into the following sections:</p> <ul> <li><code>train</code> contains 5744 MIDI files in 2872 styles (exactly 2 files per style). Each file contains 252 measures following a 2 measure count-in.</li> <li><code>val</code> and <code>test</code> each contain 1200 files in 40 styles (exactly 30 files per style, 16 bars per file after the count-in). The sets of styles are disjoint from each other and from those in <code>train</code>.</li> <li><code>itest</code> is generated from the same chord charts as <code>test</code>, but in 40 styles from the training set.</li> </ul> <p>Chord charts for all MIDI files are provided in the ABC format and the Band-in-a-Box (MGU) format. Each chord chart corresponds to at least 2 MIDI files in different styles.</p> <p>The code used to automate Band-in-a-Box is available in the <a href="https://github.com/cifkao/pybiab">pybiab</a> package.</p> <p>If you use the data in your research, please reference the paper (not just the Zenodo record):</p> <pre><code>@article{groove2groove, author={Ond\v{r}ej C\'{i}fka and Umut \c{S}im\c{s}ekli and Ga\"{e}l Richard}, title={{Groove2Groove}: One-Shot Music Style Transfer with Supervision from Synthetic Data}, journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}, publisher={IEEE}, year={2020}, volume={28}, pages={2638--2650}, doi={10.1109/TASLP.2020.3019642}, url={https://doi.org/10.1109/TASLP.2020.3019642} }</code></pre>
Los Angeles MIDI Dataset
<p>SOTA kilo-scale MIDI dataset for MIR and Music AI purposes</p> <p> </p> <p>Main Features:</p> <p>1) ~600000 MIDIs to explore :)</p> <p>2) Each MIDI file was read-checked and 100% de-duped</p> <p>3) Extensive meta-data for each MIDI file</p> <p>4) Detailed features matrix for each MIDI file</p> <p>5) Helper Python code</p> <p> </p> <p>Project Los Angeles</p> <p>Tegridy Code 2023</p>
UMD-350MB: Refined MIDI Dataset for Symbolic Music Generation
<p><strong>UMD-350MB</strong></p> <p>The Universal MIDI Dataset 350MB (UMD-350MB) is a proprietary collection of 85,618 MIDI files curated for research and development within our organization. This collection is a subset sampled from a larger dataset developed for pretraining symbolic music models.</p> <p>The field of symbolic music generation is constrained by limited data compared to language models. Publicly available datasets, such as the Lakh MIDI Dataset, offer large collections of MIDI files sourced from the web. While the sheer volume of musical data might appear beneficial, the actual amount of valuable data is less than anticipated, as many songs contain less desirable melodies with erratic and repetitive events.</p> <p>The UMD-350MB employs an attention-based approach to achieve more desirable output generations by focusing on human-reviewed training examples of single-track melodies, chord progressions, leads and arpeggios with an average duration of 8 bars. This was achieved by refining the dataset over 24 months, ensuring consistent quality and tempo alignment. Moreover, the dataset is normalized by setting the timing information to 120 BPM with a tick resolution (PPQ) of 96 and transposing the musical scales to C major and A minor (natural scales).</p> <p><strong>Melody Styles</strong></p> <p>A major portion of the dataset is composed of newly produced private data to represent modern musical styles.</p> <ul> <li>Pop: 1970s to 2020s Pop music</li> <li>EDM: Trance, House, Synthwave, Dance, Arcade</li> <li>Jazz: Bebop, Ballad, Latin-Jazz, Bossa-Jazz, Ragtime</li> <li>Soul: 80s Classic, Neo-Soul, Latin-Soul</li> <li>Urban: Pop, Hip-Hop, Trap, R&B, Afrobeat</li> <li>World: Latin, Bossa Nova, European</li> <li>Other: Film, Cinematic, Game music and piano references</li> </ul> <p><em>Actual MIDI files are unlabeled for unsupervised training.</em></p> <p><strong>Dataset Access</strong></p> <p>Please note that this is a closed-source dataset with very limited access. Considerations for access include proposals for data augmentation, chord extraction and other enhancement methods, whether through scripts, algorithmic techniques, manual editing in a DAW or additional processing methods.</p> <p>For inquiries about this dataset, please email us.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.