Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2 results for “drums demixing”

Learn how ShareScore rates datasets ↗
zenodo40/100

StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 1

<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed &nbsp;in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269&ndash;2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p>&nbsp;</p> <p><strong>*** THIS IS PART 1 OF 2 ***</strong></p> <p><strong>Download part 2 here:</strong> <a href="../records/7882857">https://zenodo.org/records/7882857&nbsp;</a> (now available!)&nbsp;</p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p>&nbsp;</p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.</p> <p>____________________________</p> <p>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters.&nbsp;</p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>

opencc-by-4.0Apr 2023View details →
zenodo40/100

StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 2

<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed &nbsp;in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269&ndash;2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p>&nbsp;</p> <p><strong>*** THIS IS PART 2 OF 2 ***</strong></p> <p><strong>Download part 1 here:</strong> <a href="../records/7860223">https://zenodo.org/records/7860223&nbsp;</a></p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p>&nbsp;</p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.<br><br>_____________________________<br><br>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters.&nbsp;</p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record