Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
112
datasets available to search
ShareScore release 0.9.0
Dataset results
112 results for “music dataset”
Music Genre fMRI Dataset - Derivatives
<p>This dataset contains preprocessed data from the Music Genre fMRI Dataset (<a href="https://openneuro.org/datasets/ds003720/versions/1.0.0">https://openneuro.org/datasets/ds003720/versions/1.0.0</a>). Experimental stimuli can be generated using GTZAN_Preprocess.py.</p> <p>References:</p> <p>1. Nakai, Koide-Majima, and Nishimoto (2021). Correspondence of categorical and feature-based representations of music in the human brain. Brain and Behavior. 11(1), e01936. https://doi.org/10.1002/brb3.1936</p> <p>2. Nakai, Koide-Majima, and Nishimoto (2022). Music genre neuroimaging dataset. Data in Brief. 40, 107675. https://doi.org/10.1016/j.dib.2021.107675</p>
MedleyDB Audio: A Dataset of Multitrack Audio for Music Research
<p>Audio files for the MedleyDB multitrack dataset. <strong>Annotation and Metadata files are version controlled and are available in the <a href="https://github.com/marl/medleydb">MedleyDB github</a> repository: </strong><em>Metadata</em> can be found <a href="https://github.com/marl/medleydb/tree/master/medleydb/data/Metadata">here</a>, <em>Annotations</em> can be found <a href="https://github.com/marl/medleydb/tree/master/medleydb/data/Annotations">here</a>.</p> <p>For detailed information about the dataset, please visit MedleyDB's <a href="http://medleydb.weebly.com">website</a>.</p> <p> </p> <p>If you make use of MedleyDB for academic purposes, please cite the following publication:<br> <br> <em>R. Bittner, J. Salamon, M. Tierney, M. Mauch, C. Cannam and J. P. Bello, "<a href="http://marl.smusic.nyu.edu/medleydb_webfiles/bittner_medleydb_ismir2014.pdf">MedleyDB: A Multitrack Dataset for Annotation-Intensive MIR Research</a>", in 15th International Society for Music Information Retrieval Conference, Taipei, Taiwan, Oct. 2014.</em></p>
Human Dataset of musical patterns and repetitions with artistic variations
<p>Dataset of musical patterns and pattern repetitions with artistic variations recorded by keyboard players and guitar players.</p> <p> </p> <p>NOTE: There are some corrupted files in the dataset. Please find the latest version at: https://zenodo.org/doi/10.5281/zenodo.10818616</p>
Hindustani Music Rhythm Dataset
<p>CompMusic Hindustani Rhythm Dataset is a rhythm annotated test corpus for automatic rhythm analysis tasks in Hindustani Music. The collection consists of audio excerpts from the CompMusic Hindustani research corpus, manually annotated time aligned markers indicating the progression through the taal cycle, and the associated taal related metadata. A brief description of the dataset is provided below. </p> <p>For a brief overview and audio examples of taals in Hindustani music, please see</p> <p><a href="http://compmusic.upf.edu/examples-taal-hindustani">http://compmusic.upf.edu/examples-taal-hindustani</a></p> <p><strong>THE DATASET </strong></p> <p><strong>Audio music content </strong></p> <p>The pieces are chosen from the CompMusic Hindustani music collection. The pieces were chosen in four popular taals of Hindustani music, which encompasses a majority of Hindustani khyal music. The pieces were chosen include a mix of vocal and instrumental recordings, new and old recordings, and to span three lays. For each taal, there are pieces in dhrut (fast), madhya (medium) and vilambit (slow) lays (tempo class). All pieces have Tabla as the percussion accompaniment. The excerpts are two minutes long. Each piece is uniquely identified using the MBID of the recording. The pieces are stereo, 160 kbps, mp3 files sampled at 44.1 kHz. The audio is also available as wav files for experiments. </p> <p><strong>Annotations</strong></p> <p>There are several annotations that accompany each excerpt in the dataset.</p> <p><strong>Sam, vibhaag and the maatras</strong>: The primary annotations are audio synchronized time-stamps indicating the different metrical positions in the taal cycle. The sam and matras of the cycle are annotated. The annotations were created using Sonic Visualizer by tapping to music and manually correcting the taps. Each annotation has a time-stamp and an associated numeric label that indicates the position of the beat marker in the taala cycle. The annotations and the associated metadata have been verified for correctness and completeness by a professional Hindustani musician and musicologist. The long thick lines show vibhaag boundaries. The numerals indicate the matra number in cycle. In each case, the sam (the start of the cycle, analogous to the downbeat) are indicated using the numeral 1. </p> <p><strong>Taal related metadata</strong>: For each excerpt, the taal and the lay of the piece are recorded. Each excerpt can be uniquely identified and located with the MBID of the recording, and the relative start and end times of the excerpt within the whole recording. A separate 5 digit taal based unique ID is also provided for each excerpt as a double check. The artist, release, the lead instrument, and the raag of the piece are additional editorial metadata obtained from the release. There are optional comments on audio quality and annotation specifics. </p> <p><strong>Data subsets</strong></p> <p>The dataset consists of excerpts with a wide tempo range from 10 MPM (matras per minute) to 370 MPM. To study any effects of the tempo class, the full dataset (HMDf) is also divided into two other subsets - the long cycle subset (HMDl) consisting of vilambit (slow) pieces with a median tempo between 10-60 MPM, and the short cycle subset (HMDs) with madhyalay (medium, 60-150 MPM) and the drut lay (fast, 150+ MPM). </p> <p><strong>Possible uses of the dataset</strong></p> <p>Possible tasks where the dataset can be used include taal, sama and beat tracking, tempo estimation and tracking, taal recognition, rhythm based segmentation of musical audio, audio to score/lyrics alignment, and rhythmic pattern discovery. </p> <p><strong>Dataset organization</strong></p> <p>The dataset consists of audio, annotations, an accompanying spreadsheet providing additional metadata, a MAT-file that has identical information as the spreadsheet, and a dataset description document.</p> <p><strong>Using this dataset</strong></p> <p>Please cite the following publication if you use the dataset in your work:</p> <blockquote> <p>Ajay Srinivasasmurthy, Andre Holzapfel, Ali Taylan Cemgil, Xavier Serra, "A generalized Bayesian model for tracking long metrical cycles in acoustic music signals", in Proc. of the 41st IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2016), Shanghai, China, March 2016</p> </blockquote> <p><a href="http://hdl.handle.net/10230/32090">http://hdl.handle.net/10230/32090</a></p> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p><strong>Contact</strong></p> <p>If you have any questions or comments about the dataset, please feel free to write to us.</p> <p>Ajay Srinivasamurthy<br> Music Technology Group<br> Universitat Pompeu Fabra,<br> Barcelona, Spain<br> ajays.murthy@upf.edu</p> <p>Kaustuv Kanti Ganguli<br> DAP lab, Dept. of Electrical Engineering,<br> Indian Institute of Technology Bombay<br> Mumbai, India<br> kaustuvkanti@ee.iitb.ac.in</p> <p> </p> <p><a href="http://compmusic.upf.edu/hindustani-rhythm-dataset">http://compmusic.upf.edu/hindustani-rhythm-dataset</a></p>
Carnatic Music Rhythm Dataset
<p>CompMusic Carnatic Rhythm Dataset is a rhythm annotated test corpus for automatic rhythm analysis tasks in Carnatic Music. The collection consists of audio excerpts from the CompMusic Carnatic research corpus, manually annotated time aligned markers indicating the progression through the taala cycle, and the associated taala related metadata. A brief description of the dataset is provided below. For a brief overview and audio examples of taalas in Carnatic music, please see <a href="http://compmusic.upf.edu/examples-taala-carnatic">http://compmusic.upf.edu/examples-taala-carnatic</a></p> <p><strong>THE DATASET</strong></p> <p><strong>Audio music content </strong></p> <p>The pieces are chosen from the CompMusic Carnatic music collection. The pieces were chosen in four popular taalas of Carnatic music, which encompasses a majority of Carnatic music. The pieces were chosen include a mix of vocal and instrumental recordings, new and old recordings, and to span a wide variety of forms. All pieces have a percussion accompaniment, predominantly Mridangam. The excerpts are full length pieces or a part of the full length pieces. There are also several different pieces by the same artist (or release group), and multiple instances of the same composition rendered by different artists. Each piece is uniquely identified using the MBID of the recording. The pieces are stereo, 160 kbps, mp3 files sampled at 44.1 kHz.</p> <p><strong>Annotations</strong></p> <p>There are several annotations that accompany each excerpt in the dataset.</p> <p><strong>Sama and beats:</strong> The primary annotations are audio synchronized time-stamps indicating the different metrical positions in the taala cycle. The annotations were created using Sonic Visualizer by tapping to music and manually correcting the taps. Each annotation has a time-stamp and an associated numeric label that indicates the position of the beat marker in the taala cycle. The marked positions in the taala cycle are shown with numbers, along with the corresponding label used. In each case, the sama (the start of the cycle, analogous to the downbeat) are indicated using the numeral 1.</p> <p><strong>Taala related metadata</strong>: For each excerpt, the taala of the piece, edupu (offset of the start of the piece, relative to the sama, measured in aksharas) of the composition, and the kalai (the cycle length scaling factor) are recorded. Each excerpt can be uniquely identified and located with the MBID of the recording, and the relative start and end times of the excerpt within the whole recording. A separate 5 digit taala based unique ID is also provided for each excerpt as a double check. The artist, release, the lead instrument, and the raaga of the piece are additional editorial metadata obtained from the release. A flag indicates if the excerpt is a full piece or only a part of a full piece. There are optional comments on audio quality and annotation specifics. </p> <p><strong>Possible uses of the dataset</strong></p> <p>Possible tasks where the dataset can be used include taala, sama and beat tracking, tempo estimation and tracking, taala recognition, rhythm based segmentation of musical audio, structural segmentation, audio to score/lyrics alignment, and rhythmic pattern discovery.</p> <p><strong>Dataset organization</strong></p> <p>The dataset consists of audio, annotations, an accompanying spreadsheet providing additional metadata. For a detailed description of the organization, please see the README in the dataset.</p> <p><strong>Data Subset</strong></p> <p>A subset of this dataset consisting of 118 two minute excerpts of music is also available. The content in the subset is equaivalent and is separately distributed for a quicker testing of algorithms and approaches.</p> <p><strong>Using this dataset</strong></p> <p>Please cite the following publications if you use the dataset in your work:</p> <blockquote> <p>Srinivasamurthy, A., Holzapfel, A., Cemgil, A. T., & Serra, X. (2015, October). Particle Filters for Efficient Meter Tracking with Dynamic Bayesian Networks. In Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR 2015) (pp. 197–203). Malaga, Spain. (Subset)</p> </blockquote> <p><a href="http://hdl.handle.net/10230/34998">http://hdl.handle.net/10230/34998</a> </p> <blockquote> <p>Srinivasamurthy, A., & Serra, X. (2014, May). A Supervised Approach to Hierarchical Metrical Cycle Tracking from Audio Music Recordings. In Proceedings of the 39th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2014) (pp. 5237–5241). Florence, Italy. (Full dataset)</p> </blockquote> <p><a href="http://doi.org/10.1109/ICASSP.2014.6854598">http://doi.org/10.1109/ICASSP.2014.6854598</a></p> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p><strong>Contact</strong></p> <p>If you have any questions or comments about the dataset, please feel free to write to us.</p> <p>Ajay Srinivasamurthy</p> <p>Music Technology Group</p> <p>Universitat Pompeu Fabra, </p> <p>Barcelona, Spain</p> <p>ajays.murthy@upf.edu</p> <p> </p> <p><a href="http://compmusic.upf.edu/carnatic-rhythm-dataset">http://compmusic.upf.edu/carnatic-rhythm-dataset</a></p>
CloserMusicDB: A Modern Multipurpose Dataset of High Quality Music
<pre>CloserMusicDB - a collection of full length studio quality tracks annotated by a team of human experts.</pre>
A dataset for Emotion Recognition using EEG and MUsical Stimuli
<p># EREMUS: A Dataset for Emotion Recognition Using EEG and Musical Stimuli</p> <p>## Description</p> <p>EREMUS is a dataset designed for the study of emotion recognition using electroencephalography (EEG) data collected during musical stimuli exposure. The dataset includes EEG recordings from 34 young subjects in a controlled laboratory setting. Each subject participated in 16 trials, with a duration of approximately 90 seconds each. Eight trials featured songs from the subject's personal music playlist, while the other eight were randomly selected from the preferences of other participants. Following each trial, subjects self-assessed their emotional responses using the Geneva Emotion Wheel, which categorizes emotions into 20 families arranged in a circular format based on valence and dominance.</p> <p>### Data Collection</p> <p>EEG data was recorded using a 32-channel EPOC Flex EEG system with saline sensors, at a sampling rate of 128 Hz. Electrode placements adhered to the international 10-20 system.</p> <p>### Training Dataset</p> <p>The training set consists of 294 trials from 26 subjects, with each subject contributing approximately 12 labeled trials. Each trial includes the following metadata:</p> <p>- **session_type**: Indicates whether the trial is from a *personal* or *other* session.<br>- **subject_id**: Identifier for the subject.<br>- **spotify_track_id**: Spotify identifier for the song.<br>- **song_title**: Title of the song played.<br>- **song_author**: List of song authors.<br>- **emotion**: Self-assessed emotion based on the Geneva Emotion Wheel.<br>- **label**: Emotion label in a Valence-Dominance space.<br>- **id**: Trial identifier.</p> <p>### Test Dataset</p> <p>The test dataset is divided into two parts: held-out trials and held-out subjects, both released without emotion labels or subject identifiers.</p> <p>- **Held-out Trials**: Comprises 104 trials from the 26 subjects included in the training dataset.<br>- **Held-out Subjects**: Contains 122 trials from 8 subjects not present in the training dataset.</p> <p>Additionally, the held-out trials dataset includes 44 extra trials for the subject identification task, which may contain or omit stimulation information.</p> <p>### Preprocessing</p> <p>The dataset is available in two versions:</p> <p>1. **Raw EEG Data**: Unprocessed data.<br>2. **Pruned Data**: Preprocessed using EEGLab, which includes a FIR filter applied between 0.5 and 40 Hz and artifact removal via Independent Component Analysis (ICA).</p> <p>### Files Structure</p> <p>The dataset is organized in a hierarchical structure according to data type (raw or pruned) and split type (training, held-out trials, held-out subjects):</p> <p>```<br>dataset<br>├── splits_subject_identification.json<br>├── splits_emotion_recognition.json<br>├── README.md<br>├── raw<br>│ ├── train<br>│ │ ├── 1135903657_eeg.fif<br>│ │ └── ...<br>│ ├── test_trial<br>│ └── test_subject<br>└── pruned<br> ├── train<br> │ ├── 1135903657_eeg.fif<br> │ └── ...<br> ├── test_trial<br> └── test_subject<br>```</p> <p>The IDs for the test sets are specified in two JSON files:<br>- `splits_subject_identification.json` for Task 1.<br>- `splits_emotion_recognition.json` for Task 2.</p>
UMD-350MB: Refined MIDI Dataset for Symbolic Music Generation
<p><strong>UMD-350MB</strong></p> <p>The Universal MIDI Dataset 350MB (UMD-350MB) is a proprietary collection of 85,618 MIDI files curated for research and development within our organization. This collection is a subset sampled from a larger dataset developed for pretraining symbolic music models.</p> <p>The field of symbolic music generation is constrained by limited data compared to language models. Publicly available datasets, such as the Lakh MIDI Dataset, offer large collections of MIDI files sourced from the web. While the sheer volume of musical data might appear beneficial, the actual amount of valuable data is less than anticipated, as many songs contain less desirable melodies with erratic and repetitive events.</p> <p>The UMD-350MB employs an attention-based approach to achieve more desirable output generations by focusing on human-reviewed training examples of single-track melodies, chord progressions, leads and arpeggios with an average duration of 8 bars. This was achieved by refining the dataset over 24 months, ensuring consistent quality and tempo alignment. Moreover, the dataset is normalized by setting the timing information to 120 BPM with a tick resolution (PPQ) of 96 and transposing the musical scales to C major and A minor (natural scales).</p> <p><strong>Melody Styles</strong></p> <p>A major portion of the dataset is composed of newly produced private data to represent modern musical styles.</p> <ul> <li>Pop: 1970s to 2020s Pop music</li> <li>EDM: Trance, House, Synthwave, Dance, Arcade</li> <li>Jazz: Bebop, Ballad, Latin-Jazz, Bossa-Jazz, Ragtime</li> <li>Soul: 80s Classic, Neo-Soul, Latin-Soul</li> <li>Urban: Pop, Hip-Hop, Trap, R&B, Afrobeat</li> <li>World: Latin, Bossa Nova, European</li> <li>Other: Film, Cinematic, Game music and piano references</li> </ul> <p><em>Actual MIDI files are unlabeled for unsupervised training.</em></p> <p><strong>Dataset Access</strong></p> <p>Please note that this is a closed-source dataset with very limited access. Considerations for access include proposals for data augmentation, chord extraction and other enhancement methods, whether through scripts, algorithmic techniques, manual editing in a DAW or additional processing methods.</p> <p>For inquiries about this dataset, please email us.</p>
Indian Art Music Tonic Datasets (audio)
<p>This repository includes the audio for the Indian Art Music Tonic Datasets (<a href="https://zenodo.org/record/1257114">https://zenodo.org/record/1257114</a>).</p> <p>Refer to the Indian Art Music Tonic Datasets page to get the features and annotations for the dataset. A complete presentation of the datasets is available there.</p>
Dataset related to the paper "The vibe of musical and social reward: Listening to beat-based music acts as a surrogate for socioemotional support during the Covid-19 pandemic"
<p>The set includes data related to the paper "The vibe of musical and social reward: Listening to beat-based music acts as a surrogate for socioemotional support during the Covid-19 pandemic"</p>
(Tiny) Music Score Generation Dataset
<p>This dataset is a much smaller version of the Music Score Generation Dataset from the thesis on automatic score-to-score music generation</p>
Music Score Generation Dataset
<p>This dataset was created for automatic score-to-score generation for piano.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.