Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
112
datasets available to search
ShareScore release 0.9.0
Dataset results
112 results for “music dataset”
Twitter Mental Disorder Tweets and Musics Dataset
<p>https://www.kaggle.com/datasets/rrmartin/twitter-mental-disorder-tweets-and-musics?resource=download</p>
YM2413-MDB: A Multi-Instrumental FM Video Game Music Dataset with Emotion Annotations
<p>YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations. It includes 669 audio and MIDI files of music from Sega and MSX PC games in the 80s using YM2413, a programmable sound generator based on FM. The collected game music is arranged with a subset of 15 monophonic instruments and one drum instrument. They were converted from binary commands of the YM2413 sound chip. Each song was labeled with 19 emotion tags by two annotators and validated by three verifiers to obtain refined tags</p> <p>For more detailed information about the dataset, please refer to our paper: <a href="https://arxiv.org/abs/2211.07131">YM2413-MDB: A Multi-Instrumental FM Video Game Music Dataset with Emotion Annotations</a>.</p> <p><strong>File Description</strong></p> <p><strong>1) Pure data</strong></p> <p>- original_vgms: crawled vgm files from <a href="https://www.smspower.org/">SMS POWER</a> and <a href="https://vgmrips.net/packs/">VGMRIPs</a></p> <p>- wav: rendered vgm files using <a href="https://github.com/vgmrips/vgmplay">VGMPlay</a></p> <p> </p> <p><strong>2) MIDI data</strong></p> <p>- midi/vgmplay_log_to_midi: converted midi files</p> <p>- midi/adjust_tempo: add postprocessing(metrically aligned using wav_downbeat files) after midi conversion</p> <p>- midi/adjust_tempo_remove_delayed_inst: add postprocessing(metrically aligned using wav_downbeat files, remove delayed instrument) after midi conversion</p> <p> </p> <p><strong>3) Metadata</strong></p> <p>- emotion_annotation/verified_annotation.csv: contains emotion annotation for each songs</p> <p>- tags_kor_eng.txt: Korean <-> English tag dictionary</p> <p> </p> <p><strong>4) Useful middle-time step data</strong></p> <p>- wav_downbeat: extracted downbeat values using TCNBeatTracker of <a href="https://github.com/CPJKU/madmom">madmom</a></p> <p>- vgm_txts: disassembled vgm files as txt using <a href="https://github.com/vgmrips/vgmtools#vgm-text-writer-vgm2txt">vgm2txt</a></p> <p>- ydr: YM2413 Disassembly Raw(YDR). command list of vgm files. generated by reading vgm_txts</p> <p> </p> <p><strong>Update Log</strong></p> <p>- version 1.0.1: Fix ticks per beat value adjust to tempo where tempo values are not 150. Also, madmom downbeat files are updated from DBNBeatTracker(ISMIR, 2015) to TCNBeatTracker(Newer one EUSIPCO, 2019).</p> <p>- version 1.0.2: <strong><a href="https://github.com/jech2/YM2413-MDB/issues/2">Wrong emotion tag issue in the verification annotation file was fixed.</a></strong></p>
China traditional music instrument dataset
<p>The FolkMusic dataset is a Chinese traditional music dataset mainly used for training instrument recognition models and performance evaluation. The dataset covers 15 traditional Chinese musical instruments, including Ba, Flute, Dongxiao, Erhu, Guqin, Guzheng, Hulusi, Liuqin, Pipa, Sanxian, Sheng, Suona, Yangqin, Zhongruan, and Falling Qin. The music clips in each instrument are saved as .mp3 files, which are recorded via two channels with a sampling rate of 44100Hz. The duration of these music clips are 3s, and a single instrument plays each music clip.</p>
Data from: Creating a multi-track classical music performance dataset for multi-modal music analysis: challenges, insights, and applications
Open the record for dataset details and reuse information.
Dataset from the EMNLP 2020 article "Modeling the Music Genre Perception across Language-Bound Cultures"
<p>We release the data required to reproduce the experiments from the article <em>Modeling the Music Genre Perception across Language-Bound Cultures</em> presented at the <a href="https://2020.emnlp.org">EMNLP 2020</a> conference.</p> <p>More information about this data and how it should be used in the experiments can be found in the GitHub repository <a href="https://github.com/deezer/CrossCulturalMusicGenrePerception">deezer/CrossCulturalMusicGenrePerception</a>.</p> <p>Please cite our paper if you use the code or data in your work.</p>
Age-related neural changes underlying long-term recognition of musical sequences - Communications Biology - Leonardo Bonetti - Dataset
Open the record for dataset details and reuse information.
FMAKv2: A Dataset of Key and Mode Annotations for the Free Music Archive
<p>We present FMAKv2, a deriavative work of <a href="../records/10719860">FMAK</a>, a dataset containing song-level key and mode annotations of 5489 songs, spread across 17 genres, released and used in the paper <a href="https://arxiv.org/abs/2407.07408"><strong>STONE: Self-supervised Tonality Estimator</strong></a>, accpeted at <strong>ISMIR 2024</strong>.</p> <p>About FMAK:</p> <blockquote> <p><a href="../records/10719860">FMAK</a> is a an expert-labeled dataset for the evaluation of key detection. The curation and annotations of 5489 songs were all<strong> </strong>created by <strong>Stella Wong</strong> (co-author of STONE) and <strong>Gandalf Hernandez</strong>. The FMAK metadata is made freely available for public use under a <strong>Creative Commons Attribution 4.0 International License. </strong>The work was presented as an LBD at ISMIR 2023 at first, later published with STONE, at ISMIR 2024.</p> <p>The DOI of FMAK is <code>10.5281/zenodo.10719860</code> and the link is <a href="../records/10719860">https://zenodo.org/records/10719860</a></p> </blockquote> <p>The difference between FMAK and FMAKv2 is <strong>only</strong> the modification of around 200 songs' annotations. Other annotations remain the same as FMAK, therefore <strong>created</strong>, <strong>curated</strong>, and <strong>annotated</strong> by <strong>Stella Wong</strong> and <strong>Gandalf Hernandez</strong>. FMA track id and Spotify URI remain unchanged from FMAK. Authors of FMAK did not verify the modifications of annotations of FMAKv2 and should <strong>not</strong> be held liable for potential mislabelings in FMAKv2.</p> <p>For each song, we provide identical information from FMAK of:</p> <ul> <li>FMA track id (6 digits)</li> <li>Spotify URI (when available)</li> <li>Key and mode</li> </ul> <p>All the audios in FMAKv2 are identical as FMAK, and can be downloaded from <a href="../records/10719860">FMAK</a> repository.<br><br></p> <p>If you use annotations from fmakv2, please cite the following papers:</p> <pre>@article{kong2024stone, title={STONE: Self-supervised Tonality Estimator}, author={Kong, Yuexuan and Lostanlen, Vincent and Meseguer-Brocal, Gabriel and Wong, Stella and Lagrange, Mathieu and Hennequin, Romain}, journal={Proceedings of International Society for Music Information Retrieval Conference (ISMIR 2024)}, year={2024} }</pre> <pre>@inproceedings{wong2023fmak, title={FMAK: A DATASET OF KEY AND MODE ANNOTATIONS FOR THE FREE MUSIC ARCHIVE--EXTENDED ABSTRACT}, author={Wong, Stella and Hernandez, Gandalf}, booktitle={International Society for Music Information Retrieval Late-Breaking/Demo Session (ISMIR-LBD)}, year={2023} }</pre>
Self-Assembly and Synchronization: Crafting Music with Multi-Agent Embodied Oscillators - DATASET
<p>Dataset for amalysis replication</p>
Passive Head-Mounted Display Music-Listening EEG dataset
<p><strong>Summary:</strong></p> <p>This dataset contains electroencephalographic recordings of 12 subjects listening to music with and without a passive head-mounted display, that is, a head-mounted display which does not include any electronics at the exception of a smartphone. The electroencephalographic headset consisted of 16 electrodes. A full description of the experiment is available at <a href="https://hal.archives-ouvertes.fr/hal-02085118">https://hal.archives-ouvertes.fr/hal-02085118</a>. Data were recorded during a pilot experiment taking place in the GIPSA-lab, Grenoble, France, in 2017 (Cattan and al, 2018). Python code for manipulating the data is downloadable at <a href="https://github.com/plcrodrigues/py.PHMDML.EEG.2017-GIPSA">https://github.com/plcrodrigues/py.PHMDML.EEG.2017-GIPSA</a>. The ID of this dataset is <em>PHMDML.EEG.2017-GIPSA.</em></p> <p> </p> <p><strong>Full description of the experiment and dataset: </strong><a href="https://hal.archives-ouvertes.fr/hal-02085118">https://hal.archives-ouvertes.fr/hal-02085118</a></p> <p> </p> <p><strong><em>Principal Investigator</em>:</strong> Eng. Grégoire Cattan</p> <p> </p> <p><strong><em>Technical Supervisors</em>:</strong> Eng. Pedro L. C. Rodrigues</p> <p> </p> <p><strong><em>Scientific Supervisor:</em></strong> Dr. Marco Congedo</p> <p> </p> <p><strong>ID of the dataset: </strong><em>PHMDML.EEG.2017-GIPSA</em></p>
Datasets for MuSiC Compare Tutorial
<p>These datasets are used for the MuSiC compare tutorial in Galaxy.</p>
Datasets for Deconvolution with MuSiC Tutorial
<p>These datasets are used in the MuSiC deconvolution tutorial in Galaxy. They were retrieved from the EBI's Array Express platform, originally published by:</p> <p>Segerstolpe Å, Palasantza A, Eliasson P, Andersson EM, Andréasson AC, Sun X, Picelli S, Sabirsh A, Clausen M, Bjursell MK, Smith DM, Kasper M, Ämmälä C, Sandberg R. Single-Cell Transcriptome Profiling of Human Pancreatic Islets in Health and Type 2 Diabetes. Cell Metab. 2016 Oct 11;24(4):593-607. doi: 10.1016/j.cmet.2016.08.020. Epub 2016 Sep 22. PMID: 27667667; PMCID: PMC5069352.</p>
Datasets for MuSiC Deconvolution benchmarking tutorial suite.
<p>These are the datasets for the MuSiC deconvolution benchmarking tutorial suite.</p>
Bioacoustic dataset assessing the impact of festival music on bat activity
<p>Sound is a critical component of an animal's habitat, where it is used to glean important environmental information from their surroundings. The modification of natural soundscapes due to the global rise in anthropogenic noise pollution over recent decades can have serious negative impacts on species fitness and survival.</p> <p>Nocturnal species such as bats are reliant on sound for many aspects of their life history and are therefore highly sensitive to anthropogenic noise. Music festivals are a source of unregulated and potentially harmful, acute noise pollution however they have become ubiquitous across our landscapes throughout the summer months and are increasingly being held in settings important for wildlife.</p> <p>Using an experimental approach, we provide the first evidence of negative impacts of music festivals on bat activity in a habitat that represents a typical festival setting i.e. woodland edge. We found that loud music playback alone can reduce the activity of bats even in the absence of other anthropogenic factors commonly associated with music festivals such as lighting and habitat disturbance. Activity of <em>Nyctalus/Eptesicus</em> spp. was reduced along woodland edge habitats exposed to loud music whereas no effect was recorded for <em>Myotis</em> spp., <em>P. pygmaeus</em> and <em>P. pipistrellus</em> compared to quiet nights. We also provide the first evidence of the spatial scale of negative effects from festival music on activity for <em>P.pygmaeus</em> as well as highlighting differential responses between cryptic species.</p> <p>In light of the paucity of research or guidance into acute noise impacts on nocturnal biodiversity, we outline the potential negative impacts of music festivals for bats. We show that music alone can reduce the activity of bats even in the absence of other anthropogenic factors commonly associated with music festivals which could potentially fragment important habitats for certain species, leading to a degradation of functional connectivity across the landscape.</p>
mshoxxDB - a Versioned Dataset for Electronic Music
<p><strong>Description:</strong> Version 1(.1) consists of 18 full-length pieces of music in the genre of Electronic Music, totaling 61 minutes of audio. The dataset spans a number of sub-genres of Electronic Music: Video Game, 8-Bit (Chiptune), EDM, Pop, House, Chillout/Dreamy. The dataset is suited for a variety of tasks in the field of Music Information Retrieval (MIR), such as Source Separation, Multi-Pitch Estimation, Beat Detection, Tempo Estimation. It is particularly interesting for instrument-agnostic methods and evaluations of model generalization due to the wide variety of synthetic and traditional timbres.</p> <p><strong>Contents:</strong><br>- mixtures and multi-tracks in FLAC format (44.1 kHz, 16-Bit, Mono, compression level 6)<br>- track-wise MIDI files<br>- CSV metadata with genre, tempo, time signature, and artist information</p> <p><strong>Technical Properties:</strong> Not all mixtures are exact sums of their respective multi-tracks. The mixtures may contain additional processing in the form of limiters and compression (on the full mix or side-chain compression from one track to another). <strong><em>No harmonic effects were added onto the mixtures</em></strong> (= effects such as reverbs, echos, and delays that add more harmonic information and would result in mismatches between MIDI and audio).</p> <p><strong>License:</strong> All contents distributed under Creative Commons BY-NC-SA 4.0.</p> <p><strong>Demo Page:</strong> For a few listening examples, please visit this dataset's github page at <a href="https://mic-tae.github.io/mshoxxdb/">https://mic-tae.github.io/mshoxxdb/</a>.</p> <p><strong>Repo: </strong>The mshoxxDB repo is located at <a href="https://github.com/mic-tae/mshoxxdb">https://github.com/mic-tae/mshoxxdb</a>, and may contain more up-to-date information (README.md).</p> <p><strong>Citation:</strong> Should you use this dataset in your work, please cite it the following way (bibtex):</p> <blockquote> <p>@misc {taenzer:mshoxxDB:2024,<br> author = {Taenzer, Michael},<br> title = {{mshoxxDB - a Versioned Dataset for Electronic Music}},<br> booktitle = {{Late-Breaking and Demo Session of the 25th International Conference on Music Information Retrieval (ISMIR)}},<br> address = {{San Francisco, CA, USA}},<br> year = {2024},<br>}</p> </blockquote> <p><strong>Version 1.1 (16 July 2025)</strong><br>- all files now reflect main DB version number v1 (previous numbers were personal track session numbers)<br>- removed umlaut from Güte --> Guete<br>- added ms12 and ms14 dataset splits (json files), a license and readme file</p>
Bioacoustic dataset assessing the impact of festival music on bat activity
Open the record for dataset details and reuse information.
Dataset from the article, "Interference by linguistic processes in the occurrence of lyrics in involuntary musical imagery"
<p>Dataset from the article, "Interference by linguistic processes in the occurrence of lyrics in involuntary musical imagery", in Journal of Cognitive Psychology.</p>
The dataset used in the article "Listener Modeling and Context-aware Music Recommendation Based on Country Archetypes"
<p>This is the dataset used in the study "Listener Modeling and Context-aware Music Recommendation Based on Country Archetypes". The dataset is a subset of the LFM-1b LastFM dataset (http://www.cp.jku.at/datasets/LFM-1b/), which contains country-specific music listening events. </p> <p> </p>
Instrumental background music mitigates self-control fatigue and improves performance in prolonged mental work (Dataset)
<p>This repository contains the raw data used for a research study titled "Instrumental Background Music Facilitates Task Endurance and Cognitive Performance: Insights from Ego Depletion". This repository contains two Microsoft Excel files, as follows:</p> <ul> <li><em>ed_music_feature_analysis.csv</em></li> <li><em>ed_raw_data.xlsx</em></li> <li><em>ed_appendices [Anonymous].pdf</em></li> </ul> <p><strong>Files Description</strong></p> <p><em><strong>File: ed_music_feature_analysis.csv</strong></em></p> <p>This file contains the specific feature analysis for each music stimulus (N = 17) used in the experiment. These features (retrieved from Spotify's "Get Track's Audio Features" API) are:</p> <ul> <li>energy level</li> <li>valence</li> <li>tempo</li> </ul> <p>The Spotify ID of each music track is also recorded in this file, to facilitate future replication of the experiment.</p> <p><em><strong>File: ed_raw_data.xlsx</strong></em></p> <p>This file contains the raw data from the experiment (N = 41). There are two tabs in this file. The "data" tab records all raw data, whilst the "index" tab contains all the abbreviations (with their explanations) used in the document, and when relevant, a brief description of how the outcome measures were calculated.</p> <p>In the "data" tab, each row corresponds to a single participant and each column to a single variable. The variables about sample characteristics include:</p> <ul> <li>Age</li> <li>Self-identified gender</li> <li>Area of study</li> <li>Level of English proficiency</li> <li>Whether the participant has specific neurodiversity</li> <li>Personality traits</li> <li>Musicianship</li> <li>Self-reported everyday music listening habits</li> <li>Self-reported perceived (positive and negative) impact of music listening while studying</li> </ul> <p>The variables about experimental conditions and outcomes measures include:</p> <ul> <li>Each participant's assigned experimental condition (music or control)</li> <li>Reading comprehension performance (measured in terms of the number of questions answered and a final score)</li> <li>Pre- and post-reading positive affect, negative affect, and motivation for cognition</li> <li>(for the music condition) personal perception of the music they heard during the experiment</li> <li>Level of ego depletion (measured based on Stroop interference)</li> <li>Verbal working memory (measured based on Reading Span Task performance)</li> <li>Divided attention capacity (measured based on Category Switch Task performance)</li> </ul> <p><em>Notes.</em></p> <p><em>1. Empty cells in between data are missing data (due to participant absentees).</em></p> <p><em>2. See the "index" tab in <strong>ed_raw_data.xlsx </strong>for </em>a <em>detailed presentation and explanation of each recorded data.</em></p> <p><strong><em>File: </em><em>ed_appendices [Anonymous].pdf</em></strong></p> <p>This file contains the four appendices mentioned in the main article.</p>
MOSA: Music mOtion and Semantic Annotation dataset
<p>MOSA dataset is a large-scale music dataset containing 742 professional piano and violin solo music performances with 23 musicians (> 30 hours, and > 570 K notes). This dataset features following types of data:</p> <ul> <li><strong>High-quality 3-D motion capture data</strong></li> <li><strong>Audio recordings</strong></li> <li><strong>Manual semantic annotations</strong></li> </ul> <p>This is the dataset of the paper: Huang et al. (2024) MOSA: Music Motion with Semantic Annotation Dataset for Multimedia Anaysis and Generation. IEEE/ACM Transactions on Audio, Speech and Language Processing. DOI: 10.1109/TASLP.2024.3407529<br>https://arxiv.org/abs/2406.06375</p> <p> </p> <p>The description of dataset is avaiable on Github: https://github.com/yufenhuang/MOSA-Music-mOtion-and-Semantic-Annotation-dataset/blob/main/MOSA-dataset/dataset.md</p> <p> </p> <p>To request the access of full dataset, please sign in with Zenodo and submit the request from.</p>
Indian Art Music Raga Recognition Dataset (audio)
<p>The <strong>Rāga Recognition Datasets (audio)</strong> comprise two sizable datasets, one for each music tradition: the <strong>Carnatic Music Dataset (CMD)</strong> and the <strong>Hindustani Music Dataset (HMD)</strong>. This repository contains audio files that complement the features provided in the <strong><a href="https://zenodo.org/records/7278506">Rāga Recognition (features</a>)</strong> dataset. As the audio files are copyrighted, they must be <strong>requested through this repository</strong>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.