Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

5 results for “Million song dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

ESSENTIA analysis of audio snippets from the Million Song Dataset Taste Profile subset

<p>This upload includes the ESSENTIA analysis output of (a subset of) song snippets from the Million Song Dataset, namely those included in the Taste Profile subset. The audio snippets were collected from 7digital.com and were subsequently analyzed with ESSENTIA 2.1-beta3. Pre-trained SVM models provided by the ESSENTIA authors on their website were applied.</p> <p>The file <strong>msd_song_jsons.rar </strong>contains the ESSENTIA analysis output after applying the SVM models for highlevel feature extraction. Please note that these are 204317 files.</p> <p>The file <strong>msd_played_songs_essentia.csv.gz </strong>contains all one-dimensional real-valued fields of the jsons merged into one csv file with 204317 rows.</p> <p>The full procedure and subsequent analysis is described in</p> <p>Fricke, K. R., Greenberg, D. M., Rentfrow, P. J., &amp; Herzberg, P. Y. (2019). Measuring musical preferences from listening behavior: Data from one million people and 200,000 songs. <em>Psychology of Music</em>, 0305735619868280.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

MSD-A: Million Song Dataset for Artists

<p>The MSD-A is a dataset related to the Million Song Dataset (MSD). It is a collection of artist tags and biographies gathered from Last.fm for all the artists that have songs in the MSD. In addition, the MSD Taste Profile (recommendation dataset) is adapted to artists.</p> <p>We provide the biographies, tags, data splits, and feature embeddings to reproduce the experiments from the paper:</p> <p>Oramas S., Nieto O., Sordo M., &amp; Serra X. (2017) A Deep Multimodal Approach for Cold-start Music Recommendation. https://arxiv.org/abs/1706.09739</p> <p>Source code is available at https://github.com/sergiooramas/tartarus</p> <p>The file dlrs-data.tar.gz in this zenodo version is corrupted. You can download the good file in this link:</p> <p>https://drive.google.com/open?id=0B-oq_x72w8NUbUpkMzZSc1JPd28</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

MSD-I: Million Song Dataset with Images for Multimodal Genre Classification

<p>The Million Song Dataset (https://labrosa.ee.columbia.edu/millionsong/) is a collection of metadata and precomputed audio features for 1 million songs. Along with this dataset, a dataset with annotations of 15 top-level genres with a single label per song was released. In our work, we combine the CD2c version of this genre datase (http://www.tagtraum.com/msd_genre_datasets.html) with a collection of album cover images.&nbsp;</p> <p><br> The final dataset contains 30,713 tracks from the MSD and their related album cover images, each annotated with a unique genre label among 15 classes. Based on an initial analysis on the images, we identified that this set of tracks is associated to 16,753 albums, yielding an average of 1.8 songs per album.</p> <p>We randomly divide the dataset into three parts: 70% for training, 15% for validation, and 15% for test, with no artist and album overlap across these sets. This is crucial to avoid possible overfitting, as the classifier may learn to predict the artist instead of the genre.&nbsp;</p> <p>&nbsp;</p> <p>Content:</p> <p>MSD-I dataset (mapping, metadata, annotations and links to images)<br> Data splits and feature vectors for TISMIR single-label classification experiments&nbsp;</p> <p>These data can be used together with the Tartarus deep learning python module&nbsp;https://github.com/sergiooramas/tartarus.</p> <p>&nbsp;</p> <p>Scientific References:</p> <p>Please cite the following paper if using MSD-I dataset or Tartarus software.</p> <p>Oramas, S., Barbieri, F., Nieto, O., and Serra, X (2018). Multimodal Deep Learning for Music Genre Classification, Transactions of the International Society for Music Information Retrieval,&nbsp;V(1).</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Hit Song Prediction (Million Song Dataset and Audio Features)

<p><strong>Hit Song Prediction Dataset</strong></p> <p>This dataset is based on the Million Song Dataset (MSD), which contains one million songs that are representative for western commercial music released between 1922 and 2011. The dataset contains release year information for 515,576 of the MSD songs. Please refer to http://millionsongdataset.com/ for further information on the million song dataset.</p> <p>For our hit song prediction experiments, we extract high- and low-level audio features using the Essentia toolkit (cf. https://essentia.upf.edu/). For the high-level features, we make use of the pre-trained classifiers as provided by Essentia. For a detailed description of the features, please visit the Essentia documentation.</p> <p><br> The dataset hence contains:</p> <ul> <li><strong>Audio features</strong>: the compressed msd_audio_features.tar.gz file contains the low- and high-level features for each track, stored as json files. Please note that we organize all MSD audio feature files based on the track&#39;s identifier with one folder holding all tracks with the same first letter of the track identifier to keep the files manageable. For each track, we provide two files: one containing the high-level and one containing the low-level features extracted by Essentia.</li> <li><strong>Billboard data:</strong> the folder billboard_data contains two files: msd_bb_matches.csv&nbsp;contains information about the MSD tracks that were also featured in the Billboard Hot 100 charts. Here, we provide the MSD id, Echo Nest id, artist name, track title, release year, peak position in Billboard charts and the number of weeks in the charts. The second file, msd_bb_non_matches.csv&nbsp;contains meta-information about the tracks of the MSD that were not featured in the Billboard Hot 100 and hence were used as negative samples. Here, we provide the MSD id, Echo Nest id, artist name, track title and the release year.</li> </ul> <p><br> If you make use of the dataset, please kindly cite the following paper:</p> <p>Eva Zangerle, Michael V&ouml;tter, Ramona Huber, and Yi-Hsuan Yang. Hit Song Prediction: Leveraging Low- and High-Level Audio Features. In Proceedings of the 20th International Society for Music Information Retrieval Conference 2019 (ISMIR 2019), 2019.</p> <p><br> @inproceedings{zangerle_ismir19,<br> title = {{Hit Song Prediction: Leveraging Low- and High-Level Audio Features}},<br> author = {Eva Zangerle and Ramona Huber and Michael V\&quot;{o}tter and Yi-Hsuan Yang},<br> year = {2019},<br> booktitle = {{Proceedings of the 20th International Society for Music Information Retrieval Conference 2019 (ISMIR 2019)}},<br> }</p>

opencc-by-4.0Jun 2019View details →
dryad32/100

The Million Song Dataset

Open the record for dataset details and reuse information.

publicOct 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record