Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
112
datasets available to search
ShareScore release 0.9.0
Dataset results
112 results for “music dataset”
A dataset recorded during development of an affective brain-computer music interface: calibration session
Open the record for dataset details and reuse information.
A dataset recorded during development of an affective brain-computer music interface: testing session
Open the record for dataset details and reuse information.
A dataset recorded during development of an affective brain-computer music interface: training sessions
Open the record for dataset details and reuse information.
The Italian Music Dataset
<p><strong>Overview</strong></p> <p>The dataset is built by exploiting the Spotify and SoundCloud APIs. It is composed of over 14,500 different songs of both famous and less famous Italian musicians. Each song in the dataset is identified by its Spotify id and its title. Tracks' metadata include also lemmatized and POS-tagged lyrics and, in the most of cases, ten musical features directly gathered from Spotify. Musical features include acousticness (float), danceability (float), duration_ms (int), energy (float), instrumentalness (float), liveness (float), loudness (float), speechiness (float), tempo (float) and valence (float). All features range from 0.0 to 1.0 except for loudness that typically ranges between -60 and 0 db, the tempo that represents beats per minute (BPM) and the duration that represents the track in milliseconds. For further information refer to the Spotify's documentation at <a href="https://developer.spotify.com/documentation/web-api/reference/tracks/get-audio-features/">Spotify Documentation</a></p> <p>For further information regarding the dataset and the related project visit <a href="https://bit.ly/2MUUwEx">SoBigData Catalogue</a></p>
Medley-solos-DB: a cross-collection dataset for musical instrument recognition
<p>Medley-solos-DB<br> =============<br> Version 1.2 March 2019.<br> </p> <p> </p> <p>Created By<br> --------------</p> <p>Vincent Lostanlen (1), Carmine-Emanuele Cella (2), Rachel Bittner (3), Slim Essid (4).<br> <br> (1): New York University<br> (2): UC Berkeley<br> (3): Spotify, Inc.<br> (4): Télécom ParisTech</p> <p> </p> <p><br> Description<br> ---------------</p> <p> </p> <p>Medley-solos-DB is a cross-collection dataset for automatic musical instrument recognition in solo recordings. It consists of a training set of 3-second audio clips, which are extracted from the MedleyDB dataset of Bittner et al. (ISMIR 2014) as well as a test set set of 3-second clips, which are extracted from the solosDB dataset of Essid et al. (IEEE TASLP 2009). Each of these clips contains a single instrument among a taxonomy of eight: clarinet, distorted electric guitar, female singer, flute, piano, tenor saxophone, trumpet, and violin.</p> <p>The Medley-solos-DB dataset is the dataset that is used in the benchmarks of musical instrument recognition in the publications of Lostanlen and Cella (ISMIR 2016) and Andén et al. (IEEE TSP 2019).</p> <p> </p> <p>[1] V. Lostanlen, C.E. Cella. Deep convolutional networks on the pitch spiral for musical instrument recognition. Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2016.</p> <p>[2] J. Andén, V. Lostanlen, and S. Mallat. Joint time-frequency scattering. IEEE Transactions in Signal Processing, vol. 67, no. 14, pp. 3704-3718, 2019. doi: 10.1109/TSP.2019.2918992</p> <p> </p> <p><br> Data Files<br> --------------</p> <p>The Medley-solos-DB contains 21571 audio clips as WAV files, sampled at 44.1 kHz, with a single channel (mono), at a bit depth of 32. Every audio clip has a fixed duration of 2972 milliseconds, that is, 65536 discrete-time samples.</p> <p>Every audio file has a name of the form:</p> <p>Medley-solos-DB_SUBSET-INSTRUMENTID_UUID.wav</p> <p> </p> <p>For example:</p> <p>Medley-solos-DB_test-0_0a282672-c22c-59ff-faaa-ff9eb73fc8e6.wav</p> <p>corresponds to the snippet whose universally unique identifier (UUID) is 0a282672-c22c-59ff-faaa-ff9eb73fc8e6, contains clarinet sounds (clarinet has instrument id equal to 0), and belongs to the test set.</p> <p> </p> <p><br> Metadata Files<br> -------------------</p> <p>The Medley-solos-DB_metadata is a CSV file containing 21572 rows (one for each audio clip) and five columns:</p> <p>1. subset: either "training", "validation", or "test"</p> <p>2. instrument: tag in Medley-DB taxonomy, such as "clarinet", "distorted electric guitar", etc.</p> <p>3. instrument id: integer from 0 to 7. There is a one-to-one between "instrument" (string format) and "instrument id" (integer). We provide both for convenience.</p> <p>4. song id: integer from 0 to 226. The track and artist names are anonymized.</p> <p>5. UUID4: universally unique identifier. Assigned and random, and different for every row.</p> <p> </p> <p>The list of instrument classes is:</p> <p>0. clarinet</p> <p>1. distorted electric guitar</p> <p>2. female singer</p> <p>3. flute</p> <p>4. piano</p> <p>5. tenor saxophone</p> <p>6. trumpet</p> <p>7. violin</p> <p> </p> <p><br> Please acknowledge Medley-solos-DB in academic research<br> ---------------------------------------------------------------------------------</p> <p>When Medley-solos-DB is used for academic research, we would highly appreciate it if scientific publications of works partly based on this dataset cite the following publication:</p> <p>V. Lostanlen, C.E. Cella. Deep convolutional networks on the pitch spiral for musical instrument recognition. Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2016.</p> <p>The creation of this dataset was supported by ERC InvariantClass grant 320959.</p> <p> </p> <p><br> Conditions of Use<br> ------------------------</p> <p>Dataset created by Vincent Lostanlen, Rachel Bittner, and Slim Essid, as a derivative work of Medley-DB and solos-Db.</p> <p>The Medley-solos-DB dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> https://creativecommons.org/licenses/by/4.0/</p> <p>The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the authors are not liable for, and expressly exclude all liability for, loss or damage however and whenever caused to anyone by any use of the Medley-solos-DB dataset or any part of it.</p> <p> </p> <p><br> Feedback<br> -------------</p> <p>Please help us improve Medley-solos-DB by sending your feedback to:<br> vincent.lostanlen@nyu.edu</p> <p>In case of a problem, please include as many details as possible.</p> <p> </p> <p> </p> <p>Acknowledgement<br> -------------------------<br> We thank all artists, recording engineers, curators, and annotators of both MedleyDB and solosDb.</p>
MAD-EEG: an EEG dataset for decoding auditory attention to a target instrument in polyphonic music
<p>The <em><strong>MAD-EEG Dataset</strong></em> is a research corpus for studying EEG-based auditory attention decoding to a target instrument in polyphonic music. </p> <p>The dataset consists of 20-channel EEG responses to music recorded from 8 subjects while attending to a particular instrument in a music mixture. </p> <p>For further details, please refer to the paper: <em><a href="https://hal.archives-ouvertes.fr/hal-02291882/document">MAD-EEG: an EEG dataset for decoding auditory attention to a target instrument in polyphonic music</a>.</em></p> <p>If you use the data in your research, please reference the paper (not just the Zenodo record):</p> <pre><code>@inproceedings{Cantisani2019, author={Giorgia Cantisani and Gabriel Trégoat and Slim Essid and Gaël Richard}, title={{MAD-EEG: an EEG dataset for decoding auditory attention to a target instrument in polyphonic music}}, year=2019, booktitle={Proc. SMM19, Workshop on Speech, Music and Mind 2019}, pages={51--55}, doi={10.21437/SMM.2019-11}, url={http://dx.doi.org/10.21437/SMM.2019-11} }</code></pre> <p> </p>
A dataset recorded during development of a tempo-based brain-computer music interface
Open the record for dataset details and reuse information.
A dataset recording joint EEG-fMRI during affective music listening
Open the record for dataset details and reuse information.
Culture-Aware Music Recommendation Dataset
<p><strong>LFM-1b dataset extended by acoustic track features and cultural cues describing users</strong></p> <p> </p> <p>This dataset is based on the LFM-1b dataset (cf. <a href="http://www.cp.jku.at/datasets/LFM-1b/">http://www.cp.jku.at/datasets/LFM-1b/</a>), however, adds acoustic features describing the tracks to the original dataset as well as cultural aspects describing users (taken from Hofstede's six dimension model and the World Happiness Report) on the country-level.</p> <p>For the creation of the dataset, we extract all users for which the original dataset contains country information for. We extract the listening events of these users and match the tracks against the Spotify API to subsequently retrieve the acoustic features of these tracks (cf. [Spotify Audio Feature Description](https://developer.spotify.com/documentation/web-api/reference/object-model/#audio-features-object)). The final dataset contains only events of users with country information and tracks with acoustic features, which can be matched with the country-level data of the World Happiness Report and Hofstede's cultural dimensions to add cultural and socio-economic aspects for users.</p> <p>This new dataset contains</p> <ul> <li>55,190 users</li> <li>3,471,884 tracks including acoustic features</li> <li>351,469,333 listening events of those users for tracks we have obtained acoustic features for</li> <li>Hofstede's cultural dimensions for 47 countries</li> <li>World Happiness Report (WHR) data for 164 countries</li> </ul> <p> </p> <p><strong>Files</strong><br> All files are tab-separated, with no quoting of strings. The dataset contains the following files, whose content we describe in more detail in the following parts.</p> <p>* acoustic_features_lfm_id.tsv: acoustic features for all tracks in the dataset, identified by their LFM track identifier<br> * events.tsv: listening events for all users<br> * hofstede.tsv: Hofstede's cultural dimensions<br> * users.tsv: user metadata<br> * world_happiness_report_2018.tsv: World Happiness Report data</p> <p>For further information on the contents of these files, please cf. the Readme file.</p> <p> </p> <p>Please cite the following paper when using the dataset:<br> Zangerle, E., Pichl, M. and Schedl, M., 2020. User Models for Culture-Aware Music Recommendation: Fusing Acoustic and Cultural Cues. <em>Transactions of the International Society for Music Information Retrieval</em>, 3(1), pp.1–16. DOI: <a href="http://doi.org/10.5334/tismir.37">http://doi.org/10.5334/tismir.37</a></p>
P4KxSpotify: A Dataset of Pitchfork Music Reviews and Spotify Musical Features
<p>18,403 music reviews scraped from Pitchfork, including relevant metadata such as author, review date, record release year, score, and genre, along with those album's audio features pulled from Spotify's API.</p>
OrchideaSOL: an audio dataset of isolated musical notes, including mutes and extended playing techniques
<p>OrchideaSOL<br> ==========<br> Version 2.0, April 2020.<br> </p> <p> </p> <p>Created By<br> --------------</p> <p>Carmine-Emanuele Cella (1), Daniele Ghisi (1), Vincent Lostanlen (2), Fabien Lévy (3), Joshua Fineberg (4), Yan Maresz (5)<br> <br> (1): UC Berkeley<br> (2): New York University<br> (3): Columbia University<br> (4): Boston University<br> (5): Conservatoire de Paris</p> <p> </p> <p>Description<br> ---------------</p> <p><br> OrchideaSOL is a dataset of 13265 samples, each containing a single musical note from one of 14 different instruments:</p> <ol> <li>Bass Tuba</li> <li>French Horn</li> <li>Trombone</li> <li>Trumpet in C</li> <li>Accordion</li> <li>Contrabass</li> <li>Violin</li> <li>Viola</li> <li>Violoncello</li> <li>Bassoon</li> <li>Clarinet in B-flat</li> <li>Flute</li> <li>Oboe</li> <li>Alto Saxophone</li> </ol> <p> </p> <p>These sounds were originally recorded at Ircam in Paris (France) between 1996 and 1999, as part of a larger project named Studio On Line (SOL). One asset of OrchideaSOL is that it contains many combinations of mutes and extended playing techniques.<br> <br> The OrchideaSOL audio data can be used for creative purposes insofar at the use complies with the Ircam Forum License. Please visit: https://forum.ircam.fr/legal/contrat-de-licence-forum-ircam/</p> <p><br> The OrchideaSOL metadata can be used for creative purposes insofar at the use complies with the Creative Commons Attribution 4.0 International license (see below).<br> <br> OrchideaSOL can be used for education and research purposes. In particular, it can be employed as a dataset for training and/or evaluating music information retrieval (MIR) systems, for tasks such as instrument recognition, playing technique recognition, or fundamental frequency estimation. For this purpose, we provide an official 5-fold split of OrchideaSOL. This split has been carefully balanced in terms of instrumentation, pitch range, and dynamics. For the sake of research reproducibility, we encourage users of OrchideaSOL to adopt this split and report their results in terms of average performance across folds.</p> <p> </p> <p>Data Files<br> --------------</p> <p>OrchideaSOL contains 13265 audio clips as WAV files, sampled at 44.1 kHz, with a single channel (mono), at a bit depth of 16. This is equivalent to the audio quality of a compact disc. Audio clips vary in duration between two and ten seconds.</p> <p>Every audio file has a file path of the form:<br> <FAMILY>/<INSTRUMENT><+MUTE>/<TECHNIQUE>/<INSTR><+M>-<TECH>-<PITCH>-<DYN>-<INSTANCE>-<MISC>.wav</p> <p><br> where:</p> <ul> <li><FAMILY> corresponds to the instrument family: "Brass", "Keyboards" (includes accordion), "Strings", and "Winds" (i.e., woodwinds).</li> <li><INSTRUMENT> is the full name of the instrument.</li> <li><+MUTE> is the type of mute being used, such as "wah", "harmon", "piombo", or "sordina". If there is no mute, this field is absent.</li> <li><TECHNIQUE> is the type of playing technique.</li> <li><INSTR> is the abbreviation of the instrument.</li> <li><+M> is the abbreviation of the type of mute, if applicable.</li> <li><TECH> is the abbreviation of playing technique.</li> <li><PITCH> denotes the pitch of the musical note. This pitch is encoded in the American standard pitch notation: pitch class (C means "do") followed by pitch octave. According to this convention, A4 has a fundamental frequency of 440 Hz.</li> <li><DYN> denotes the intensity dynamics, ranked from pp (pianissimo) to ff (fortissimo).</li> <li><INSTANCE> contains additional information, when applicable. For example, for bowed string instruments, the same pitch may sometimes be achieved on different positions and different strings, resulting in small timbre differences. In this case the label "1c", "2c", "3c", or "4c" denotes the string which is being bowed. (The letter c originates from the word "corde", which means string in French.) By convention, the first string is the one with the highest pitch when played as an open string. Furthermore, on some wind instruments, the same note was played multiple times, e.g. at multiple durations. In this case, we use the label "alt1", "alt2", etc. to denote alternative instances of the note. If none of these tags apply, the <INSTANCE> field becomes "N", which stands for "Not Applicable".</li> <li><MISC> contains additional information, if applicable. In OrchideaSOL, some pitches were never recorded, and thus missing from the chromatic scale. In this case, the <MISC> tag contains a letter "R", to denote the fact that the corresponding WAV file has been obtained by transforming a different audio clip via some digital frequency transposition (similar to Auto-Tune). The letter "R" stands for "resampled". Furthermore, some pitches were slightly out of tune in comparison with the A440 tuning standard. Again, we applied some digital frequency transposition to correct them and put them exactly in tune. The amount of frequency transposition is measured in "cents" of an equal-tempered semitone. The letter "T" stands for "tuned". Because we employed a high-fidelity algorithm for frequency transposition, and because the amount of digital frequency transposition is small, the timbre of pitch-corrected notes remains faithful to the instrument. If none of these tags apply, the <MISC> field becomes "N", which stands for "natural"; in this case, the note is distributed exactly as it was recorded in the studio.</li> </ul> <p>For example, "Strings/Violin+sordina/tremolo/Vn+S-trem-A4-mf-4c-T13d_R200d.wav" corresponds to:</p> <ul> <li>a violin sound;</li> <li>equipped with a sordina mute;</li> <li>played in the tremolo playing technique;</li> <li>at pitch A4 (440 Hz);</li> <li>with mezzoforte dynamics;</li> <li>on the fourth string (i.e. the lowest);</li> <li>resampled from a B4 by lowering pitch by a semitone, i.e. 100 cents (R100d)</li> <li>lowered by 13 cents (T22d) to match the A440 tuning standard.</li> </ul> <p> </p> <p>The audio data for OrchideaSOL is not directly downloadable on Zenodo. Rather, it can be downloaded for free after registering to the Ircam forum. Please visit: https://forum.ircam.fr/</p> <p> </p> <p>Metadata File<br> -------------------</p> <p>The OrchideaSOL_metadata.csv file contains 13265 rows, one for each audio clip. It can be opened by a text editor or by a spreadsheet software application. It contains 13 columns:</p> <ol> <li>Path to the WAV file, in UNIX filesystem format. For Windows compatibility, replace the slashes ("/") by backslashes ("\"). Ex: "Strings/Violin+sordina/tremolo/Vn+S-trem-A4-mf-4c-T13d_R200d.wav"</li> <li>Fold ID. Either equal to 0, 1, 2, 3, or 4.</li> <li>Family. Ex: "Brass"</li> <li>Instrument abbreviation. Ex: "BTb"</li> <li>Instrument name in full. Ex: "Bass Tuba"</li> <li>Technique abbreviation.</li> <li>Technique name in full.</li> <li>Pitch. Ex: "A#1"</li> <li>Pitch ID in MIDI format. Ex: 34. Integer in the range 0-127.</li> <li>Dynamics. Ex: "ff".</li> <li>Dynamics ID. Integer. pp maps to 0 and ff maps to 4. The higher, the louder.</li> <li>Instance ID. Integer in the range 0-4</li> <li>String ID. Equal to 1, 2, 3, 4, or empty if not applicable.</li> <li>"Needed digital retuning". TRUE if the file has been pitch-shifted with digital audio effects; FALSE otherwise.</li> </ol> <p> </p> <p>Conditions of Use<br> ------------------------</p> <p>OrchideaSOL was created in 2020 by Carmine-Emanuele Cella, Daniele Ghisi, Vincent Lostanlen, Fabien Lévy, Joshua Fineberg, and Yan Maresz.</p> <p>OrchideaSOL is a derivative of SOL. We wish to thank Hugues Vinet, Greg Beller, and all coordinators of the Ircam Forum for their authorization to upload the metadata of OrchideaSOL to Zenodo.</p> <p>The audio samples in OrchideaSOL are offered free of charge under the Ircam Forum License. Please visit: https://forum.ircam.fr/legal/contrat-de-licence-forum-ircam/</p> <p>The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the authors are not liable for, and expressly exclude all liability for, loss or damage however and whenever caused to anyone by any use of the OrchideaSOL dataset or any part of it.</p> <p> </p> <p>Versions<br> -----------<br> 1.0 was released on February 24th, 2020.<br> 2.0 was released on April 4th, 2020. It fixes a bug in the instance IDs of oboe sounds in the "blow without reed" technique.</p> <p> </p> <p>Feedback<br> -------------</p> <p>Please help us improve OrchideaSOL by sending your feedback to:<br> carmine.cella@berkeley.edu</p> <p>For issues regarding the metadata encoding, the five-fold split, or the OrchideaSOL module in mirdata, please write to:<br> vincent.lostanlen@nyu.edu</p> <p>In case of a problem, please include as many details as possible.</p>
GUITAR-FX-DIST: A Dataset of Processed Guitar Recordings for Music Research - (Poly Continuous)
<p><strong>GUITAR-FX-DIST</strong> is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects. It was developed for research in guitar effects detection, classification and parameters estimation. The dataset is also useful for research on automatic music transcription, intelligent music production, signal processing or effects modelling. It contains both unprocessed and processed recordings.</p> <p>The dataset is split into 4 sub-datasets: Mono Continuous, Mono Discrete, Poly Continuous, Poly Discrete</p> <p> </p> <p><strong>Authors:</strong></p> <p>Marco Comunità - <a href="http://c4dm.eecs.qmul.ac.uk/">Centre for Digital Music</a>, Queen Mary University of London</p> <p> </p> <p><strong>Reference:</strong></p> <p>If you make use of GUITAR-FX-DIST, please cite the following publication:</p> <pre><code>@article{comunità2021guitar, title={Guitar Effects Recognition and Parameter Estimation with Convolutional Neural Networks}, author={Comunità, Marco and Stowell, Dan and Reiss, Joshua D.}, journal={Journal of the Audio Engineering Society}, year={2021}, volume={69}, number={7/8}, pages={594-604}, doi={}, month={July} }</code></pre> <p> </p> <p><strong>Dataset Snapshot:</strong></p> <ul> <li><strong>Size:</strong> ~550k samples (~305 hours) + 550k mel spectrograms</li> <li><strong>Audio Format:</strong> WAV - 44.1kHz, 16bit, mono, -6dBFS</li> <li><strong>Mel-Spectrogram Format:</strong> NPY - 128 frequency bands, sample rate 22050Hz, window length 1024, hop size 512,</li> <li><strong>Effects:</strong> 14 between overdrive, distortion and fuzz</li> <li><strong>Unprocessed recordings</strong> <ul> <li>624 monophonic notes</li> <li>420 polyphonic (2, 3 and 4 notes intervals and chords)</li> <li>2 guitars, with up to 2 pick-up settings and up to 3 plucking styles (finger pluck - hard, finger pluck - soft, pick) <ul> <li>Schecter Diamond C-1 Classic</li> <li>Chester Stratocaster</li> </ul> </li> </ul> </li> <li><strong>Samples length:</strong> 2 sec</li> </ul> <p> </p> <p><strong>Unprocessed Recordings:</strong></p> <p>The original (unprocessed) recordings are from the <a href="https://www.idmt.fraunhofer.de/en/business_units/m2d/smt/audio_effects.html">IDMT-SMT-Audio-Effects</a> dataset.</p> <p>For details please refer to the website and the accompagning publication:</p> <p><em>Stein, Michael; Abeßer, Jakob; Dittmar, Christian; Schuller, Gerald: Automatic Detection of Audio Effects in Guitar and Bass Recordings. Proceedings of the AES 128th Convention, 2010.</em></p> <p> </p> <p><strong>Processed Recordings:</strong></p> <p>The processed recordings are divided into 4 sub-datasets which are named depending on the unprocessed recordings used (monophonic or polyphonic) and on the settings' values (discrete or continuous).</p> <p>The sub-datasets are called: Mono Discrete, Poly Discrete, Mono Continuous, Poly Continuous</p> <p>Mono Discrete and Poly Discrete use a discrete set of combinations selected as the most common and representative settings a person might use (see README file for details).</p> <p>For Mono Continuous and Poly Continuous both unprocessed samples as well as settings’ values are drawn from a uniform distribution (10000 samples for each effect).</p> <p>Samples:</p> <ul> <li>Mono Discrete: ~160k</li> <li>Poly Discrete: ~110k</li> <li>Mono Continuous: 140k</li> <li>Poly Continuous: 140k</li> </ul> <p> </p> <p><strong>Scripts:</strong></p> <p>The dataset includes the MATLAB scripts used to generate the samples</p>
GUITAR-FX-DIST: A Dataset of Processed Guitar Recordings for Music Research - (Mono Discrete)
<p><strong>GUITAR-FX-DIST</strong> is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects. It was developed for research in guitar effects detection, classification and parameters estimation. The dataset is also useful for research on automatic music transcription, intelligent music production, signal processing or effects modelling. It contains both unprocessed and processed recordings.</p> <p>The dataset is split into 4 sub-datasets: Mono Continuous, Mono Discrete, Poly Continuous, Poly Discrete</p> <p> </p> <p><strong>Authors:</strong></p> <p>Marco Comunità - <a href="http://c4dm.eecs.qmul.ac.uk/">Centre for Digital Music</a>, Queen Mary University of London</p> <p> </p> <p><strong>Reference:</strong></p> <p>If you make use of GUITAR-FX-DIST, please cite the following publication:</p> <pre><code>@article{comunità2021guitar, title={Guitar Effects Recognition and Parameter Estimation with Convolutional Neural Networks}, author={Comunità, Marco and Stowell, Dan and Reiss, Joshua D.}, journal={Journal of the Audio Engineering Society}, year={2021}, volume={69}, number={7/8}, pages={594-604}, doi={}, month={July} }</code></pre> <p> </p> <p><strong>Dataset Snapshot:</strong></p> <ul> <li><strong>Size:</strong> ~550k samples (~305 hours) + 550k mel spectrograms</li> <li><strong>Audio Format:</strong> WAV - 44.1kHz, 16bit, mono, -6dBFS</li> <li><strong>Mel-Spectrogram Format:</strong> NPY - 128 frequency bands, sample rate 22050Hz, window length 1024, hop size 512,</li> <li><strong>Effects:</strong> 14 between overdrive, distortion and fuzz</li> <li><strong>Unprocessed recordings</strong> <ul> <li>624 monophonic notes</li> <li>420 polyphonic (2, 3 and 4 notes intervals and chords)</li> <li>2 guitars, with up to 2 pick-up settings and up to 3 plucking styles (finger pluck - hard, finger pluck - soft, pick) <ul> <li>Schecter Diamond C-1 Classic</li> <li>Chester Stratocaster</li> </ul> </li> </ul> </li> <li><strong>Samples length:</strong> 2 sec</li> </ul> <p> </p> <p><strong>Unprocessed Recordings:</strong></p> <p>The original (unprocessed) recordings are from the <a href="https://www.idmt.fraunhofer.de/en/business_units/m2d/smt/audio_effects.html">IDMT-SMT-Audio-Effects</a> dataset.</p> <p>For details please refer to the website and the accompagning publication:</p> <p><em>Stein, Michael; Abeßer, Jakob; Dittmar, Christian; Schuller, Gerald: Automatic Detection of Audio Effects in Guitar and Bass Recordings. Proceedings of the AES 128th Convention, 2010.</em></p> <p> </p> <p><strong>Processed Recordings:</strong></p> <p>The processed recordings are divided into 4 sub-datasets which are named depending on the unprocessed recordings used (monophonic or polyphonic) and on the settings' values (discrete or continuous).</p> <p>The sub-datasets are called: Mono Discrete, Poly Discrete, Mono Continuous, Poly Continuous</p> <p>Mono Discrete and Poly Discrete use a discrete set of combinations selected as the most common and representative settings a person might use (see README file for details).</p> <p>For Mono Continuous and Poly Continuous both unprocessed samples as well as settings’ values are drawn from a uniform distribution (10000 samples for each effect).</p> <p>Samples:</p> <ul> <li>Mono Discrete: ~160k</li> <li>Poly Discrete: ~110k</li> <li>Mono Continuous: 140k</li> <li>Poly Continuous: 140k</li> </ul> <p> </p> <p><strong>Scripts:</strong></p> <p>The dataset includes the MATLAB scripts used to generate the samples</p>
GUITAR-FX-DIST: A Dataset of Processed Guitar Recordings for Music Research - (Poly Discrete)
<p><strong>GUITAR-FX-DIST</strong> is a dataset of electric guitar recordings processed with overdrive, distortion and fuzz audio effects. It was developed for research in guitar effects detection, classification and parameters estimation. The dataset is also useful for research on automatic music transcription, intelligent music production, signal processing or effects modelling. It contains both unprocessed and processed recordings.</p> <p>The dataset is split into 4 sub-datasets: Mono Continuous, Mono Discrete, Poly Continuous, Poly Discrete</p> <p> </p> <p><strong>Authors:</strong></p> <p>Marco Comunità - <a href="http://c4dm.eecs.qmul.ac.uk/">Centre for Digital Music</a>, Queen Mary University of London</p> <p> </p> <p><strong>Reference:</strong></p> <p>If you make use of GUITAR-FX-DIST, please cite the following publication:</p> <pre><code>@article{comunità2021guitar, title={Guitar Effects Recognition and Parameter Estimation with Convolutional Neural Networks}, author={Comunità, Marco and Stowell, Dan and Reiss, Joshua D.}, journal={Journal of the Audio Engineering Society}, year={2021}, volume={69}, number={7/8}, pages={594-604}, doi={}, month={July} }</code></pre> <p> </p> <p><strong>Dataset Snapshot:</strong></p> <ul> <li><strong>Size:</strong> ~550k samples (~305 hours) + 550k mel spectrograms</li> <li><strong>Audio Format:</strong> WAV - 44.1kHz, 16bit, mono, -6dBFS</li> <li><strong>Mel-Spectrogram Format:</strong> NPY - 128 frequency bands, sample rate 22050Hz, window length 1024, hop size 512,</li> <li><strong>Effects:</strong> 14 between overdrive, distortion and fuzz</li> <li><strong>Unprocessed recordings</strong> <ul> <li>624 monophonic notes</li> <li>420 polyphonic (2, 3 and 4 notes intervals and chords)</li> <li>2 guitars, with up to 2 pick-up settings and up to 3 plucking styles (finger pluck - hard, finger pluck - soft, pick) <ul> <li>Schecter Diamond C-1 Classic</li> <li>Chester Stratocaster</li> </ul> </li> </ul> </li> <li><strong>Samples length:</strong> 2 sec</li> </ul> <p> </p> <p><strong>Unprocessed Recordings:</strong></p> <p>The original (unprocessed) recordings are from the <a href="https://www.idmt.fraunhofer.de/en/business_units/m2d/smt/audio_effects.html">IDMT-SMT-Audio-Effects</a> dataset.</p> <p>For details please refer to the website and the accompagning publication:</p> <p><em>Stein, Michael; Abeßer, Jakob; Dittmar, Christian; Schuller, Gerald: Automatic Detection of Audio Effects in Guitar and Bass Recordings. Proceedings of the AES 128th Convention, 2010.</em></p> <p> </p> <p><strong>Processed Recordings:</strong></p> <p>The processed recordings are divided into 4 sub-datasets which are named depending on the unprocessed recordings used (monophonic or polyphonic) and on the settings' values (discrete or continuous).</p> <p>The sub-datasets are called: Mono Discrete, Poly Discrete, Mono Continuous, Poly Continuous</p> <p>Mono Discrete and Poly Discrete use a discrete set of combinations selected as the most common and representative settings a person might use (see README file for details).</p> <p>For Mono Continuous and Poly Continuous both unprocessed samples as well as settings’ values are drawn from a uniform distribution (10000 samples for each effect).</p> <p>Samples:</p> <ul> <li>Mono Discrete: ~160k</li> <li>Poly Discrete: ~110k</li> <li>Mono Continuous: 140k</li> <li>Poly Continuous: 140k</li> </ul> <p> </p> <p><strong>Scripts:</strong></p> <p>The dataset includes the MATLAB scripts used to generate the samples</p>
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
<p>This upload contains the supplementary material for our <a href="https://arxiv.org/abs/2312.09207" target="_blank" rel="noopener">paper</a> presented at the <a href="https://mmm2024.org/" target="_blank" rel="noopener">MMM2024 conference</a>.</p> <h2>Dataset</h2> <p>The dataset contains rich text descriptions for music audio files collected from Wikipedia articles.</p> <p>The audio files are freely accessible and available for download through the URLs provided in the dataset.</p> <h3>Example</h3> <p>A few hand-picked, simplified examples of the dataset. </p> <table> <tbody> <tr> <td> <p><strong>file</strong></p> </td> <td> <p><strong>aspects</strong></p> </td> <td> <p><strong>sentences</strong></p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/7/7a/Bongo_sound.wav" target="_blank" rel="noopener"><strong>🔈 Bongo sound.wav</strong></a></p> </td> <td> <p>['bongoes', 'percussion instrument', 'cumbia', 'drums']</p> </td> <td> <p>['a loop of bongoes playing a cumbia beat at 99 bpm']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/4/46/Example_of_double_tracking_in_a_pop-rock_song_%283_guitar_tracks%29.ogg" target="_blank" rel="noopener"><strong>🔈 Example of double tracking in a pop-rock song (3 guitar tracks).ogg</strong></a></p> </td> <td> <p>['bass', 'rock', 'guitar music', 'guitar', 'pop', 'drums']</p> </td> <td> <p>['a pop-rock song']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/6/62/OriginalDixielandJassBand-JazzMeBlues.ogg" target="_blank" rel="noopener"><strong>🔈 OriginalDixielandJassBand-JazzMeBlues.ogg</strong></a></p> </td> <td> <p>['jazz standard', 'instrumental', 'jazz music', 'jazz']</p> </td> <td> <p>['Considered to be a jazz standard', 'is an jazz composition']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/5/58/Colin_Ross_-_Etherea.ogg" target="_blank" rel="noopener"><strong>🔈 Colin Ross - Etherea.ogg</strong></a></p> </td> <td> <p>['chirping birds', 'ambient percussion', 'new-age', 'flute', 'recorder', 'single instrument', 'woodwind']</p> </td> <td> <p>['features a single instrument with delayed echo, as well as ambient percussion and chirping birds', 'a new-age composition for recorder']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/8/8b/Belau_rekid_%28instrumental%29.oga" target="_blank" rel="noopener"><strong>🔈 Belau rekid (instrumental).oga</strong></a></p> </td> <td> <p>['instrumental', 'brass band']</p> </td> <td> <p>['an instrumental brass band performance']</p> </td> </tr> <tr> <td> <p><strong>...</strong></p> </td> <td> <p>...</p> </td> <td> <p>...</p> </td> </tr> </tbody> </table> <h3>Dataset structure</h3> <p>We provide three variants of the dataset in the <code>data</code> folder.</p> <p>All are described in the paper.</p> <ol> <li><code>all.csv</code> contains all the data we collected, without any filtering.</li> <li><code>filtered_sf.csv</code> contains the data obtained using the <em>self-filtering</em> method.</li> <li><code>filtered_mc.csv</code> contains the data obtained using the <em>MusicCaps</em> dataset method.</li> </ol> <h3>File structure</h3> <p>Each CSV file contains the following columns:</p> <ul> <li><code>file</code>: the name of the audio file</li> <li><code>pageid</code>: the ID of the Wikipedia article where the text was collected from</li> <li><code>aspects</code>: the short-form (tag) description texts collected from the Wikipedia articles</li> <li><code>sentences</code>: the long-form (caption) description texts collected from the Wikipedia articles</li> <li><code>audio_url</code>: the URL of the audio file</li> <li><code>url</code>: the URL of the Wikipedia article where the text was collected from</li> </ul> <h3>Citation</h3> <div> <p>If you use this dataset in your research, please cite the following paper:</p> <div> <pre><code>@inproceedings{wikimute,</code><br><code> title = {WikiMuTe: {A} Web-Sourced Dataset of Semantic Descriptions for Music Audio},</code><br><code> author = {Weck, Benno and Kirchhoff, Holger and Grosche, Peter and Serra, Xavier},</code><br><code> booktitle = "MultiMedia Modeling",</code><br><code> year = "2024",</code><br><code> publisher = "Springer Nature Switzerland",</code><br><code> address = "Cham",</code><br><code> pages = "42--56",</code><br><code> doi = {10.1007/978-3-031-56435-2_4},</code><br><code> url = {https://doi.org/10.1007/978-3-031-56435-2_4},</code><br><code>}</code></pre> </div> </div> <h3>License</h3> <p>The data is available under the <a href="https://creativecommons.org/licenses/by-sa/3.0/" target="_blank" rel="noopener">Creative Commons Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) license</a>.</p> <p>Each entry in the dataset contains a URL linking to the article, where the text data was collected from.</p>
N20EMv2 dataset for automatic music transcription from multimodal singing
<p>N20EMv2 dataset for multimodal automatic music transcription from multimodal singing, presented in our TOMM 2024 paper, Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing. This dataset contains recordings of two modalities: audio and video. </p> <p>Our paper is available at: https://dl.acm.org/doi/10.1145/3651310.</p> <p>Code is available at: https://github.com/guxm2021/SVT_SpeechBrain</p> <p>Please cite our work as:</p> <pre>@article{gu2024automatic, title={Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing}, author={Gu, Xiangming and Ou, Longshen and Zeng, Wei and Zhang, Jianan and Wong, Nicholas and Wang, Ye}, journal={ACM Transactions on Multimedia Computing, Communications and Applications}, publisher={ACM New York, NY}, year={2024} }</pre>
Music Data Sharing Platform for Computational Musicology Research (CCMUSIC DATASET)
<p>This platform is a multi-functional music data sharing platform for Computational Musicology research. It contains many music datas such as the sound information of Chinese traditional musical instruments and the labeling information of Chinese pop music, which is available for free use by computational musicology researchers.</p> <p>This platform is also a large-scale music data sharing platform specially used for Computational Musicology research in China, including 3 music databases: Chinese Traditional Instrument Sound Database (CTIS), Midi-wav Bi-directional Database of Pop Music and Multi-functional Music Database for MIR Research (CCMusic). All 3 databases are available for free use by computational musicology researchers. For the contents contained in the database, we will provide audio files recorded by the professional team of the conservatory of music, as well as corresponding labelled files, which have no commodity copyright problem and facilitate large-scale promotion. We hope that this music data sharing platform can meet the one-stop data needs of users and contribute to the research in the field of Computational Musicology.</p> <p> </p> <p>If you want to know more information or obtain complete files, please go to the official website of this platform:</p> <p><a href="https://ccmusic-database.github.io/en/">Music Data Sharing Platform for Academic Research</a></p> <p> </p> <ul> <li> <p><strong>Chinese Traditional Instrument Sound Database (CTIS)</strong></p> </li> </ul> <p>This database is developed by Prof. Han Baoqiang's team for many years, which collects sound information about Chinese traditional musical instruments. The database includes 287 Chinese national musical instruments, including traditional musical instruments, improved musical instruments and ethnic minority musical instruments.</p> <ul> <li> <p><strong>Multi-functional Music Database for MIR Research</strong></p> </li> </ul> <p>This database collects sound materials of pop music, folk music and hundreds of national musical instruments, and makes comprehensive annotation to form a multi-purpose music database for MIR researchers.</p> <ul> <li><strong>Midi-wav Bi-directional Database of Pop Music</strong></li> </ul> <p>This database contains hundreds of Chinese pop songs, and each song contains the corresponding midi-audio-lyric information. Among them, recording the vocal part and accompaniment part of audio independently is helpful to study the MIR task under the ideal situation. In addition, the information of singing techniques consistent with vocal part (such as breath sound, falsetto, breathing, vibrato, mute, slide, etc.) is marked in MuseScore, which constitutes a Midi-Wav bi-direction corresponding pop music database.</p>
Harmonized Cultural Access & Participation Dataset for Music
<p>Changes since the last version: in the .csv export there was a naming problem.</p> <p>- `visit_concert`: This is a standard CAP variables about visiting frequencies, in numeric form. <br> - `fct_visit_concert`: This is a standard CAP variables about visiting frequencies, in categorical form. <br> - `is_visit_concert`: binary variable, 0 if the person had not visited concerts in the previous 12 months.<br> - `artistic_activity_played_music`: A variable of the frequency of playing music as an amateur or professional practice, in some surveys we have only a binary variable (played in the last 12 months or not) in other we have frequencies. We will convert this into a binary variable. <br> - `fct_artistic_activity_played_music`: The `artistic_activity_played_music` in categorical representation.<br> - `artistic_activity_sung`: A variable of the frequency of singing as an amateur or professional practice, like played_muisc. Because of the liturgical use of singing, and the differences of religious practices among countries and gender, this is a significantly different variable from played_music.<br> - `fct_artistic_activity_sung`: The `artistic_activity_sung` variable in categorical representation.<br> - `age_exact`: The respondent’s age as an integer number. <br> - `country_code`: an ISO country code<br> - `geo`: an ISO code that separates Germany to the former East and West Germany, and the United Kingdom to Great Britain and Northern Ireland, and Cyprus to Cyprus and the Turiksh Cypriot community.[we may leave Turkish Cyprus out for practical reasons.]<br> - `age_education`: This is a harmonized education proxy. Because we work with the data of more than 30 countries, education levels are difficult to harmonize, and we use the Eurobarometer standard proxy, age of leaving education. It is a specially coded variable, and we will re-code them into two variables, `age_education` and `is_student`. <br> - `is_student`: is a dummy variable for the special coding in age_education for “still studying”, i.e. the person does not have yet a school leaving age. It would be tempting to impute `age` in this case to `age_education`, but we will show why this is not a good strategy.<br> - `w`, `w1`: Post-stratification weights for the 15+ years old population of each country. Use `w1` for averages of `geo` entities treating Northern Ireland, Great Britain, the United Kingdom, the former GDR, the former West Germany, and Germany as geographical areas. Use `w` when treating the United Kingdom and Germany as one territory.<br> - `wex`: Projected weight variable. For weighted average values, use `w`, `w1`, for projections on the population size, i.e., use with sums, use `wex`.<br> - `id`: The identifier of the original survey.<br> - `rowid``: A new unique identifier that is unique in all harmonized surveys, i.e., remains unique in the harmonized dataset.</p>
Hindustani Music Alankar Dataset
<p>This dataset is a collection of CSV files that include the time interval annotations of Alankars (ornamentations) present in the compositions of Hindustani classical music. The selected music files comprise the files from the Dunya Corpus [1].</p> <p>The annotations are for the Alankar segments in the truncated music files (MBID included). In total, there are 361 Alankar segments. There are 17 recordings in the dataset, with a total duration of approximately 25 minutes. Annotations for the Alankar segments include middle and end sections.</p> <p>References:<br> [1]. Gulati, S., Serrà, J., Ganguli, K. K., & Serra, X. (2014). Landmark detection in Hindustani music melodies. In Proceedings of the International Computer Music Conference / Sound and Music Computing Conference (ICMC-SMC), pp. 1062- 1068. Athens, Greece.<br> _________________________________________________________________________________________________________<br> This project was funded under the grant number: ECR/2018/000204 by the Science & Engineering Research Board (SERB).</p>
Hindustani Music Nyas Dataset
<p>This dataset is a collection of CSV files that include the time interval annotations of Nyas swara present in the compositions of Hindustani classical music. The selected music files comprise the files from the Nyas dataset [1] of the Dunya Corpus. The annotations are for the Nyas segments in the truncated music files (MBID included). In total, there are 2269 Nyaas segments. Overall, there are 67 recordings in the dataset, with a total duration of approx 100 minutes. Annotations for the Nyas segments comprise Alap sections, middle sections (medium tempo), and end sections (fast tempo).</p> <p>References:<br> [1]. Gulati, S., Serrà, J., Ganguli, K. K., & Serra, X. (2014). Landmark detection in Hindustani music melodies. In Proceedings of the International Computer Music Conference / Sound and Music Computing Conference (ICMC-SMC), pp. 1062- 1068. Athens, Greece.<br> _________________________________________________________________________________________________________<br> This project was funded under the grant number: ECR/2018/000204 by the Science & Engineering Research Board (SERB).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.