Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “piano music”
An Annotated Corpus of Tonal Piano Music from the Long 19th Century
<p>This corpus has been created within the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and employs the <a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p><strong>Version 1</strong> has been released for submitting it as part of the data report <code>Hentschel, J., Rammos, Y., Neuwirth, M., Rohrmeier, M. (forthcoming). An Annotated Corpus of Tonal Piano Music from the Long 19th Century</code> that accompanies nine corpora grouped under the DOI <a href="https://doi.org/10.5281/zenodo.7483349">10.5281/zenodo.7483349</a>.</p> <p><strong>Version 1.1</strong> comes with a complete set of metadata and score headers. Among more accurate composition dates, the metadata now include URIs that identify the compositions in terms of the <a href="https://viaf.org/">Virtual International Authority File (VIAF)</a>, <a href="https://www.wikidata.org/">Wikidata</a>, <a href="https://imslp.org/">IMSLP</a> and <a href="https://musicbrainz.org/">MusicBrainz</a>. The data has been re-extracted from the scores using <a href="https://pypi.org/project/ms3/">ms3 1.1.1</a>.</p> <p>The publication covers the following corpora (the DOI links always point at the latest version respectively):</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7473560">Ludwig van Beethoven - Piano Sonatas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473566">Frédéric Chopin - Mazurkas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473568">Claude Debussy - Suite Bergamasque</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473576">Antonín Dvořák - Silhouettes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473580">Franz Liszt - Années de Pèlerinage</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473528">Nikolai Medtner - Tales</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473582">Robert Schumann - Kinderszenen</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473586">Pyotr Tchaikovsky - The Seasons</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473578">Edvard Grieg - Lyric Pieces</a></li> </ul> <p> </p>
Saarland Music Data: MIDI-Audio Piano Music
<p>This is an improved version of the dataset originally referred to as <strong>SMD MIDI-Audio Piano Music.</strong> For more details, please visit the website: <a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi">https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi</a></p> <p>Saarland Music Data provides audio recordings along with perfectly synchronized MIDI files for various piano pieces. The pieces were performed by students of the <a href="http://www.hfm.saarland.de">Hochschule für Musik Saar</a> on a hybrid acoustic/digital piano <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a>. The Disklavier allows for capturing key and pedal movements of the piano while playing. This information, which can be stored in a MIDI file, yields an accurate annotation of the corresponding audio recording in form of a symbolic description of all played musical note events. The SMD MIDI-Audio pairs constitute a valuable dataset for various music analysis tasks such as music transcription, performance analysis, music synchronization, audio alignment, or source separation. </p> <p>All performances were recorded in the studios of the <a href="http://www.hfm.saarland.de">Hochschule für Musik Saar</a>, played by students of piano classes of different levels, on a <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a> model <a href="http://www.yamaha.com/yamahavgn/CDA/ContentDetail/ModelSeriesDetail.html?CNTID=556850&CNTYP=PRODUCT">DCFIIISM4PRO</a>. Using two cardioid-condenser microphones fixed over the resonating body of the piano, all performances were directly recorded into Steinberg Cubase 4. Except for trimming the beginnings and ends of the recordings, no further post-processing (filters, effects) was applied to the musical material. From each Cubase project, an audio file (44.1 kHz, stereo) as well as a synchronized standard MIDI file (SMF) were exported. Besides these files, we also provide the audio files as WAV (22.05 kHz, mono) and the MIDI files encoded as CSV files and as WAV files (22.05 kHz, mono) rendered using the Software synthesizer <a href="https://www.fluidsynth.org/">FluidSynth</a>.</p> <p>SMD MIDI-Audio Piano Music (V1) contains the following data:</p> <ul> <li>wav_44100_stereo: Audio file (44.1 kHz, stereo)</li> <li>wav_22050_mono: Audio file (22.05 kHz, mono)</li> <li>midi: MIDI file</li> <li>csv: Export of note events from MIDI file into CSV format</li> <li>midi_wav_22050_mono: MIDI file rendered as audio file (22.05 kHz, mono)</li> </ul> <p>If you publish results obtained using this dataset, please cite:</p> <p>Meinard Müller, Verena Konz, Wolfgang Bogler, Vlora Arifi-Müller: Saarland Music Data (SMD). In Late-Breaking and Demo Session of the 12th International Conference on Music Information Retrieval (ISMIR), 2011. [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/2011_MuellerKonzBoglerArifi_SaarlandMusicData_ISMIR-LateBreaking.pdf">pdf</a>] [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/bibtex.html">bib</a>]</p>
JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting
<p>Release of the MIDI data pairs that constitute the JAZZVAR dataset. See below for the abstract of the publication.</p> <p>The data is also available transposed to C/Am and subsequently, to all keys, with accompanying metadata.</p> <p>Abstract:</p> <p>Jazz pianists often uniquely interpret jazz standards. Passages from these interpretations can be viewed as sections of variation. We manually extracted such variations from solo jazz piano performances. The JAZZVAR dataset is a collection of 502 pairs of Variation and Original MIDI segments. Each Variation in the dataset is accompanied by a corresponding Original segment containing the melody and chords from the original jazz standard. Our approach differs from many existing jazz datasets in the music information retrieval (MIR) community, which often focus on improvisation sections within jazz performances. In this paper, we outline the curation process for obtaining and sorting the repertoire, the pipeline for creating the Original and Variation pairs, and our analysis of the dataset. We also introduce a new generative music task, Music Overpainting, and present a baseline Transformer model trained on the JAZZVAR dataset for this task. Other potential applications of our dataset include expressive performance analysis and performer identification.</p>
EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation
<p>EMOPIA (pronounced ‘yee-mò-pi-uh’) dataset is a shared multi-modal (audio and MIDI) database focusing on perceived emotion in <strong>pop piano music</strong>, to facilitate research on various tasks related to music emotion. The dataset contains <strong>1,087</strong> music clips from 387 songs and <strong>clip-level</strong> emotion labels annotated by four dedicated annotators. </p> <p>For more detailed information about the dataset, please refer to our paper: <a href="https://arxiv.org/abs/2108.01374"><strong>EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation</strong></a>. </p> <p><strong>File Description</strong></p> <ul> <li><em><strong>midis/</strong></em>: midi clips transcribed using GiantMIDI. <ul> <li>Filename `Q1_xxxxxxx_2.mp3`: Q1 means this clip belongs to Q1 on the V-A space; xxxxxxx is the song ID on YouTube, and the `2` means this clip is the 2nd clip taken from the full song.</li> </ul> </li> <li><em><strong>metadata/</strong></em>: metadata from YouTube. (Got when crawling)</li> <li> <p><em><strong>songs_lists/</strong></em>: YouTube URLs of songs.</p> </li> <li> <p><em><strong>tagging_lists/</strong></em>: raw tagging result for each sample.</p> </li> <li> <p><em><strong>label.csv</strong></em>: metadata that records filename, 4Q label, and annotator.</p> </li> <li> <p><em><strong>metadata_by_song.csv</strong></em>: list all the clips by the song. Can be used to create the train/val/test splits to avoid the same song appear in both train and test.</p> </li> <li> <p><em><strong>scripts/prepare_split.ipynb:</strong></em> the script to create train/val/test splits and save them to csv files.</p> </li> </ul> <p>------</p> <p><strong>2.2 Update</strong></p> <ul> <li>Add tagging files in <em><strong>tagging_lists/</strong></em> that are missing in the previous version.</li> <li>Add <em><strong>timestamps.json</strong></em> for easier usage. It records all the timestamps in dict format. You can see <em><strong>scripts/load_timestamp.ipynb</strong></em> for the format example.</li> <li>Add <em><strong>scripts/timestamp2clip.py</strong></em>: After the raw audio are crawled and put in <em><strong>audios/raw</strong></em>, you can use this script to get audio clips. The script will read <em><strong>timestamps.json</strong></em> and use the timestamp to extract clips. The clips will be saved to <em><strong>audios/seg</strong> </em>folder.</li> <li>remove 7 midi files that were added by mistake, and also corrected the number in <em><strong>metadata_by_song.csv</strong></em>.</li> </ul> <p> </p> <p><strong>2.1 Update</strong></p> <p>Add one file and one folder:</p> <ul> <li><em><strong>key_mode_tempo.csv</strong></em>: key, mode, and tempo information extracted from files.</li> <li><strong><em>CP_events/</em></strong>: CP events used in our paper. Extracted using this <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/representations/uncond/cp/corpus2events.py">script</a>, and add the emotion event to the front.</li> </ul> <p>Modify one folder:</p> <ul> <li>The <strong><em>REMI_events/</em></strong> files in version 2.0 contain some information that is not related to the paper, so remove it.</li> </ul> <p> </p> <p><strong>2.0 Update</strong></p> <p>Add two new folders:</p> <ul> <li><strong><em>corpus/</em></strong>: processed data that following <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/Dataset.md">the preprocessing flow</a>. (Please notice that although we have <code>1078</code> clips in our dataset, we lost some clips during steps 1~4 of the flow, so the final number of clips in this <strong><code>corpus</code></strong> is <code>1052</code>, and that's the number we used for training the generative model.)</li> <li><strong><em>REMI_events/</em></strong>: REMI event for each midi file. They are generated using this <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/representations/uncond/remi/corpus2events.py">script</a>.</li> </ul> <p>-------- </p> <p> </p> <p> </p> <p> </p> <p><strong>Cite this dataset</strong></p> <pre><code>@inproceedings{{EMOPIA}, author = {Hung, Hsiao-Tzu and Ching, Joann and Doh, Seungheon and Kim, Nabin and Nam, Juhan and Yang, Yi-Hsuan}, title = {{MOPIA}: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation}, booktitle = {Proc. Int. Society for Music Information Retrieval Conf.}, year = {2021} }</code></pre>
jazznet: A Dataset of Fundamental Piano Patterns for Music Audio Machine Learning Research
<p>Jazznet is a dataset of piano patterns for music audio machine learning research. The dataset comprises chords, arpeggios, scales, and chord progressions in all keys of an 88-key piano and in all the inversions, for a total of 162520 labeled piano patterns, resulting in 95GB of data and more than 26k hours of audio. The data is also accompanied by Python scripts to enable the easy generation of new piano patterns beyond those present in the dataset. The data is broken down into small, medium, and large subsets, comprising 21516, 30328, and 52360 patterns, respectively (with all the chords, arpeggios, and scales being present in all subsets). </p> <p>The GitHub page of the dataset, containing details of the dataset and scripts for generating new data is https://github.com/tosiron/jazznet.</p>
Towards Musically Informed Evaluation of Piano Transcription Models
<p>We provide here the evaluation set employed in our experiments described in "Towards Musically Informed Evaluation of Piano Transcription Models", published in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), San Francisco, United States, 2024.</p> <p>In this work, we demonstrate musically informed piano transcription metrics using transcriptions derived from three state-of-the-art transcriptions ([1], [2], [3]). To this end, we create an evaluation set that includes (1) a subset of the original audio recordings from the MAESTRO dataset [1], (2) a re-recorded version that subset, and (3) a perturbed version of recordings from both (1) and (2). In this data repository, we provide components (2) and (3).</p> <p>[1] Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations, 2019. </p> <p>[2] Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Wan, and Yuxan Wang, “High-resolution piano transcription with pedals by regressing onset and offset times,” IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 29, pp. 3707–3717, 2021. </p> <p>[3] Curtis Hawthorne, Ian Simon, Rigel Swavely, Ethan Manilow, and Jesse Engel. “Sequence-to-sequence piano transcription with transformers,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021.</p>
Field Report on 3D Audio Capture of Solo Piano for Classical Music Productions - Audio Files
<p>This online repository contains audio files related to research on the side surround loudspeakers on 3D audio recordings of solo piano in classical music productions, undertaken by Emre Ekici, Will Howie, and Toru Kamekawa at Tokyo University of the Arts, March 2023. Please find the guidelines for usage below: </p> <p>The archive (3DPIANO_TRACKS.zip) contains 23 channels of mono audio tracks for ITU 4+7+0 solo piano recording, played by Yamaha Disklavier. Files are recorded at 96 kHz / 24-bit.</p> <p>File naming convention:<br> ##_ProjectTitle_Loudspeaker_Position/Variable_Position/Variable_Option (if applicable).</p> <p>Anyone is free to download and listen to these files for reference.</p> <p>If you wish to use these files for your research, please contact Emre Ekici (mrekici@outlook.com) or Will Howie (wghowie@gmail.com) for permission.</p>
Stimuli and Results for "Investigating the Perceptual Validity of Evaluation Metrics for Automatic Piano Music Transcription"
<p>This contains the stimuli and the participants data for the listening tests presented in the paper:</p> <p>Adrien Ycart, Lele Liu, Emmanouil Benetos, Marcus T. Pearce. "Investigating the Perceptual Validity of Evaluation Metrics for Automatic Piano Music Transcription". <em>Transactions of the International Society for Music Information Retrieval</em>, 3(1):68-81, 2020 .</p> <p>More precisely, it contains:</p> <ul> <li>MAPS_midi_cut.zip: The MIDI files used to create the stimuli </li> <li>cut_points_seconds.zip: The points in seconds at which the MAPS music pieces were cut to make the stimuli. These correspond to manually-selected 5 to 10 seconds chunks, roughly corresponding to musical phrases.</li> <li>listening_test_results.zip: The data gathered during the listening test: <ul> <li>user_data.csv contains data about participants</li> <li>answers_data.csv contains the answers given by all participants</li> <li>comments.txt contains the comments left by the participants.</li> </ul> </li> </ul> <p>For any enquiries, please contact Adrien Ycart (a.ycart@qmul.ac.uk) or Emmanouil Benetos (emmanouil.benetos@qmul.ac.uk).</p> <p> </p>
Dataset for Evaluating Sustain-Pedal Detection from Polyphonic Piano Music
<p>To evaluate methods of sustain-pedal detection from polyphonic piano music, we built a dataset consisting of ten well-known passages of Chopin's music. Ground-truth annotations of this dataset represent sustain-pedal on/off states at every 0.1 second. This annotation was based on the sustain-pedal movement tracked by a dedicated measurement system.</p> <p>Music scores of the ten passages were saved in PDF files. They were performed by a pianist using a Yamaha baby grand piano situated in the studios at Queen Mary University of London. The audio were recorded at 44.1 kHz and 24 bits using the spaced-pair stereo microphone technique. A pair of Earthworks QTC40 omnidirectional condenser microphones was positioned about 50 cm above the strings. </p> <p>We have developed a transfer learning method such that the on/off state of the sustain pedal can be detected at every 0.1 second. The ground-truth annotation and our detection results were saved in <em>transfer-learning-y_segment.npz</em>, which can be loaded using <em>numpy.load </em>in Python. The key for passage name, ground-truth annotation and our detection results is <em>filename_record</em>, <em>y_true</em> and <em>y_pred</em>, respectively.</p>
Recordings of classical music (voice with piano, piano solo) with smartphones and professional audio equipment
<p>To investigate differences of performance rating and perception of classical music demo videos, two demos were recorded with six devices (four Smartphones with main cam video function, one field recorder, one professional setup). Video and Audio were separated for each file, only the audio files were used in the study and are thus presented here. The music of the voice demo is from the genres of romantic lied and romantic and modern opera, the piano solo music was composed in the 20th century.</p> <p>The documentation follows the recommendations of the German Society for Acoustics (DEGA), the information is as following:</p> <p>A 3D model of the recording situation.<br>The audio files, normalised (as used in the corresponding study) and with original level.<br>Geometric measurements of room dimensions.<br>Pictures of the recording.<br>A list of the equipment and recording system.<br>Render statistics (provided by the DAW used, Reaper).<br>Scores of two of the three pieces.<br>A Takelist.</p>
Musical Performance Critique Documents for Piano
<p>CROCUS (CRitique dOCUmentS): Dataset of Musical Performance Critique Documents (in Japanese) CC BY-NC-ND 4.0</p> <p>This open dataset contains 23 piano performances and 144 critiques of those performances.</p> <p>For more information, please visit the project page below.<br><a href="https://masaki-cb.github.io/crocus/">https://masaki-cb.github.io/crocus/</a></p> <p>Please access CrestMusePEDB (<a href="http://www.crestmuse.jp/pedb/" target="_blank" rel="noopener">http://www.crestmuse.jp/pedb/</a>) for performance recordings and scores.</p> <p> </p> <p>n01 F. Chopin “Tristesse”, Op. 10-3<br>n08 F. Chopin “24 Préludes”, Op. 28-7<br>n17 J. S. Bach Invention No. 1 in C major, BWV 772<br>n22 J. S. Bach Invention No. 15 in B minor, BWV 786<br>n26 L. v. Beethoven Sonata No. 8 in A flat major, Op. 13, 2nd Mov.<br>n28 L. v. Beethoven Sonata No. 8 in C minor, Op. 13, 3rd Mov.<br>n31 R. Schumann “Traumerai”, Kinderszenen No.7, Op. 15<br>n32 W.A. Mozart Sonata No.32 in A major, KV. 331<br>n46 C. Debussy La Fille aux Cheveux de Lin</p> <p>n48 C. Debussy Rêverie</p>
Art of Chokin Vintage Piano Music Box
Another vintage music box for our collection! This piano shaped music box was handmade in Japan and features a sweet sercet note on the bottom. Interested in a different unique music box model? Check out this one! https://skfb.ly/op6st Source: Objaverse 1.0 / Sketchfab
Listening Samples for Paper "Score to Audio: Integrating Systems to Transform Music Scores into Expressive Piano Audio"
<p>Listening Samples for Paper "Score to Audio: Integrating Systems to Transform Music Scores into Expressive Piano Audio".</p> <p>Please check the README.md for more details.</p>
Investigating the Effect of Simulated Live Piano Music on Preoperative Cancer Patients, Health Care Providers and Hospital Volunteers Using Validated Questionnaires and Proteomic Analysis
ClinicalTrials.gov study NCT03239587. IPD Sharing: Not stated. Countries: 0. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.