Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

76

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

76 results for “piano”

Learn how ShareScore rates datasets ↗
zenodo52/100

BiVib - Audio-Tactile Piano Sample Library

<p><strong>BiVib</strong> is an extensive piano sample library consisting of <strong>bi</strong>naural sounds and keyboard <strong>vib</strong>ration signals.<br>Samples were acquired with high-quality audio and vibration measurement equipment on two <a href="https://en.wikipedia.org/wiki/Disklavier">Yamaha Disklavier pianos</a> (one grand and one upright model) by means of computer-controlled playback of each key at ten different MIDI velocity values.<br>Project files (<em>instruments</em> and <em>multis</em>) are provided for use with the software sampler <a href="https://www.native-instruments.com/en/products/komplete/samplers/kontakt-6/">Native Instruments Kontakt</a> (version 5 and above, available for Windows and Mac OS).<br>The nominal specifications of the equipment used in the acquisition chain are reported in a companion document, allowing researchers to calculate physical quantities (e.g. acoustic pressure, vibration acceleration) from the recordings.<br>The library is especially suited for acoustic and vibration research on the piano, as well as for research on multimodal interaction with musical instruments.</p>

opencc-by-nc-sa-4.0Jan 2019View details →
zenodo52/100

LTER-Italy site Piano Limina CAL1 figure

<p>Geographical representation of the LTER-Italy site Piano Limina CAL1 (LTER_EU_IT_032) - DEIMS-ID <a href="https://deims.org/d35d5417-d167-4137-97d1-c62ae4bc580b">https://deims.org/d35d5417-d167-4137-97d1-c62ae4bc580b</a></p>

opencc-by-sa-4.0Aug 2021View details →
zenodo48/100

ATEPP: A Dataset of Automatically Transcribed Expressive Piano Performance

<p>ATEPP is a dataset of expressive piano performances by virtuoso pianists. The dataset contains 11742&nbsp;11677 performances (~1000 hours) by 49 pianists and covers 1580 movements by 25 composers. All of the MIDI files in the dataset come from the piano transcription of existing audio recordings of piano performances. Scores in MusicXML format are also available for around half of the tracks. The dataset is organized and aligned by compositions and movements for comparative studies. For more details, please check <a href="https://github.com/BetsyTang/ATEPP">here</a>.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo48/100

Dataset accompanying the publication: Acoustic cues of keyboard mechanics enable auditory localization of upright piano tones

<p>Dataset accompanying the publication: Acoustic cues of keyboard mechanics enable auditory localization of upright piano tones (in J. Acoust. Soc. Am., 2024)</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Piano sonatas and Beethoven works in Hans Georg Nägelis catalogues

<p>This dataset (in two csv files) collects the references to piano sonatas, and to works by Ludwig van Beethoven, in the catalogues published by Hans Georg N&auml;geli as a bookseller in Zurich between 1792 and 1805.</p> <p>The data was used by the author in his paper &quot;Hans Georg N&auml;geli as Publisher and Bookseller of Piano Music&quot;, presented at the conference&nbsp;<a href="https://www.hkb-interpretation.ch/beethoven2020">Beethoven and the Piano</a>,&nbsp;4&ndash;7&nbsp;November 2020.</p> <p>The printed version of the article is at present in preparation.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

An Annotated Corpus of Tonal Piano Music from the Long 19th Century

<p>This corpus has been created within the&nbsp;<a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a>&nbsp;and employs&nbsp;the&nbsp;<a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p><strong>Version 1</strong>&nbsp;has been released for submitting it as part of the data&nbsp;report&nbsp;<code>Hentschel, J., Rammos, Y., Neuwirth, M., Rohrmeier, M. (forthcoming). An Annotated Corpus of Tonal Piano Music from the Long 19th Century</code>&nbsp;that accompanies nine corpora grouped under the DOI&nbsp;<a href="https://doi.org/10.5281/zenodo.7483349">10.5281/zenodo.7483349</a>.</p> <p><strong>Version 1.1</strong>&nbsp;comes with a complete set of metadata and score headers.&nbsp;Among more accurate composition dates,&nbsp;the&nbsp;metadata now include URIs that identify the compositions in terms of&nbsp;the&nbsp;<a href="https://viaf.org/">Virtual International Authority File (VIAF)</a>,&nbsp;<a href="https://www.wikidata.org/">Wikidata</a>,&nbsp;<a href="https://imslp.org/">IMSLP</a>&nbsp;and&nbsp;<a href="https://musicbrainz.org/">MusicBrainz</a>.&nbsp;The data has been re-extracted from the scores&nbsp;using&nbsp;<a href="https://pypi.org/project/ms3/">ms3 1.1.1</a>.</p> <p>The publication covers the following corpora (the DOI links always point at the latest version respectively):</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7473560">Ludwig van Beethoven - Piano Sonatas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473566">Fr&eacute;d&eacute;ric Chopin - Mazurkas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473568">Claude Debussy - Suite Bergamasque</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473576">Anton&iacute;n Dvoř&aacute;k - Silhouettes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473580">Franz Liszt - Ann&eacute;es de P&egrave;lerinage</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473528">Nikolai Medtner - Tales</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473582">Robert Schumann - Kinderszenen</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473586">Pyotr Tchaikovsky - The Seasons</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473578">Edvard Grieg - Lyric Pieces</a></li> </ul> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2022View details →
zenodo44/100

PIANO (Penetration and Interruption of Alpine Foehn) – flux station data set

<p>ABSTRACT</p> <p>This resource comprises meteorological and turbulence data from four flux stations operated during the PIANO (Penetration and Interruption of Alpine Foehn) field campaign. The campaign took place in and around Innsbruck, Austria, during autumn and early winter 2017. The goal of the PIANO campaign was to study south foehn events, in particular the interaction between cold air pools and foehn, the mechanisms by which foehn can break through to reach the valley floor and the processes affecting the subsequent breakdown of foehn. This dataset provides near-surface turbulence observations (including surface fluxes obtained using the eddy covariance technique), along with radiation and soil measurements, as well as meteorological information.</p> <p>DATA SET DESCRIPTION</p> <p>1. Spatial coverage and locations</p> <p>Three eddy covariance (EC) stations were operated at grassland sites during the PIANO campaign. One station (&lsquo;EC_South&rsquo;) was installed in the Wipp Valley near to the village of Patsch, south of the city of Innsbruck. Two stations were installed in the Inn Valley, one to the east of Innsbruck in the region of Thaur (&lsquo;EC_East&rsquo;) and one to the west of Innsbruck at Innsbruck Airport (&lsquo;EC_West&rsquo;). Data from a fourth EC station at the Innsbruck Atmospheric Observatory (IAO, Karl et al. (2020)) in the centre of Innsbruck (&lsquo;EC_Centre&rsquo;) was also used. Precise station co-ordinates are provided in the data files.</p> <p>Three of the stations were located on grassland surrounded by mixed agricultural fields: the two stations in the Inn Valley (EC_East, EC_West) were installed on the fairly flat valley floor, while the site in the Wipp Valley (EC_South) gently sloped downwards to the west. During the campaign the vegetation was generally short at 5-10 cm. As far as possible, sites were selected to have a clear fetch for at least a few hundred metres. All three grassland sites experienced snow cover during winter. The urban station (EC_Centre) is a long-term site installed above roof level and representative of the surrounding neighbourhood close to the city centre of Innsbruck.</p> <p>2. Temporal coverage</p> <p>The temporal coverage of the datasets for the PIANO campaign are as follows:</p> <p>&bull; EC_West: 15 Sep 2017 - 31 Dec 2017<br> &bull; EC_South: 08 Sep 2017 &ndash; 15 Dec 2017<br> &bull; EC_East: 13 Oct 2017 &ndash; 15 Dec 2017<br> &bull; EC_Centre: 1 Sep 2017 &ndash; 31 Dec 2017</p> <p>The timeseries for EC_East begins later than the other sites because electrical interference thought to be from a nearby transmitter meant there was no useable flux data for the first month. The site was relocated on 13 October 2017 (no data is included before this date). Repeated theft of the batteries at EC_East resulted in gaps for the last few days of the dataset in December 2017. Due to issues with remote data collection, data availability at EC_West is low in September 2017. The PIANO campaign took place during autumn and early winter 2017 but the EC_West station was operated for longer (until 22 May 2018 after which use of the site was no longer permitted) as it provided a useful rural comparison station for the urban measurements (Karl et al., 2020; Ward et al., submitted). Data for 1 January &ndash; 22 May 2018 are available from the first author on request. Data collection at the long-term EC_Centre/IAO site began in spring 2017 and is ongoing.</p> <p>3. Instrument details</p> <p>At EC_West a closed-path eddy covariance system (CPEC200, Campbell Scientific) provided fast response measurements of the three wind components, temperature, water vapour mixing ratio and carbon dioxide mixing ratio. At EC_East and EC_South a sonic anemometer (CSAT3B, Campbell Scientific) and krypton hygrometer (KH20, Campbell Scientific) provided fast response measurements of the three wind components, temperature and water vapour. These fast data were logged at 20 Hz (CR6, Campbell Scientific). All three stations were equipped with a four-component radiometer (CNR4, Kipp and Zonen) to provide incoming and outgoing shortwave and longwave radiation. Meteorological measurements included air temperature and humidity (Rotronic HC2A-S3, mounted in an actively ventilated radiation shield Rotronic RS12T), atmospheric pressure (Campbell CS100, mounted inside the logger box) and precipitation (ARG100 tipping bucket gauge, Campbell Scientific). Soil instruments comprised two soil heat flux plates at 0.05 m depth (HFP01, Hukseflux), two soil temperature sensors (107, Campbell Scientific) at 0.02 and 0.04 m depth and a soil probe (ACC-SEN-SDI, Acclima) providing soil moisture and soil temperature at 0.05 m depth. At each site, the fast-response anemometer and gas analyser were mounted on a tripod at around 2.5 m above ground, while the radiometer and temperature-humidity probe were slightly lower, at around 2.0 m (exact sensor heights are provided in the data files).</p> <p>At EC_Centre a closed-path eddy covariance system (CPEC200, Campbell Scientific) provided fast response measurements of the three wind components, temperature, water vapour mixing ratio and carbon dioxide mixing ratio at 10 Hz (CR3000, Campbell Scientific) measured at 42.8 m above ground level on a lattice mast installed on top of a university building. A four-component radiometer (CNR4, Kipp and Zonen) provided incoming and outgoing shortwave and longwave radiation and air temperature and humidity are also measured (Rotronic HC2A-S3, mounted in a ventilated radiation shield). Atmospheric pressure is measured by a pressure sensor mounted inside one of the electronics boxes supplied as part of the CPEC200 (EC100, Campbell Scientific). No soil or precipitation measurements were made at the urban station.</p> <p>4. Data processing</p> <p>The fast-response eddy covariance data were processed to 30-min statistics following standard procedures using EddyPro version 7.0.7 (LI-COR Biosciences, 2021). These include despiking of raw data, time-lag compensation using maximum covariance, double coordinate rotation (meaning the 30-min mean vertical wind speed is forced to zero),&nbsp;simple block averaging (i.e. no filtering was applied), humidity correction of sonic temperature (Schotanus et al., 1983), and spectral corrections at low frequencies (Moncrieff et al., 2004) and high frequencies (after Fratini et al. (2012) for the closed-path CPEC200 data and Moncrieff et al. (1997) for the krypton hygrometer data). Oxygen (Tanner et al., 1993; van Dijk et al., 2003) and density (Webb et al., 1980) corrections were also applied at the sites with krypton hygrometers. Automated calibration (zero and span for carbon dioxide and zero for water vapour) was performed for the CPEC instruments once per day at EC_West and twice per day at EC_Centre.</p> <p>In addition to the standard processing described above, gust speeds were calculated from the sonic data. First the instantaneous horizontal wind speed was calculated (neglecting any vertical component). A 3-s running mean of the horizontal wind speed was then obtained, and the gust speed taken as the maximum of this 3-s running mean over a 1-min averaging interval.</p> <p>The dissipation rate of turbulent kinetic energy was obtained from the fast-response measurements of the three wind components (u, v, w) as follows. First, spectra were calculated for u, v and w using evenly spaced logarithmic frequency bins. The inertial subrange was identified as the region around 1 Hz where a local linear fit to the spectral slope was within &plusmn;20% of the expected -5/3 slope. The dissipation rate was calculated for each frequency bin in the identified inertial subrange according to Kolmogorov theory (e.g. Kaimal and Finnigan, 1994), using a value of 0.55 for u and 0.73 for v and w for the Kolmogorov inertial subrange constants, and the mean value over the frequency bins was used to provide the dissipation rate for u, v, and w for each 30-min period. Further discussion can be found in Ward et al. (in prep.).</p> <p>Quality control removed data during times of power outage and instrument malfunction and data adversely affected by rainfall (all KH20 data during rainfall were removed). To exclude any potential effects of turbulence distortion, data were removed when the wind direction was within &plusmn;10&deg; of the mounting structure. Data falling outside physically reasonable thresholds were removed, including times when the rotation angle exceeded 45&deg;. Stationarity tests following Foken and Wichura (1996) were applied with a threshold of 100 (i.e. data were excluded when the difference between 5-min and 30-min statistics exceeded 100%).</p> <p>For the meteorological, radiation and soil data, quality control removed data during times of power outage and instrument malfunction (including when dew on the radiometer adversely affected readings).</p> <p>5. Data file structure</p> <p>Two files in netCDF format are provided containing processed and quality-controlled data:</p> <p>&bull; PIANO_EC_MetData_QC_1min_v1-00.nc containing the meteorological, radiation and soil data for each site at 1-min resolution. This file also contains horizontal wind speed (before co-ordinate rotation), wind direction and gust speed for each site at 1-min resolution.</p> <p>&bull; PIANO_EC_FluxData_QC_30min_v1-00.nc containing processed statistics and fluxes for each site at 30-min resolution.</p> <p>There are also quicklook plots (provided in PNG format, monthly and for the whole period) showing the data contained in these files.</p> <p>Four sets of files in ASCII format are provided containing the fast (10/20 Hz) eddy covariance data for each site for every 30-minute period. These files are timestamped with the time corresponding to the end of the period and are named:</p> <p>&bull; PIANO_EC_FastData_SITENAME_yyyymmdd_HHMM.csv.</p> <p>These sets of files are provided as a single .zip folder for each site which is named according to the site.</p> <p>All timestamps are given in UTC (in seconds since 00:00 UTC 01 January 1970) and denote the end of the averaging period.</p> <p>The following variables can be found in the MetData file: air temperature (ta), relative humidity (rh), atmospheric pressure (pa), precipitation (prec), soil temperature (ts1, ts2, ts3), soil volumetric water content (vwc), soil heat flux from each heat flux plate (shf1, shf2), incoming shortwave radiation (swin), outgoing shortwave radiation (swout), incoming longwave radiation (lwin), outgoing longwave radiation (lwout), wind speed (wspeed, i.e. vector average horizontal wind speed before double rotation), wind direction (wdir) and gust speed (gust).</p> <p>The following variables can be found in the FluxData file: friction velocity (ustar), sensible heat flux (h), latent heat flux (le), carbon dioxide flux (fco2), stability parameter (zeta), turbulent kinetic energy (tke), wind speed (wspeed, i.e. vector average wind speed after double rotation), wind direction (wdir), unrotated vertical wind velocity (wunrot, i.e. before double rotation), the standard deviation of the wind components and temperature (sigu, sigv, sigw, sigt), and dissipation rate of turbulent kinetic energy calculated from u, v and w spectra (epu, epv, epw).</p> <p>The following variables can be found in the RawData files: unrotated lateral, longitudinal and vertical wind components (in m s-1), temperature (in degree C), water vapour concentration (supplied for EC_West and EC_Centre as the mixing ratio (in mmol m-1) and supplied for EC_South and EC_East as the absolute humidity (g m-3) and carbon dioxide mixing ratio (in &mu;mol mol-1) for EC_West and EC_Centre. Note that the absolute value of the water vapour concentration from the krypton hygrometers should not be used. These lateral, longitudinal and vertical wind components are as measured in the co-ordinate system of the sonic anemometers and the angle of installation of the sonic needed to convert to north-south east-west co-ordinates is given in the FluxData file.</p> <p>6. Publications</p> <p>Data from these flux stations have been included in multiple publications as part of the PIANO project (Haid et al., 2020; Haid et al., 2021; Muschinski et al., 2021; Umek et al., 2021; Umek et al., submitted) as well as publications as part of a related study on turbulent exchange in complex environments (Ward et al., in prep.; Ward et al., submitted).</p> <p>7. Contact</p> <p>Contact helen.ward(at)uibk.ac.at for any questions regarding the data set.</p> <p>8. Acknowledgements</p> <p>The PIANO campaign was supported by the Austrian Science Fund (FWF) and the Weiss Science Foundation under Grant P29746-N32. Collection of this dataset was also supported by an FWF Lise Meitner project (M2244-N32) and a research stipend from Innsbruck University. Measurements at IAO are supported by the Bundesministerium f&uuml;r Wissenschaft, Forschung und Wirtschaft (Hochschulraum-Strukturmittel grant), the European Commission for funding ALP-AIR within FP7-PEOPLE and the FWF (P30600_NBL, P33701-N). The PIANO campaign was also supported by KIT IMK-IFU, Austro Control GmbH, Zentralanstalt f&uuml;r Meteorologie und Geodynamik (ZAMG), the Hydrographic Service of Tyrol, Innsbrucker Kommunalbetriebe AG (IKB), Bergisel Betriebsgesellschaft m.b.H., Innsbrucker Nordkettenbahnen Betriebs GmbH, T-Mobile Austria GmbH, Unser Lagerhaus Warenhandelsgesellschaft, PEMA Immobilien GmbH, HTL Anichstra&szlig;e, Hilton Innsbruck, TINETZ-Tiroler Netze GmbH, Land Tirol, and the communities Patsch and V&ouml;ls.</p> <p>9. References</p> <p>Foken T, Wichura B (1996) Tools for quality assessment of surface-based flux measurements. Agric. For. Meteorol. 78: 83-105 doi: 10.1016/0168-1923(95)02248-1</p> <p>Fratini G, Ibrom A, Arriga N, Burba G, Papale D (2012) Relative humidity effects on water vapour fluxes measured with closed-path eddy-covariance systems with short sampling lines. Agric. For. Meteorol. 165: 53-63 doi: 10.1016/j.agrformet.2012.05.018</p> <p>Haid M, Gohm A, Umek L, Ward HC, Muschinski T, Lehner L, Rotach MW (2020) Foehn&ndash;cold pool interactions in the Inn Valley during PIANO IOP2. Q. J. R. Meteorol. Soc. 146: 1232-1263 doi: 10.1002/qj.3735</p> <p>Haid M, Gohm A, Umek L, Ward HC, Rotach MW (2021) Cold-air pool processes in the Inn Valley during foehn: A comparison of four cases during PIANO. Boundary Layer Meteorology doi: 10.1007/s10546-021-00663-9</p> <p>Kaimal JC, Finnigan JJ (1994) Atmospheric Boundary Layer Flows: Their structure and management. Oxford University Press, 289 pp.</p> <p>Karl T et al. (2020) Studying urban climate and air quality in the Alps - The Innsbruck Atmospheric Observatory. Bull. Amer. Meteorol. Soc. doi: 10.1175/BAMS-D-19-0270.1</p> <p>LI-COR Biosciences (2021) Eddy Covariance Processing Software - version 7.0.7, Available at www.licor.com/EddyPro.</p> <p>Moncrieff JB, Clement R, Finnigan JJ, Meyers T (2004) Averaging, detrending and filtering of eddy covariance time series. In: X Lee,</p> <p>Massman WJ and Law BE (Editors), Handbook of Micrometeorology: a guide for surface flux measurements.</p> <p>Moncrieff JB et al. (1997) A system to measure surface fluxes of momentum, sensible heat, water vapour and carbon dioxide. Journal of Hydrology 188-199: 589-611</p> <p>Muschinski T, Gohm A, Haid M, Umek L, Ward HC (2021) Spatial heterogeneity of the Inn Valley Cold Air Pool during south foehn: Observations from an array of temperature. Meteorol. Z. 30: 153-168 doi: 10.1127/metz/2020/1043</p> <p>Schotanus P, Nieuwstadt FTM, Bruin HAR (1983) Temperature measurement with a sonic anemometer and its application to heat and moisture fluxes. Bound.-Layer Meteor. 26: 81-93 doi: 10.1007/bf00164332</p> <p>Tanner B, Swiatek E, Greene J (1993) Density fluctuations and use of the krypton hygrometer in surface flux measurements. Management of irrigation and drainage systems: integrated perspectives. American Society of Civil Engineers, New York, NY: 945-952</p> <p>Umek L, Gohm A, Haid M, Ward HC, Rotach MW (2021) Large eddy simulation of foehn-cold pool interactions in the Inn Valley during PIANO IOP2. Quart J Roy Meteorol Soc 147: 944-982 doi: 10.1002/qj.3954</p> <p>Umek L, Gohm A, Haid M, Ward HC, Rotach MW (submitted) Influence of grid resolution of large-eddy simulations on foehn-cold pool interaction. Quart J Roy Meteorol Soc</p> <p>van Dijk A, Kohsiek W, de Bruin HAR (2003) Oxygen Sensitivity of Krypton and Lyman-&alpha; Hygrometers. J. Atmos. Ocean. Technol. 20: 143-151 doi: 10.1175/1520-0426(2003)020&lt;0143:osokal&gt;2.0.co;2</p> <p>Ward HC, Rotach MW, Gohm A, Graus M, Karl T, Haid M, Umek L, Muschinski T (submitted) Energy and mass exchange at an urban site in mountainous terrain &ndash; the Alpine city of Innsbruck. Atmos. Chem. Phys.</p> <p>Ward HC, Rotach MW, Graus M, Karl T, Gohm A, Umek L, Haid M (in prep.) Turbulence characteristics at an urban site in highly complex terrain.</p> <p>Webb EK, Pearman GI, Leuning R (1980) Correction of flux measurements for density effects due to heat and water-vapor transfer. Q. J. R. Meteorol. Soc. 106: 85-100</p> <p></p> <p></p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

PIANO-AUG - Piano Multipitch Estimation with Augmentation Dataset

<p>This dataset includes a 29 minutes long piano recording with a total of 4584 notes. The recording includes sections of single notes as well as groups of simultaneously sounding notes (intervals, chords), which ascend chromatically within the MIDI pitch range [21,92].</p> <p>In particular, the segments cover single notes, intervals (two-note groups), as well as common three- voiced and four-voiced chords. It must be noted that we do not investigate chord inversions here, all chords are used in their root position.</p> <p>The audio is created by rendering the MIDI file using the &ldquo;Grand Piano&rdquo; plugin in Ableton Live 8 with default settings except from the &ldquo;Reverb Amount&rdquo; being set to zero.</p> <p>The&nbsp;<a href="https://github.com/spotify/pedalboard">pedalboard</a>&nbsp;python library (0.4.1) can be used to create multiple augmented versions of the unprocessed piano recording. In particular, we create pairs of mild (-) and heavy (+) augmentations for each of the compression (comp), gain (gain), low-pass filter (lpf), and reverb (rev) effects using <a href="https://github.com/jakobabesser/piano_aug/blob/main/create_augmented_versions.py">this Python script</a></p> <p>Further details are provided at <a href="https://github.com/jakobabesser/piano_aug">https://github.com/jakobabesser/piano_aug</a>.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

An Automatically Transcribed Piano Performer Dataset

<p>Data-driven approaches for performers&#39; style analysis need large corpora of music performances to derive expressive performance parameters. One good reason for Deep Neural Networks (DNNs) not being used for performer identification is the lack of large-scale datasets with overlapping performances by different performers. Hence, to bridge this&nbsp;gap, we created a score-aligned&nbsp;automatically transcribed performer dataset.&nbsp;There are a total of 474 performances in the dataset, played by 6 pianists, spanning 35 movements by 2 composers&nbsp;for a total of 474 Western classical piano recordings in MIDI format.&nbsp;Every single MIDI file that&#39;s included in this collection was derived from a piano transcription of an already-existing audio recording of a piano performance. Score files in musicXML format and the alignment results are also included in the dataset.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Influence of Playing on the Tonal Characteristics of a Concert Piano - Dataset

<p>Dataset supporting conference proceedings&nbsp;publication:</p> <p><strong>## Abstract</strong></p> <p>Well-maintained pianos are said to &rdquo;mature&rdquo; and to &rdquo;change for the better&rdquo; over the first years. When auditioning concert<br> pianos for purchase, technicians do often not choose the best sounding instrument but the one with the greatest potential<br> for future development. The present work addresses the following questions: Are structural changes measurable on a<br> piano after one year of operation in a concert house? Are these changes perceivable by listeners?<br> Measurements are performed on two occasions: First, on a brand new instrument prepared for sale. Second, on the same<br> piano after having been played for one year in a concert hall. Single notes are recorded with dummy-head-microphones<br> in player position in an anechoic chamber. An extended ABX listening test engaging approx. 100 players, tuners,<br> and builders, addresses the questions whether a variation in tonal quality is audible and if so, what sound properties<br> could lead to a perceived difference. Semantic sub-grouping allows for indication on the vocabulary listeners of varying<br> expertise use to verbalize their sensation. The statements give hints on what could have changed over the year and<br> are used as a basis for the analysis of corresponding physical properties and psychoacoustic parameters related to the<br> described sensations.</p> <p><strong>## Data Structure</strong></p> <p><strong>### Naming Rules:</strong></p> <p><strong>Example:</strong> &#39;D1_PROD07_01__0000_M0000.csv&#39;</p> <ul> <li>Ignore the <em>D1</em></li> <li><em>PROD07&nbsp;</em>is before the year in a concert hall, <em>PROD08&nbsp;</em>is after the year.</li> <li><em>01</em> is the key (range: <em>01-88</em>)</li> <li>Ignore the&nbsp;<em>0000.</em></li> <li><em>M0000 </em>is the take (range: <em>M0000-M0004</em>)</li> </ul> <p>.csv files contain the data, the corresponding&nbsp;.txt files are the headers.</p> <p><strong>### Columns in&nbsp;.csv file:</strong></p> <ul> <li><strong>Time [s]</strong>; Time vector</li> <li><strong>Kistler 9722A500 [N]</strong>; Force sensor, measures the key / key bed impact force</li> <li><strong>PCB 352C23 [m/s^2]</strong>; Acceleration sensor, measures the acceleration at the bridge (at the hitch pins of the corresponding string) in direction normal to the soundboard.&nbsp;</li> <li><strong>HSU 3.2 left [Pa]</strong>; Left ear channel of dummy head, measures sound pressure in player position.</li> <li><strong>HSU 3.2 right [Pa]</strong>;&nbsp;Right ear channel of dummy head, measures sound pressure in player position.</li> </ul> <p>sample rate = 50000</p> <p>The listening test utilizing this dataset is still online (however, your input will not be evaluated):&nbsp;<a href="http://www.culturalheritage.digital/listeningtest/">http://www.culturalheritage.digital/listeningtest/</a></p> <p>Feel free to contact me for comments or questions: niko.plath@uni-hamburg.de</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Saarland Music Data: MIDI-Audio Piano Music

<p>This is an improved version of the dataset originally referred to as <strong>SMD MIDI-Audio Piano Music.</strong> For more details, please visit the website: <a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi">https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi</a></p> <p>Saarland Music Data provides audio recordings along with perfectly synchronized MIDI files for various piano pieces. The pieces were performed by students of the <a href="http://www.hfm.saarland.de">Hochschule f&uuml;r Musik Saar</a> on a hybrid acoustic/digital piano <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a>. The Disklavier allows for capturing key and pedal movements of the piano while playing. This information, which can be stored in a MIDI file, yields an accurate annotation of the corresponding audio recording in form of a symbolic description of all played musical note events. The SMD MIDI-Audio pairs constitute a valuable dataset for various music analysis tasks such as music transcription, performance analysis, music synchronization, audio alignment, or source separation.&nbsp;</p> <p>All performances were recorded in the studios of the <a href="http://www.hfm.saarland.de">Hochschule f&uuml;r Musik Saar</a>, played by students of piano classes of different levels, on a <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a> model <a href="http://www.yamaha.com/yamahavgn/CDA/ContentDetail/ModelSeriesDetail.html?CNTID=556850&amp;CNTYP=PRODUCT">DCFIIISM4PRO</a>. Using two cardioid-condenser microphones fixed over the resonating body of the piano, all performances were directly recorded into Steinberg Cubase 4. Except for trimming the beginnings and ends of the recordings, no further post-processing (filters, effects) was applied to the musical material. From each Cubase project, an audio file (44.1 kHz, stereo) as well as a synchronized standard MIDI file (SMF) were exported. Besides these files, we also provide the audio files as WAV (22.05 kHz, mono) and the MIDI files encoded as CSV files and as WAV files (22.05 kHz, mono) rendered using the Software synthesizer <a href="https://www.fluidsynth.org/">FluidSynth</a>.</p> <p>SMD MIDI-Audio Piano Music (V1) contains the following data:</p> <ul> <li>wav_44100_stereo: Audio file (44.1 kHz, stereo)</li> <li>wav_22050_mono: Audio file (22.05 kHz, mono)</li> <li>midi: MIDI file</li> <li>csv: Export of note events from MIDI file into CSV format</li> <li>midi_wav_22050_mono: MIDI file rendered as audio file (22.05 kHz, mono)</li> </ul> <p>If you publish results obtained using this dataset, please cite:</p> <p>Meinard M&uuml;ller, Verena Konz, Wolfgang Bogler, Vlora Arifi-M&uuml;ller: Saarland Music Data (SMD). In Late-Breaking and Demo Session of the 12th International Conference on Music Information Retrieval (ISMIR), 2011. [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/2011_MuellerKonzBoglerArifi_SaarlandMusicData_ISMIR-LateBreaking.pdf">pdf</a>] [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/bibtex.html">bib</a>]</p>

opencc-by-3.0Mar 2024View details →
zenodo44/100

Claude Debussy – Other Piano Pieces

<p>This dataset originates from the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and contains musicological research data. For more information, please refer to its documentation page <a href="https://dcmlab.github.io/debussy_other_piano_pieces">https://dcmlab.github.io/debussy_other_piano_pieces</a></p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0May 2023View details →
zenodo44/100

Claude Debussy – Pour le Piano

<p>This dataset originates from the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and contains musicological research data. For more information, please refer to its documentation page <a href="https://dcmlab.github.io/debussy_pour_le_piano">https://dcmlab.github.io/debussy_pour_le_piano</a></p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0May 2023View details →
zenodo40/100

JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting

<p>Release of the MIDI data pairs that constitute the JAZZVAR dataset. See below for the abstract of the publication.</p> <p>The data is also available transposed to C/Am and subsequently, to all keys, with accompanying metadata.</p> <p>Abstract:</p> <p>Jazz pianists often uniquely interpret jazz standards. Passages from these interpretations can be viewed as sections of variation. We manually extracted such variations from solo jazz piano performances. The JAZZVAR dataset is a collection of 502 pairs of Variation and Original MIDI segments. Each Variation in the dataset is accompanied by a corresponding Original segment containing the melody and chords from the original jazz standard. Our approach differs from many existing jazz datasets in the music information retrieval (MIR) community, which often focus on improvisation sections within jazz performances. In this paper, we outline the curation process for obtaining and sorting the repertoire, the pipeline for creating the Original and Variation pairs, and our analysis of the dataset. We also introduce a new generative music task, Music Overpainting, and present a baseline Transformer model trained on the JAZZVAR dataset for this task. Other potential applications of our dataset include expressive performance analysis and performer identification.</p>

opencc-by-nc-sa-2.0May 2024View details →
zenodo40/100

Audio Piano Triad Dataset

<p>Created by: Agust&iacute;n Macaya Valladares<br> Date: May 5th, 2021</p> <p>- Dataset contains 43.200 examples of piano triads in .wav format.<br> - Second Version: The audios are the same as the first version, but the octave number in the names were&nbsp;corrected from (2, 3, 4) to (3, 4, 5), respectively.</p> <p>Details:<br> - Sample rate: 16000 Hz.<br> - Data type: 16-bit PCM (int16).<br> - File size: Each example has a file size of 128 kB (5.53 GB for complete dataset).<br> - Duration: 4 seconds.<br> - Sound: Piano (digital).<br> - Chords were played by a human on a velocity-sensitive piano keyboard.<br> - 3 seconds pressed, 1 second released.</p> <p>- 3 octaves (3, 4, 5).<br> - 12 base notes per octave: Cn, Df, Dn, Ef, En, Fn, Gf, Gn, Af, An, Bf, Bn. (n is natural, f is flat).<br> - 4 triad types per note: major (j), minor (n), diminished (d), augmented (a). No inversions.<br> - 3 volumes per triad: forte (f), metsoforte (m), piano (p).</p> <p><br> - 10 original examples per combination of octave, base note, triad type, and volume. (10*3*12*4*3 = 4.320 examples).<br> - x10 data augmentation for each example (4.320 * 10 = 43.200 total examples).<br> - Data augmentation through random temporal and amplitude shifts.<br> - Metadata is in&nbsp;the name of the chord. For example: &quot;piano_4_Af_d_m_45.wav&quot; is a piano chord, (4) 4th&nbsp;octave, (Af) A flat base note, (d) diminished, (m) metsoforte, 45th example.</p> <p>Note:<br> - The audios are in 16-bit PCM (int16) data type to reduce the file size. This means that the dynamic range of values in the array is -32768 to 32768, integers. To normalize the audios in the range -1 to 1 (float) just divide by 32768.</p> <p>Second Version: The audios are the same as the first version, but the octave number in the names were corrected from (2, 3, 4) to (3, 4, 5), respectively.</p>

opencc-by-4.0May 2021View details →
zenodo40/100

EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

<p>EMOPIA (pronounced &lsquo;yee-m&ograve;-pi-uh&rsquo;) dataset is a shared multi-modal (audio and MIDI) database focusing on perceived emotion in&nbsp;<strong>pop piano music</strong>, to facilitate research on various tasks related to music emotion. The dataset contains&nbsp;<strong>1,087</strong>&nbsp;music clips from 387 songs and&nbsp;<strong>clip-level</strong>&nbsp;emotion labels annotated by four dedicated annotators.&nbsp;</p> <p>For more detailed information about the dataset, please refer to our paper:&nbsp;<a href="https://arxiv.org/abs/2108.01374"><strong>EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation</strong></a>.&nbsp;</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>midis/</strong></em>:&nbsp;midi clips transcribed using GiantMIDI. <ul> <li>Filename `Q1_xxxxxxx_2.mp3`: Q1 means this clip belongs to Q1 on the V-A space; xxxxxxx is the song ID on YouTube, and the `2` means this clip is the 2nd clip taken from the full song.</li> </ul> </li> <li><em><strong>metadata/</strong></em>:&nbsp;metadata from YouTube. (Got when crawling)</li> <li> <p><em><strong>songs_lists/</strong></em>:&nbsp;YouTube URLs of songs.</p> </li> <li> <p><em><strong>tagging_lists/</strong></em>:&nbsp;raw tagging result for each sample.</p> </li> <li> <p><em><strong>label.csv</strong></em>: metadata that records filename, 4Q label, and annotator.</p> </li> <li> <p><em><strong>metadata_by_song.csv</strong></em>: list all the clips by the song. Can be used to create the train/val/test splits to avoid the same song appear in both train and test.</p> </li> <li> <p><em><strong>scripts/prepare_split.ipynb:</strong></em> the script to create train/val/test splits and save them to csv files.</p> </li> </ul> <p>------</p> <p><strong>2.2 Update</strong></p> <ul> <li>Add tagging files in <em><strong>tagging_lists/</strong></em> that are missing in the previous version.</li> <li>Add <em><strong>timestamps.json</strong></em>&nbsp;for easier usage. It records all the timestamps in dict format. You can see <em><strong>scripts/load_timestamp.ipynb</strong></em>&nbsp;for the format example.</li> <li>Add&nbsp;<em><strong>scripts/timestamp2clip.py</strong></em>:&nbsp;After the raw audio are crawled and put in <em><strong>audios/raw</strong></em>, you can use this script to get audio clips. The script will read <em><strong>timestamps.json</strong></em>&nbsp;and use the timestamp to extract clips. The clips will be saved to <em><strong>audios/seg</strong>&nbsp;</em>folder.</li> <li>remove 7 midi files that were added by mistake, and also corrected the number in <em><strong>metadata_by_song.csv</strong></em>.</li> </ul> <p>&nbsp;</p> <p><strong>2.1 Update</strong></p> <p>Add one file and one folder:</p> <ul> <li><em><strong>key_mode_tempo.csv</strong></em>: key, mode, and tempo information extracted from files.</li> <li><strong><em>CP_events/</em></strong>:&nbsp; CP events used in our paper. Extracted using this <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/representations/uncond/cp/corpus2events.py">script</a>, and add the emotion event to the front.</li> </ul> <p>Modify one folder:</p> <ul> <li>The <strong><em>REMI_events/</em></strong> files in version 2.0 contain&nbsp;some information that is not related to the paper, so remove it.</li> </ul> <p>&nbsp;</p> <p><strong>2.0 Update</strong></p> <p>Add two new folders:</p> <ul> <li><strong><em>corpus/</em></strong>:&nbsp; processed data that following <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/Dataset.md">the&nbsp;preprocessing flow</a>. (Please notice that although we have&nbsp;<code>1078</code>&nbsp;clips in our dataset, we lost some clips during steps&nbsp;1~4 of&nbsp;the flow, so the final number of clips in this&nbsp;<strong><code>corpus</code></strong>&nbsp;is&nbsp;<code>1052</code>, and that&#39;s the number we&nbsp;used for training the generative model.)</li> <li><strong><em>REMI_events/</em></strong>: REMI event for each midi file. They are generated using this <a href="https://github.com/YatingMusic/compound-word-transformer/blob/main/dataset/representations/uncond/remi/corpus2events.py">script</a>.</li> </ul> <p>--------&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Cite this dataset</strong></p> <pre><code>@inproceedings{{EMOPIA}, author = {Hung, Hsiao-Tzu and Ching, Joann and Doh, Seungheon and Kim, Nabin and Nam, Juhan and Yang, Yi-Hsuan}, title = {{MOPIA}: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation}, booktitle = {Proc. Int. Society for Music Information Retrieval Conf.}, year = {2021} }</code></pre>

opencc-by-4.0Jul 2021View details →
zenodo40/100

jazznet: A Dataset of Fundamental Piano Patterns for Music Audio Machine Learning Research

<p>Jazznet is a&nbsp;dataset of&nbsp;piano patterns for music audio machine learning research. The dataset comprises chords, arpeggios, scales, and chord progressions in all keys of an 88-key piano and in all the inversions, for a total of&nbsp;162520 labeled piano patterns, resulting in 95GB of data and more than 26k hours of audio. The data is also accompanied by Python scripts to enable the easy generation of new piano patterns beyond those present in the dataset. The data is broken down into small, medium, and large subsets, comprising 21516, 30328, and 52360 patterns, respectively (with all the chords, arpeggios, and scales being present in all subsets).&nbsp;</p> <p>The GitHub page of the dataset, containing details of the dataset and scripts for generating new data is&nbsp;https://github.com/tosiron/jazznet.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

IDMT-PIANO-MM Dataset

<p>The dataset IDMT-PIANO-MM includes a total of 432 piano recordings (around four hours), which cover nine music pieces recorded in eight different rooms using six different recording devices. The pieces cover classical music (B. Bart&oacute;k, W. A. Mozart, J. Pachelbel, and L. v. Beethoven) as well as jazz (S. Joplin as well as own compositions) and range from simple to medium difficulty. All music pieces are in the public domain. The recording locations range from small rooms to a large lecture hall. Information about the room geometries, piano position within the room, as well as wall materials are documented. The rooms include four different grand pianos, three upright pianos, and one stage-piano. At each location, audio recordings were made with three mobile phones (iPhone 6S Plus, Redmi Note 8, LG G6), two tablets (iPad Air 2, Amazon Fire tablet), and one stereo setup using two high-quality Oktava MK 012 microphones in an AB recording setup.</p>

opencc-by-nc-nd-4.0Jan 2023View details →
zenodo40/100

PFVN-synth: Synthesized Violin-Piano Ensemble Dataset

<p>A PFVN-synth dataset contains realistic instrumental triplet audios of piano, violin, and their mixture by rendering MIDI files with virtual instruments using musical scores of 45 different pieces by 23 classical composers with a total duration of 7 hours. All MIDI files were collected on the <a href="https://musescore.com">MuseScore website</a>.&nbsp;All tracks are rendered into monaural audio files with the standard CD quality: 44.1kHz, 16-bit. We used commercial virtual instruments to synthesize piano and violin. Specifically, we used &lsquo;B&ouml;sendorfer Grand Piano&rsquo; in Apple Logic Pro and &lsquo;SWAM Violin V3&rsquo; by Audio Modeling.</p> <p>It is divided into a train set containing 32 pieces with a duration of 5.8 hours, a validation set containing 3 pieces with a duration of 30 minutes, and a test set containing 10 pieces with a duration of 50 minutes so that the composers are not biased to the split sets.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

The Claude Debussy Solo Piano Corpus

<p>This dataset originates from the <a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a> and contains musicological research data. For more information, please refer to its documentation page <a href="https://dcmlab.github.io/debussy_piano">https://dcmlab.github.io/debussy_piano</a></p> <p>Please cite this dataset as</p> <pre><code>Laneve, S., Schaerf, L., Cecchetti, G., Hentschel, J., &amp; Rohrmeier, M. (2023). The diachronic development of Debussy's musical style: A corpus study with Discrete Fourier Transform. Humanities and Social Sciences Communications, 10(1), 289. https://doi.org/10.1057/s41599-023-01796-7</code></pre> <p>The folders in the ZIP file that can be downloaded here from Zenodo are empty because they are submodules in <a href="https://github.com/DCMLab/debussy_piano">the original Git repository</a> which Zenodo is not including. However, the submodules have their own records:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7920474">debussy_childrens_corner</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963617">debussy_deux_arabesques</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963639">debussy_estampes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963636">debussy_etudes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963649">debussy_images</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963676">debussy_other_piano_pieces</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963656">debussy_pour_le_piano</a></li> <li><a href="https://doi.org/10.5281/zenodo.7963660">debussy_preludes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473568">debussy_suite_bergamasque</a></li> </ul> <p>&nbsp;</p>

opencc-by-nc-sa-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record