Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Musicology”
Music Data Sharing Platform for Computational Musicology Research (CCMUSIC DATASET)
<p>This platform is a multi-functional music data sharing platform for Computational Musicology research. It contains many music datas such as the sound information of Chinese traditional musical instruments and the labeling information of Chinese pop music, which is available for free use by computational musicology researchers.</p> <p>This platform is also a large-scale music data sharing platform specially used for Computational Musicology research in China, including 3 music databases: Chinese Traditional Instrument Sound Database (CTIS), Midi-wav Bi-directional Database of Pop Music and Multi-functional Music Database for MIR Research (CCMusic). All 3 databases are available for free use by computational musicology researchers. For the contents contained in the database, we will provide audio files recorded by the professional team of the conservatory of music, as well as corresponding labelled files, which have no commodity copyright problem and facilitate large-scale promotion. We hope that this music data sharing platform can meet the one-stop data needs of users and contribute to the research in the field of Computational Musicology.</p> <p> </p> <p>If you want to know more information or obtain complete files, please go to the official website of this platform:</p> <p><a href="https://ccmusic-database.github.io/en/">Music Data Sharing Platform for Academic Research</a></p> <p> </p> <ul> <li> <p><strong>Chinese Traditional Instrument Sound Database (CTIS)</strong></p> </li> </ul> <p>This database is developed by Prof. Han Baoqiang's team for many years, which collects sound information about Chinese traditional musical instruments. The database includes 287 Chinese national musical instruments, including traditional musical instruments, improved musical instruments and ethnic minority musical instruments.</p> <ul> <li> <p><strong>Multi-functional Music Database for MIR Research</strong></p> </li> </ul> <p>This database collects sound materials of pop music, folk music and hundreds of national musical instruments, and makes comprehensive annotation to form a multi-purpose music database for MIR researchers.</p> <ul> <li><strong>Midi-wav Bi-directional Database of Pop Music</strong></li> </ul> <p>This database contains hundreds of Chinese pop songs, and each song contains the corresponding midi-audio-lyric information. Among them, recording the vocal part and accompaniment part of audio independently is helpful to study the MIR task under the ideal situation. In addition, the information of singing techniques consistent with vocal part (such as breath sound, falsetto, breathing, vibrato, mute, slide, etc.) is marked in MuseScore, which constitutes a Midi-Wav bi-direction corresponding pop music database.</p>
Erkomaishvili Dataset: A Curated Corpus of Traditional Georgian Vocal Music for Computational Musicology
<p><strong>Abstract</strong></p> <p>The analysis of recorded audio material using computational methods has received increased attention in ethnomusicological research. We present a curated dataset of traditional Georgian vocal music for computational musicology. The corpus is based on historic tape recordings of three-voice Georgian songs performed by the the former master chanter Artem Erkomaishvili. In this article, we give a detailed overview on the audio material, transcriptions, and annotations contained in the dataset. Beyond its importance for ethnomusicological research, this carefully organized and annotated corpus constitutes a challenging scenario for music information retrieval tasks such as fundamental frequency estimation, onset detection, and score-to-audio alignment. The corpus is publicly available and accessible through score-following web-players.</p> <p><strong>License</strong></p> <p>This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nc/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</p> <p><strong>Copyright of Audio (wav)</strong></p> <p>Ministry of Culture, Sports and Youth of Georgia<br> Legal Entity of Public Law<br> Vano Sarajishvili Tbilisi State Conservatoire (TSC)<br> 8-10, GRIBOEDOV St, TBILISI 0108, GEORGIA Tel. / fax :(+995 32) 2 999 144,<br> www.tsc.edu.ge E-mail: info@tsc.edu.ge; inter@tsc.edu.ge</p> <p>We thank the rector of TSC, Nana Sharikadze, for the permission to publish the recordings along with our annotations on Zenodo.</p> <p><strong>Copyright of Annotations (csv)</strong></p> <p>Sebastian Rosenzweig^1, Frank Scherbaum^2, David Shugliashvili^3, Vlora Arifi-Müller^1, and Meinard Müller^1<br> ^1: International Audio Laboratories Erlangen, Germany<br> ^2: University of Potsdam, Germany<br> ^3: Tbilisi State Conservatoire, Georgia</p> <p>The provided digital sheet music in MusicXML-format is based on the transcriptions by David Shugliashvili as published in the book:</p> <p>David Shugliashvili<br> Georgian Church Hymns, Shemokmedi School<br> Georgian Chanting Foundation, 2014.</p> <p><strong>References</strong></p> <p>If you use the Erkomaishvili dataset in your research, please cite:</p> <p>Sebastian Rosenzweig, Frank Scherbaum, David Shugliashvili, Vlora Arifi-Müller, and Meinard Müller<br> Erkomaishvili Dataset: A Curated Corpus of Traditional Georgian Vocal Music for Computational Musicology<br> Transactions of the International Society for Music Information Retrieval (TISMIR), 3(1): 31–41, 2020.</p>
Computational Musicology – Distant Reading of Sheet Music
<p>2<sup>nd </sup>Lecture</p>
TrustMus benchmark: The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
<p>TrustMus is an initial, rigorously validated benchmark designed to assess the accuracy and reliability of large language models (LLMs) in the domain of musicology. This dataset includes a collection of 400 human-validated multiple-choice questions, categorized into four thematic areas: People (Ppl), Instruments and Technology (I&T), Genres, Forms, and Theory (Thr), and Culture and History (C&H).</p> <p>The questions are derived from <em>The Grove Dictionary Online</em> using a semi-automated methodology. The process involves generating initial questions with a fine-tuned retrieval-augmented generation (RAG) model, filtering them through a series of automated checks, and finally validating them through expert human annotation. TrustMus is introduced in an initial paper, providing a critical resource for researchers and developers aiming to evaluate and improve LLM performance in this specialized field of musicology.</p> <p>This benchmark is discussed in the paper : </p> <p><strong>BibTeX Citation:</strong></p> <div> <div><code>@inproceedings{ramoneda2024trustmus,</code><br><code> title={The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?},</code><br><code> author={Ramoneda, Pedro and Parada-Cabaleiro, Emilia and Weck, Benno and Serra, Xavier},</code><br><code> booktitle={Proceedings of the 3rd Workshop on NLP for Music and Audio (NLP4MusA)},</code><br><code> year={2024},</code><br><code> month={November},</code><br><code> address={San Francisco, USA},</code><br><code> organization={Co-located with ISMIR'2024}</code><br><code>}</code><br><br> <div> </div> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.