Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “lyrics”

Learn how ShareScore rates datasets ↗
zenodo48/100

Cadenza Challenge (CAD2): databases for lyric intelligibility task

<h2>Cadenza</h2> <p>This is the training and validation data for the lyric intelligibility task from the <a href="https://cadenzachallenge.org/">Second Cadenza Machine Learning Challenge (CAD2).</a></p> <p>The Cadenza Challenges are improving music production and processing for people with a hearing loss. According to The World Health Organization, 430 million people worldwide have a disabling hearing loss. Studies show that not being able to understand lyrics is an important problem to tackle for those with hearing loss. Consequently, this task is about improving the intelligibility of lyrics when listening to pop/rock over headphones. But this needs to be done without losing too much audio quality - you can't improve intelligibility just by turning off the rest of the band! We will be using one metric for intelligibility and another metric for audio quality, and giving you different targets to explore the balance between these metrics.</p> <p>Please see the <a href="https://cadenzachallenge.org/">Cadenza website</a> for a full description of the data</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Arab-Andalusian music lyrics dataset

<p>The dataset contains lyrics for the songs in the&nbsp;<a href="https://musicbrainz.org/collection/142ea0d7-7fdf-4ea5-9b04-219f68023d01">Arab-Anadalusian music collection curated within the CompMusic project</a>, that belong to the nawbas &quot;Isbahan&quot;, &quot;Maya&rdquo;, &ldquo;Raml Maya&rdquo;, &ldquo;Gharibat al-Husayn&rdquo;, &ldquo;Hijaz Kabir&rdquo;, &ldquo;Hijaz Msharqi&rdquo;, &ldquo;Istihlal&rdquo;, &ldquo;Rasd&rdquo;, and &rdquo;Rasd Dayl&rdquo;.</p> <p>Lyrics are stored in two formats: as Tab Separated Values (TSV) files and as JSON files.</p> <p>Each file is identified by its MusicBrainz recording ID (MBID).</p> <p>The lyrics are stored both in their original Arabic script (folder &#39;original&#39;) and a romanized/transliterated version (folder &#39;transliterated&#39;) using the American Library of Congress (ALA-LC standard).</p> <p>Corresponding audio files are available from&nbsp;<a href="https://zenodo.org/record/1291776#.WyeYnZ9fjCI">the Arab-Andalusian music corpus</a>, as well as the Internet Archive URL included in the metadata file (&#39;metadata.csv&#39;).</p> <p>For more information about the exact format and contents of the dataset, please consult the README provided in the archive.</p> <p>For more information, please refer to&nbsp;<a href="http://compmusic.upf.edu/corpora">http://compmusic.upf.edu/corpora</a>.</p>

opencc-by-nc-4.0Jun 2018View details →
zenodo40/100

MUSDB18 lyrics extension

<p>This is a set of annotated lyrics transcripts for songs belonging to&nbsp;the <a href="https://zenodo.org/record/1117372">MUSDB18 dataset</a>.&nbsp;The set comprises lyrics of&nbsp;all songs&nbsp;which have English lyrics, i.e. 96 out of 100 songs for the training set and 45 out of 50 songs&nbsp;for the test set. MUSDB18 is a dataset for music source separation and provides the following separated tracks for each song: vocals, bass, drums, other (rest of the accompaniment), mixture.</p> <p>The lyrics transcripts, together with the audio files of MUSDB18, are a valuable resource for research on tasks such as text-informed singing voice separation,&nbsp;automatic lyrics alignment, automatic lyrics transcription, and&nbsp;singing voice synthesis and analysis. The provided data should be used for research purposes only.&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>The lyrics were transcribed manually by the authors who are not native English speakers. It is likely that the transcriptions&nbsp;are not 100% correct. The composers of the songs are the copyright holders of the original lyrics.</p> <p>The songs were divided into sections of lengths between 3 and 12&nbsp;seconds. The priority when choosing the section boundaries was that they correspond to natural pauses and do not cut vocal sounds. The sections do not necessarily correspond to lyrically meaningful&nbsp;lines.&nbsp;Most of the sections do not overlap, some have an overlap of 1 second.&nbsp;In some difficult cases, e.g. shouting in metal songs or mumbled words, where the words are barely intelligible, we made an effort to make the transcriptions as accurate as possible phonetically and did not prioritize semantically meaningful phrases.</p> <p><strong>Citation</strong></p> <p>The dataset was built&nbsp;for the paper</p> <p><a href="https://hal.telecom-paris.fr/hal-03255334">Schulze-Forster, K., Doire, C., Richard, G., &amp; Badeau, R. &quot;Phoneme Level Lyrics Alignment and Text-Informed Singing Voice Separation.&quot;&nbsp;<em>IEEE/ACM Transactions on Audio, Speech and Language Processing&nbsp;</em>(2021).</a></p> <p>If you use the data for your research, please cite the corresponding paper:</p> <pre><code>@article{schulze2021phoneme, title={Phoneme Level Lyrics Alignment and Text-Informed Singing Voice Separation}, author={Schulze-Forster, Kilian and Doire, Clement and Richard, Ga{\"e}l and Badeau, Roland}, journal={IEEE/ACM Transactions on Audio, Speech and Language Processing}, year={2021}, publisher={IEEE} }</code></pre> <p><strong>Annotations</strong></p> <p>For each section, the annotations comprise:&nbsp;the start and end time, the corresponding lyrics, and a label indicating one of the following four properties:</p> <p>(a) &nbsp;only one person is singing<br> (b) &nbsp;several singers are pronouncing the same phonemes at the&nbsp;same time (possibly singing different notes)<br> (c) &nbsp;several singers are pronouncing different phonemes simultaneously (possibly singing different notes)<br> (d) &nbsp;no singing</p> <p>Segments that are labelled with the property (b) or (c) do not necessarily have this property over the whole segment duration. As soon as somewhere in a segment several singers are present, label (b) was assigned; as soon as they sung different phonemes somewhere at the same time, label (c) was assigned. Property (a) and (d) are valid for the entire segment. Furthermore, segments with property (c) can contain either some (lead) singer(s) singing some words in the presence of background singers singing long vowels such as &rsquo;ah&rsquo; or &rsquo;oh&rsquo; or they can contain multiple singers who sing different words at the same time. In the latter case, it was very difficult to recognise the sung words&nbsp;and to decide in which order to transcribe words or phrases sung simultaneously. These segments are marked with a &#39;*&#39; and it&nbsp;is recommended to reject them for most use cases.</p> <p>The annotations have the following format:<br> &lt;start time&gt; &lt;end time&gt; &lt;vocals property&gt; &lt;lyrics&gt;</p> <p><em>Example:</em><br> 00:18 00:23 a i know the reasons why &nbsp;--&gt; starts&nbsp;at 18 sec., ends&nbsp;at 23 sec., vocals type (a), lyrics: i know the reasons why</p> <p>The Python script <em>musdb_lyrics_cut_audio.py</em>&nbsp;is provided to automatically cut the MUSDB songs into the annotated segments. The script requires the <em>musdb </em>and <em>soundfile</em> package. The user needs to update the paths and select the desired sources and vocals types in lines 19-26. The script saves wav-files for each selected source for each annotated segment as well as the corresponding lyrics as txt-file. The MUSDB training partition is divided into a training and validation set. The tracks for the validation set can be changed below line 29.</p> <p>The file&nbsp;<em>words_and_phonemes.txt&nbsp;</em>contains a list of all words and their&nbsp;decomposition into phonemes. The phonemes are written in <a href="https://en.wikipedia.org/wiki/ARPABET">2-letter ARPABET style</a>&nbsp;and obtained with the&nbsp;<a href="http://www.speech.cs.cmu.edu/tools/lextool.html">LOGIOS&nbsp;Lexicon Tool</a>.</p> <p><strong>License</strong></p> <p>The data is licensed under the terms of the&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</a>.&nbsp;To view a copy of this license, read the provided LICENSE.txt file,&nbsp;visit&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a> or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</p> <p>The creators&nbsp;of MUSDB18 lyrics extension<em>&nbsp;</em>and their corresponding affiliation institutes are not liable for, and expressly exclude, all liability for loss or damage however and whenever caused to anyone by any use of MUSDB18 lyrics extension&nbsp;or any part of it.&nbsp;</p> <p><strong>Acknowledgment</strong></p> <p>The authors would like to thank Olumide Okubadejo and Sinead Namur for their help with transcribing and correcting&nbsp;part of the&nbsp;lyrics.</p> <p>&nbsp;</p>

openother-ncMar 2021View details →
zenodo40/100

N20EM dataset for multimodal lyric transcription

<p>N20EM dataset for multimodal lyric transcription, proposed in our ACM MM 2022 paper, MM-ALT: A Multimodal Automatic Lyric Transcription System. This dataset contains recordings of three modalities: audio, video, and IMU motion signal.&nbsp;</p> <p>Our paper's camera ready version:&nbsp;https://arxiv.org/abs/2207.06127</p> <p>Project website:&nbsp;https://n20em.github.io/</p> <p><strong>Note:&nbsp;</strong></p> <ol> <li><strong>Once you download the dataset, we assume you have read and agreed with the <a href="https://drive.google.com/file/d/1te7AxPTSGAdyqqNtkfjFbtCv4ydwgcOF/view?usp=sharing">Terms and Conditions</a>.</strong></li> <li><strong>Commercial usage is strictly prohibited.</strong></li> </ol> <p>Please cite our work as:</p> <p>@inproceedings{gu2022mm, &nbsp;title={MM-ALT: A multimodal automatic lyric transcription system}, &nbsp;author={Gu, Xiangming and Ou, Longshen and Ong, Danielle and Wang, Ye}, &nbsp;booktitle={Proceedings of the 30th ACM International Conference on Multimedia}, &nbsp;pages={3328--3337}, &nbsp;year={2022} }</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Oct 2022View details →
zenodo40/100

Archaic Lyric Agora (ALyrA) - Ontoterminology of Archaic Lyric Poets - version 1.0

<p>This is version 1.0 of the ALyrA ontoterminology (i.e., 'a terminology whose conceptual system is a formal ontology', Roche 2007), which focuses on modelling Archaic Lyric poets (799-430 BC&Epsilon;). AlyrA stands for "Archaic Lyrical Agora".&nbsp;</p> <p><br>The ontoterminology was created by Rafail Giannadakis (Department of Philology, University of Crete) under the supervision of Dr. Maria Papadopoulou (Department of Philology, University of Crete).</p> <p><br>We gratefully acknowledge the generous funding of the European Commission in the context of the TALOS AI for SSH project (Grant agreement ID: 101087269).</p> <p>Contact: giannadakis.uni@gmail.com&nbsp;</p>

openAug 2024View details →
zenodo36/100

Edvard Grieg – Lyric Pieces (A corpus of annotated scores)

https://dcmlab.github.io/grieg_lyric_pieces/

opencc-by-nc-sa-4.0Dec 2022View details →
zenodo36/100

Emotion4MIDI: A Lyrics-Based Emotion-Labeled Symbolic Music Dataset

<p>This dataset includes emotion labels for the publicly available MIDI dataset, namely Lakh MIDI Dataset and Reddit MIDI dataset. The values represent the probability of containing a particular emotion. For a single song, more than one emotion can be present, hence the values don't add up to 1.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems

<p>FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems. Dataset for the task force "Vulcan" from ISWS 2024 led by Aldo Gangemi and Andrea Poltronieri</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Corpus of song lyrics in Spanish labeled for gender-based violence against women

<pre><strong>Content:</strong> Labeled corpus including expressions of gender-based violence extracted from song lyrics in Spanish.<br><strong>Tags:</strong> [0: no gender-based violence; 1: gender-based violence] <strong>Description:</strong> It consists of 1000 labeled expressions, of which 778 correspond to expressions without content of gender-based violence and 222 data <br>that contain expressions of gender-based violence collected from song lyrics in Spanish. <strong>Language:</strong> Spanish <strong>Size:</strong> 549KB</pre>

opencc-by-4.0Aug 2024View details →
zenodo36/100

The Lyric Poetry, Angers Theater (1871)

The Lyric Poetry is one of the 4 sculptures of the Angers Theater facade who represents the principles of theater. **free download but do not forget to credit me if you use it publicly.** **DSLR scan of only 17 shots with Sony a300, complicated to shot ( 4 meters from the ground ), process with Reality Capture ** Source: Objaverse 1.0 / Sketchfab

opencc-byJul 2018View details →
zenodo32/100

Jingju Lyrics Datasets

<p>In order to study the expressive functions of jingju metrical patterns according to its lyrics, a series of different datasets have been created from the <a href="https://github.com/MTG/Jingju-Lyrics-Collection">Jingju Lyrics Collection</a>, that has been collected through scraping the online repository of jingju libretti <a href="http://www.xikao.com/"><em>Zhongguo jingju xikao</em> 中国京剧戏考</a>. These datasets have been created for the analysis of lyrics of the <em>banshi yuanban</em>, <em>manban</em>, <em>kuaiban </em>and <em>yaoban </em>both in the <em>shengqiang xipi </em>and <em>erhuang </em>(<em>kuaiban </em>is not used in <em>erhuang</em>) by applying NLP techniques, namely topic modelling and document classification.</p> <p><strong>Using this dataset</strong></p> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p><a href="http://compmusic.upf.edu/jingju-lyrics-datasets">http://compmusic.upf.edu/jingju-lyrics-datasets</a></p>

opencc-by-nc-nd-4.0Jul 2017View details →
zenodo28/100

THE GENESIS OF THE SYMBOL OF THE "SOUL" IN LYRIC POETRY

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

vocadito: A dataset of solo vocals with f0, note, and lyric Annotations

<p>vocadito is a dataset&nbsp;of 40 short excerpts of solo, monophonic singing. The excerpts are sung in 7 different languages by singers with varying of levels of training, and are recorded on a variety of devices. For a detailed description, see the <a href="https://arxiv.org/pdf/2110.05580.pdf">technical report</a>.</p> <p>Annotations are labeled&nbsp;by trained musicians. For each excerpt,&nbsp;we provide:</p> <ul> <li>frame-level f0 annotations</li> <li>2 versions of note annotations (from 2 different annotators)</li> <li>lyrics</li> <li>language</li> </ul> <p>Python code for loading this dataset is included in <a href="https://github.com/mir-dataset-loaders/mirdata">mirdata</a>.</p>

opencc-by-4.0Oct 2021View details →
geo24/100

Human tongue squamous cell carcinoma cell line SAS-Control vs. SAS-LYRIC knockdown stable clones global gene profiling

GEO Series GSE44766. Homo sapiens. 4 samples. Type: Expression profiling by array.

openGEO-OpenFeb 2016View details →
zenodo24/100

Dataset from the article, "Interference by linguistic processes in the occurrence of lyrics in involuntary musical imagery"

<p>Dataset from the article, &quot;Interference by linguistic processes in the occurrence of lyrics in involuntary musical imagery&quot;, in Journal of Cognitive Psychology.</p>

opencc-by-4.0Mar 2020View details →
zenodo24/100

CCogS-Mx/Spanish-lyrics-dataset-for-LGBTQ-phobia-screening: Spanish lyrics dataset for LGBTQ+phobia screening

<p>The development of this corpus includes songs written in Spanish that may contain LGBTQ+phobic text or not, contributor update</p>

restrictedcc-by-4.0Jul 2024View details →
ClinicalTrials.gov24/100

"Fun.Feel.Share" Lyrics-writing and Singing Show

ClinicalTrials.gov study NCT03368014. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Lyric Self-replacement Clinical Investigation

ClinicalTrials.gov study NCT05349981. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record