Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “Indian Music”

Learn how ShareScore rates datasets ↗
zenodo44/100

Indian Art Music Raga Recognition Dataset (features)

<p>The <strong>Rāga Recognition Datasets (features)</strong> comprise two sizable datasets, one for each music tradition: the <strong>Carnatic Music Dataset (CMD)</strong> and the <strong>Hindustani Music Dataset (HMD)</strong>. Each dataset entry includes features such as <strong>pitch</strong>, <strong>tonic</strong>, and <strong>nyas</strong> and <strong>tani</strong> segments. These datasets can be used to develop and evaluate approaches for automatic rāga recognition in Indian art music. To the best of our knowledge, they are the largest and most comprehensive datasets (in terms of available metadata) ever used for studying this task.</p> <p>This repository only contains the metadata and computed features for the dataset, and shared in open access. To get the audio, please refer&nbsp;<a href="https://zenodo.org/records/7278511" target="_blank" rel="noopener">to this zenodo entry</a> and submit your request.</p> <p>&nbsp;</p> <p>Please cite the following publications if you use the material shared here in your research work.</p> <blockquote> <p>Gulati, S., Serr&agrave;, J., Ganguli, K. K., &cedil;Sent&uuml;rk, S., &amp; Serra, X. (2016). Time-delayed melody surfaces for raga recognition. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), pp. 751&ndash;757. New York, USA. [<a href="http://hdl.handle.net/10230/33117">Postprint PDF</a>]</p> </blockquote> <blockquote> <p>Gulati, S., Serr&agrave;, J., Ishwar, V., &cedil;Sent&uuml;rk, S., &amp; Serra, X. (2016). Phrase-based raga recognition using vector space modeling. In Proceedings of the 41st IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 66&ndash;70. Shanghai, China. [<a href="http://hdl.handle.net/10230/32879">Postprint PDF</a>]</p> </blockquote> <p>&nbsp;</p> <h2>Annotation Format</h2> <p>We provide both tsv files and json files that contain information about each audio recording in terms of its mbid, the path of the audio/feature files and the associated rāga identifier. Each rāga is assigned a unique identifier by Dunya, which is similar to the mbid in terms of purpose. We also provide a mapping of the rāga id to its transliterated name.</p> <h2>Mirdata</h2> <p>This dataset is included in <a href="https://github.com/mir-dataset-loaders/mirdata">mirdata</a>. Use the following code snippet to access the dataset in mirdata.</p> <pre><code># Import midata import mirdata # Initialize dataset dataset_name = 'compmusic_raga' data_home = 'mirdata/dataset' dataset = mirdata.initialize(dataset_name, data_home=data_home) # Download dataset dataset.download() # Validate dataset dataset.validate() # Load dataset as a dictionary with track ids as keys and track objects as values data = dataset.load_tracks()</code></pre> <p>In order to load the audio files in mirdata, they must be requested beforehand and placed in the data home directory.</p> <h2>Contact&nbsp;</h2> <p>If you have any questions or comments about the dataset, please feel free to email:</p> <p><a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a></p> <p>&nbsp;</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

Indian Regional Music Dataset

<p>This dataset is a collection of mel-spectrogram features&nbsp;extracted from&nbsp;Indian regional music&nbsp;containing the following languages:<br> Hindi, Gujarati, Marathi, Konkani, Bengali, Oriya, Kashmiri, Assamese, Nepali, Konyak, Manipuri, Khasi &amp; Jaintia, Tamil, Malayalam, Punjabi, Telugu, Kannada.</p> <p>Five recordings are collected for each language for four artists (2Male + 2Female) each.&nbsp;2 artists out of 4 for each language are old veteran performers, and the remaining 2 are contemporary performers. Overall, the dataset includes 17 languages and 68 artists (34 Males and 34 Females). There are 340 recordings in the dataset, with a total duration of 29.3 hrs.</p> <p>Mel-spectrogram is extracted from a 3-second segment with a 1/2 second sliding window for each song. Extracted mel-spectrogram for each segment is annotated with&nbsp;language, location,&nbsp;local_song_index,&nbsp;global_song_index, language_id, location_id,&nbsp;artist_id, gender_id&nbsp;and no_of_artists.</p> <p>_________________________________________________________________________________________________________</p> <p>This project was funded under the grant number: ECR/2018/000204 by the Science &amp; Engineering Research Board (SERB).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Indian Folk Music Dataset

<p>This dataset is a collection of mel-spectrogram features extracted from Indian folk music containing the following 15 folk styles:<br> Bauls, Bhavageethe, Garba, Kajri, Maand, Sohar, Tamang Selo, Veeragase, Bhatiali, Bihu, Gidha, Lavani, Naatupura Paatu, Sufi, Uttarakhandi.</p> <p>The number of recordings varies from 16 to 50 in the mentioned folk styles representing the scarcity of availability of given folk styles on the Internet. There are at least 4 artists and a maximum of 22. Overall there are 125 artists (34 female + 91 male) in these 15 folk styles.&nbsp;</p> <p>There is a total of 606 recordings in the dataset, with a total duration of 54.45 hrs.<br> Mel-spectrogram is extracted from a 3-second segment with each song&#39;s 1/2 second sliding window. Extracted mel-spectrogram for each segment is annotated with folk_style, state, artist, gender, song, source, no_of_artists, folk_style_id, state_id, artist_id, gender_id.<br> _________________________________________________________________________________________________________<br> This project was funded under the grant number: ECR/2018/000204 by the Science &amp; Engineering Research Board (SERB).</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Indian Semi-Classical Music Dataset

<p>This dataset is a collection of mel-spectrogram features extracted from Indian semi-classical music containing the following 9 semi-classical styles:<br> Bhajan, Chaiti, Dadra, Ghazal, Kajri, Natya Sangeet, Qawwali, Tappa, Thumri.</p> <p>The number of recordings varies from 25 (for Chaiti) to 50 in the mentioned styles representing the scarcity of availability of given folk styles on the Internet. There are at least 5 artists and a maximum of 13. Overall there are 48 artists (36 female + 12 male) in these 9 semi-classical styles.&nbsp;<br> There is a total of 425 recordings in the dataset, with a total duration of 54.69 hrs.<br> Mel-spectrogram is extracted from a 3-second segment with each song&#39;s 1/2 second sliding window. Extracted mel-spectrogram for each segment is annotated with the genre, artist, gender, song, source, no_of_artists, genre_id, artist_id,&nbsp;&nbsp; &nbsp;gender_id.<br> _________________________________________________________________________________________________________<br> This project was funded under the grant number: ECR/2018/000204 by the Science &amp; Engineering Research Board (SERB).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Indian Art Music Tonic Datasets

<p><strong>Introduction</strong></p> <p>These datasets comprise audio excerpts and manually done&nbsp;annotations&nbsp;of the tonic pitch of the lead artist for each audio excerpt. Each excerpt is accompanied by its associated editorial metadata. These datasets can be used to develop and evaluate computational approaches for automatic tonic identification in Indian art music. These datasets have been used in several articles mentioned below.&nbsp;A majority of these datasets come from the&nbsp;<a href="http://compmusic.upf.edu/corpora">CompMusic corpora</a>&nbsp;of Indian art music, for which each recording is associated with a&nbsp;<a href="https://musicbrainz.org/doc/MusicBrainz_Identifier">MBID</a>. With the MBID other information can be obtained using the&nbsp;<a href="http://dunya.compmusic.upf.edu/developers/">Dunya API</a>. We here provide an overview of the tonic identification datasets.&nbsp;</p> <p><strong>Datasets&nbsp;</strong></p> <p>The statistics about the datasets for tonic identification is listed in the table below. These six datasets are used in [1] for a comparative evaluation. To the best of our knowledge these are the largest datasets available for tonic identification for Indian art music. These datases vary in terms of the audio quality, recording period (decade), the number of recordings for Carnatic, Hindustani, male and female singers and instrumental and vocal excerpts. For a detailed information about these datasets we refer to Chapter 3 of this&nbsp;<a href="http://mtg.upf.edu/node/3592">thesis</a>.</p> <p>All the datasets (annotations) are version controlled. To know how the features are extracted visit the companion page for the&nbsp;<a href="http://compmusic.upf.edu/node/323">publication</a>.</p> <p>The audio files corresponding to these datsets are made available on request for only research purposes. To obtain the files, please refer to <a href="https://zenodo.org/record/7342372">this Zenodo entry</a>.</p> <p><strong>Annotation Format&nbsp;</strong></p> <p>The tonic annotations are availabe both in tsv and json format.&nbsp;</p> <p>TSV: &lt;relative path to audio&gt;&lt;tab&gt;&lt;tonic(Hz)&gt;&lt;tab&gt;&lt;Carnatic or Hindustani&gt;&lt;tab&gt;&lt;artist_name&gt;&lt;tab&gt;&lt;gender of the singer&gt;&lt;vocal or instrumental&gt;&nbsp;</p> <p>JSON:&nbsp;{<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;artist&#39;: &lt;name of the lead artist if available&gt;,&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;filepath&#39;: &lt;relative path to the audio file&gt;,</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;gender&#39;: &lt;gender of the lead singer if available&gt;,</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;mbid&#39;: &lt;musicbrainz id when available&gt;,</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;tonic&#39;: &lt;tonic in Hz&gt;,</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;tradition&#39;: &lt;Hindustani or Carnatic&gt;,</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;type&#39;: &lt;vocal or instrumental&gt;<br> &nbsp;&nbsp;&nbsp;&nbsp; }</p> <p><br> where keys of the main dictionary are the filepaths to the audio files (feature path is exactly the same with a different extension of the file name).</p> <p><strong>Using this dataset</strong></p> <p>If you use this dataset in a publication, please cite:</p> <blockquote> <p>Gulati, S., Bellur, A., Salamon, J., Ranjani, H. G., Ishwar, V., Murthy, H. A., &amp; Serra, X. (2014). Automatic Tonic Identification in Indian Art Music: Approaches and Evaluation.&nbsp;<em>Journal of New Music Research</em>,&nbsp;<em>43</em>(01), 55&ndash;73.&nbsp;</p> </blockquote> <p><a href="http://hdl.handle.net/10230/25675">http://hdl.handle.net/10230/25675</a></p> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p><strong>Contact&nbsp;</strong></p> <p>If you have any questions or comments about the dataset, please feel free to email: [sankalp (dot) gulati (at) gmail (dot) com], or&nbsp;[sankalp (dot) gulati (at) upf (dot) edu]</p> <p>&nbsp;</p> <p><a href="http://compmusic.upf.edu/iam-tonic-dataset">http://compmusic.upf.edu/iam-tonic-dataset</a></p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Feb 2014View details →
zenodo32/100

Saraga: research datasets of Indian Art Music

<p><strong>Dataset introduction</strong></p> <p>This repository contains time aligned melody, rhythm, and structural annotations for two large open corpora of Indian Art Music (Carnatic and Hindustani music).</p> <p>The repository contains Carnatic and Hindustani collections in separated zip files, and each collection is organized by songs grouped by artist concerts/live performances. This organization follows the structure generated by downloading the data using the scripts available at the dataset Github repository: <a href="https://github.com/MTG/saraga">https://github.com/MTG/saraga</a>.</p> <p>Moreover, there is a part of the Carnatic collection, 168 tracks to be specific, that counts with multitrack audio files apart from the mix audio. The considered instruments are: Ghatam, Mridangam, Violin, Voice and Secondary Voice.</p> <p>&nbsp;</p> <p><strong>Annotations in the dataset</strong></p> <p>Section and tempo annotations stored as start and end timestamps together with the name of the section and tempo during the section (in a separate file). Sama annotations referring to rhythmic cycle boundaries stored as timestamps. Phrase annotations stored as timestamps and transcription of the phrases using solf&egrave;ge symbols ({S, r, R, g, G, m, M, P, d, D, n, N}). Audio features automatically extracted and stored: pitch and tonic.</p> <p>For more information about the dataset tracks and annotations, please refer to the Saraga website: <a href="https://mtg.github.io/saraga/">https://mtg.github.io/saraga/</a></p> <p>&nbsp;</p> <p><strong>Using this dataset</strong></p> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p>*Please note that you can also use this dataset through the MIRDATA library (<a href="https://github.com/mir-dataset-loaders/mirdata">https://github.com/mir-dataset-loaders/mirdata</a>), where this dataset is in the list of available datasets.</p>

opencc-by-nc-sa-4.0May 2018View details →
zenodo24/100

Indian Art Music Raga Recognition Dataset (audio)

<p>The <strong>Rāga Recognition Datasets (audio)</strong> comprise two sizable datasets, one for each music tradition: the <strong>Carnatic Music Dataset (CMD)</strong> and the <strong>Hindustani Music Dataset (HMD)</strong>. This repository contains audio files that complement the features provided in the <strong><a href="https://zenodo.org/records/7278506">Rāga Recognition (features</a>)</strong> dataset. As the audio files are copyrighted, they must be <strong>requested through this repository</strong>.</p>

restrictedcc-by-4.0Aug 2016View details →
ClinicalTrials.gov24/100

Indian Instrumental Music in Hypertension

ClinicalTrials.gov study NCT02147366. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov24/100

Music Therapy for Rehabilitation in Post-stroke Non-fluent Aphasia: the Indian Adaptation

ClinicalTrials.gov study NCT06323330. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo12/100

Indian Art Music Tonic Datasets (audio)

<p>This repository includes the audio for the Indian Art Music Tonic Datasets (<a href="https://zenodo.org/record/1257114">https://zenodo.org/record/1257114</a>).</p> <p>Refer to the Indian Art Music Tonic Datasets page to get the features and annotations for the dataset. A complete presentation of the datasets is available there.</p>

restrictedFeb 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record