Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “auto-tagging”
Contextual Tags for music auto-tagging
<p>The dataset is composed of 15 contextual tags extracted based on user's usage through created playlists in the Deezer catalog. The tags are: " car, chill, club, dance, gym, happy, night, party, relax, running, sad, sleep, summer, work, workout". For each track one or multiple tags are associated with it indicating that users listen to the track in the associated context. </p> <p>The creation of the dataset and the initial baseline of an auto-tagging model is described in the paper: Ibrahim, Karim M., Jimena Royo-Letelier, Elena V. Epure, Geoffroy Peeters, and Gaël Richard. "AUDIO-BASED AUTO-TAGGING WITH CONTEXTUAL TAGS FOR MUSIC." <em>2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</em>. IEEE, 2020.</p> <p>The dataset is composed of the SONG_ID which is the ID of the track in the Deezer catalog. Each track is labeled with each tag as either 1 (indicating a track's presence in the context) or 0 (indicating a track's absence). The 30 seconds track previews used to train the model in the paper can be accessed through the Deezer API: <a href="https://developers.deezer.com/api">https://developers.deezer.com/api</a> </p> <p>.</p>
User-aware music auto-tagging with contextual tags
<p>This is a user-aware music dataset labeled with the contextual use of each track according to each user. The dataset is composed of 10 contextual tags extracted based on user's usage through created playlists in the Deezer catalog. The tags are: " car, gym, happy, night, relax, running, sad, summer, work, workout". For each track/user pair, a contextual tag is associated with it indicating that the user listens to the track in the associated context. Additionally, the users are represented as embeddings based on their listening history computed through the matrix factorization of the user/track matrix.</p> <p>The creation of the dataset and the baseline of our auto-tagging model is described in the paper: Karim M. Ibrahim, Elena V. Epure, Geoffroy Peeters, and Gaël Richard. "Should we consider the users in contextual music auto-tagging models?" <em>21st International Society for Music Information Retrieval Conference (ISMIR)</em>. 2020. The source code of the paper is available here: <a href="https://github.com/KarimMibrahim/user-aware-music-autotagging">https://github.com/KarimMibrahim/user-aware-music-autotagging</a></p> <p>The dataset is composed of the SONG_ID which is the ID of the track in the Deezer catalog. Each track/user pair is labeled with each tag as either 1 (indicating a track's presence in the context) or 0 (indicating a track's absence). The 30 seconds track previews used to train the model in the paper can be accessed through the Deezer API: <a href="https://developers.deezer.com/api">https://developers.deezer.com/api</a>. Each user is represented with an anonymized USER_ID which is associated with the user embedding available in the user_embeddings.csv file. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.