Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “multilabel classification”
BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings
<p>BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings<br> =============================================================================================<br> Version 1.0, September 2019.</p> <p> </p> <p>Created By<br> -------------</p> <p>Elizabeth Mendoza (1), Vincent Lostanlen (2, 3, 4), Justin Salamon (3, 4), Andrew Farnsworth (2), Steve Kelling (2), and Juan Pablo Bello (3, 4).</p> <p> </p> <p>(1): Forest Hills High School, New York, NY, USA<br> (2): Cornell Lab of Ornithology, Cornell University, Ithaca, NY, USA<br> (3): Center for Urban Science and Progress, New York University, New York, NY, USA<br> (4): Music and Audio Research Lab, New York University, New York, NY, USA</p> <p>https://wp.nyu.edu/birdvox</p> <p> </p> <p>Description<br> --------------</p> <p>The BirdVox-scaper-10k dataset contains 9983 artificial soundscapes. Each soundscape lasts exactly ten seconds and contains one or several avian flight calls from up to 30 different species of New World warblers (Parulidae). Alongside each audio file, we include an annotation file describing the start time and end time of each flight call in the corresponding soundscape, as well as the species of warbler it belongs to.</p> <p>In order to synthesize soundscapes in BirdVox-scaper-10k, we mixed natural sounds from various pre-recorded sources. First, we extracted isolated recordings of flight calls containing little or no background noise from the CLO-43SD dataset [1]. Secondly, we extracted 10-second "empty" acoustic scenes from the BirdVox-DCASE-20k dataset [2]. These acoustic scenes contain various sources of real-world background noise, including biophony (insects) and anthropophony (vehicles), yet are guaranteed to be devoid of any flight calls. Lastly, we "fill" each acoustic scene by mixing it with flight calls sampled at random.</p> <p>Although the BirdVox-scaper-10k does not consist of natural recordings, we have taken several measures to ensure the plausibility of each synthesized soundscape, both from qualitative and quantitative standpoints.<br> <br> The BirdVox-scaper-10k dataset can be used, among other things, for the research, development, and testing of bioacoustic classification models.</p> <p>For details on the hardware of ROBIN recording units, we refer the reader to [2].</p> <p>[1] J. Salamon, J. Bello. Fusing shallow and deep learning for bioacoustic bird species classification. Proc. IEEE ICASSP, 2017.</p> <p>[2] V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, and J. Bello. BirdVox-full-night: a dataset and benchmark for avian flight call detection. Proc. IEEE ICASSP, 2018.</p> <p>[3] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016.</p> <p> </p> <p> </p> <p> </p> <p>@inproceedings{lostanlen2018icassp,<br> title = {BirdVox-full-night: a dataset and benchmark for avian flight call detection},<br> author = {Lostanlen, Vincent and Salamon, Justin and Farnsworth, Andrew and Kelling, Steve and Bello, Juan Pablo},<br> booktitle = {Proc. IEEE ICASSP},<br> year = {2018},<br> published = {IEEE},<br> venue = {Calgary, Canada},<br> month = {April},<br> }</p>
MN-DS: A Multilabeled News Dataset for News Articles Hierarchical Classification
<p><strong>Overview</strong></p> <p>This dataset contains 10,917 news articles with hierarchical news categories collected between January 1st 2019, and December 31st 2019 classified by using NewsCodes Media Topic taxonomy. We manually labelled the articles based on a hierarchical taxonomy with 17 first-level and 109 second-level categories.</p> <p>This dataset can be used to train machine learning models for automatically classifying news articles by topic. This dataset can be helpful for researchers working on news structuring, classification, and predicting future events based on released news.</p> <p><strong>Reproducibility of results</strong></p> <p>The results presented in the research paper "MN-DS: A Multilabeled News Dataset for News Articles Hierarchical Classification", technical validation can be reproduced using functions in <a href="https://github.com/alinapetukhova/mn-ds-news-classification">github repository</a>.</p> <p><strong>Licenses</strong></p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a> license (see `LICENSE_DATA.txt`).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.