Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “multilabel dataset”
A multilabel dataset for distinguishing Bosnian, Croatian, Montenegrin, and Serbian
<p>This dataset contains files used in the VarDial 2024 Shared Task on Distinguishing Between Similar Languages - Multiple Labels for the Bosnian - Croatian - Montenegrin - Serbian (BCMS) subtask.</p> <p>The starting point for this dataset is the one published by Rupnik et al. (2023). It contains geolocated data from the BCMS linguistic area collected from Twitter (rebranded as X).<br>Each instance contains the full tweet production of a single user, which was manually annotated for the user's country.<br>The original annotation was single-label, and it was produced by a single annotator. In the version of the data produced here, the test and dev sets were reannotated by multiple annotators, in a multi-label setting. For the details on the reannotation process, please see Miletić and Miletić (2024). We have also excluded retweets from the original data, as these represent reproduced content from a different user account and may not be representative of the language use of the user themselves.</p> <p>For details on the shared task, we refer you to Chifu et al. (2024).</p>
BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings
<p>BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings<br> =============================================================================================<br> Version 1.0, September 2019.</p> <p> </p> <p>Created By<br> -------------</p> <p>Elizabeth Mendoza (1), Vincent Lostanlen (2, 3, 4), Justin Salamon (3, 4), Andrew Farnsworth (2), Steve Kelling (2), and Juan Pablo Bello (3, 4).</p> <p> </p> <p>(1): Forest Hills High School, New York, NY, USA<br> (2): Cornell Lab of Ornithology, Cornell University, Ithaca, NY, USA<br> (3): Center for Urban Science and Progress, New York University, New York, NY, USA<br> (4): Music and Audio Research Lab, New York University, New York, NY, USA</p> <p>https://wp.nyu.edu/birdvox</p> <p> </p> <p>Description<br> --------------</p> <p>The BirdVox-scaper-10k dataset contains 9983 artificial soundscapes. Each soundscape lasts exactly ten seconds and contains one or several avian flight calls from up to 30 different species of New World warblers (Parulidae). Alongside each audio file, we include an annotation file describing the start time and end time of each flight call in the corresponding soundscape, as well as the species of warbler it belongs to.</p> <p>In order to synthesize soundscapes in BirdVox-scaper-10k, we mixed natural sounds from various pre-recorded sources. First, we extracted isolated recordings of flight calls containing little or no background noise from the CLO-43SD dataset [1]. Secondly, we extracted 10-second "empty" acoustic scenes from the BirdVox-DCASE-20k dataset [2]. These acoustic scenes contain various sources of real-world background noise, including biophony (insects) and anthropophony (vehicles), yet are guaranteed to be devoid of any flight calls. Lastly, we "fill" each acoustic scene by mixing it with flight calls sampled at random.</p> <p>Although the BirdVox-scaper-10k does not consist of natural recordings, we have taken several measures to ensure the plausibility of each synthesized soundscape, both from qualitative and quantitative standpoints.<br> <br> The BirdVox-scaper-10k dataset can be used, among other things, for the research, development, and testing of bioacoustic classification models.</p> <p>For details on the hardware of ROBIN recording units, we refer the reader to [2].</p> <p>[1] J. Salamon, J. Bello. Fusing shallow and deep learning for bioacoustic bird species classification. Proc. IEEE ICASSP, 2017.</p> <p>[2] V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, and J. Bello. BirdVox-full-night: a dataset and benchmark for avian flight call detection. Proc. IEEE ICASSP, 2018.</p> <p>[3] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016.</p> <p> </p> <p> </p> <p> </p> <p>@inproceedings{lostanlen2018icassp,<br> title = {BirdVox-full-night: a dataset and benchmark for avian flight call detection},<br> author = {Lostanlen, Vincent and Salamon, Justin and Farnsworth, Andrew and Kelling, Steve and Bello, Juan Pablo},<br> booktitle = {Proc. IEEE ICASSP},<br> year = {2018},<br> published = {IEEE},<br> venue = {Calgary, Canada},<br> month = {April},<br> }</p>
SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network
<p><strong>SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network</strong></p> <p>Version 2.3, September 2020</p> <p> </p> <p><strong>Created by</strong></p> <p>Mark Cartwright (1,2,3), Jason Cramer (1), Ana Elisa Mendez Mendez (1), Yu Wang (1), Ho-Hsiang Wu (1), Vincent Lostanlen (1,2,4), Magdalena Fuentes (1), Graham Dove (2), Charlie Mydlarz (1,2), Justin Salamon (5), Oded Nov (6), Juan Pablo Bello (1,2,3)</p> <ol> <li>Music and Audio Research Lab, New York University</li> <li>Center for Urban Science and Progress, New York University</li> <li>Department of Computer Science and Engineering, New York University</li> <li>Cornell Lab of Ornithology</li> <li>Adobe Research</li> <li>Department of Technology Management and Innovation, New York University</li> </ol> <p> </p> <p><strong>Publication</strong></p> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Cartwright, M., Cramer, J., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In <em>Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE)</em>, 2020.<br> <a href="https://arxiv.org/abs/2009.05188">[pdf]</a></p> <p> </p> <p><strong>Description</strong></p> <p>SONYC Urban Sound Tagging (SONYC-UST) is a dataset for the development and evaluation of machine listening systems for realistic urban noise monitoring. The audio was recorded from the <a href="https://wp.nyu.edu/sonyc">SONYC</a> acoustic sensor network. Volunteers on the <a href="https://zooniverse.org">Zooniverse</a> citizen science platform tagged the presence of 23 classes that were chosen in consultation with the New York City Department of Environmental Protection. These 23 fine-grained classes can be grouped into 8 coarse-grained classes. The recordings are split into three sets: training, validation, and test. The training and validation sets are disjoint with respect to the sensor from which each recording came, and the test set is displaced in time. For increased reliability, three volunteers annotated each recording. In addition, members of the SONYC team subsequently created a subset of verified, ground-truth tags using a two-stage annotation procedure in which two annotators independently tagged and then collectively resolved any disagreements. This subset of recordings with verified annotations intersects with all three recording splits. All of the recordings in the test set have these verified annotations. In v2 version of this dataset, we have also included coarse spatiotemporal context information to aid in tag prediction when time and location is known. For more details on the motivation and creation of this dataset see the <a href="http://dcase.community/challenge2020/task-urban-sound-tagging-with-spatiotemporal-context">DCASE 2020 Urban Sound Tagging with Spatiotemporal Context Task website</a>.</p> <p> </p> <p><strong>Audio data</strong></p> <p>The provided audio has been acquired using the SONYC acoustic sensor network for urban noise pollution monitoring. Over 60 different sensors have been deployed in New York City, and these sensors have collectively gathered the equivalent of over 50 years of audio data, of which we provide a small subset. The data was sampled by selecting the nearest neighbors on VGGish features of recordings known to have classes of interest. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings. To maintain privacy, we quantized the spatial information to the level of a city block, and we quantized the temporal information to the level of an hour. We also limited the occurrence of recordings with positive human voice annotations to one per hour per sensor.</p> <p> </p> <p><strong>Label taxonomy</strong></p> <p>The label taxonomy is as follows:</p> <ol> <li>engine<br> 1: small-sounding-engine<br> 2: medium-sounding-engine<br> 3: large-sounding-engine<br> X: engine-of-uncertain-size</li> <li>machinery-impact<br> 1: rock-drill<br> 2: jackhammer<br> 3: hoe-ram<br> 4: pile-driver<br> X: other-unknown-impact-machinery</li> <li>non-machinery-impact<br> 1: non-machinery-impact</li> <li>powered-saw<br> 1: chainsaw<br> 2: small-medium-rotating-saw<br> 3: large-rotating-saw<br> X: other-unknown-powered-saw</li> <li>alert-signal<br> 1: car-horn<br> 2: car-alarm<br> 3: siren<br> 4: reverse-beeper<br> X: other-unknown-alert-signal</li> <li>music<br> 1: stationary-music<br> 2: mobile-music<br> 3: ice-cream-truck<br> X: music-from-uncertain-source</li> <li>human-voice<br> 1: person-or-small-group-talking<br> 2: person-or-small-group-shouting<br> 3: large-crowd<br> 4: amplified-speech<br> X: other-unknown-human-voice</li> <li>dog<br> 1: dog-barking-whining</li> </ol> <p>The classes preceded by an <code>X</code> code indicate when an annotator was able to identify the coarse class, but couldn’t identify the fine class because either they were uncertain which fine class it was or the fine class was not included in the taxonomy. <code>dcase-ust-taxonomy.yaml</code> contains this taxonomy in an easily machine-readable form.</p> <p> </p> <p><strong>Data splits</strong></p> <p>This release contains a training subset (13538 recordings from 35 sensors), and validation subset (4308 recordings from 9 sensors), and a test subset (669 recordings from 48 sensors). The training and validation subsets are disjoint with respect to the sensor from which each recording came. The sensors in the test set will not disjoint from the training and validation subsets, but the test recordings are displaced in time, occurring after any of the recordings in the training and validation subset. The subset of recordings with verified annotations (1380 recordings) intersects with all three recording splits. All of the recordings in the test set have these verified annotations.</p> <p> </p> <p><strong>Annotation data</strong></p> <p>The annotation data are contained in <code>annotations.csv</code>, and encompass the training, validation, and test subsets. Each row in the file represents one multi-label annotation of a recording—it could be the annotation of a single citizen science volunteer, a single SONYC team member, or the agreed-upon ground truth by the SONYC team (see the <em>annotator_id</em> column description for more information). Note that since the SONYC team members annotated each class group separately, there may be multiple annotation rows by a single SONYC team annotator for a particular audio recording.</p> <p> </p> <p> </p> <p><strong>Columns</strong></p> <p><em>split</em></p> <p>The data split. (<em>train</em>, <em>validate, test</em>)</p> <p><em>sensor_id</em></p> <p>The ID of the sensor the recording is from.</p> <p><em>audio_filename</em></p> <p>The filename of the audio recording</p> <p><em>annotator_id</em></p> <p>The anonymous ID of the annotator. If this value is positive, it is a citizen science volunteer from the Zooniverse platform. If it is negative, it is a SONYC team member. If it is <code>0</code>, then it is the ground truth agreed-upon by the SONYC team.</p> <p><em>year</em></p> <p>The year the recording is from.</p> <p><em>week</em></p> <p>The week of the year the recording is from.</p> <p><em>day</em></p> <p>The day of the week the recording is from, with Monday as the start (i.e. <code>0</code>=Monday).</p> <p><em>hour</em></p> <p>The hour of the day the recording is from</p> <p><em>borough</em><br> The NYC borough in which the sensor is located (<code>1</code>=Manhattan, <code>3</code>=Brooklyn, <code>4</code>=Queens). This corresponds to the first digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>block</em></p> <p>The NYC block in which the sensor is located. This corresponds to digits 2—6 digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>latitude</em></p> <p>The latitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em>longitude</em></p> <p>The longitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_presence</em></p> <p>Columns of this form indicate the presence of fine-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset.</p> <p><em><coarse_id>_<coarse_name>_presence</em></p> <p>Columns of this form indicate the presence of a coarse-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset. These columns are computed from the fine-level class presence columns and are presented here for convenience when training on only coarse-level classes.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_proximity</em></p> <p>Columns of this form indicate the proximity of a fine-level class. After indicating the presence of a fine-level class, citizen science annotators were asked to indicate the proximity of the sound event to the sensor. Only the citizen science volunteers performed this task, and therefore this data is not included in the verified annotations. This column may take on one of the following four values: (<code>near</code>, <code>far</code>, <code>notsure</code>, <code>-1</code>). If <code>-1</code>, then the proximity was not annotated because either the annotation was not performed by a citizen science volunteer, or the citizen science volunteer did not indicate the presence of the class.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>Dataset created by Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez, Yu Wang, Ho-Hsiang Wu, Vincent Lostanlen, Magdalena Fuentes, Graham Dove, Charlie Mydlarz, Justin Salamon, Oded Nov, and Juan Pablo Bello</p> <p>The SONYC-UST dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-UST dataset or any part of it.</p> <p> </p> <p><strong>Feedback</strong></p> <p>Please help us improve SONYC-UST by sending your feedback to:</p> <ul> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>We would like to thank all the Zooniverse volunteers who continue to contribute to our project. This work is supported by <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">National Science Foundation award 1544753</a>.</p> <p> </p> <p><strong>Change log</strong></p> <ul> <li>2.3 Added the ground truth annotations for the test set, and regrouped the audio files for upload to Zenodo.</li> <li>2.2 Added the audio for the test set (audio-eval.tar.gz).</li> <li>2.1 The DCASE 2020 development dataset. 14778 new recordings added along with coarse spatiotemporal context information.</li> <li>1.0 Data is the same as v0.4. Publication added to README.</li> <li>0.4 Fixed error in annotations. Previously, the coarse class "machinery-impact" was accidentally indicated as present whenever "non-machinery-impact" was present regardless of the presence of "machinery-impact". This error has been fixed.</li> <li>0.3 Test set annotations added</li> <li>0.2 Test set audio files added</li> </ul>
multilabel datasets
<p>multilabel datasets</p>
MN-DS: A Multilabeled News Dataset for News Articles Hierarchical Classification
<p><strong>Overview</strong></p> <p>This dataset contains 10,917 news articles with hierarchical news categories collected between January 1st 2019, and December 31st 2019 classified by using NewsCodes Media Topic taxonomy. We manually labelled the articles based on a hierarchical taxonomy with 17 first-level and 109 second-level categories.</p> <p>This dataset can be used to train machine learning models for automatically classifying news articles by topic. This dataset can be helpful for researchers working on news structuring, classification, and predicting future events based on released news.</p> <p><strong>Reproducibility of results</strong></p> <p>The results presented in the research paper "MN-DS: A Multilabeled News Dataset for News Articles Hierarchical Classification", technical validation can be reproduced using functions in <a href="https://github.com/alinapetukhova/mn-ds-news-classification">github repository</a>.</p> <p><strong>Licenses</strong></p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a> license (see `LICENSE_DATA.txt`).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.