Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “audio classification”
An Open-set Recognition and Few-Shot Learning Dataset for Audio Event Classification in Domestic Environments
<p>The problem of training a deep neural network with a small set of positive samples is known as few-shot learning (FSL). It is widely known that traditional deep learning (DL) algorithms usually show very good performance when trained with large datasets. However, in many applications, it is not possible to obtain such a high number of samples. In the image domain, typical FSL applications are those related to face recognition. In the audio domain, music fraud or speaker recognition can be clearly benefited from FSL methods. This paper deals with the application of FSL to the detection of specific and intentional acoustic events given by different types of sound alarms, such as door bells or fire alarms, using a limited number of samples. These sounds typically occur in domestic environments where many events corresponding to a wide variety of sound classes take place. Therefore, the detection of such alarms in a practical scenario can be considered an open-set recognition (OSR) problem. To address the lack of a dedicated public dataset for audio FSL, researchers usually make modifications on other available datasets. This paper is aimed at providing the audio recognition community with a carefully annotated dataset for FSL and OSR comprised of 1360 clips from 34 classes divided into pattern sounds and unwanted sounds. To facilitate and promote research in this area, results with two baseline systems (one trained from scratch and another based on transfer learning), are presented.</p> <p> </p>
BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings
<p>BirdVox-scaper-10k: a synthetic dataset for multilabel species classification of flight calls from 10-second audio recordings<br> =============================================================================================<br> Version 1.0, September 2019.</p> <p> </p> <p>Created By<br> -------------</p> <p>Elizabeth Mendoza (1), Vincent Lostanlen (2, 3, 4), Justin Salamon (3, 4), Andrew Farnsworth (2), Steve Kelling (2), and Juan Pablo Bello (3, 4).</p> <p> </p> <p>(1): Forest Hills High School, New York, NY, USA<br> (2): Cornell Lab of Ornithology, Cornell University, Ithaca, NY, USA<br> (3): Center for Urban Science and Progress, New York University, New York, NY, USA<br> (4): Music and Audio Research Lab, New York University, New York, NY, USA</p> <p>https://wp.nyu.edu/birdvox</p> <p> </p> <p>Description<br> --------------</p> <p>The BirdVox-scaper-10k dataset contains 9983 artificial soundscapes. Each soundscape lasts exactly ten seconds and contains one or several avian flight calls from up to 30 different species of New World warblers (Parulidae). Alongside each audio file, we include an annotation file describing the start time and end time of each flight call in the corresponding soundscape, as well as the species of warbler it belongs to.</p> <p>In order to synthesize soundscapes in BirdVox-scaper-10k, we mixed natural sounds from various pre-recorded sources. First, we extracted isolated recordings of flight calls containing little or no background noise from the CLO-43SD dataset [1]. Secondly, we extracted 10-second "empty" acoustic scenes from the BirdVox-DCASE-20k dataset [2]. These acoustic scenes contain various sources of real-world background noise, including biophony (insects) and anthropophony (vehicles), yet are guaranteed to be devoid of any flight calls. Lastly, we "fill" each acoustic scene by mixing it with flight calls sampled at random.</p> <p>Although the BirdVox-scaper-10k does not consist of natural recordings, we have taken several measures to ensure the plausibility of each synthesized soundscape, both from qualitative and quantitative standpoints.<br> <br> The BirdVox-scaper-10k dataset can be used, among other things, for the research, development, and testing of bioacoustic classification models.</p> <p>For details on the hardware of ROBIN recording units, we refer the reader to [2].</p> <p>[1] J. Salamon, J. Bello. Fusing shallow and deep learning for bioacoustic bird species classification. Proc. IEEE ICASSP, 2017.</p> <p>[2] V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, and J. Bello. BirdVox-full-night: a dataset and benchmark for avian flight call detection. Proc. IEEE ICASSP, 2018.</p> <p>[3] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016.</p> <p> </p> <p> </p> <p> </p> <p>@inproceedings{lostanlen2018icassp,<br> title = {BirdVox-full-night: a dataset and benchmark for avian flight call detection},<br> author = {Lostanlen, Vincent and Salamon, Justin and Farnsworth, Andrew and Kelling, Steve and Bello, Juan Pablo},<br> booktitle = {Proc. IEEE ICASSP},<br> year = {2018},<br> published = {IEEE},<br> venue = {Calgary, Canada},<br> month = {April},<br> }</p>
Auditory Scene Analysis dataset (Multichannel universal sound separation & polyphonic audio classification)
<p>We constructed a new dataset for <strong>multichannel universal sound separation</strong> and <strong>polyphonic audio classification</strong> tasks.</p> <p>We constructed a new dataset for multichannel USS and polyphonic audio classification tasks. The proposed dataset is designed to reflect various conditions, including moving sources with temporal onsets and offsets. For foreground sound sources, signals from 13 audio classes were selected from open-source databases (Pixabay and FSD50K, Librispeech, MUSDB18, Vocalsound). These signals were resampled to 16 kHz and pre-processed by either padding zeros or cropping to 4 seconds. Each sound source has a 75% probability of being a moving source, with speeds ranging from 0 to 3 m/s. The dataset features between 2 to 4 foreground sound sources, along with one background noise from the diffused TAU-SNoise dataset with a signal-to-noise ratio (SNR) ranging from 6 to 30 dB. The simulations were conducted using gpuRIR. Room dimensions were set to a width and length between 5 and 8 meters, and a height between 3 and 4 meters, with reverberation times ranging from 0.2 to 0.6 seconds. These parameters were sampled from uniform distributions. We simulated spatialized sound sources using a 4-channel tetrahedral microphone array with a radius of 4.2 cm. The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p> <p>The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p>
AVECL-UMONS DATABASE FOR AUDIO-VISUAL EVENT CLASSIFICATION AND LOCALIZATION
<p>AVECL-UMONS is a dataset for audio-visual event classification and localization in the context of office environments. The dataset is composed of 11 event classes recorded at several realistic positions in two different rooms. The dataset comprises two types of sequences according to the number of events in the sequence. 2662 unilabel sequences and 2724 multilabel sequences are recorded corresponding to a total of 5.24 hours.</p> <p>The dataset is complementary to the <em>SECL-UMons </em>dataset<em> </em>(also available on zenodo). The <em>SECL-UMons</em> dataset includes recordings made with a microphone array while <em>AVECL-UMONS</em> includes recordings made with 4 webcams.</p>
Database for Automatic Spatial Audio Scene Classification in Binaural Recordings of Music
<p>This repository contains supplementary material for the paper titled ‘<em>Automatic Spatial Audio Scene Classification in Binaural Recordings of Music</em>.’</p> <p>The database consists of the five following folders:</p> <ol> <li>Binaural recordings used for training</li> <li>Binaural recordings used for testing</li> <li>Extracted features</li> <li>Classification algorithm</li> <li>Music credits</li> </ol>
Cough Audio Classification as a TB Triage Test
ClinicalTrials.gov study NCT05317247. IPD Sharing: YES. Countries: 2. Publications: 0.
Audio Data Collection for Identification and Classification of Coughing
ClinicalTrials.gov study NCT04326309. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.