Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “Sound Classification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset-AOB: urban sounds events classification

<p>The dataset Dataset-AOB is an audio dataset collected and manually edited for urban sounds events classification using Convolutional Neural Networks for the Master Thesis:&nbsp;</p> <p>Ospina, A. &quot;Audio Event Classification using Deep Learning. Use case: Urban Sounds Events classification with Convolutional Neural Networks,&quot; M.Eng. thesis, Beuth University of Applied Sciences, Berlin, 2020.</p> <p>- 10 audio events:&nbsp;alarm-siren, children playing, dog bark, engine, footsteps, glass breaking, gun shot, metro train, rain and screams.</p> <p>- duration: &lt; 4 seconds</p> <p>- format: (.wav)</p> <p>- sampling rate: 22KHz - 44KHz</p> <p>- files: Dataset-AOB: development dataset (4831 samples), DatasetEVAL-AOB: evaluation (218 samples)</p> <p>- metadata: (.csv)</p> <p>- sources per class: (.png)</p> <p>Contact: aospinab@gmail.com</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Data from: Crowdsourcing training material for automated bird sound classification – a pilot study

<p>Data from the manuscript &quot;Crowdsourcing training material for automated bird sound classification &ndash; a pilot study&quot; by&nbsp;Petteri Lehikoinen, Meeri Rannisto, Ulisses Camargo, Aki Aintila, Patrik Lauha, Esko Piirainen, Panu Somervuo &amp; Otso Ovaskainen</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Classification of Phonocardiograms with Convolutional Neural Networks-Figure 2. PCGs for sample sound file of 103-1305031931979-B

<p>Heart sounds that provide valuable diagnostic information in clinical examinations are among the most important physiological signals in the human body. However, heart sounds include noise, such as external sounds and lung sounds, caused by signal recording conditions. Noisy heart sound signal negatively affects the diagnosis of the doctor (Denga &amp; Hanb, 2018). Digital filters are often used to filter biomedical signal. Digital filtering is defined as the acquisition of desired frequency values according to the characterization of the desired filter in order to improve the signal according to the intended use (Shenoi, 2005; Thede, 1995). Based on the experience gained from previous studies, an elliptic filter was used in this study (Deperlioglu, 2018; Guraksin et. al., 2009) . The heart sound signal covering the first two steps is given in Figure 2 for sample sound file of 103_1305031931979_B in the PASCAL Btraining data set.</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

Semi-Supervised Active Learning for Sound Classification in Hybrid Learning Environments

<p>There are 16,930 sound instances in our database with durations ranging 242 from 1 to 10 seconds, which correspond to (approximately) 15 hours of environmental 243 sounds. All sound files were converted into raw 16 bit encoding, mono-channel, and 16 244 kHz sampling rate, as various formats and rates were used in the original versions 245 retrieved from the web.</p>

opencc-zeroAug 2016View details →
zenodo36/100

Semi-Supervised Active Learning for Sound Classification in Hybrid Learning Environments

<p>There are 16,930 sound instances in our database with durations ranging 242 from 1 to 10 seconds, which correspond to (approximately) 15 hours of environmental 243 sounds. All sound files were converted into raw 16 bit encoding, mono-channel, and 16 244 kHz sampling rate, as various formats and rates were used in the original versions 245 retrieved from the web.</p>

opencc-by-4.0Aug 2016View details →
zenodo36/100

Real-Life Indoor Sound Event Dataset (ReaLISED) for Sound Event Classification (SEC)

<p>The Real-Life Indoor Sound Event Dataset (ReaLISED) offers&nbsp;the scientific community the possibility of testing Sound Event Classification (SEC) algorithms with new real indoor&nbsp;audio event recordings. The full set is made up of 2479 sound recordings of 18 events. The 18 event classes are the following:&nbsp;beater, cooking, cupboard/wardrobe,&nbsp;dishwasher, door, drawer, furniture movement, microwave, object falling, smoke extractor, speech, switch, television, vacuum cleaner, walking, washing machine, water tap, and window. There are 2479 clips of isolated sounds, which result in 3624.51 seconds.&nbsp;The number of events in each class is between 104 for the &quot;Window&quot; class and 190 for the &ldquo;Speech&rdquo; class, with a mean value of 138 events and a standard deviation of 25.</p> <p>Four Olympus LS-100 recorders&nbsp;were used. The sampling frequency was set to 44.1 kHz and 24 bits per sample. The stereo mode was used, and a medium sensitivity of the microphone was set. The distance between the recorder and the sound source was set to approximately 30-40 cm.</p> <p>Apart from the labels related to the class of event, extra information for each recording is provided in order to be exploited if necessary in the future, with other research purposes. This extra information completes the description of the sound source.</p> <p>The dataset is introduced to the scientific community by providing all the .flac files which composed it. The name of the files is built with 5 pieces of information, separated with underscores (&ldquo;_&rdquo;), with the format &ldquo;abc_123_45_67_8.flac&prime;&prime;:</p> <ul> <li> <p>&ldquo;abc&rdquo;: the first three letters indicate the source that produces the sound. This segment can take 18 different values: &lsquo;bea&rsquo; (beater), &lsquo;coo&rsquo; (cooking), &lsquo;cup&rsquo; (cupboard/wardrobe), &lsquo;dis&rsquo; (dishwasher), &lsquo;doo&rsquo; (door), &lsquo;dra&rsquo; (drawer), &lsquo;fur&rsquo; (furniture movement), &lsquo;mic&rsquo; (microwave), &lsquo;obj&rsquo; (object falling), &lsquo;smo&rsquo; (smoke extractor), &lsquo;spe&rsquo; (speech), &lsquo;swi&rsquo; (switch), &lsquo;tel&rsquo; (television), &lsquo;vac&rsquo; (vacuum cleaner), &lsquo;wal&rsquo; (walking), &lsquo;was&rsquo; (washing machine), &lsquo;wat&rsquo; (water tap), win&rsquo; (window).</p> </li> <li> <p>&ldquo;123&rdquo;: this set of digits identifies the event among the number of events produced by the source identified with &ldquo;abc&rdquo;. This segment can take all the values between &lsquo;001&rsquo; and &lsquo;190&rsquo;, which is the maximum number of events of a particular class we can find in the dataset (speech).</p> </li> <li> <p>&ldquo;45&rdquo;: this set of digits identifies the action that produce the sound. This segment can take 11 different values: &lsquo;01&rsquo; (close), &lsquo;02&rsquo; (open), &lsquo;03&rsquo; (throw), &lsquo;04&rsquo; (turn on), &lsquo;05&rsquo; (turn off), &lsquo;06&rsquo; (move), &lsquo;07&rsquo; (plug), &lsquo;08&rsquo; (unplug), &lsquo;09&rsquo; (raise), &lsquo;10&rsquo; (lower), and &lsquo;00&rsquo; (there is no information about the action).</p> </li> <li> <p>&ldquo;67&rdquo;: this set of digits identifies the material the sound source is made of. This segment can take 14 different values: &rsquo;01&rsquo; (wood), &rsquo;02&rsquo; (glass), &rsquo;03&rsquo; (metal), &rsquo;04&rsquo; (plastic), &rsquo;05&rsquo; (ceramic), &rsquo;06&rsquo; (synthetic), &rsquo;07&rsquo; (cardboard), &rsquo;08&rsquo; (marble), &rsquo;09&rsquo; (floating platform), &rsquo;10&#39;&nbsp;(platelet), &rsquo;11&rsquo; (wicker), &rsquo;12&rsquo; (carpet), &rsquo;13&rsquo; (medium-density fibreboard MDF), and &rsquo;00&rsquo; (there is no information about the material).</p> </li> <li> <p>&ldquo;8&rdquo;: the last digit gives approximate information about the intensity of the recorded sound. It can take 4 different values: &rsquo;1&rsquo; (low intensity), &rsquo;2&rsquo; (medium intensity), &rsquo;3&rsquo; (high intensity), &rsquo;0&rsquo; (there ir no information about the intensity).</p> <p>For clarity, some examples of audio file&nbsp;names with this code are shown hereunder:</p> </li> <li> <p>&ldquo;doo_040_02_00_3.flac&rdquo; is the name of the 40th file in the Door class, described as &ldquo;opening a door of unknown material with high intensity&rdquo;.</p> </li> <li> <p>&ldquo;fur_058_06_01_2.flac&rdquo; is the name of the 58th file in the furniture movement class, described as &ldquo;moving a wooden furniture with medium intensity&rdquo;.</p> </li> <li> <p>&ldquo;vac_001_00_00_0.flac&rdquo; is the name of the 1st audio file in the vacuum cleaner class, described as &ldquo;using the vacuum cleaner, without information about the action, neither the material or the intensity&rdquo;.</p> </li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Auditory Scene Analysis dataset (Multichannel universal sound separation & polyphonic audio classification)

<p>We constructed a new dataset for <strong>multichannel universal sound separation</strong> and <strong>polyphonic audio classification</strong> tasks.</p> <p>We constructed a new dataset for multichannel USS and polyphonic audio classification tasks. The proposed dataset is designed to reflect various conditions, including moving sources with temporal onsets and offsets. For foreground sound sources, signals from 13 audio classes were selected from open-source databases (Pixabay and FSD50K, Librispeech, MUSDB18, Vocalsound). These signals were resampled to 16 kHz and pre-processed by either padding zeros or cropping to 4 seconds. Each sound source has a 75% probability of being a moving source, with speeds ranging from 0 to 3 m/s. The dataset features between 2 to 4 foreground sound sources, along with one background noise from the diffused TAU-SNoise dataset with a signal-to-noise ratio (SNR) ranging from 6 to 30 dB. The simulations were conducted using gpuRIR. Room dimensions were set to a width and length between 5 and 8 meters, and a height between 3 and 4 meters, with reverberation times ranging from 0.2 to 0.6 seconds. These parameters were sampled from uniform distributions. We simulated spatialized sound sources using a 4-channel tetrahedral microphone array with a radius of 4.2 cm. The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p> <p>The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

SLoClas: A Database for Joint Sound Localization and Classification

<p>We present a new database namely Sound Localization and Classification (SLoClas) corpus, for studying and analyzing&nbsp; sound localization and classification. The corpus contains a total of 23.27 hours of data recorded using a 4-channel microphone array. 10 classes of sounds are played over a loudspeaker at 1.5 meters distance from the array by varying the DoA from 1 degree to 360 degree at an interval of 5 degree. To facilitate the study of noise robustness,&nbsp; 6&nbsp; types of outdoor noise are recorded at 4 DoAs, using the same devices.</p> <p>We release this database for research purpose only. If you use this corpus, cite the following paper</p> <p>Qian Xinyuan, Bidisha Sharma, Amine El Abridi, and Haizhou Li. &quot;SLoClas: A Database for Joint Sound Localization and Classification.&quot; <em>arXiv preprint arXiv:2108.02539</em> (2021).</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

<p>Occupancy modeling is used to evaluate avian distributions and habitat associations, yet it typically requires extensive survey effort because a minimum of three repeat samples are required for accurate parameter estimation. Autonomous recording units (ARUs) can reduce the need for surveyors on site, yet ARUs utility were limited by hardware costs and the time required to manually annotate recordings. Software that identifies bird vocalizations may reduce expert time needed, if classification is sufficiently accurate. We assessed the performance of BirdNET – an automated classifier capable of identifying vocalizations from &gt;900 North American and European bird species – by comparing automated to manual annotations of recordings of 13 breeding bird species collected in northwestern California. We compared the parameter estimates of occupancy models evaluating habitat associations supplied with manually annotated data (9 min recording segments) to output from models supplied with BirdNET detections. We used three sets of BirdNET output to evaluate the duration of automatic annotation needed to approach manually annotated model parameter estimates: 9-min, 87-min, and 87-min of high-confidence detections. We incorporated 100 3-sec manually validated BirdNET detections per species to estimate true and false positive rates within an occupancy model. BirdNET correctly identified 90% and 65% of the bird species a human detected when data were restricted to detections exceeding a low or high confidence score threshold, respectively. Occupancy estimates, including habitat associations, were similar regardless of method. Precision (proportion of true positives to all detections) was &gt;0.70 for 9 of 13 species, and a low of 0.29. However, processing of longer recordings was needed to rival manually annotated data. We conclude that BirdNET is suitable for annotating multispecies recordings for occupancy modeling when extended recording durations are used. Together, ARUs and BirdNET may benefit monitoring and, ultimately, conservation of bird populations by greatly increasing monitoring opportunities.   </p>

opencc-zeroFeb 2022View details →
zenodo32/100

Ambient sound classification dataset - UC CRLAB

<p>This dataset contains three classes: Alarmed, Social, and Disengaged. It contains audio file of a number of 925 for Disengaged, 941 for Social, and 942 for Alarmed.</p>

restrictedcc-by-4.0Sep 2024View details →
zenodo32/100

SOUND-BASED DRONE FAULT CLASSIFICATION USING MULTI-TASK LEARNING

<p>arxiv :&nbsp;https://arxiv.org/abs/2304.11708</p> <p>Accepted at 29th International Congress on Sound and Vibration (ICSV29).&nbsp;</p> <p>The drone has been used for various purposes including military applications, aerial photography, and pesticide spraying. However, the drone is vulnerable to external disturbances, and malfunction in propellers and motors can easily occur. To improve the safety of drone operations, early detection of mechanical faults should be made in real-time. In this paper, we propose a sound-based deep neural network (DNN) fault classifier and drone sound dataset. The dataset was constructed by collecting the operating sounds of drones from microphones mounted on three different drones in an anechoic chamber. The dataset includes various operating conditions of drones, such as flight directions (front, back, right, left, clockwise, counter clockwise) and faults on propellers and motors. The drone sounds were then mixed with noises recorded in five different spots on the university campus, with a signal-to-noise ratio (SNR) varying from 10 dB to 15 dB. Using the acquired dataset, we train a DNN classifier, 1DCNN-ResNet, that classifies the types of mechanical faults and their locations from short-time input waveforms. We employ multitask learning (MTL) and incorporate the direction classification task as an auxiliary task to make the classifier learn more general audio features. The test over unseen data reveals that the proposed multitask model can successfully classify faults in drones and outperforms single-task models even with less training data.&nbsp;</p> <p>&nbsp;</p> <p>please reorganize the file directory like below</p> <p>drone</p> <p>ㄴA</p> <p>ㄴB</p> <p>ㄴC</p> <p>&nbsp;</p> <p>For each drone type A, B, and C have 54000*2 files. (Here, *2 means stereo channel, you can find mic1 and mic2 in subdirectory)&nbsp;They are divided into train, valid, and test by a 6:2:2 ratio. For each file, recording information is labeled below.</p> <p>{model_type}_{maneuvering_direction}_{fault}_{drone_file_index}_{background}_{background_file_index}_{SNR}</p> <p>model_type: A, B, C</p> <p>maneuvering_direction: F(Front), B(Back), R(Right), L(Left), C(Clockwise), CC(Counter-clockwise)</p> <p>fault: N (Normal), MF1~4&nbsp;(Moter Failure), PC1~4 (Propeller Cut) -&gt; 1~4 means each motor/propeller of the quadcopter.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

Open the record for dataset details and reuse information.

publicFeb 2022View details →
zenodo28/100

Soundscape Datasets for Few-Shot Bird Sound Classification

<p>This repository provides easy access to open-source soundscape datasets of bird sounds, specifically optimized for few-shot classification.</p> <p><code>soundscapes.zip</code> contains evaluation soundscape datasets from the BIRB benchmark (https://arxiv.org/abs/2312.07439), downsampled to 16kHz, preprocessed using CNN14 from PANNs (https://arxiv.org/abs/1912.10211), to select a 6-second window with the highest bird activation, and converted to Pytorch (.pt) format to facilitate usability for evaluating deep neural networks.&nbsp;</p> <p>These preprocessed datasets are employed in the work "<em>Domain-Invariant Representation Learning of Bird Sounds</em>" (https://arxiv.org/abs/2409.08589), which evaluates the few-shot learning capabilities of deep learning models trained on focal recordings (e.g., Xeno-Canto) and tested on soundscape recordings.</p> <h2>Dataset Structure</h2> <h3>Validation Dataset</h3> <ul> <li><strong>POW (</strong><code>pow.pt</code><strong>):&nbsp;</strong>The validation dataset consists of 16,047 examples across 43 classes and is organized as a dictionary with&nbsp;<code>'data'</code> and<code> 'label'</code> keys representing bird sounds and their corresponding labels. Storing the entire validation dataset in a single tensor enables rapid loading and efficient processing, significantly accelerating the validation process. Classes with only one example are removed, as they are insufficient for one-shot classification tasks. Source: https://zenodo.org/records/4656848#.Y7ijhOxudhE</li> </ul> <h3>Test Datasets&nbsp;</h3> <p>Each test dataset is structured with multiple subfolders, each labeled with an eBird species code to represent data for a specific bird species.</p> <ul> <li><strong>SSW (</strong><code>ssw/</code><strong>):</strong> Contains 50,760 examples across 96 classes. Source: https://zenodo.org/records/7079380#.Y7ijHOxudhE</li> <li><strong>NES (</strong><code>coffee_farms/</code><strong>):</strong> Contains 6,952 examples across 89 classes. Source: https://zenodo.org/records/7525349#.ZB8z_-xudhE</li> <li><strong>UHH (</strong><code>hawaii/</code><strong>):</strong> Contains 59,583 examples across 27 classes. Source: https://zenodo.org/records/7078499#.Y7ijPuxudhE</li> <li><strong>HSN (</strong><code>high_sierras/</code><strong>):</strong> Contains 10,296 examples across 19 classes. Source: https://zenodo.org/records/7525805#.ZB8zsexudhE</li> <li><strong>SNE (</strong><code>sierras_kahl/</code><strong>):</strong> Contains 20,147 examples across 56 classes. Source: https://zenodo.org/records/7050014#.Y7ijWexudhE</li> <li><strong>PER (</strong><code>peru/</code><strong>):</strong> Contains 14,768 examples across 132 classes. Source: https://zenodo.org/records/7079124#.Y7iis-xudhE</li> </ul> <p>Code and detailed instructions, including data loading, model implementation, and few-shot evaluation, can be found at: https://github.com/ilyassmoummad/ProtoCLR</p>

opencc-by-nc-4.0Oct 2024View details →
zenodo24/100

Convolutional Neural Networks for Scops Owl Sound Classification

<p>These are the training, validation, and testing&nbsp;datasets used in a publication titled &quot;Convolutional Neural Networks for Scops Owl Sound Classification&quot; . Each set consists&nbsp;of WAV&nbsp;files of seven&nbsp;Indonesian scops owl species. If you use this dataset for your research publication please cite the following paper:&nbsp;<a href="https://doi.org/10.1016/j.procs.2020.12.010">https://doi.org/10.1016/j.procs.2020.12.010</a></p>

opencc-by-4.0Nov 2022View details →
ClinicalTrials.gov24/100

A Study of Breathing Sound-based Classification of Patients With Breathing Disorders

ClinicalTrials.gov study NCT05868694. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Feasibility of AI-based Classification of Normal, Wheeze and Crackle Sounds From Stethoscope in Clinical Settings

ClinicalTrials.gov study NCT05268263. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
zenodo20/100

SECL-UMONS DATABASE FOR SOUND EVENT CLASSIFICATION AND LOCALIZATION

<p>SECL-UMons is a dataset for sound event classification and localization in the context of office environments. The multichannel dataset is composed of 11 event classes recorded at several realistic positions in two different rooms. The dataset comprises two types of sequences according to the number of events in the sequence. 2662 unilabel sequences and 2724 multilabel sequences are recorded corresponding to a total of 5.24 hours.</p>

openother-ncDec 2019View details →
ClinicalTrials.gov20/100

Clinical Evaluation of AI-aided Auscultation With Automatic Classification of Respiratory System Sounds

ClinicalTrials.gov study NCT04208360. IPD Sharing: NO. Countries: 0. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record