Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

37 results for “sound event”

Learn how ShareScore rates datasets ↗
zenodo32/100

TUT Sound events 2016, Development dataset

<p>TUT Sound events 2016, development dataset consists of 22 audio recordings from two acoustic scenes:</p> <ul> <li>Home (indoor), 10 recordings, totaling 36:16</li> <li>Residential area (outdoor), 12 recordings, totalling 42:00</li> </ul>

openother-ncFeb 2016View details →
zenodo32/100

TUT Sound events 2017, Development dataset

<p>TUT Sound events 2017, development dataset consists of 24 audio recordings from a single acoustic scene:</p> <ul> <li>&nbsp;Street (outdoor), totaling 1:32:08</li> </ul>

openother-ncMar 2017View details →
zenodo32/100

TAU Sound Events and Speech Privacy Preservation

<div>The TAU Sound Events and Speech Privacy Preservation Dataset is a collection of audio data used in the work "Adversarial Representation Learning for Robust Privacy Preservation in Audio" by S. Gharib, M. Tran, D. Luong, K. Drossos, and T. Virtanen. The dataset is created by merging subsets of the <a href="../records/4060432">Freesound 50k Dataset (FSD50K)</a> and the <a href="https://www.openslr.org/12">LibriSpeech corpus</a>. Both FSD50K and LibriSpeech are licensed under the Creative Commons license.</div> <div>The dataset contains of ~5000 one-second sound event samples with or without speech content provided in WAV and NumPy array format (approximately half of the samples contains speech). The creation of the dataset ensures an equal number of samples for male and female speakers across each sound event class. The sound event classes included in this dataset are:</div> <ul> <li>dog barking</li> <li>glass breaking</li> <li>gun shot</li> <li>cough</li> <li>slam</li> <li>applause</li> <li>dished pot pan</li> <li>toilet flush</li> <li>cat meowing</li> <li>doorbell</li> <li>crying</li> <li>drill</li> </ul> <p>Please check the README for better understanding of the dataset.</p>

opencc-by-nc-4.0Dec 2023View details →
zenodo32/100

TUT Sound events 2016, Evaluation dataset

<p>TUT Sound events 2016, evaluation dataset consists of 10 audio recordings from two acoustic scenes:</p> <ul> <li>Home (indoor), 5 recordings, totaling 17:49</li> <li>Residential area (outdoor), 5 recordings, totaling 17:49</li> </ul>

openother-ncSep 2017View details →
zenodo32/100

TUT Sound events 2017, Evaluation dataset

<p>TUT Sound events 2017, evaluation dataset consists of 8 audio recordings from a single acoustic scene:</p> <ul> <li>&nbsp;Street (outdoor), totaling 29:09</li> </ul>

openother-ncOct 2017View details →
zenodo32/100

TUT Sound Events 2018 - Circular array, Reverberant and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Circular array, Reverberant and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated, reverberant, and circular-array format recordings with&nbsp;stationary point sources each associated with a spatial coordinate.&nbsp;The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with&nbsp;reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz&ndash;4 kHz band center frequencies.</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least&nbsp;1 meter away from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Anechoic and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Anechoic, and Synthetic Impulse Response Dataset&nbsp;</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of [1 10] meters from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Circular array, Anechoic and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Circular array, Anechoic and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated anechoic circular-array format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of [1 10] meters from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated reverberant first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate.&nbsp;The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with&nbsp;reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz&ndash;4 kHz band center frequencies.</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least&nbsp;1 meter away from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT)&nbsp;Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset</strong></p> <p>This dataset consists of real-life first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset was&nbsp;generated by collecting impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing<br> a maximum length sequence around the array in circular trajectory in one elevation at a time. The playback volume was set to be 30 dB greater than the ambient sound level. The recording was done in a corridor inside the university with classrooms around it during work hours.The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20&nbsp;to 20&nbsp;with 10-degree increments at 2 m.&nbsp;</p> <p>The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://serv.cusp.nyu.edu/projects/urbansounddataset/urbansound8k.html">urbansound8k dataset</a>.&nbsp;This dataset consists of 10 sound event classes such as air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. We do not consider the air_conditioner and children_playing sound events. Further, we only include the sound event examples marked as foreground in the dataset. We used the splits 1, 8 and 9 provided in the urbansound8k as the three CV splits. These splits were chosen as they had a good number of examples for all the chosen sound event classes after selecting only the foreground examples.&nbsp;During the sound scene synthesis, we randomly chose a sound event example and associated it with a random distance among the collected ones, azimuth and elevation angle. The sound event example was then convolved with the respective IR for the given distance, azimuth and elevation to spatially position it.</p> <p>The metadata.zip folder consists of the license and the metadata for the complete dataset. The rest of the nine zip files consists dataset for given split and overlap. For example, the&nbsp;wav_ov3_split1_30db.zip file consists of training and testing recordings for the case of maximum three temporally overlapping sound events (ov3) for the first cross-validation split (split1). Within each audio folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p><strong>Data collector (s): </strong>Fagerlund, Eemi;&nbsp;Koskimies, Aino</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Development Datasets

<p>This package consists of two development datasets, <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;providing four-channel First-Order Ambisonic (FOA) recordings while <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These development datasets are part of the <a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>.</p> <p>Both the development set consists of 400, one minute long recordings sampled at 48000 Hz, and divided into four cross-validation splits of 100 recordings each. These recordings were synthesized using spatial room impulse response (IRs) collected from five indoor locations, at 504 unique combinations of azimuth-elevation-distance. Furthermore, in order to synthesize the recordings, the collected IRs were convolved with <a href="http://www.cs.tut.fi/sgn/arg/dcase2016/task-sound-event-detection-in-synthetic-audio#audio-dataset">isolated sound events dataset from DCASE 2016 task 2</a>. Finally, to create a realistic sound scene recording, natural ambient noise collected in the IR recording locations was added to the synthesized recordings such that the average SNR of the sound events was 30 dB.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <p><strong>Download instructions</strong></p> <p>The three files, &nbsp;<strong><em>foa_dev.z01</em></strong>,<strong><em> foa_dev.z02</em></strong>&nbsp;and <strong><em>foa_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;development dataset.<br> The two files, <strong><em>mic_dev.z01</em></strong>&nbsp;and, <strong><em>mic_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development dataset.<br> The <strong><em>metadata_dev.zip</em></strong>&nbsp;is the common metadata for both <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development datasets.</p> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> &nbsp;</p>

openother-ncFeb 2019View details →
zenodo32/100

TAU Moving Sound Events 2019 - Ambisonic, Anechoic, Synthetic IR and Moving Source Dataset

<p><strong>Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Anechoic and Synthetic Impulse Response (IR) and Moving Source Dataset</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location and directional spatial location in azimuth and elevation angles (in degrees), angular velocity of motion, and distance from the microphone (in meters).</p> <p>The isolated sound events were taken from the DCASE 2016 task 2 dataset. This dataset consists of 11 sound event classes such as Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. Every event is assigned a spatial trajectory on an arc with a constant distance from the microphone (in the range 1-10 m) and moving with a constant angular velocity for its duration. Due to the choice of the ambisonic spatial recording format, the steering vectors for a plane wave source or point source in the far field are frequency-independent. Hence, there is no need for a time-variant convolution or impulse response interpolation scheme as the source is moving; the spatial encoding of the monophonic signal was done sample-by-sample using instantaneous ambisonic encoding vectors for the respective DOA of the moving source. The synthesized trajectories in the dataset vary in both azimuth and elevation and are simulated to have a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the &#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks&#39;</a> work.</p>

openother-ncApr 2019View details →
zenodo32/100

Sound event localization and detection (SELDnet) results

<p>This package is part of the work -&nbsp;<a href="https://github.com/sharathadavanne/seld-metric">Joint Measurement of Localization and Detection of Sound Events</a>&nbsp;presented in WASPAA 2019.</p> <p>This package consists of results from the <a href="https://arxiv.org/abs/1905.08546">SELDnet method</a> for joint&nbsp;sound event localization and detection. The results corresponding to different training states of 5, 25 and 75 epochs are provided here.&nbsp; These results are for the four cross-validation splits of the&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset. The sound events in this dataset&nbsp;consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle.&nbsp;This <strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset is part of the&nbsp;<a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>&nbsp;and can be downloaded <a href="https://zenodo.org/record/2599196#.XT_RmHUzaCg">here</a>.</p> <p>Each of the results folders consists&nbsp;of 400 files, corresponding to the results of the individual recordings of the&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset.</p> <p>This data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <p><strong>Download instructions</strong></p> <p>The three files, &nbsp;<strong><em>mic_dev_5</em></strong>,<strong><em>&nbsp;mic_dev_25,&nbsp;</em></strong>and&nbsp;<strong><em>mic_dev_75</em></strong>, correspond to SELDnet results at 5, 25 and 75 epochs.</p> <p>Download the zip files&nbsp;and use your favorite compression tool to unzip these split zip files.</p>

openother-ncJul 2019View details →
zenodo28/100

A large joint sound scene and sound event dataset for source separation of foreground sound events

<p>This large scale data set contains 10000 samples of sound scenes generated from real world recordings, and the original source recordings. It includes 10 different backgrounds with 6-9 appropriate foreground sound events. Strong labels (timed annotations) are provided in four formats for all samples. The original sourceids to identify the class type, a two source method to simply separate foreground and backgrounds, a 32 source annotation for all distinct foregrounds, and a by background (scene) type annotation where sources are according to the background.&nbsp;</p> <p>Baseline results will be presented later in 2020. Further evolutions of this dataset will also be produced with more complex, polyphonic foreground sound events. Please email h.bear@qmul.ac.uk with any questions.</p> <p>Data is free to use for Research purposes only.&nbsp;&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo28/100

WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System

<p>A new large language model (LLM)-powered dataset namely wild domestic environment sound event detection (WildDESED). It is crafted as an extension to the original DESED dataset to reflect diverse acoustic variability and complex noises in home settings. We leveraged LLMs to generate eight different domestic scenarios based on target sound categories of the DESED dataset. Then we enriched the scenarios with a carefully tailored mixture of noises selected from AudioSet and ensured no overlap with target sound. We consider widely popular convolutional neural recurrent network to study WildDESED dataset, which depicts its challenging nature. We then apply curriculum learning by gradually increasing noise complexity to enhance the model's generalization capabilities across various noise levels.</p>

opencc-by-4.0Aug 2024View details →
zenodo24/100

TAU Moving Sound Events 2019 - Ambisonic, Reverberant, Real-life IR and Moving Source Dataset

<p><strong>Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Reverberant and Real-life Impulse Response and Moving Source Dataset</strong></p> <p>This dataset consists of real-life first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset was generated by collecting impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing<br> a maximum length sequence around the array in circular trajectory in one elevation at a time. The playback volume was set to be 30 dB greater than the ambient sound level. The recording was done in a corridor inside the university with classrooms around it during work hours. The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20 to 20 with 10-degree increments at 2 m.&nbsp;</p> <p>The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. All sound events in this dataset are moving only along azimuth with a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location in azimuth and elevation angles (in degrees), the angular velocity of motion and distance from the microphone (in meters).</p> <p>The isolated sound events were taken from the urbansound8k dataset. This dataset consists of 10 sound event classes such as air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. We do not consider air_conditioner and children_playing sound events. Further, we only include the sound event examples marked as foreground in the dataset. We used the splits 1, 8 and 9 provided in the urbansound8k as the three CV splits. These splits were chosen as they had a good number of examples for all the chosen sound event classes after selecting only the foreground examples. During the sound scene synthesis, every sound event is assigned a spatial trajectory on an arc with a constant distance from the microphone and moving with a constant angular velocity for its duration.</p> <p>Other than the license file, there are nine zip files that consist of the dataset and corresponding metadata for given split and overlap. For example, the ov3_split1.zip file consists of training and testing recordings and metadata for the case of a maximum of three temporally overlapping sound events (ov3) for the first cross-validation split (split1). Within each folder, the filenames for training split have the &#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks&#39;</a> work.</p> <p>Data collector (s): Fagerlund, Eemi; Koskimies, Aino; Hakala, Aapo</p>

openother-ncApr 2019View details →
zenodo20/100

SECL-UMONS DATABASE FOR SOUND EVENT CLASSIFICATION AND LOCALIZATION

<p>SECL-UMons is a dataset for sound event classification and localization in the context of office environments. The multichannel dataset is composed of 11 event classes recorded at several realistic positions in two different rooms. The dataset comprises two types of sequences according to the number of events in the sequence. 2662 unilabel sequences and 2724 multilabel sequences are recorded corresponding to a total of 5.24 hours.</p>

openother-ncDec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record