Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

37 results for “sound event”

Learn how ShareScore rates datasets ↗
zenodo44/100

TAU-NIGENS Spatial Sound Events 2020

<p><strong>DESCRIPTION:</strong></p> <p>The <strong>TAU-NIGENS Spatial Sound Events 2020</strong> dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and distances as seen from the recording position.&nbsp;The spatialization of all sound events is based on filtering through real spatial room impulse responses (RIRs), captured in multiple rooms of various shapes, sizes, and acoustical absorption properties. Furthermore, each scene recording is delivered in two spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). The sound events are spatialized as either stationary sound sources in the room, or moving sound sources, in which case time-variant RIRs are used. Each sound event in the sound scene is associated with a trajectory of its direction-of-arrival (DoA) to the recording point, and a temporal onset and offset time. The isolated sound event recordings used for the synthesis of the sound scenes are obtained from the <a href="https://doi.org/10.5281/zenodo.2535878">NIGENS general sound events database</a>. These recordings serve as the development dataset for the <a href="http://dcase.community/challenge2020/task-sound-event-localization-and-detection">DCASE 2020 Sound Event Localization and Detection Task</a> of the <a href="http://dcase.community/challenge2020/">DCASE 2020 Challenge</a>.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the corresponding DCASE2020 task setup:</p> <p>Politis., Archontis, Adavanne, Sharath, &amp; Virtanen, Tuomas (2020). A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection. In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020)</em>, Tokyo, Japan.</p> <p>A longer version with more detailed information can be also found <a href="https://arxiv.org/pdf/2006.01919.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset includes a large number of mixtures of sound events with realistic spatial properties under different acoustic conditions, and hence it is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>600 one-minute long sound scene recordings (development dataset).</li> <li>200 one-minute long sound scene recordings (evaluation dataset).</li> <li>Sampling rate 24kHz.</li> <li>About 700 sound event samples spread over 14 classes (see <a href="http://doi.org/10.5281/zenodo.2535878">here</a> for more details).</li> <li>8 provided cross-validation splits of 100 recordings each, with unique sound event samples and rooms in each of them.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics&nbsp;(<strong>FOA</strong>) and tetrahedral microphone array.</li> <li>Realistic spatialization and reverberation through RIRs collected in 15 different enclosures.</li> <li>From about 1500 to 3500 possible RIR positions across the different rooms.</li> <li>Both static reverberant and moving reverberant sound events.</li> <li>Up to two overlapping sound events allowed, temporally and spatially.</li> <li>Realistic spatial ambient noise collected from each room is added to the spatialized sound events, at varying signal-to-noise ratios (SNR) ranging from noiseless (30dB) to noisy (6dB).</li> </ul> <p>The IRs were collected in Finland by staff of Tampere University between 12/2017 - 06/2018, and between 11/2019 - 1/2020. The older measurements from five rooms were also used for the earlier <a href="https://doi.org/10.5281/zenodo.2580091">development</a> and <a href="https://doi.org/10.5281/zenodo.3066124">evaluation</a> datasets&nbsp;<strong>TAU Spatial Sound Events 2019</strong>, while ten additional rooms were added for this dataset. The data collection received funding from the European Research Council, grant agreement <a href="https://cordis.europa.eu/project/id/637422">637422 EVERYSOUND</a>.</p> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2020">here</a>. This implementation serves as the baseline method in the <a href="http://dcase.community/challenge2020/task-sound-event-localization-and-detection">DCASE 2020 Sound Event Localization and Detection Task</a>.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>Version 1.0 of the dataset included only the 600 development audio recordings and labels, used by the participants of Task 3 of DCASE2020 Challenge to train and validate their submitted systems. Version 1.1 included additionally the 200 evaluation audio recordings without labels, for the evaluation phase of DCASE2020. The latest version 1.2, published after the completion of the challenge, includes also the labels for the evaluation files.</p> <p>If researchers wish to compare their system against the submissions of DCASE2020 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The three files, <strong><em>foa_dev.z01</em></strong>,<strong><em> foa_dev.z02</em></strong>, and <strong><em>foa_dev.zip</em></strong>, correspond to audio data of the <strong>FOA </strong>recording format.<br> The three files, <strong><em>mic_dev.z01</em></strong>,<strong><em> mic_dev.z02</em></strong>, and <strong><em>mic_dev.zip</em></strong>, correspond to audio data of the <strong>MIC</strong> recording format.<br> The <strong><em>metadata_dev.zip</em></strong>&nbsp;is the common metadata for both formats.</p> <p>The file, <em><strong>foa_eval.zip</strong></em>, corresponds to audio data of the <strong>FOA</strong> recording format for the evaluation dataset.<br> The file, <em><strong>mic_eval.zip</strong></em>, corresponds to audio data of the <strong>MIC</strong> recording format for the evaluation dataset.<br> The <em><strong>metadata_eval.zip</strong></em> is the common metadata for both formats. An info file is included (<em>metadata_eval_info.txt</em>) which specifies which of the two evaluation folds the mix file belongs to, and what is its number of overlapping events.</p> <p>Download the zip files corresponding to the format of interest and use your favorite compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal:</p> <ol> <li>Combine the split archive to a single archive: <pre>zip -s 0 split.zip --out single.zip</pre> </li> <li>Extract the single archive using unzip: <pre>unzip single.zip</pre> </li> </ol>

opencc-by-nc-4.0Apr 2020View details →
zenodo44/100

Dataset-AOB: urban sounds events classification

<p>The dataset Dataset-AOB is an audio dataset collected and manually edited for urban sounds events classification using Convolutional Neural Networks for the Master Thesis:&nbsp;</p> <p>Ospina, A. &quot;Audio Event Classification using Deep Learning. Use case: Urban Sounds Events classification with Convolutional Neural Networks,&quot; M.Eng. thesis, Beuth University of Applied Sciences, Berlin, 2020.</p> <p>- 10 audio events:&nbsp;alarm-siren, children playing, dog bark, engine, footsteps, glass breaking, gun shot, metro train, rain and screams.</p> <p>- duration: &lt; 4 seconds</p> <p>- format: (.wav)</p> <p>- sampling rate: 22KHz - 44KHz</p> <p>- files: Dataset-AOB: development dataset (4831 samples), DatasetEVAL-AOB: evaluation (218 samples)</p> <p>- metadata: (.csv)</p> <p>- sources per class: (.png)</p> <p>Contact: aospinab@gmail.com</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Phlorest phylogeny derived from Hruschka et al. 2015 'Detecting regular sound changes in linguistics as events of concerted evolution'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Hruschka, D. J., Branford, S., Smith, E. D., Wilkins, J., Meade, A., Pagel, M., &amp; Bhattacharya, T. (2015). Detecting regular sound changes in linguistics as events of concerted evolution. Current Biology, 25(1), 1-9.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

USM Dataset - A Dataset for Polyphonic Sound Event Tagging in Urban Sound Monitoring Scenarios

<p>This dataset includes 24,000 5-seconds-long polyphonic stereo soundscapes composed of sounds taken from the FSD50k dataset:</p> <p>- Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra. FSD50K: an Open Dataset of Human-Labeled Sound Events (<a href="https://arxiv.org/abs/2010.00475">https://arxiv.org/abs/2010.00475</a>)</p> <p>FSD50k samples used in the USM dataset were selected to allow for commercial usage.</p> <p>Find more details about the USM dataset at&nbsp;<a href="https://github.com/jakobabesser/USM">https://github.com/jakobabesser/USM</a></p>

openmit-licenseApr 2022View details →
zenodo44/100

TAU-NIGENS Spatial Sound Events 2021

<p><strong>DESCRIPTION:</strong></p> <p>The&nbsp;<strong>TAU-NIGENS Spatial Sound Events 2021</strong>&nbsp;dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and distances as seen from the recording position.&nbsp;The spatialization of all sound events is based on filtering through real spatial room impulse responses (RIRs), captured in multiple rooms of various shapes, sizes, and acoustical absorption properties. Furthermore, each scene recording is delivered in two spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). The sound events are spatialized as either stationary sound sources in the room, or moving sound sources, in which case time-variant RIRs are used. Each sound event in the sound scene is associated with a single direction-of-arrival (DoA) if static, a&nbsp;trajectory DoAs if moving, and a temporal onset and offset time. The isolated sound event recordings used for the synthesis of the sound scenes are obtained from the&nbsp;<a href="https://doi.org/10.5281/zenodo.2535878">NIGENS general sound events database</a>. These recordings serve as the development dataset for the&nbsp;<a href="http://dcase.community/challenge2021/task-sound-event-localization-and-detection">DCASE 2021 Sound Event Localization and Detection Task</a>&nbsp;of the&nbsp;<a href="http://dcase.community/challenge2021/">DCASE 2021 Challenge</a>.</p> <p>This&nbsp;dataset is the third iteration of spatialized&nbsp;sound event datasets based on&nbsp;real room responses and ambient noise from multiple spaces, with each iteration introducing more challenging conditions closer to real-life. Those iterations, including the present one, are:</p> <ul> <li><strong>TAU Spatial Sound Events 2019,&nbsp;</strong><a href="https://doi.org/10.5281/zenodo.2580091">development</a>&nbsp;and&nbsp;<a href="https://doi.org/10.5281/zenodo.3066124">evaluation</a>&nbsp;datasets.<br> 5 rooms, high direct-to-reverberant&nbsp;ratios (DRR), static sources only, minimum DoA separation 10&deg;, discrete grid of DoAs, high SNR for ambient noise, maximum polyphony of 2 simultaneous events</li> <li><strong><a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a></strong>, development and evaluations datasets.<br> 13 rooms, low-to-high DRRs, static and moving sources, continuous DoAs, low-to-high SNR for ambient noise,<br> maximum polyphony of 2 simultaneous events</li> <li><strong>TAU-NIGENS Spatial Sound Events 2021</strong>.<br> Same as 2020, with the following exceptions: a more natural temporal distribution of sound events,<br> maximum polyphony of 3&nbsp;target events, <strong>inclusion of additional out-of-target-classes directional interference events</strong></li> </ul> <p>The inclusion of directional interferences is the main new challenging property of the new dataset. They are spatialized in the scene in the same way as the target events, and can be either static or moving. The interfering events are sourced from the &quot;engine&quot;, &quot;fire&quot;, and &quot;general&quot; classes of the NIGENS sound event database. The interferers are considered unknown and no activity or directional labels of them are provided with the training datasets.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the corresponding DCASE2020 task setup:</p> <p>Archontis Politis, Sharath Adavanne, Daniel Krause, Antoine Deleforge, Prerak Srivastava, Tuomas Virtanen (2021).<br> A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection.&nbsp;<br> In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2021)</em>, Barcelona, Spain.</p> <p>available <a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Politis_43.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset includes a large number of mixtures of sound events with realistic spatial properties under different acoustic conditions, and hence it is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>600 one-minute long sound scene recordings with metadata (development dataset).</li> <li>200 one-minute long sound scene recordings without metadata (evaluation dataset).</li> <li>Sampling rate 24kHz.</li> <li>About 500 sound event samples distributed over the 12 target classes (see [here](http://doi.org/10.5281/zenodo.2535878) for more details).</li> <li>About 400 sound event samples used as interference events (see [here](http://doi.org/10.5281/zenodo.2535878) for more details).</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array.</li> <li>Realistic spatialization and reverberation through multichannel RIRs collected in 13 different enclosures.</li> <li>From 1184 to 6480 possible RIR positions across the different rooms.</li> <li>Both static reverberant and moving reverberant sound events.</li> <li>Three possible angular speeds for moving sources of approximately 10, 20, or 40deg/sec.</li> <li>Up to three overlapping sound events possible, temporally and spatially.</li> <li>Simultaneous directional interfering sound events with their own temporal activities, static or moving.</li> <li>Realistic spatial ambient noise collected from each room is added to the spatialized sound events, at varying signal-to-noise ratios (SNR) ranging from noiseless (30dB) to noisy (6dB) conditions.</li> </ul> <p>The IRs were collected in Finland by staff of Tampere University between 12/2017 - 06/2018, and between 11/2019 - 1/2020.&nbsp;The data collection received funding from the European Research Council, grant agreement&nbsp;<a href="https://cordis.europa.eu/project/id/637422">637422 EVERYSOUND</a>.</p> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset will be provided soon. That&nbsp;implementation will serve as the baseline method in the&nbsp;<a href="http://dcase.community/challenge2021/task-sound-event-localization-and-detection">DCASE 2021 Sound Event Localization and Detection Task</a>.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current and final version (Version 1.2) of the dataset includes the 600 development audio recordings and labels, used by the participants of Task 3 of DCASE2021&nbsp;Challenge to train and validate their submitted systems, and the 200 evaluation audio recordings including their&nbsp;labels, used in the evaluation phase of DCASE2021.</p> <p>If researchers wish to compare their system against the submissions of DCASE2021 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The three files,&nbsp;<strong><em>foa_dev.z01</em></strong>, and&nbsp;<strong><em>foa_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>FOA&nbsp;</strong>recording format.<br> The three files,&nbsp;<strong><em>mic_dev.z01</em></strong>,&nbsp;and&nbsp;<strong><em>mic_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>MIC</strong>&nbsp;recording format.<br> The&nbsp;<strong><em>metadata_dev.zip</em></strong>&nbsp;is&nbsp;the common metadata for both formats.</p> <p>The file,<strong>&nbsp;<em>foa_eval.zip</em></strong>, corresponds to audio data of the&nbsp;FOA&nbsp;recording format for the evaluation dataset.<br> The file,&nbsp;<strong><em>mic_eval.zip</em></strong>, corresponds to audio data of the&nbsp;MIC&nbsp;recording format for the evaluation dataset.<br> The&nbsp;<strong><em>metadata_eval.zip</em></strong>&nbsp;is the common metadata for both formats.</p> <p>Download the zip files corresponding to the format of interest and use your favorite compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal:</p> <ol> <li>Combine the split archive to a single archive: <pre>zip -s 0 split.zip --out single.zip</pre> </li> <li>Extract the single archive using unzip: <pre>unzip single.zip</pre> </li> </ol>

opencc-by-nc-4.0Feb 2021View details →
zenodo40/100

UNS-Exterior Spatial Sound Events 2023 (UNS-ESSE2023)

<p><span>This dataset was generated within the UNS2: Localising Audio Events in Crowds use case of the H2020 MARVEL project. The purpose of the dataset is to offer audio samples collected outdoors in an urban area that could be used for development of sound event localisation and detection (SELD) models for acoustic monitoring in urban environments.</span></p> <p><span>The dataset was generated within a staged recording process. Audio samples were collected by utilising Infineon Audiohub Nano 8-channel microphone array board with a sampling rate of 48&nbsp;kHz. The audio was synthetised by mixing target sound events from the FSD50K dataset &ldquo;gunshot&rdquo; and &ldquo;gunfire&rdquo;, &ldquo;boom&rdquo;, and &ldquo;shatter&rdquo; with samples of the class &ldquo;chatter&rdquo;, that was used as a background noise, extracted also from the FSD50k dataset. The scenario included different SNR values for the events in the mixtures and the measurement of the Sound Pressure Level (SPL) before and after the recording. Sound events were reproduced using eight JBL VP7212MDP10 speakers, which were positioned circularly around the microphone array board in equidistant positions. Data collection was performed for two different distances: 5 m and 10 m. One of the speakers was used to reproduce the target events (overlaid on the background noise), where the selected speaker varied during the data collection, while the others were used to reproduce background noise only. The recording setup was placed outdoors, involving additional ambient noise of the urban city area, which is a case closer to the real-world scenario comparing to the datasets recorded in the laboratory conditions. Together with the audio files, spatiotemporal annotation of the sound events is provided. These annotations include temporal onset and offset of the target events, azimuth, source distance, SNR level and SPL values for background ambience noise. </span></p> <p><span>This work was funded by the European Union&rsquo;s Horizon 2020 research and innovation program MARVEL under grant agreement No 957337. This publication reflects the authors&rsquo; views only. The European Commission is not responsible for any use that may be made of the information it contains.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Compositional discovery of architecture-aware and sound process models from event logs of multi-agent systems: experimental data.

<p>This repository contains the experimental data used for the evaluation of the compositional approach to the discovery of process models from event logs of multi-agent systems, where agents interact according to specific patterns of synchronous and asynchronous interactions.</p> <p>According to the experiment plan, there is the folder for each interface pattern containing:</p> <ol> <li>The reference model (Petri net encoded in PNML-file)</li> <li>The event log obtained by simulating the behavior of the reference model (XES-file)</li> <li>The model discovered directly from the generated event log (Petri net encoded in PNML-file)</li> <li>The model discovered by composing the agent model w.r.t. the interface pattern (Petri net encoded in&nbsp;PNML-file)</li> </ol>

opencc-by-4.0May 2021View details →
zenodo40/100

TUT Rare sound events, Evaluation dataset

<p>TUT Rare Sound events 2017, evaluation dataset consists of source files for creating mixtures of rare sound events (classes baby cry, gun shot, glass break) with background audio, as well a set of readily generated mixtures and recipes for generating them.</p> <p>The &quot;source&quot; part of the dataset consists of two subsets:</p> <ul> <li>background recordings from 15 different acoustic scenes,</li> <li>recordings with the target rare sound events from three classes, accompanied by annotations of their temporal occurrences.</li> </ul> <p>The mixture set consists of two 1500 mixtures (500 per target class, with half of the mixtures not containing any target class events).&nbsp;</p> <p>The collection of the background recording data has been financially supported by European Research Council under the European Unions H2020 Framework Programme through ERC Grant Agreement 637422 EVERYSOUND.</p>

openother-ncJan 2018View details →
zenodo40/100

TUT Rare sound events, Development dataset

<p>TUT Rare Sound events 2017, development dataset consists of source files for creating mixtures of rare sound events (classes baby cry, gun shot, glass break) with background audio, as well a set of readily generated mixtures and recipes for generating them.</p> <p>The &quot;source&quot; part of the dataset consists of two subsets:</p> <ul> <li>background recordings from 15 different acoustic scenes,</li> <li>recordings with the target rare sound events from three classes, accompanied by annotations of their temporal occurrences,</li> <li>a set of meta files providing the cross-validation setup: lists of background and target event recordings split into training and test subsets (called &quot;devtrain&quot; and &quot;devtest&quot;, respectively, indicating they are provided as the development dataset, as opposed to the evaluation dataset released separately).&nbsp;</li> </ul> <p>The mixture set consists of two subsets (training and testing), each containing ~1500 mixtures (~500 per target class in each subset, with half of the mixtures not containing any target class events).&nbsp;</p> <p>&nbsp;</p> <p>The collection of the background recording data has been financially supported by European Research Council under the European Unions H2020 Framework Programme through ERC Grant Agreement 637422 EVERYSOUND.</p>

openother-ncMar 2017View details →
zenodo40/100

Joint sound scene and event dataset

<p>A dataset of synthentic sound scenes created in scaper [1]using real-world recordings. Synthesized scenes are&nbsp;split into train and test such that no original recordings are in both splits.&nbsp;</p> <p>Dataset created for the task of performing both acoustic scene classification jointly with sound event detection. Time stamped annotations and jams files included.&nbsp;</p> <p>10 acoustic scene classes, 32 sound event classes.</p> <p>filename syntax is [sceneLabel_recordingID_polyphonylevel]</p> <p>Full credits / citation details are below,&nbsp;please contact Yogi at h.bear@qmul.ac.uk with any queries.&nbsp;</p> <p>Citation:</p> <p>Helen L Bear, Ines Nolasco, and Emmanouil Benetos,&nbsp;<em>Towards joint sound scene and polyphonic sound event detection.</em>&nbsp;Interspeech 2019. Graz.&nbsp;</p> <p>References</p> <ol> <li>Salamon, Justin, et al. &quot;Scaper: A library for soundscape synthesis and augmentation.&quot;&nbsp;<em>Applications of Signal Processing to Audio and Acoustics (WASPAA), 2017 IEEE Workshop on</em>. IEEE, 2017.</li> </ol>

opencc-by-4.0Feb 2019View details →
zenodo40/100

NIGENS general sound events database

<p>NIGENS (<strong>N</strong>eural <em><strong>I</strong></em>nformation Processing group <em><strong>GEN</strong></em>eral sounds) is a database provided for sound-related modeling in the field of computational auditory scene analysis, particularly for sound event detection, that has emerged from the <a href="http://twoears.eu">Two!Ears project</a>.</p> <p>It contains 1017 wav files of various lengths (between 1s and 5mins), in total comprising 4h:46m of sound material. Mostly, sounds are provided with 32-bit precision and 44100 Hz sampling rate. The files contain sound events in isolation, i.e. without superposition of ambient or other foreground sources.&nbsp;&nbsp; &nbsp;&nbsp;</p> <p>Fourteen distinct sound classes are included: <em>alarm</em>, <em>crying baby</em>, <em>crash</em>, <em>barking dog</em>, <em>running engine</em>, <em>burning fire</em>, <em>footsteps</em>, <em>knocking on door</em>, <em>female</em>&nbsp;and <em>male speech</em>, <em>female</em>&nbsp;and <em>male scream</em>, <em>ringing phone</em>, <em>piano</em>. Additionally, there is the <em>general</em>&nbsp;(&ldquo;anything else&rdquo;) class.&nbsp;Care has been taken to select sound classes representing different features, like noise-like or pronounced, discrete or continuous.</p> <p>The <em>general</em>&nbsp;class is a pool of sound events different than the 14 distuingished target sound classes, containing as heterogeneous sounds as possible (303 in total). For example, it includes nature sounds such as wind, rain, or animals, sounds from human-made environments such as honks, doors, or guns, as well as human sounds like coughs.&nbsp;These sounds are intended both as ``disturbance&#39;&#39; sound events (superposing) and as counterexamples to target sound classes.<br> <br> Wav&nbsp;files are accompanied by annotation (.txt) files that include&nbsp;perceptual on- and offset times of the&nbsp;file&#39;s sound events.&nbsp;</p> <p>You are free to use this database non-commercially under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0&nbsp;license.</p> <p><strong>If you use this data set, please cite as:</strong></p> <p><strong>Ivo Trowitzsch, Jalil Taghia, Youssef Kashef, and Klaus Obermayer (2019). <em>The NIGENS general sound events database</em>. Technische Universit&auml;t Berlin, Tech. Rep. arXiv:1902.08314 [cs.SD] &nbsp;</strong></p> <p>In [1], we have developed and analyzed a robust binaural sound event detection training scheme using NIGENS. In [2], we have extended it to join sound event detection and localization through spatial segregation.</p> <p>[1]&nbsp;Trowitzsch, I., Mohr, J., Kashef, Y., Obermayer, K. (2017). <em>Robust detection of environmental sounds in binaural auditory scenes</em>. IEEE/ACM Transactions on Audio, Speech, and Language Processing 25(6).</p> <p>[2]&nbsp;Trowitzsch, I., Schymura, C., Kolossa, D., Obermayer, K. (2019). <em>Joining Sound Event Detection and Localization Through Spatial Segregation</em>. accepted for publication in&nbsp;IEEE/ACM Transactions on Audio, Speech, and Language Processing. DOI: 10.1109/TASLP.2019.2958408. E-Preprint:&nbsp;<a href="https://arxiv.org/abs/1904.00055">arXiv:1904.00055</a>&nbsp;[cs.SD].</p>

opencc-by-nc-nd-4.0Feb 2019View details →
zenodo40/100

Sound Events for Surveillance Applications

<p>The Sound Events for Surveillance Applications (SESA) dataset files were obtained from Freesound. The dataset was divided between train (480 files) and test (105 files) folders. All audio files are WAV, Mono-Channel, 16 kHz, and 8-bit with up to 33 seconds. # Classes: 0 - Casual (not a threat) 1 - Gunshot 2 - Explosion 3 - Siren (also contains alarms)</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

CLDF dataset derived from Hruschka et al.'s "Detecting regular sound changes in linguistics as events of concerted evolution" from 2015

<p>Cite the source of the dataset as:</p> <blockquote> <p>Hruschka, D. J., Branford, S., Smith, E. D., Wilkins, J., Meade, A., Pagel, M., &amp; Bhattacharya, T. (2015). Detecting regular sound changes in linguistics as events of concerted evolution. Current Biology, 25(1), 1-9.</p> </blockquote>

opencc-by-nc-4.0Jul 2023View details →
zenodo40/100

"I was the class teacher at that time. It was a class trip, usually organized near the end of the schoolterm in summer. The pupils went there by bike to have a barbecue at the sandy banks of the river Rhine near Dusseldorf. The landscape around is mostly dominated by agriculture and glasshouse cultures. You find a mixture of former villages nowadays completely suburbanized. The population finds jobs in the nearby urban centers like Dusseldorf, Neuss and other big cities. The reason why Irecorded the scene is simply because Iam interested in collecting sounds in general by doing recordings in different surroundings like nature, cities and everything between. My memories about the event are that it was a relaxing and funny atmosphere, which is not always the case while teaching in a classroom" [Reinhard/reinsamba]15 in Collecting Sounds. Online Sharing of Field Recordings as Cultural Practice

"I was the class teacher at that time. It was a class trip, usually organized near the end of the schoolterm in summer. The pupils went there by bike to have a barbecue at the sandy banks of the river Rhine near Dusseldorf. The landscape around is mostly dominated by agriculture and glasshouse cultures. You find a mixture of former villages nowadays completely suburbanized. The population finds jobs in the nearby urban centers like Dusseldorf, Neuss and other big cities. The reason why Irecorded the scene is simply because Iam interested in collecting sounds in general by doing recordings in different surroundings like nature, cities and everything between. My memories about the event are that it was a relaxing and funny atmosphere, which is not always the case while teaching in a classroom" [Reinhard/reinsamba]15

opencc-by-4.0Dec 2019View details →
dryad40/100

Selected large model output files and Buffalo sounding data from: Lake Huron enhances snowfall downwind of Lake Erie: a modeling study of the 2010 near year’s Lake-effect snowfall event

Open the record for dataset details and reuse information.

publicDec 2024View details →
zenodo36/100

Real-Life Indoor Sound Event Dataset (ReaLISED) for Sound Event Classification (SEC)

<p>The Real-Life Indoor Sound Event Dataset (ReaLISED) offers&nbsp;the scientific community the possibility of testing Sound Event Classification (SEC) algorithms with new real indoor&nbsp;audio event recordings. The full set is made up of 2479 sound recordings of 18 events. The 18 event classes are the following:&nbsp;beater, cooking, cupboard/wardrobe,&nbsp;dishwasher, door, drawer, furniture movement, microwave, object falling, smoke extractor, speech, switch, television, vacuum cleaner, walking, washing machine, water tap, and window. There are 2479 clips of isolated sounds, which result in 3624.51 seconds.&nbsp;The number of events in each class is between 104 for the &quot;Window&quot; class and 190 for the &ldquo;Speech&rdquo; class, with a mean value of 138 events and a standard deviation of 25.</p> <p>Four Olympus LS-100 recorders&nbsp;were used. The sampling frequency was set to 44.1 kHz and 24 bits per sample. The stereo mode was used, and a medium sensitivity of the microphone was set. The distance between the recorder and the sound source was set to approximately 30-40 cm.</p> <p>Apart from the labels related to the class of event, extra information for each recording is provided in order to be exploited if necessary in the future, with other research purposes. This extra information completes the description of the sound source.</p> <p>The dataset is introduced to the scientific community by providing all the .flac files which composed it. The name of the files is built with 5 pieces of information, separated with underscores (&ldquo;_&rdquo;), with the format &ldquo;abc_123_45_67_8.flac&prime;&prime;:</p> <ul> <li> <p>&ldquo;abc&rdquo;: the first three letters indicate the source that produces the sound. This segment can take 18 different values: &lsquo;bea&rsquo; (beater), &lsquo;coo&rsquo; (cooking), &lsquo;cup&rsquo; (cupboard/wardrobe), &lsquo;dis&rsquo; (dishwasher), &lsquo;doo&rsquo; (door), &lsquo;dra&rsquo; (drawer), &lsquo;fur&rsquo; (furniture movement), &lsquo;mic&rsquo; (microwave), &lsquo;obj&rsquo; (object falling), &lsquo;smo&rsquo; (smoke extractor), &lsquo;spe&rsquo; (speech), &lsquo;swi&rsquo; (switch), &lsquo;tel&rsquo; (television), &lsquo;vac&rsquo; (vacuum cleaner), &lsquo;wal&rsquo; (walking), &lsquo;was&rsquo; (washing machine), &lsquo;wat&rsquo; (water tap), win&rsquo; (window).</p> </li> <li> <p>&ldquo;123&rdquo;: this set of digits identifies the event among the number of events produced by the source identified with &ldquo;abc&rdquo;. This segment can take all the values between &lsquo;001&rsquo; and &lsquo;190&rsquo;, which is the maximum number of events of a particular class we can find in the dataset (speech).</p> </li> <li> <p>&ldquo;45&rdquo;: this set of digits identifies the action that produce the sound. This segment can take 11 different values: &lsquo;01&rsquo; (close), &lsquo;02&rsquo; (open), &lsquo;03&rsquo; (throw), &lsquo;04&rsquo; (turn on), &lsquo;05&rsquo; (turn off), &lsquo;06&rsquo; (move), &lsquo;07&rsquo; (plug), &lsquo;08&rsquo; (unplug), &lsquo;09&rsquo; (raise), &lsquo;10&rsquo; (lower), and &lsquo;00&rsquo; (there is no information about the action).</p> </li> <li> <p>&ldquo;67&rdquo;: this set of digits identifies the material the sound source is made of. This segment can take 14 different values: &rsquo;01&rsquo; (wood), &rsquo;02&rsquo; (glass), &rsquo;03&rsquo; (metal), &rsquo;04&rsquo; (plastic), &rsquo;05&rsquo; (ceramic), &rsquo;06&rsquo; (synthetic), &rsquo;07&rsquo; (cardboard), &rsquo;08&rsquo; (marble), &rsquo;09&rsquo; (floating platform), &rsquo;10&#39;&nbsp;(platelet), &rsquo;11&rsquo; (wicker), &rsquo;12&rsquo; (carpet), &rsquo;13&rsquo; (medium-density fibreboard MDF), and &rsquo;00&rsquo; (there is no information about the material).</p> </li> <li> <p>&ldquo;8&rdquo;: the last digit gives approximate information about the intensity of the recorded sound. It can take 4 different values: &rsquo;1&rsquo; (low intensity), &rsquo;2&rsquo; (medium intensity), &rsquo;3&rsquo; (high intensity), &rsquo;0&rsquo; (there ir no information about the intensity).</p> <p>For clarity, some examples of audio file&nbsp;names with this code are shown hereunder:</p> </li> <li> <p>&ldquo;doo_040_02_00_3.flac&rdquo; is the name of the 40th file in the Door class, described as &ldquo;opening a door of unknown material with high intensity&rdquo;.</p> </li> <li> <p>&ldquo;fur_058_06_01_2.flac&rdquo; is the name of the 58th file in the furniture movement class, described as &ldquo;moving a wooden furniture with medium intensity&rdquo;.</p> </li> <li> <p>&ldquo;vac_001_00_00_0.flac&rdquo; is the name of the 1st audio file in the vacuum cleaner class, described as &ldquo;using the vacuum cleaner, without information about the action, neither the material or the intensity&rdquo;.</p> </li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Artificial sound mixes with event insertions

<p><strong>Contains artificial sound mixes and meta data that were created for the task of<br> sound event detection.</strong></p> <p>The mixes were created using background and event audio recordings<br> from Tampere University&#39;s Detection and Classification of Acoustic Scenes and<br> Events (DCASE) Community. More information on the source data can be found at<br> http://www.cs.tut.fi/sgn/arg/dcase2017/challenge/task-rare-sound-event-detection#audio-dataset.</p> <p>Source data credits: Diment, Aleksandr et al (2017, 2018)</p> <p><strong>Created artificial sound mixes are 10 seconds long and contain:</strong></p> <p>- background audio from diverse scenes from start to end<br> - 0 to 4 event insertions</p> <p>With this formula, two datasets were created separately: a training, and an<br> evaluation dataset. These were created separately so that the backgrounds and<br> event recordings used for the evaluation dataset were not used in any of the<br> training audio mixes. Thus, keeping them unseen by the system during development.</p> <p><strong>Audio data:</strong><br> <strong>mixes_train.zip</strong>: Contains the audio mixes created for training.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Audio format: .wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Count: 1000 tracks</p> <p><strong>mixes_eval.zip</strong>: Contains the audio mixes created for evaluation.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Audio format: .wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Count: 500 tracks</p> <p><strong>Meta-data:</strong><br> Meta-data was maintained documenting the source background, the overlaid source<br> events, and the time onset and offset of each sound event or confusing sound.</p> <p><strong>meta_track_info_train.csv and meta_track_info_eval.csv columns:</strong></p> <p>- trackID: the unique ID of a mix<br> - class_dummy: whether a mix contain a glas break event (other events can be<br> &nbsp; determined using meta_clip_insertions_train and meta_clip_insertions_eval)<br> - background_file: the unique file reference used for background sound to the<br> &nbsp; source data (i.e. DCASE original audio data set)<br> - background_t0: the second in the original background sound recording in which<br> &nbsp; the 10 second background starts.</p> <p><strong>meta_clip_insertions_train.csv and meta_clip_insertions_eval.csv columns:</strong></p> <p>- trackID: the unique ID of a mix<br> - event: the type of event insertion<br> - event_start: at which time in the mix the event start<br> - event_end: at which time in the mix the event ends<br> - event_file: the unique file_number of event reference to the<br> &nbsp; source data (i.e. DCASE original audio data set)<br> &nbsp; an event_file with value 345584_4.wav means the event comes from file<br> &nbsp; 345584.wav in the DCASE audio set and is the 4th event in that audio file.</p> <p>To see the project for which this data set was created visit the github<br> repository at https://github.com/reyvaz/sound-event-detection</p> <p><strong>References:</strong><br> Diment, Aleksandr, Mesaros, Annamaria, Heittola, Toni, &amp; Virtanen, Tuomas. (2017).<br> TUT Rare sound events, Development dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.401395</p> <p>Aleksandr Diment, Annamaria Mesaros, Toni Heittola, &amp; Tuomas Virtanen. (2018).<br> TUT Rare sound events, Evaluation dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.1160455<br> &nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Evaluation Datasets

<p>This package consists of two evaluation datasets,&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;providing four-channel First-Order Ambisonic (FOA) recordings while&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These evaluation datasets are part of the&nbsp;<a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>.&nbsp;The corresponding development datasets can be downloaded <a href="https://doi.org/10.5281/zenodo.2599196">here</a>.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <ul> <li>The <strong>foa_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;evaluation dataset.</li> <li>The <strong>mic_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;evaluation dataset.</li> </ul> <p>-- Version 2 updates --</p> <p>The<a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection-results"> DCASE 2019 sound event localization and detection task has now ended</a>. Hence we are releasing the reference labels for the evaluation dataset in this version.</p> <ul> <li>The&nbsp;<strong><em>metadata_eval.zip</em></strong>&nbsp;is the common metadata for both&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;evaluation datasets.&nbsp; &nbsp;</li> <li>The <strong>short2longnames.txt</strong>&nbsp;file consists of the corresponding names for each recording in the dataset in the <a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection#development-dataset">development-set format</a>, i.e., including the information of the impulse response location and the maximum number of overlapping sound events in the recording.</li> </ul> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

openother-ncMay 2019View details →
zenodo36/100

TAU-SEBin Binaural Sound Events 2021

<p><strong>TAU-SEBin Binaural Sound Events 2021 </strong>is a dataset of synthetic binaural audio recordings, which consist of sound events spaced in simulated shoebox rooms. The data is suitable for experiments with several acoustic scene analysis tasks such as sound source localization, sound distance estimation or sound event detection.</p> <p>&nbsp;</p> <p>Data is created using isolated sound events derived from several datasets: NIGENS [1], DESED [2] and TUT Rare Sound Events 2017 [3], containing 18 total sound classes, namely: alarm, baby, blender, cat, crash, dishes, dog, engine, fire, footsteps, glassbreak, gunshot, knock, phone, piano, scream, speech, water. The data is split into two subsets, one of which&nbsp;(bin_prox_dir) contains up to two overlapping sound events, whereas the other one consists of single sources only (bin_prox_dir_one). Each subset contains 400 audio files, divided into 4 equal splits for fold-wise cross-validation.</p> <p>&nbsp;</p> <p>The metadata provides the following information:</p> <p><strong>sound_event_recording</strong> - sound event class</p> <p><strong>start_time, end_time - </strong>onset and offset times of the sound events (in seconds)</p> <p><strong>azi, ele - </strong>the azimuth and elevation angle of the sound source (in degrees)</p> <p><strong>dist</strong> - sound source to receiver distance (in metres)</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>References:</p> <p>[1] I. Trowitzsch, J. Taghia, Y. Kashef, and K. Obermayer, NIGENS general sound events database. Zenodo, 2019.<br> [2] N. Turpault, R. Serizel, A. Parag Shah, and J. Salamon, &ldquo;Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,&rdquo; in Workshop on Detection and Classification of Acoustic Scenes and Events, 2019.<br> [3] A. Mesaros, T. Heittola, A. Diment, B. Elizalde, A. Shah, E. Vincent, B. Raj, and T. Virtanen, &ldquo;DCASE 2017 challenge setup: Tasks, datasets and baseline system,&rdquo; in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), 2017, pp. 85&ndash;92.</p>

openother-openJul 2021View details →
zenodo32/100

OPEN-WINDOW: SOUND EVENT DATABASE FOR RESEARCH AND DEVELOPMENT

<p><strong>(1) Background:</strong></p> <p>Situated in the domain of urban sound scene classification by humans and machines, the research in this project will be a first step towards mapping urban noise pollution experienced indoors and finding ways to reduce its negative impact in peoples&#39; homes. The acoustic distinction between outdoor and indoor scenes is an active research field and can be automated with some success. A much subtler difference is the change in the indoor soundscape induced by an open window. Being able to determine this, however, would allow applications in warning systems and be a prerequisite for an app-based urban sound mapping project.</p> <p>Acoustic detection requires neither line of sight nor sensors at the window frame or knowledge of the number of windows or their size. The task, however, varies substantially in difficulty with the amount of sound inside and outside. From the point of machine classification, the lack of specificity is the most problematic aspect: Very few sounds if any can be assumed to originate exclusively from outside <em>and</em> be present at all times to aid automatic detection. The required generalisation ability, however, can be assumed for humans, who might also use very subtle cues in the change of reverberations.</p> <p>&nbsp;</p> <p><strong>(2) Dataset</strong></p> <p><em>(a) Recording locations</em></p> <p>The recordings have been made at three different locations.&nbsp;</p> <ul> <li>Farm: A farm in Brook, Surrey, United Kingdom. The recordings were made in an open-plan studio flat area in the centre of the farm. The recordings in this location have the lowest levels of background noise, due mainly to a quiet environmental surrounding.</li> <li>Office 1: An office at the University of Surrey, Guildford, United Kingdom. The recordings were made in an open-plan&nbsp;office located on the first floor, at the Centre for Vision, Speech and Signal Processing (CVSSP). Since this office accommodates 16 researchers, recordings in this location have the highest level of background noise</li> <li>Office 2: An office at the University of Surrey, Guildford, United Kingdom. The recordings were made in a small size open-plan office at the CVSSP. This office accommodates 8 researchers and the recordings made in this office considered to have a medium level of background noise.</li> </ul> <p><em>(b) Recording equipment</em></p> <p>The recordings made at the two offices and a studio flat in a farm used a dedicated laptop, Focusrite Clarett 4pre USB external sound card (44,100 Hz sample rate at 16 bits per sample) 1, and a Behringer ECM 8000 microphone.</p> <p><em>(c)&nbsp;Recording setup</em></p> <p>The Behringer ECM 8000 microphone is connected to the External Line Return (XLR) input of the Focusrite Clarett external sound<br> card via an XLR cable. The external sound card is connected to the dedicated laptop and controlled using Ableton Live 10&nbsp;software for setting configurations and exporting the recorded audio files. The microphone is located approximately 10 cm away from the<br> window and fixed using a microphone holder. At each location 90 audio sessions are recorded; 60 one minute recordings for static state setup and 30 fifteen seconds recordings for transitional state setup.</p> <p><em>(d)&nbsp;File naming conventions</em></p> <p>The naming convention for audio recording is as follows:<br> [Location] [State] [Time] [IDX]<br> [State] will be one of the following: &ldquo;O stands for open, C stands&nbsp;for Close, OC means a transition from Open to Close and CO stands for a transition from Close to Open.&rdquo; [Time] stamp will be one of the following: &ldquo;AM stands for morning between 9:00 to 12:00, N stands for noon which is between 13:00 to 15:00 and PM which stands for an afternoon which is between 17:00 to 20:00.&rdquo; [IDX] is<br> representing the file ID number. For example, &ldquo;Farm C PM 01.wav&rdquo;, means this file is recorded at the farm and in the afternoon when the window is closed and the file ID is 01.</p> <p><em>(e) Dataset acquisition:</em></p> <p>A recording kit consisting of a dedicated laptop and microphone will be given to volunteers. Custom-programmed software will remind the user to specify the window state (establishing the so-called ground truth).</p> <p><em>(f)&nbsp;Specifications</em><br> &nbsp; - Open-Window contains 270 audio recordings totalling 3.37 hours of audio.<br> &nbsp; - Each audio recording belongs to one of the four classes representing the window states; two stationary states (Open, Close) and two transitional states (Open-Close, Close-Open).<br> &nbsp; - The recordings were carried out in different locations and at different times of the day.<br> &nbsp; &nbsp; &nbsp; - Three locations: Office1, Office2, Farm<br> &nbsp; &nbsp; &nbsp; - Three periods of the day: Morning, Afternoon, Evening<br> &nbsp; - The recordings are split into six-folds.<br> &nbsp; &nbsp; &nbsp; - Fold 1 is the test set.<br> &nbsp; &nbsp; &nbsp; - Fold 2 is the validation set.<br> &nbsp; &nbsp; &nbsp; - Folds 3-6 comprise the training set.<br> &nbsp; &nbsp; Each fold is balanced in terms of the class and location distribution.<br> &nbsp; - The annotations/metadata can be found in annotations.csv.<br> &nbsp; - The recordings for the stationary states are approximately 60 seconds, while the recordings for the transitional states are approximate 15 seconds.<br> &nbsp; - The format of the recordings is 2-channel 16-bit PCM sampled at 44.1 kHz.</p>

opencc-by-4.0Nov 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record