Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “ambisonic”

Learn how ShareScore rates datasets ↗
zenodo52/100

PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures

<h1>PAN-AR</h1> <p>This is <strong>PAN-AR</strong> (Panoramas, Ambient Noise &amp; Ambisonics RIRs), a dataset described in the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <blockquote> <p>Filippo Denti, Davide Fantini, Federico Avanzini and Giorgio Presti. PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures. In <em>Proceedings of the 19th International Audio Mostly Conference</em>, Milan, Italy, September 2024.</p> </blockquote> <p>The dataset includes Spatial Room Impulse Responses (SRIRs) in second-order Ambisonics format, ambient noise recordings, and spherical photos. These data have been captured in four environments with different configurations of the source and listener positions:</p> <ol> <li>Printer room</li> <li>Meeting room</li> <li>Classroom</li> <li>Underground parking area</li> </ol> <p>Panoramas and planimetries are provided in a temporary version. The final version with post-processed panoramas and complete planimetries will be available soon. An example of the final panoramas is provided for position A of the printer room, while an example of complete planimetry is provided for the printer and the meeting rooms.</p> <h2>SOFA</h2> <p>The SRIRs are also provided in SOFA format&nbsp;<a href="https://sofacoustics.org/data/database/pan-ar/" target="_blank" rel="noopener">here</a>.</p> <h2>How to cite</h2> <p>If you use the PAN-AR dataset, please cite the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <pre><code>@inproceedings{denti2024panar,</code><br><code> title = {{PAN-AR}: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures},</code><br><code> author = {Denti, Filippo and Fantini, Davide and Avanzini, Federico and Presti, Giorgio},</code><br><code> year = {2024},</code><br><code> month = {September},</code><br><code> booktitle = {Proceedings of the 19th International Audio Mostly Conference (AM '24)},</code><br><code> location = {Milan, Italy},</code><br><code> publisher = {ACM},</code><br><code> isbn = {979-8-4007-0968-5/24/09},</code><br><code> doi = {10.1145/3678299.3678332}</code><br><code>}</code></pre>

opencc-by-sa-4.0Dec 2024View details →
zenodo40/100

Ambisonic Room Impulse Responses

<p>Datasets of Ambisonic room impulse responses from 3 european museums/touristic sites:</p> <ul> <li>La Fundaci&oacute;&nbsp;Miro, Barcelona, Spain</li> <li>Die Alte Pinakotheke, Munich, Germany</li> <li>St Andrews Castle, St Andrews, Scotland</li> </ul> <p>The datasets are in SOFA format&nbsp;of convention AmbisonicsDRIR, recently proposed by the author.</p>

opencc-by-nc-sa-4.0Sep 2018View details →
zenodo40/100

Ambisonic tracks

<p>Ambisonic tracks for both users.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Ambisonic Recordings of Typical Environments (ARTE) Database

<p><strong>Note:</strong> The sound recordings are only to be used for non-commercial personal, educational or research purposes.</p> <p><strong>Overview:&nbsp;</strong>The ARTE database is described in detail in Weisser et al. (2019) and was designed primarily to:</p> <ol> <li>Provide to the research community accessible multichannel recordings of a range of realistic acoustic everyday scenes that can be used in a large variety of auditory perception tests with improved ecological validity and played back in loudspeaker arrays of different geometries as well as on headphones.</li> <li>Enable standardization and replication of auditory perception tests that utilize realistic noisy environments.</li> <li>Complement the multichannel recordings with measured multichannel Room Impulse Responses (RIRs) as well as basic derived acoustic data.</li> </ol> <p>The ARTE database, so far, contains 13 acoustic environments that were recorded with a purpose-built 62-channel microphone array in various locations around Sydney (Australia), and was decoded into the higher-order Ambisonics (HOA) format.</p> <p>For each acoustic environment the following files are provided:</p> <p><strong>HOA environment files:</strong> The recorded environments were decoded into 31mixed-order HOA channels and saved as WAV-files with a sampling frequency of 44.1 kHz and 32 bits per sample. Thereby, channels 1-25 refer to the 3D HOA periphonic (horizontal) components up to the order of M = 4, and channels 26-31 refer to additional sectorial 2D components (i.e., m = n) up to the order of M = 7.</p> <p><strong>HOA RIR files: </strong>In each environment, Room Impulse Responses (RIRs) were measured with a Tannoy V8 dual-concentric loudspeaker at a number of positions relative to the microphone array. Currently, only a single RIR is provided in each environment which was measured with a loudspeaker in front of the microphone array (0 degree azimuth) at a distance of 1.3 m. Similar to the noise files, the RIRs are provided as 31-channel WAV-files with a sampling frequency of 44.1 kHz and 32 bits per sample. In addition to the &ldquo;standard&rdquo; RIR, a second version is provided in which the RIR was split into a direct sound (DS) component as well as a reverberation component (REV). The separated version of the RIR can be useful for enhancing the directionality (and frequency response) of the direct sound by decoding it into a single loudspeaker channel (i.e., a loudspeaker at an azimuth angle of 0 degrees) and then adding it back to the reverberant component, which is decoded normally. This process has been shown to be particularly useful when evaluating the benefit provided by directional signal enhancement methods (e.g., beamformers) in hearing aids.</p> <p><strong>Binaural environment files:</strong> The HOA noise files were transformed into binaural headphone signals by simulating their playback via a 41-channel loudspeaker array to the in-ear microphones of a calibrated Bruel &amp; Kjaer Head and Torso Simulator (HATS type 4128C). These binaural signals are provided in two versions: (a) an unprocessed version that needs to be presented via headphones that are equalized using an artificial ear and (b) a version that can be directly played back via any diffuse-field equalized headphones.&nbsp;</p> <p><strong>Binaural RIRs:</strong> The HOA RIRs were transformed into binaural RIRs in the same way as the HOA noise files (see above) and were saved both unequalized and diffuse-field equalized.</p> <p><strong>Basic acoustic measures:</strong> A number of basic acoustic measures are provided by a separate PDF-file for each environment, including: (a) unweighted sound pressure levels (dB SPL), (b) A-weighted sound pressure levels (dBA), (c) reverberation time (RT60), (d) third-octave power spectra in dB SPL, (e) temporal envelopes, (f) amplitude modulation spectra, and (g) directional characteristics in the horizontal plane. The acoustic measures were derived by simulating the playback of the MOA noise files (and RIRs) via a 41-channel loudspeaker array to a calibrated omni-directional 1/4&rdquo; GRAS microphone (Type 46BL).</p> <p>Apart from the acoustic environment specific files, the ARTE database includes a number of Matlab<sup>TM</sup> functions that help decoding the provided HOA files into a format that can be played back via a given loudspeaker array, and includes a number of examples.</p> <p>Further technical details are described in Weisser, et al. (2019).</p> <p><strong>Supporting material:&nbsp;</strong>The provided Matlab<sup>TM</sup> scripts and examples assume that the downloaded files are organized in a specific directory structure. This structure is generated automatically when downloading (and unzipping) the main zip-file (ARTE database downloas.7z). Note that this zip-file contains all required functions except the MOA and binaural sound files and RIRs. Due to their file size (about 10 GB in total), these sound files should be downloaded, one by one, from the individual links provided below.</p> <p><strong>Some notes on calibration:&nbsp;</strong>All HOA noise files were normalized in the same way such that they correctly maintain their original differences in sound pressure level. Hence, once the sensitivity of the loudspeaker playback system is known, the same playback gain must be applied to all noise files. Even though this playback gain can be derived using any of the provided noise files, the easiest noise file for calibrating the loudspeaker playback system is the provided diffuse noise due to its steady-state behavior. Given that most playback environments contain significant low-frequency background noise, and loudspeakers have different low-frequency roll-offs, the provided A-weighted sound pressure levels should be best used for calibration. Also, it is assumed here that all loudspeakers in the playback array have the same distance to the listener, identical sensitivity, and a flat frequency response. If this is not the case the loudspeakers need to be equalized individually. Also, reverberation of the playback room should be as low as possible.</p> <p><strong>Acknowledgement:&nbsp;</strong>The research related to&nbsp;the ARTE database was financially supported by the Oticon foundation as well as the HEARing CRC, established and supported under the Cooperative Research Centres Program &ndash; an initiative of the Australian Government, and the Oticon foundation.</p> <p><strong>References:&nbsp;</strong>Weisser, A., Buchholz, J. M., Oreinos, C., Badajoz-Davila, J., Galloway, J., Beechey, T., Keidser, G. (2019). The Ambisonics Recordings of Typical Environments (ARTE) database. Acta Acustica united with Acustica. (see provided pdf-file)</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Ambisonic Stimuli Files and Binaural Renders

<p>This repository hosts the audio files used to render the layer-based stimuli for the experiment described in 'A Study on Loudspeaker SPL Decays for Envelopment and Engulfment across an Extended Audience' (2024 AES International Conference on Acoustics and Sound Reinforcement).&nbsp;</p><p>Binaural renders of the 7th-order Ambisonic stimuli are included for headphone listening.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Ambisonic-Binaural

<p>A dataset collected for ambisonic-based binaural rendering. The recordings are collected at ByteDance in Nov. 2021.&nbsp;We recorded 31 minutes and 18 minutes of audios for training and testing respectively. These audios are paired ambisonics and binaurals.</p>

openmit-licenseOct 2022View details →
zenodo36/100

BINCI 360º demo video with 2nd Ambisonics audio

<p>BINCI - Binaural Tools for the Creative Industries</p> <p>There are multiple use cases as well as interpretations about immersive audio. As an aid for understanding BINCI approach of immersive audio we created  this first 360 demo video with a basi introduction.  You can download this video and watch and hear with Google Cardboard or even better, with Samsung Gear. </p>

opencc-by-nc-nd-4.0Nov 2017View details →
zenodo36/100

TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Evaluation Datasets

<p>This package consists of two evaluation datasets,&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;providing four-channel First-Order Ambisonic (FOA) recordings while&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These evaluation datasets are part of the&nbsp;<a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>.&nbsp;The corresponding development datasets can be downloaded <a href="https://doi.org/10.5281/zenodo.2599196">here</a>.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <ul> <li>The <strong>foa_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;evaluation dataset.</li> <li>The <strong>mic_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;evaluation dataset.</li> </ul> <p>-- Version 2 updates --</p> <p>The<a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection-results"> DCASE 2019 sound event localization and detection task has now ended</a>. Hence we are releasing the reference labels for the evaluation dataset in this version.</p> <ul> <li>The&nbsp;<strong><em>metadata_eval.zip</em></strong>&nbsp;is the common metadata for both&nbsp;<strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;evaluation datasets.&nbsp; &nbsp;</li> <li>The <strong>short2longnames.txt</strong>&nbsp;file consists of the corresponding names for each recording in the dataset in the <a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection#development-dataset">development-set format</a>, i.e., including the information of the impulse response location and the maximum number of overlapping sound events in the recording.</li> </ul> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

openother-ncMay 2019View details →
zenodo32/100

Ambisonics Binaural Rendering via Masked Magnitude Least Squares - Supplemental Material

<p>Refer to the readme file for more information.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

FOA-MEIR Dataset: multi-environment impulse response recordings with a first-order ambisonic microphone

<p>FOA-MEIR is an impulse response (IR) dataset recorded in over 100 environments for use in sound event localization and detection (SELD) tasks. This dataset is set up to develop a robust SELD system in an unknown environment, and the IRs for the inferred environment are recorded at a different location from that of training data. The dataset also contains dry source recordings that can be combined with IR recordings to generate audio clips for training the SELD task.</p> <p>License: see the file named LICENSE.pdf</p> <p>Further information is available at [1] and Github: https://github.com/nttrd-mdlab/seld-foa-meir<br> <br> [1] Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, &ldquo;Echo-aware Adaptation of Sound Event Localization and Detection in Unknown Environments,&rdquo; in IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), 2022.</p>

openother-ncFeb 2022View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Anechoic and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Anechoic, and Synthetic Impulse Response Dataset&nbsp;</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of [1 10] meters from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated reverberant first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate.&nbsp;The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with&nbsp;reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz&ndash;4 kHz band center frequencies.</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least&nbsp;1 meter away from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT)&nbsp;Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset</strong></p> <p>This dataset consists of real-life first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset was&nbsp;generated by collecting impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing<br> a maximum length sequence around the array in circular trajectory in one elevation at a time. The playback volume was set to be 30 dB greater than the ambient sound level. The recording was done in a corridor inside the university with classrooms around it during work hours.The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20&nbsp;to 20&nbsp;with 10-degree increments at 2 m.&nbsp;</p> <p>The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://serv.cusp.nyu.edu/projects/urbansounddataset/urbansound8k.html">urbansound8k dataset</a>.&nbsp;This dataset consists of 10 sound event classes such as air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. We do not consider the air_conditioner and children_playing sound events. Further, we only include the sound event examples marked as foreground in the dataset. We used the splits 1, 8 and 9 provided in the urbansound8k as the three CV splits. These splits were chosen as they had a good number of examples for all the chosen sound event classes after selecting only the foreground examples.&nbsp;During the sound scene synthesis, we randomly chose a sound event example and associated it with a random distance among the collected ones, azimuth and elevation angle. The sound event example was then convolved with the respective IR for the given distance, azimuth and elevation to spatially position it.</p> <p>The metadata.zip folder consists of the license and the metadata for the complete dataset. The rest of the nine zip files consists dataset for given split and overlap. For example, the&nbsp;wav_ov3_split1_30db.zip file consists of training and testing recordings for the case of maximum three temporally overlapping sound events (ov3) for the first cross-validation split (split1). Within each audio folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p><strong>Data collector (s): </strong>Fagerlund, Eemi;&nbsp;Koskimies, Aino</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Tietotalo Ambisonic Impulse Response

<p><strong>Tampere University of Technology (TUT) Tietotalo Ambisonic Impulse Response</strong></p> <p>This dataset consists of impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The recordings were done in a fairly large spaced corridor inside the university (Tietotalo building) with classrooms around it. The IR acquisition was done using a maximum length sequence (MLS). The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing the MLS around the Eigenmike in a circular trajectory. The playback volume was set to be 30 dB greater than the ambient sound level. The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20 to 20 with 10-degree increments at 2 m.&nbsp;</p> <p>The moving-source IRs were obtained by a freely available tool from CHiME challenge which estimates the time-varying responses in STFT domain by forming a least-squares regression between the known measurement signal and the far-field recording independently at each frequency. The IR for any azimuth within one trajectory can be analyzed by assuming block-wise stationarity of acoustic channel. The CHiME IR estimation tool was applied independently on all 32 channels of the Eigenmike. For the dataset creation, we analyzed the DOA of each time frame using MUSIC and extracted IRs for azimuthal angles at 10&deg; resolution (36 IRs for each elevation).</p> <p>The IR file is in .mat format and can be read both in Matlab and Python. The details of the IR file are as following,</p> <p>Size: (2, 9, 1025, 36, 4, 32) = (distance_wrt_mic, elevation_wrt_mic, FFT, &nbsp;azimuth_wrt_mic, blocks, channels).</p> <p>where,</p> <p>distance_wrt_mic = two distances (1m and 2m)<br> elevation_wrt_mic = 9 elevation angles (-40:10:40) at distance 1m, and 5 elevations angles (-20:10:20) at distance 2m.<br> azimuth_wrt_mic = 36 azimuth angles (-180:10:180) for all distance-elevation combination<br> The IRs were extracted assuming block-wise stationarity (four blocks) for each frequency bin (1025 bins).</p> <p>During synthesis, after convolving the IR with a&nbsp;sound event, the 32 channel audio will have to be transformed to Ambisonic format using the transformation matrix of Eigenmike.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources using convolutional recurrent neural network</a>&#39; work, more details about this IR dataset can be found in this work.</p> <p><strong>Data collector (s):</strong> Fagerlund, Eemi; Koskimies, Aino</p>

openother-ncOct 2018View details →
zenodo32/100

TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Development Datasets

<p>This package consists of two development datasets, <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;providing four-channel First-Order Ambisonic (FOA) recordings while <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These development datasets are part of the <a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>.</p> <p>Both the development set consists of 400, one minute long recordings sampled at 48000 Hz, and divided into four cross-validation splits of 100 recordings each. These recordings were synthesized using spatial room impulse response (IRs) collected from five indoor locations, at 504 unique combinations of azimuth-elevation-distance. Furthermore, in order to synthesize the recordings, the collected IRs were convolved with <a href="http://www.cs.tut.fi/sgn/arg/dcase2016/task-sound-event-detection-in-synthetic-audio#audio-dataset">isolated sound events dataset from DCASE 2016 task 2</a>. Finally, to create a realistic sound scene recording, natural ambient noise collected in the IR recording locations was added to the synthesized recordings such that the average SNR of the sound events was 30 dB.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <p><strong>Download instructions</strong></p> <p>The three files, &nbsp;<strong><em>foa_dev.z01</em></strong>,<strong><em> foa_dev.z02</em></strong>&nbsp;and <strong><em>foa_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;development dataset.<br> The two files, <strong><em>mic_dev.z01</em></strong>&nbsp;and, <strong><em>mic_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development dataset.<br> The <strong><em>metadata_dev.zip</em></strong>&nbsp;is the common metadata for both <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development datasets.</p> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> &nbsp;</p>

openother-ncFeb 2019View details →
zenodo32/100

TAU Moving Sound Events 2019 - Ambisonic, Anechoic, Synthetic IR and Moving Source Dataset

<p><strong>Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Anechoic and Synthetic Impulse Response (IR) and Moving Source Dataset</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location and directional spatial location in azimuth and elevation angles (in degrees), angular velocity of motion, and distance from the microphone (in meters).</p> <p>The isolated sound events were taken from the DCASE 2016 task 2 dataset. This dataset consists of 11 sound event classes such as Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. Every event is assigned a spatial trajectory on an arc with a constant distance from the microphone (in the range 1-10 m) and moving with a constant angular velocity for its duration. Due to the choice of the ambisonic spatial recording format, the steering vectors for a plane wave source or point source in the far field are frequency-independent. Hence, there is no need for a time-variant convolution or impulse response interpolation scheme as the source is moving; the spatial encoding of the monophonic signal was done sample-by-sample using instantaneous ambisonic encoding vectors for the respective DOA of the moving source. The synthesized trajectories in the dataset vary in both azimuth and elevation and are simulated to have a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the &#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks&#39;</a> work.</p>

openother-ncApr 2019View details →
zenodo24/100

TAU Moving Sound Events 2019 - Ambisonic, Reverberant, Real-life IR and Moving Source Dataset

<p><strong>Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Reverberant and Real-life Impulse Response and Moving Source Dataset</strong></p> <p>This dataset consists of real-life first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset was generated by collecting impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing<br> a maximum length sequence around the array in circular trajectory in one elevation at a time. The playback volume was set to be 30 dB greater than the ambient sound level. The recording was done in a corridor inside the university with classrooms around it during work hours. The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20 to 20 with 10-degree increments at 2 m.&nbsp;</p> <p>The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. All sound events in this dataset are moving only along azimuth with a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location in azimuth and elevation angles (in degrees), the angular velocity of motion and distance from the microphone (in meters).</p> <p>The isolated sound events were taken from the urbansound8k dataset. This dataset consists of 10 sound event classes such as air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. We do not consider air_conditioner and children_playing sound events. Further, we only include the sound event examples marked as foreground in the dataset. We used the splits 1, 8 and 9 provided in the urbansound8k as the three CV splits. These splits were chosen as they had a good number of examples for all the chosen sound event classes after selecting only the foreground examples. During the sound scene synthesis, every sound event is assigned a spatial trajectory on an arc with a constant distance from the microphone and moving with a constant angular velocity for its duration.</p> <p>Other than the license file, there are nine zip files that consist of the dataset and corresponding metadata for given split and overlap. For example, the ov3_split1.zip file consists of training and testing recordings and metadata for the case of a maximum of three temporally overlapping sound events (ov3) for the first cross-validation split (split1). Within each folder, the filenames for training split have the &#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks&#39;</a> work.</p> <p>Data collector (s): Fagerlund, Eemi; Koskimies, Aino; Hakala, Aapo</p>

openother-ncApr 2019View details →
zenodo24/100

Perceptually-motivated spatial audio codec for higher-order Ambisonics compression - Examples

<p>Scene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide-range of (potentially unknown) devices. The number of channels required to deliver high spatial resolution Ambisonic audio, however, can be prohibitive for low-bandwidth applications. Therefore, in this paper, a compression codec is proposed, which is based upon the higher-order Directional Audio Coding (HO-DirAC) model. The encoder downmixes the higher-order Ambisonics (HOA) input audio into a reduced number of signals, which are accompanied by spatial parameterization metadata. The downmixed audio is coded using a perceptual audio coder, whereas the metadata is grouped into perceptual bands, quantised, and downsampled.&nbsp;On the decoder side, low Ambisonic orders are fully recovered. Whereas, not fully recoverable high Ambisonic orders are synthesized based on the spatial metadata. The results of a listening test indicate that the proposed parametric spatial audio codec can improve the adopted perceptual coder, especially at low to medium-high bitrates, when applied to fifth-order HOA signals.</p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record