Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,300

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,300 results for “Sounds”

Learn how ShareScore rates datasets ↗
zenodo32/100

Urban background sound recordings for virtual acoustics under various weather conditions at IHTApark

<p>This is an open database of calibrated background sound recordings with metadata on meteorological and acoustical parameters ready to be used in virtual reality applications, for soundsacpe studies.</p> <p>The soundscapes have been recorded at the IHTApark (green space next to the Institute for Hearing Technology and Acoustics) in Aachen (Germany) in winter and spring of 2022.</p> <p>They have been segemented into 30-seconds auralizable snippets.</p> <p>For an<strong> <a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener">interactive exploration</a></strong> of the soundscapes, please visit the follwing website:</p> <p><a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener"></a></p> <p><a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener">https://paad-group.github.io/IHTApark-ambient-recordings/</a></p> <p>The dabase is available as binaural files, first-order ambisonics files, and as omni files.</p> <p>The metadata is included as a .xlsx file.</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset of Gunshot Sounds and Koogu Model related to "Impacts of logging, hunting, and conservation on vocalizing biodiversity in Gabon" by Yoh et al. 2024

<h2>Description</h2> <p>This repository contains gunshot sound training data, testing data, and the Convolutional Neural Network (CNN) model used in our study on gunhunting patterns in Gabon. The dataset includes hundreds of audio recordings containing gunshots and other environmental sounds, their associated annotations, and the trained Koogu machine-learning model.</p> <h2>Data Summary</h2> <ul> <li><strong>Data Types:</strong> Gunshot sound recordings, non-gunshot sound recordings, annotations, model files. Data are in the format required for Koogu, where training and test data are contained in folders of audio files with associated folders containing the Raven Pro selection tables.&nbsp;</li> <li><strong>Data Format:</strong> <ul> <li>Audio files: .wav</li> <li>Annotations: .txt</li> <li>Model: TensorFlow/Keras model (.h5)</li> <li>Code: .py</li> </ul> </li> </ul> <h2>Data Details</h2> <p><em>Training and Test Data:</em></p> <ul> <li>train_annotations <ul> <li>"Training_Gunshots_Batch1.txt" is a Raven Pro selection table that contains 203 gunshot annotations that were detected through the spectral cross-correlation analysis.&nbsp;&nbsp;</li> <li>"Training_Gunshots_Batch2.txt" is a Raven Pro selection table that contains 223 gunshot annotations that were detected through the initial Koogu model that was run over the whole dataset.</li> <li>"Training_NotGunshots.txt" is a Raven Pro selection table that contains 8,614 non-gunshot annotations, including sounds from branch snaps, tree falls, calls of the putty-nosed monkey, ambient noise, etc.</li> </ul> </li> <li>train_audio <ul> <li>"audio_files" contains 184 audio files associated with the selection table "Training_Gunshots_Batch1.txt"</li> <li>"audio_files_second_batch" contains 203 clips associated with the selection table "Training_Gunshots_Batch2.txt".&nbsp;</li> <li>"other_clips" contains 9,274 clips associated with the selection table "Training_NotGunshots.txt".</li> </ul> </li> <li>test_annotations <ul> <li>"allgunshots_test_gunshots_adjusted_final_clean_20221120.txt" is a Raven Pro selection table that contains the 138 gunshot annotations associated with the files in the "test_audio" folder. This comprises the test dataset.&nbsp;&nbsp;</li> </ul> </li> <li>test_audio<br> <ul> <li>Contains the 131 recordings associated with the 138 gunshot annotations in "allgunshots_test_gunshots_adjusted_final_clean_20221120.txt"</li> </ul> </li> </ul> <p><em>Koogu Model:</em></p> <ul> <li>qDN_4x8_x4_16_Allsites_V3<br> <ul> <li>This is the Koogu model used in this study. It can be run using the script "Koogu_Run_Local_Python_Gabon_ModelV3.py" that you modify depending on the locations of your data and model.&nbsp;</li> </ul> </li> </ul> <p><em>CoLab</em>:&nbsp;An example Koogu CoLab page that contains the code used to train and test the model is available <a href="https://colab.research.google.com/drive/1TerOhzsCs9zSMj9uQ31HX7XPKG1XFYCt?usp=sharing">here.</a></p> <h2>Data Collection</h2> <ul> <li><strong>Collection Method:</strong> BAR-LT audio recorders deployed in 110 sites within Gabonese national parks, logging concessions, and community forests.</li> <li><strong>Time Period:</strong> Recordings from February 2021 to June 2022.</li> <li><strong>Geographical Information:</strong> Specific coordinates provided upon request.&nbsp;</li> </ul> <h2>Data Preparation</h2> <ul> <li><strong>Downsampling:</strong> Recordings were downsampled from 44.1 kHz to 4 kHz.</li> <li><strong>Segmentation:</strong> Audio recordings split into consecutive 2.25-second segments with 1.5 seconds of overlap.</li> <li><strong>Normalization:</strong> Waveform normalized to [-1.0, 1.0].</li> <li><strong>Spectrogram Computation:</strong> Using a 64 ms window length and 50% overlap, trimmed to 10&ndash;1200 Hz frequency range.</li> <li><strong>Training Inputs:</strong> Consisted of 584 positive class (gunshots) and 25842 negative class spectrograms.</li> </ul> <h2>Requirements</h2> <ul> <li><strong>Software Requirements:</strong> Python 3.8, TensorFlow 2.5, NumPy, Pandas.</li> <li><strong>Hardware Requirements:</strong> GPU with at least 8GB VRAM recommended.</li> </ul> <h2>References</h2> <ul> <li>For more information about this dataset and the model specifications, please consult "Impacts of logging, hunting, and conservation on vocalizing biodiversity in Gabon" (Yoh et al. 2024)</li> </ul> <h2>How to Cite this Dataset:</h2> <ul> <li>If you use this dataset in your work, please cite it: <ul> <li>Gottesman, B. (2024). Dataset of Gunshot Sounds and Koogu Model related to "Impacts of logging, hunting, and conservation on vocalizing biodiversity in Gabon" by Yoh et al. 2024 (1.0.0) [Data set]. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.11192704" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.11192704</a></li> </ul> </li> </ul>

opencc-by-nc-4.0May 2024View details →
zenodo32/100

MmWave 28 GHz MIMO channel sounding data collected in a conference room

<p>Channel sounding data sampled using a 28 GHz switched array MIMO channel sounder in a conference room.</p> <p>More information can be found in <a href="https://arxiv.org/abs/2404.19297">the corresponding publication</a> (DOI: 10.1109/TCOMM.2024.3392805), while the underlying data employs a self-explanatory structure, implemented in Matlab.</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

The Sounds of Home 2.3

<p>The Sounds of Home dataset consists of recordings from 14 microphones, each housed in its own folder. Here you will find Folder 14.<br><br>For Folders 1 to 7, access here:&nbsp;<a href="../records/12737915">https://zenodo.org/records/12737915</a></p> <p>For Folders 8 to 10, access here:&nbsp;<a href="../records/12800789">https://zenodo.org/records/12800789</a></p> <p>For Folder 11 to 13, access here: <a href="https://doi.org/10.5281/zenodo.12802632">https://doi.org/10.5281/zenodo.12802632</a></p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

The Sounds of Home 2.2

<p>The Sounds of Home dataset consists of recordings from 14 microphones, each housed in its own folder. Here you will find Folders 11 to 13.<br><br>For Folders 1 to 7, access here:&nbsp;<a href="../records/12737915">https://zenodo.org/records/12737915</a></p> <p>For Folders 8 to 10, access here: <a href="../records/12800789">https://zenodo.org/records/12800789</a></p> <p>For Folder 14, access here:&nbsp;<a href="https://doi.org/10.5281/zenodo.12802716">https://doi.org/10.5281/zenodo.12802716</a></p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

The Sounds of Home 2.1

<p>The Sounds of Home dataset consists of recordings from 14 microphones, each housed in its own folder. Here you will find Folders 8 to 10.<br><br>For Folders 1 to 7, access here: <a href="../records/12737915">https://zenodo.org/records/12737915</a></p> <p>For Folders 11 to 13, access here: <a href="https://doi.org/10.5281/zenodo.12802632">https://doi.org/10.5281/zenodo.12802632</a></p> <p>For Folder 14, access here: <a href="https://doi.org/10.5281/zenodo.12802716">https://doi.org/10.5281/zenodo.12802716</a></p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Analysis of Orbital Sounding in Context with In Situ Ground Penetrating Radar at Jezero Crater, Mars

<p>This release includes all the SHARAD data used to produce analysis and figures in the paper, "Analysis of Orbital Sounding in Context with In Situ Ground Penetrating Radar at Jezero Crater, Mars", by M.C. Raguso et al., submitted to Geophysical Research Letters in February 2024.</p> <p>The manuscript is currently under review. The dataset will be released following the completion of the review process.</p> <p>This release also includes slides (pdf format) presented during the RIMFAX Science Team meeting (09/6/22-09/09/22).</p> <p>Preferred citation (DataCite format):&nbsp;</p> <p>Raguso, M.C., &amp; Nunes, D.C. (2024). Analysis of Orbital Sounding in Context with In Situ Ground Penetrating Radar at Jezero Crater, Mars. [Dataset]. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.10681430">https://doi.org/10.5281/zenodo.10681430</a></p> <p><em>For inquiries regarding the contents of this dataset, please contact the Corresponding Author listed in the README.txt file.</em></p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Realistic urban sound mixture dataset

<p>This dataset resumes an urban sound corpus whose the realism has been proved through a perceptual test [1]. This corpus has been used in order to estimate the traffic sound level with the Non-negative Matrix Factorization formula [2].</p> <p>This dataset presents 4 folders :</p> <ul> <li><em>dictionary</em> where the traffic audio samples dedicated to the dictionary design of NMF are,</li> <li><em>recordings</em> which contains the 74 original recordings,</li> <li><em>annotation</em> which contains the annotations text files of the 74 audio files,</li> <li><em>transcribed scenes </em>which contains the 74 transcribed audio files generated with <a href="https://bitbucket.org/mlagrange/simscene"><em>SimScene</em></a> software. 4 folders composed it, according to the sound environment of the audio files (<em>park, quiet street, noisy street, very noisy street). </em> In each folder, one can find the global sound mixtures, the audio of each sound class as well as the files that include all the elements associated with the traffic and the interfering class (which contains all the other sound sources).</li> </ul> <p>[1] Gloaguen, J. R., Can, A., Lagrange, M., &amp; Petiot, J. F. (2017, June). Creation of a corpus of realistic urban sound scenes with controlled acoustic properties. In <em>173rd Meeting of the Acoustical Society of America and the 8th Forum Acusticum (Acoustics&#39; 17)</em>.</p> <p>[2] Gloaguen, J. R., Can, A., Lagrange, M., &amp; Petiot, J. F. (2018), Road traffic sound level from realistic urban sound mixtures by Non-negative Matrix Factorization, submitted for publication</p>

opencc-by-4.0Mar 2018View details →
zenodo32/100

Isolated urban sound database

<p>The Isolated urban sound database contains the audio samples used to design urban sound mixtures using <a href="https://bitbucket.org/mlagrange/simscene">SimScene</a> software.&nbsp;</p> <p>This database has already been used to design urban sound mixtures that can be found in&nbsp; <a href="https://zenodo.org/record/1145855">Estimation of the road traffic sound levels based on Non-Negative Matrix Factorization dataset</a> [1] and in <a href="https://zenodo.org/record/1184443">Realistic urban sound mixture dataset</a> [2]</p> <p>The dataset contains two folders :</p> <p>- &#39;event&#39; which includes includes 231 brief sound samples considered as salient, with a 1 to 20 seconds duration and classified among 21 sound classes (ringing bell, whistling bird, car horn, passing car, hammer, barking dog, siren, footstep, metallic noise, voice...)</p> <p>- &#39;background&#39; which includes 162 long duration sounds (~1mn30), whose acoustic properties do not vary in time. This category includes among others, whistling bird, crowd noise, rain, children playing in schoolyard, constant traffic noise ...</p> <p>More details on this sound database can be found in [3]</p> <p>&nbsp;</p> <p>[1] J.-R. Gloaguen, M. Lagrange, A. Can, J.-F. Petiot, Estimation of the road traffic sound levels in urban areas based on non-negative matrix factorization techniques, submitted for publication</p> <p>[2] J.-R. Gloaguen, A. Can, M. Lagrange, J.-F. Petiot, Road traffic sound level estimation from realistic urban sound mixtures by Non-negative Matrix Factorization, submitted for publication</p> <p>[3] J.-R. Gloaguen, A. Can, M. Lagrange, J.-F. Petiot, Creation of a corpus of realistic urban sound scenes with controlled acoustic properties, in: Acoustics &rsquo;17 Boston, Vol. 141 of The Journal of the Acoustical Society of America, Acoustical Society of America and the European Acoustics Association, Boston, United States, 2017, pp. 4044&ndash;4044.</p>

opencc-by-4.0Apr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Circular array, Reverberant and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Circular array, Reverberant and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated, reverberant, and circular-array format recordings with&nbsp;stationary point sources each associated with a spatial coordinate.&nbsp;The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with&nbsp;reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz&ndash;4 kHz band center frequencies.</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least&nbsp;1 meter away from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Anechoic and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Anechoic, and Synthetic Impulse Response Dataset&nbsp;</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of [1 10] meters from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Circular array, Anechoic and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Circular array, Anechoic and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated anechoic circular-array format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of [1 10] meters from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT) Sound Events 2018 - Ambisonic, Reverberant and Synthetic Impulse Response Dataset</strong></p> <p>This dataset consists of simulated reverberant first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate.&nbsp;The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters). The sound events are spatially placed within a room using the image source method. The room size chosen was 10x8x4 meter with&nbsp;reverberation time per octave band of [1.0, 0.8, 0.7, 0.6, 0.5, 0.4] s and 125 Hz&ndash;4 kHz band center frequencies.</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://archive.org/details/dcase2016_task2_train_dev">DCASE 2016 task 2 dataset.</a> This dataset consists of 11 sound event classes such as&nbsp;Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. The sound events are randomly placed in a spatial&nbsp;grid with 10-degree resolution in full azimuth and [-60 60) degree elevation angles. Additionally, the sound events are placed at a random distance of at least&nbsp;1 meter away from the microphone.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of&nbsp;datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

TUT Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset

<p><strong>Tampere University of Technology (TUT)&nbsp;Sound Events 2018 - Ambisonic, Reverberant and Real-life Impulse Response Dataset</strong></p> <p>This dataset consists of real-life first order Ambisonic (FOA) format recordings with&nbsp;stationary point sources each associated with a spatial coordinate. The dataset was&nbsp;generated by collecting impulse responses (IR) from a real environment using the Eigenmike spherical microphone array. The measurement was done by slowly moving a Genelec G Two loudspeaker continuously playing<br> a maximum length sequence around the array in circular trajectory in one elevation at a time. The playback volume was set to be 30 dB greater than the ambient sound level. The recording was done in a corridor inside the university with classrooms around it during work hours.The IRs were collected at elevations &minus;40 to 40 with 10-degree increments at 1 m from the Eigenmike and at elevations &minus;20&nbsp;to 20&nbsp;with 10-degree increments at 2 m.&nbsp;</p> <p>The dataset consists of three sub-datasets with a) maximum one temporally&nbsp;overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240&nbsp;recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), spatial location in azimuth and elevation angles (in degrees), and distance from the microphone (in meters).</p> <p>The isolated&nbsp;sound events were taken from the <a href="https://serv.cusp.nyu.edu/projects/urbansounddataset/urbansound8k.html">urbansound8k dataset</a>.&nbsp;This dataset consists of 10 sound event classes such as air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. We do not consider the air_conditioner and children_playing sound events. Further, we only include the sound event examples marked as foreground in the dataset. We used the splits 1, 8 and 9 provided in the urbansound8k as the three CV splits. These splits were chosen as they had a good number of examples for all the chosen sound event classes after selecting only the foreground examples.&nbsp;During the sound scene synthesis, we randomly chose a sound event example and associated it with a random distance among the collected ones, azimuth and elevation angle. The sound event example was then convolved with the respective IR for the given distance, azimuth and elevation to spatially position it.</p> <p>The metadata.zip folder consists of the license and the metadata for the complete dataset. The rest of the nine zip files consists dataset for given split and overlap. For example, the&nbsp;wav_ov3_split1_30db.zip file consists of training and testing recordings for the case of maximum three temporally overlapping sound events (ov3) for the first cross-validation split (split1). Within each audio folder, the filenames for training split have the&nbsp;&#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of&nbsp;the &#39;<a href="https://github.com/sharathadavanne/seld-net">Sound event localization and detection of overlapping sources&nbsp;using convolutional recurrent neural network</a>&#39; work.</p> <p><strong>Data collector (s): </strong>Fagerlund, Eemi;&nbsp;Koskimies, Aino</p> <p>&nbsp;</p>

openother-ncApr 2018View details →
zenodo32/100

Supplementary file 2 for "Pushes and pulls from below: anatomical variation, articulation and sound change": the data

<p>The data (as TAB-separated .TSV files and .RData standard R data files) and Rmarkdown script needed to reproduce the results reported in the main paper as a single .TAR.XZ archive.</p> <p>Please note that the actual results obtained when compiling this script may differ slightly from those reported in the paper, especially for Markov-Chain Monte Carlo (due to the heavy use of random numbers), but also for the Maximum Likelihood (due to differences between particular BLAS/LAPACK implementations).</p> <p>The first compilation of this script is quite expensive computationally (on an Intel Core i7-3770 \@ 3.40 GHz with 32 Gb RAM this took about 25 hours), but the next ones are much faster as the most expensive results are cached locally.</p>

opencc-by-4.0Nov 2018View details →
zenodo32/100

TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Development Datasets

<p>This package consists of two development datasets, <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;providing four-channel First-Order Ambisonic (FOA) recordings while <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These development datasets are part of the <a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>.</p> <p>Both the development set consists of 400, one minute long recordings sampled at 48000 Hz, and divided into four cross-validation splits of 100 recordings each. These recordings were synthesized using spatial room impulse response (IRs) collected from five indoor locations, at 504 unique combinations of azimuth-elevation-distance. Furthermore, in order to synthesize the recordings, the collected IRs were convolved with <a href="http://www.cs.tut.fi/sgn/arg/dcase2016/task-sound-event-detection-in-synthetic-audio#audio-dataset">isolated sound events dataset from DCASE 2016 task 2</a>. Finally, to create a realistic sound scene recording, natural ambient noise collected in the IR recording locations was added to the synthesized recordings such that the average SNR of the sound events was 30 dB.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <p><strong>Download instructions</strong></p> <p>The three files, &nbsp;<strong><em>foa_dev.z01</em></strong>,<strong><em> foa_dev.z02</em></strong>&nbsp;and <strong><em>foa_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;development dataset.<br> The two files, <strong><em>mic_dev.z01</em></strong>&nbsp;and, <strong><em>mic_dev.zip</em></strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development dataset.<br> The <strong><em>metadata_dev.zip</em></strong>&nbsp;is the common metadata for both <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong>&nbsp;and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>&nbsp;development datasets.</p> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> &nbsp;</p>

openother-ncFeb 2019View details →
zenodo32/100

Plane-wave propagation path data from wideband MIMO channel sounding in an urban microcellular scenario

<p>We provide plane-wave propagation path data from a wideband MIMO radio channel sounding measurement in an urban microcell scenario. The binary Matlab file includes: direction of departure (DOD: variables &quot;par.PhiTx&quot; and &quot;par.ThetaTx&quot; in [rad]), direction of arrival (DOA: &quot;par.PhiRx&quot; and &quot;par.ThetaRx&quot; in [rad]), delay (&quot;par.Tau&quot; to be multiplied with 8.3ns, the tab length), and complex polarimetric path gain (&quot;par.Alpha&quot; is a 2x2 matrix, where element [1,1]=TXtheta -&gt;RXtheta, [1,2]=TXphi -&gt;RXtheta, [2,1]=TXtheta-&gt;RXphi, and [2,2]=TXphi-&gt;RXphi), for the 30 strongest signal paths (from TX to RX) at each of the 4574 RX locations along the route described below. In the element names above, the term &quot;theta&quot; refers to the vertically polarised component, and accordingly the term &quot;phi&quot; referes to the horizontally polarised component.<br> Note #1: The exact RX location for each individual measured radio channel was NOT recorded (see route description below). &nbsp;<br> Note #2: The complex path gain (par.Alpha) is NOT calibrated, but depends on the initially fixed AGC level in the receiver, which was chosen to provide the best dynamic range for the given mobile (RX) route.<br> Both these limitations are seen reasonable since this dataset is meant for the realistic *statistical comparison* of the performance of different RX antennas in a microcell environment (and not to determine the actual received power at each exact location of the measured route).<br> Background information: The provided dataset is processed and is based on a radio channel sounder measurement at 5.3 GHz, carried out in downtown Helsinki, Finland, in April 2004. The uniform rectangular transmit (TX) array was placed at 10 m height in Aleksanterinkatu-street (an approx. 15-m wide street canyon), in front of the Nordea building, broadside pointing westwards (towards Stockmann building). The semishperical receive (RX) array was moved at 1.6-m height and for about 50 m along Aleksanterinkatu-street in line-of-sight (LOS), i.e. from in front of Kluuvi shopping centre westwards just across the crossing of Kluuvikatu-street. The TX and RX arrays cover the relevant azimuth and elevation ranges, so that this plane wave propagation path data can directly be combined with the polarimetric directional radiation pattern(s) of an antenna (array).</p>

opencc-by-nc-nd-4.0Mar 2019View details →
zenodo32/100

TAU Moving Sound Events 2019 - Ambisonic, Anechoic, Synthetic IR and Moving Source Dataset

<p><strong>Tampere University (TAU) Moving Sound Events 2019 - Ambisonic, Anechoic and Synthetic Impulse Response (IR) and Moving Source Dataset</strong></p> <p>This dataset consists of simulated anechoic first order Ambisonic (FOA) format recordings with moving point sources each in 2D spherical space represented with azimuth and elevation angles. The dataset consists of three sub-datasets with a) maximum one temporally overlapping sound events, b) maximum two temporally overlapping sound events, and c) maximum three temporally overlapping sound events. Each of the sub-datasets has three cross-validation splits, that consists of 240 recordings of about 30 seconds long for training split and 60 recordings of the same length for the testing split. For each recording, the metadata file with the same name consists of the sound event name, the temporal onset and offset time (in seconds), starting spatial location and directional spatial location in azimuth and elevation angles (in degrees), angular velocity of motion, and distance from the microphone (in meters).</p> <p>The isolated sound events were taken from the DCASE 2016 task 2 dataset. This dataset consists of 11 sound event classes such as Clearing throat, Coughing, Door knock, Door slam, Drawer, Human laughter, Keyboard, Keys (put on a table), Page turning, Phone ringing and Speech. Every event is assigned a spatial trajectory on an arc with a constant distance from the microphone (in the range 1-10 m) and moving with a constant angular velocity for its duration. Due to the choice of the ambisonic spatial recording format, the steering vectors for a plane wave source or point source in the far field are frequency-independent. Hence, there is no need for a time-variant convolution or impulse response interpolation scheme as the source is moving; the spatial encoding of the monophonic signal was done sample-by-sample using instantaneous ambisonic encoding vectors for the respective DOA of the moving source. The synthesized trajectories in the dataset vary in both azimuth and elevation and are simulated to have a constant angular velocity in the range [-90, 90]/s with 10-degree/s steps.</p> <p>The license of the dataset can be found in the LICENSE file. The rest of the nine zip files consists of datasets for a given split and overlap. For example, the ov3_split1.zip file consists of the audio and metadata folders for the case of maximum three temporally overlapping sound events (ov3) and the first cross-validation split (split1). Within each audio/metadata folder, the filenames for training split have the &#39;train&#39; prefix, while the testing split filenames have the &#39;test&#39; prefix.</p> <p>This dataset was collected as part of the &#39;<a href="https://github.com/sharathadavanne/seld-net">Localization, Detection and Tracking of Multiple Moving Sound Sources with Convolutional Recurrent Neural Networks&#39;</a> work.</p>

openother-ncApr 2019View details →
zenodo32/100

Sound event localization and detection (SELDnet) results

<p>This package is part of the work -&nbsp;<a href="https://github.com/sharathadavanne/seld-metric">Joint Measurement of Localization and Detection of Sound Events</a>&nbsp;presented in WASPAA 2019.</p> <p>This package consists of results from the <a href="https://arxiv.org/abs/1905.08546">SELDnet method</a> for joint&nbsp;sound event localization and detection. The results corresponding to different training states of 5, 25 and 75 epochs are provided here.&nbsp; These results are for the four cross-validation splits of the&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset. The sound events in this dataset&nbsp;consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle.&nbsp;This <strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset is part of the&nbsp;<a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>&nbsp;and can be downloaded <a href="https://zenodo.org/record/2599196#.XT_RmHUzaCg">here</a>.</p> <p>Each of the results folders consists&nbsp;of 400 files, corresponding to the results of the individual recordings of the&nbsp;<strong>TAU Spatial Sound Events 2019 - Microphone Array </strong>dataset.</p> <p>This data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <p><strong>Download instructions</strong></p> <p>The three files, &nbsp;<strong><em>mic_dev_5</em></strong>,<strong><em>&nbsp;mic_dev_25,&nbsp;</em></strong>and&nbsp;<strong><em>mic_dev_75</em></strong>, correspond to SELDnet results at 5, 25 and 75 epochs.</p> <p>Download the zip files&nbsp;and use your favorite compression tool to unzip these split zip files.</p>

openother-ncJul 2019View details →
zenodo32/100

Glyph sets for increased sound-shape systematicity

<p>A set of glyphs from parametric fonts, along with measurements of their sound-shape systematicity.</p>

opencc-by-4.0Aug 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record