Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,300
datasets available to search
ShareScore release 0.9.0
Dataset results
1,300 results for “Sounds”
Dataset for Sound-based Anomalies Detection in Agricultural Robotics Application
<p>This data set contains data related to a Mowing Intelligent Tool (MowIT).</p> <p>Two different microphones were used to collect the sound samples, recording the audio with just one single channel, with a sampling rate of 44100 Hz and 16 bits resolution.</p> <p>The data provided by an inertial measurement unit (IMU) was also recorded since that was already integrated into the MowIT.</p> <p>Two different data collections were performed in different open-air environments with grass to cut.</p> <p>In each collection, eight different sample sets were made, five with the machine cutting using a trimmer line and the other three using the blades. Various combinations were used in each set, and tools were or were not placed on each of the three cutting axes of the MowIT. For each group, the acquisitions were designated from 0 to 7.</p> <p>Each folder of the first collection is a combination containing two audio files, one for each microphone used, the IMU data and a photograph of the lower part of the MowIT to understand the configuration used.</p> <p>In the second collection, to improve the variety of data, three distinct sub-sets were performed for combination: the first with the MowIT turned on but not cutting grass and the next two cutting grass. </p> <p>In samples 4 and 7, there is one audio where the MowIT cuts but stops due to motor stress. In sample 6, the initial recording was not made without cutting grass, and only the two recordings were made cutting grass.</p> <p> </p> <p> </p> <p> </p>
Exploring the relation between fundamental frequency and spectral envelope in the perception of musical instrument sounds – sound files and participant responses
<p>This database contains synthesized instrument sounds (sounds.zip), participant responses (data.zip), and a key for the stimulus order (stimKey.zip) as complementary data to [1].</p> <p>Sounds include individual stimuli for two experiments. Experiment 1 contains stimuli used for sound pleasantness and sound brightness ratings. Every rating scale includes three acoustic conditions: congruent, incongruent, and fixed, corresponding to the relation of fundamental frequency (F0) and spectral envelope (SE). See [1] for further details on this matter. Experiment 2 contains stimuli for four synthesized instrument sounds: violin, alto voice, clarinet, and tuba. Sounds were synthesized using congruent spectral envelopes (all), and register fixed spectral envelopes (low, mid, high).</p> <p>All sounds are mono signals with a sampling frequency of 44100 Hz in WAV format.</p> <p>Data includes four sets of participant response data: sound pleasantness ratings (Exp. 1), sound brightness ratings (Exp. 1), sound pleasantness ratings (Exp. 2), and sound plausibility ratings (Exp. 2).</p> <p>StimKey provides values for the first and second principal components of the synthesis space for the pleasantness (1 to 81) and brightness (1 to 16) stimulus numbers for Exp. 1.</p> <p> </p> <p>Names convention for Exp. 2 sounds:</p> <p>InstrumentName_congruencyCondition_F0..wav</p> <p> </p> <p>References:</p> <p>[1] Jacobsen, S. and Siedenburg, K. (2024). Exploring the relation between fundamental frequency and spectral envelope in the perception of musical instrument sounds. Acta Acustica, 8, 48. <a href="https://doi.org/10.1051/aacus/2024038">https://doi.org/10.1051/aacus/2024038</a>.</p>
The Aliased Complex Oscillator as a Paradigm for Analog Physical Modeling Sound Synthesis --- Audio Samples
<p>Additional material to the paper with the title: "The Aliased Complex Oscillator as a Paradigm for Analog Physical Modeling Sound Synthesis"</p>
Dataset of Room Impulse Responses from Baffled Microphone Arrays and Sound Sources at Three Elevations
<p>This data set contains a collection of impulse responses (stored in SOFA format) from <em>spherical microphone arrays</em> (<strong>SMA</strong>s), <em>equatorial microphone arrays</em> (<strong>EMA</strong>s), and <em>non-spherical microphone arrays</em> (<strong>XMA</strong>s). Thereby, impulse response sets are provided for each array type at various spatial resolutions, for a loudspeaker sound source at three source elevations, and in four diverse acoustic environments (see <strong>DATA</strong> section for a full description).</p> <p>The original purpose of the microphone array data is the binaural rendering in the <em>spherical harmonics</em> (<strong>SH</strong>) domain into ear signals for high-fidelity reproduction of the acoustic scenario via headphones. Therefore, <em>binaural room impulse responses</em> (<strong>BRIR</strong>s) for 360 horizontal head orientations of a <em>G.R.A.S KEMAR</em> acoustic dummy head are provided as a reference for each scenario.</p> <p>Please contact the authors for questions or additional information regarding the room setups and utilized measurement devices.</p> <p> </p> <p><strong>======<br> DATA<br>======</strong></p> <p>This archive contains the processed impulse response sets of various measurement configurations, as described in this section.</p> <p>Directory "resources/ARIR_processed/":</p> <ul> <li>Post-processed SMA and EMA impulse responses <ul> <li><strong>"_SMA*_"</strong> or <strong>"_EMA*_"</strong> in the file name</li> <li>In SOFA format with <em>"SingleRoomSRIR"</em> convention</li> <li>From 1x <em>DPA 4060</em> microphone flush mounted in a wooden spherical scattering body with an 8.5 cm radius</li> <li>High-resolution data (measured sequentially on VariSphear turntable with two degrees-of-freedom rotations): <ul> <li>Hall: <strong>1202</strong> <strong>channels</strong> (Lebedev grid) for maximum SH order 29</li> <li>Others: <strong>2702</strong> <strong>channels</strong> (Lebedev grid) for maximum SH order 44</li> </ul> </li> <li>Lower-resolution data via subsampling in the SH domain (arbitrary sampling grids and lower target orders can be achieved): <ul> <li>SH order 29: <strong>1742 channels</strong> (t-design grid) for SMA; <strong>59</strong> <strong>channels</strong> (equiangular grid) for EMA</li> <li>SH order 12: <strong>314</strong> <strong>channels</strong> (t-design grid) for SMA; <strong>25</strong> <strong>channels</strong> (equiangular grid) for EMA</li> <li>SH order 8: <strong>146</strong> <strong>channels</strong> (t-design grid) for SMA; <strong>17</strong> <strong>channels</strong> (equiangular grid) for EMA</li> <li>SH order 4: <strong>42</strong> <strong>channels</strong> (t-design grid) for SMA; <strong>9</strong> <strong>channels</strong> (equiangular grid) for EMA</li> <li>SH order 2: <strong>14</strong> <strong>channels</strong> (t-design grid) for SMA; <strong>5</strong> <strong>channels</strong> (equiangular grid) for EMA</li> <li>SH order 1: <strong>6</strong> <strong>channels</strong> (t-design grid) for SMA; <strong>3</strong> <strong>channels</strong> (equiangular grid) for EMA</li> </ul> </li> </ul> </li> <li>Post-processed XMA impulse responses <ul> <li><strong>"_XMA*_"</strong> in the file name</li> <li>In SOFA format with <em>"SingleRoomSRIR"</em> convention</li> <li>From 18x <em>Rode Lavalier GO</em> microphone mounted in an elastic band on a wooden head-shaped scattering body (7.5 cm to 10.5 cm radius)</li> <li>High-resolution data (measured simultaneously): <ul> <li><strong>18 channels</strong> for maximum SH order 8</li> </ul> </li> <li>Lower-resolution data via integer subsets of microphones: <ul> <li>SH order 4: <strong>9 channels</strong></li> <li>SH order 2: <strong>6 channels</strong></li> </ul> </li> <li>Anechoic: For 360 horizontal scattering body orientations (measured sequentially on a VariSphear turntable with azimuth in 1-degree steps)</li> <li>Rooms: For 36 horizontal scattering body orientations (measured sequentially on VariSphear turntable with azimuth in 10-degree steps)</li> </ul> </li> <li>Generated XMA calibration filters and equalization filters <ul> <li><strong>"_x_nm_"</strong> and <strong>"_e_nm_"</strong> in the file name</li> <li>In proprietary Matlab format</li> <li>Time-domain representation of filters in the respective orders of "real" spherical harmonics</li> </ul> </li> <li>Post-processed binaural impulse responses <ul> <li><strong>"_KEMAR_"</strong> in the file name</li> <li>In SOFA format with <em>"SingleRoomSRIR"</em> convention</li> <li>From <em>G.R.A.S KEMAR</em> dummy head with large pinna</li> <li>For 360 horizontal head orientations (measured sequentially on VariSphear turntable with azimuth in 1-degree steps)</li> </ul> </li> <li>Thereby, impulse response sets are included for five acoustic environments <ul> <li><strong>"Simulation_"</strong>: Anechoic simulation of a plane wave impinging from the frontal direction on the array (SMA and EMA only)</li> <li><strong>"Anechoic_"</strong>: Anechoic measurement of a <em>Genelec 8030A</em> loudspeaker at the same height of the array</li> <li>"<strong>LabDry_"</strong>: Room measurement in an acoustically damped laboratory of a <em>Genelec 8030A</em> loudspeaker at three different source heights (the direct floor reflection is attenuated with an additional porous absorber but otherwise identical to the following condition)</li> <li><strong>"LabWet_"</strong>: Room measurement in an acoustically damped laboratory of a <em>Genelec 8030A</em> loudspeaker at three different source heights (the direct reflection is not obstructed from the hard concrete floor, but otherwise identical to the former condition)</li> <li><strong>"Hall_"</strong>: Room measurement in a very reverberant hall of a <em>Genelec 8030A</em> loudspeaker at three different source heights</li> </ul> </li> <li>Thereby, the room impulse response sets are included for three relative source elevations (from placing the loudspeaker to varying heights on the same vertical axis) <ul> <li><strong>"_SrcHigh"</strong>: The source is located above the horizon of the receiver</li> <li>"<strong>_SrcEar"</strong>: The source and receiver are located at the same height</li> <li><strong>"_SrcLow"</strong>: The source is located below the horizon of the receiver</li> </ul> </li> <li>Additionally, anechoic impulse responses of the measurement loudspeaker and the utilized microphones are included <ul> <li><strong>"Anechoic_MicSMAnoTape_"</strong>: SMA measurement microphone without the applied tape (the source was compensated)</li> <li><strong>"Anechoic_MicSMAwithTape_"</strong>: SMA measurement microphone with the applied tape (the source was compensated)</li> <li><strong>"Anechoic_MicXMAmic19_"</strong>: XMA measurement microphone (the source was compensated)</li> <li><strong>"Anechoic_SrcFreeField_"</strong>: Measurement source (on-axis) (the influence of the utilized high-quality free-field measurement microphone can be neglected)</li> <li><strong>"Anechoic_SrcFreeField+MicSMAnoTape_"</strong>: Measurement source and SMA measurement microphone without the tape applied</li> <li><strong>"Anechoic_SrcFreeField+MicSMAwithTape_"</strong>: Measurement source and SMA measurement microphone with the tape applied</li> <li><strong>"Anechoic_SrcFreeField+MicXMAmic19_"</strong>: Measurement source and XMA measurement microphone</li> <li>Overall, the resulting impulse response sets contain the following compensations (including exact compensation of the phase/time behavior): <ul> <li>Anechoic KEMAR: Source</li> <li>Anechoic SMA/EMA/XMA: Source and array microphones</li> <li>Rooms KEMAR: None</li> <li>Rooms SMA/EMA/XMA: Array microphones</li> <li>There is the option to compensate for the source's on-axis response in the room measurement data. However, the direction-dependent directivity of the loudspeaker cannot be compensated. Therefore, we decided not to compensate for the source in the room measurement data since the on-axis frequency response of the utilized loudspeaker is reasonably flat.</li> </ul> </li> </ul> </li> </ul> <p> </p> <p><strong>===========<br> DATA_RAW<br>===========</strong></p> <p><strong>This archive is too large to be uploaded to Zotero (around 77.5 GB). Please get in touch with the authors to request the data.</strong></p> <p>The archive contains the raw acoustic data of all measurement configurations captured by the measurement scripts (see section <strong>CODE_AND_PLOTS</strong>). The data yields the final impulse responses (see section <strong>DATA</strong>), as described in this section.</p> <p>Directory "resources/ARIR_raw/":</p> <ul> <li>Subdirectories by room and source position containing the raw SMA, XMA, and KEMAR acoustic measurement data</li> <li>In proprietary Matlab format, separate for every measurement position of each configuration</li> <li>Each data file contains extensive metadata, e.g., describing the utilized hardware devices, input/output ports, and descriptions.</li> <li>Each data file contains the raw utilized exponential sweep signal and the resulting captured microphone signals. Each impulse response may be recomputed with alternative deconvolution and post-processing parameters.</li> </ul> <p>Directory "resources/ARIR_raw/Logs_temp_humidity/":</p> <ul> <li>Air temperature and humidity data were captured in 5-second intervals during all acoustic measurements</li> <li>In CSV format (automatically loaded and included in the final impulse response sets as part of the measurement post-processing; see section <strong>CODE_AND_PLOTS</strong>)</li> <li>This data is not further utilized at the moment but seemed worthwhile to capture since some acoustic measurements (particularly the high-resolution SMA data sets) were conducted over multiple hours.</li> </ul> <p> </p> <p><strong>==================<br> CODE_AND_PLOTS<br>==================</strong></p> <p>This archive contains the code required to gather the raw acoustic measurement data (see section <strong>DATA_RAW</strong>), the code to post-process and yield the final impulse response data (see section <strong>DATA</strong>), and the resulting plots as described in this section.</p> <p>Directory "dependencies/":</p> <ul> <li>Matlab and Python functions that are utilized in the code</li> <li>Additional dependencies of available open-source projects may be required for certain code functions. If so, the source and setup process for the necessary dependencies are documented in the code header.</li> </ul> <p>Directory "plots/":</p> <ul> <li>Plots that were exported (and that may be regenerated) by the following scripts to validate different stages of the data simulation, measurement, and subsampling.</li> </ul> <p>Shell script "x1_Start_Jupyter.sh":</p> <ul> <li>Prepare a Python environment with the required tools described as dependencies.</li> <li>Activate the prepared Python environment to perform impulse response measurements using Jupyter Notebooks setup for different acoustic settings.</li> </ul> <p>Python Jupyter notebook "x1a_Measure_Microphones.ipynb":</p> <ul> <li>Setup and test the utilized acoustic measurement hardware.</li> <li>Perform a series of acoustic measurements of all utilized microphones in an anechoic environment.</li> <li>Export the raw acoustic data and processed impulse responses.</li> </ul> <p>Python Jupyter notebook "x1b_Measure_BRIRs.ipynb":</p> <ul> <li>Setup and test the utilized acoustic measurement hardware.</li> <li>Generate a horizontal grid of measurement orientations for the VariSphear turntable according to the desired dummy head orientations.</li> <li>Perform a series of acoustic measurements of the dummy head at the pre-defined grid in anechoic and various room environments.</li> <li>Export the raw acoustic data and processed impulse responses.</li> </ul> <p>Python Jupyter notebook "x1c_Measure_SMAs.ipynb":</p> <ul> <li>Setup and test the utilized acoustic measurement hardware.</li> <li>Generate a spherical grid of measurement orientations for the VariSphear turntable according to the desired SMA sampling grid.</li> <li>Perform a series of acoustic measurements of the SMA microphone at the pre-defined grid in anechoic and various room environments.</li> <li>Export the raw acoustic data and processed impulse responses.</li> </ul> <p>Python Jupyter notebook "x1d_Measure_XMAs.ipynb":</p> <ul> <li>Setup and test the utilized acoustic measurement hardware.</li> <li>Generate a horizontal grid of measurement orientations for the VariSphear turntable according to the desired scattering body orientations.</li> <li>Perform a series of acoustic measurements of the XMA microphones at the pre-defined grid in anechoic and various room environments.</li> <li>Export the raw acoustic data and processed impulse responses.</li> </ul> <p>Matlab script "x1e_Simulate_SMAs.m":</p> <ul> <li>Simulate a plane wave impinging from an arbitrary direction on SMAs and EMAs with a desired sampling grid in an anechoic environment.</li> <li>The simulations are helpful to evaluate the rendering method and to investigate the influence of different sampling grids and equalization methods on the rendered binaural signals.</li> </ul> <p>Matlab script "x2_Gather_And_Plot_Measurements.m":</p> <ul> <li>Gather the stored single files with individually measured impulse responses and the according metadata into a combined data set.</li> <li>The initial impulse responses can be recomputed with pre- and post-processing parameters tuned towards the specific acoustic scenario, including compensation of provided source and receiver impulse responses.</li> <li>Many plots may be generated during the processing to validate the input and output data.</li> </ul> <p>Matlab script "x2a_Compare_Measurement_Lengths.m":</p> <ul> <li>Compare the length of the resulting impulse responses of designated measurement configurations.</li> <li>This may be helpful for the tuning of pre-processing and post-processing parameters of the measured impulse responses.</li> </ul> <p>Matlab script "x3_Subsample_Measurements.m":</p> <ul> <li>Spatially subsample a high-resolution directional impulse response data set into a different (lower-resolution) sampling grid in the spherical harmonics domain.</li> <li>This is suitable for array and HRIR data sets.</li> <li>The script also compares the subsampled data against a reference set if available. In the current data set, an evaluation is performed for an anechoic simulation and a room measurement of an SMA at SH order 8.</li> </ul> <p>Matlab script "x3a_Gather_XMA_Measurements.m":</p> <ul> <li> <p>Transform anechoic XMA measurement data from SOFA into the data format required by the processing scripts to calculate the respective calibration and equalization filters.</p> </li> </ul> <p>Readme file "x3b_Generate_XMA_Filters.txt":</p> <ul> <li>The code for this functionality follows the publication [1] but is currently not polished enough for publication. Please contact Jens Ahrens (jens.ahrens@chalmers.se) for questions regarding this functionality.</li> <li>[1] J. Ahrens, H. Helmholz, D. Lou Alon, and S. V. Amengual Garí, “Spherical Harmonic Decomposition of a Sound Field Using Microphones on a Circumferential Contour Around a Non-Spherical Baffle,” <em>IEEE/ACM Trans. Audio, Speech, Lang. Process.</em>, vol. 30, pp. 3110–3119, 2022, doi: 10.1109/TASLP.2022.3209940.</li> </ul> <p>Matlab script "x3c_Gather_XMA_Filters.m":</p> <ul> <li>Rename the files containing the computed calibration and equalization filters into a suitable convention for this collection of scripts.</li> <li>The generated name includes an incremental index to track different versions of provided filter sets.</li> </ul> <p>Matlab script "x3d_Compare_XMA_Filters.m":</p> <ul> <li> <p>Generate various time domain and frequency domain plots to compare different versions of the generated XMA calibration and equalization filters.</p> </li> </ul> <p> </p> <p><strong>================<br> DOCUMENTATION<br>================</strong></p> <p>This archive contains additional documentation of the setups and processes while conducting the acoustic measurements, as described in this section.</p> <p>Directory "documentation/":</p> <ul> <li>Various photographs of the different room, source, and receiver arrangements of the data set</li> <li>The room dimensions and source and receiver positions are documented in the form of the original measurement notes (this may be improved in the future).</li> </ul> <p> </p>
Supplementary material related to the thesis "Perception of airborne sounds and vibrations in crocodiles"
<p>This repository contains all datasets, statistical codes and videos examples for the 4 different studies conducted during my thesis. </p>
Infrared-Microwave-Sounding methanol
<p>Monthly daytime methanol (ppbv) for 2008-2018 produced using the Rutherford Appleton Laboratory Infrared-Microwave-Sounding scheme. Further data description can be found in Pope et al. (2021) and the associated supplementary materials. </p> <p>This data has been used in Sands et al. (2024), currently available as a preprint: https://doi.org/10.5194/egusphere-2024-503. </p>
Compilation of existing underwater PAM repositories, libraries, and applications for sound processing
<p>Resources for passive acoustic monitoring (PAM) are continuously expanding and being developed, yet a major challenge for users is staying up-to-date and finding the best software or application for their acoustics investigation. We expand on previous efforts (Rhinehart & Nicholson, 2022; Felgate, 2023) with the aim of providing a current, comprehensive list of 1) underwater sound repositories of raw sound data without significant processing, 2) biological sound reference libraries, with species or taxa identification, and 3) sound processing tools for visualization, annotation, or analysis. This spreadsheet contains three pages, one dedicated to each of the aforementioned items, along with some descriptive information to help users identify the best resources for their needs.</p> <p>This work was done to support the Global Library of Underwater Biological Sounds (GLUBS) project and funded in part by the Richard Lounsbery Foundation and from funding to SCOR WG #169 (GLUBS) provided by national committees of the Scientific Committee on Oceanic Research (SCOR) and from a grant to SCOR from the US National Science Foundation (OCE--2140395), with support from the International Quiet Ocean Experiment.</p> <p> </p>
Sound database of Industrial Machine for Audio Anomaly Detection
<p><span>Audio anomaly detection(AAD) can seamlessly determinefaults in industrial machines and improve the efficiency of predictive maintenance systems. However, the unavailability of audio sound recordings of real industrial machines operating in their actual industrial setup has limited the efficacy of detection systems. Many different audio databases exist having collections of sounds from dummy (or real) systems operating in controlled environments but a collection of audio sounds from actual industrial machines is missing. Therefore, audio sound recordings of an Air compressor machine working in its natural industrial environment are presented. Only real sounds of an actual machine are captured. Synthetic mixing of sounds is avoided. Damaging the machine to create an anomalous state is avoided. Yet fourteen different unhealthy states are identified and their audio recordings are presented. Dataset with varied values of SNRs is also presented. Spectrograms are plotted and spectral shape parameter values of the developed corpus are calculated. The findings demonstrate the divergence in the developed database and its usefulness in building an effective AAD system for a real industrial machine.</span> </p>
A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).
<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>
To bee or not to bee: An annotated dataset for beehive sound recognition
<p><strong>-- Dataset documentation --</strong></p> <p><br> <strong>1- Introduction</strong></p> <p>The present dataset was developed in the context of our work in [1] that focus on the automatic recognition of beehive sounds. The problem is posed as the classification of sound segments in two classes: Bee and noBee. The novelty of the explored approach and the need for annotated data, dictated the construction of such dataset.</p> <p><strong>2- Description</strong></p> <p><strong>2.1- Audio recordings:</strong></p> <p>The annotated dataset was developed based on a selected set of recordings acquired in the context of two different projects: the Open Source Beehive (OSBH) project [2] and the NU-Hive project [3]. Both projects main goal is to develop a beehive monitoring system capable of identifying and predict certain events and states of the hive that are of interest to the beekeeper. Among many different variables that can be measured and that help the recognition of different states of the hive, the analysis and use of the sound the bees produce is a big focus for both projects.</p> <p>The recordings from the OSBH project were acquired through a citizen science initiative which asked people from the general public to record the sound from their beehives together with the registering of the hive state at the moment. Because of the amateur and collaborative nature of this project, the recordings from the OSBH project present great diversity due to the very different conditions in which the signals were acquired: different recording devices used, different environments where the hives were placed, and even different position for the microphones inside the hive. This variety of settings makes this dataset a very interesting tool to help evaluate and challenge the methods developed.</p> <p>The NU-Hive project is a comprehensive effort of data acquisition, concerning not only sound, but a vast amount of variables that will allow the study of bees behaviors and other unknown aspects. The selected recordings are taken from 2 hives and labeled regarding two states: queen bee is present, and queen bee not present. Contrary to the OSBH project recordings, the recordings from the NU-Hive project are from a much more controlled and homogeneous environment. Here the occurring external sounds are mainly traffic, car honks and birds.</p> <p><strong>The annotated dataset:</strong></p> <p>For each selected recording, time segments are labeled as Bee or noBee depending on the perceived source of the sound signal being from bees or external to the hive.</p> <p>The whole annotated dataset consists of 78 recordings of varying lengths which make up for a total duration of approximately 12 hours of which 25% is annotated as noBee events.</p> <p>About 60% of the recordings are from the NU-Hive dataset and represent 2 hives, the remaining are recordings from the OSBH dataset and 6 different hives. The recorded hives are from 3 main locations: North America, Australia and Europe.</p> <p> </p> <p><strong>2- Annotation procedure<a href="http://localhost:8888/notebooks/Dropbox/QMUL/BEESzzzz/Data/Annotations/readme.ipynb#2--Annotation-procedure">¶</a></strong></p> <p>The annotation procedure consists in hearing the selected recordings and marking the beginning and the end of every sound that could not be recognized as a beehive sound. The recognition of external sounds is based primarily on the perceived heard sounds, but a visual aid is also used by visualizing the log-mel frequency spectrum of the signal. All the above are functionalities offered by the Sonic Visualiser software, which was used by two volunteers that are neither bee-specialists nor specially trained in sound annotation tasks.</p> <p>By marking these pairs of moments corresponding to the beginning and end of external sound periods, we are able to get the whole recording labeled into Bee and noBee intervals. Thus in the resulting Bee intervals only pure beehive sounds, (no external sounds) should be perceived for the entirety of the segment. The noBee intervals refer to periods where an external sound can be perceived (superimposed to the bee sounds).</p> <p> </p> <p><strong>File Structure:</strong></p> <p>Each audio file is coupled with its corresponding annotation file, identified by the same name and extension <em>.lab</em>.<br> For convenience, all the annotations are collected in a single master label file named <em>beeAnnotations.mlf</em></p> <p>The <em>.lab</em> files consist of : </p> <ul> <li>First row identifies the audio file to which the annotations refer to.</li> <li>Each line after that describes an interval with starting time point, end time point and label. The time points are expressed in seconds.</li> </ul> <p>Below is an example of such an annotation file: </p> <pre><code>Hive3_20_07_2017_QueenBee_H3_audio_15_30_00 0 78.45 bee 78.46 78.95 nobee 78.96 103.92 bee 103.93 112.48 nobee 112.49 152.48 bee . </code></pre> <p>This dataset is licensed under a Creative Commons Attribution 4.0 International License.<br> When using this dataset, please cite [1]:</p> <p>[1] I. Nolasco and E. Benetos, “To bee or not to bee: Investigating machine learning approaches to beehive sound recognition,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2018, submitted.</p> <p>[2] “Open Source Beehives Project,” https://www.osbeehives.com/.</p> <p>[3] S. Cecchi, A. Terenzi, S. Orcioni, P. Riolo, S. Ruschioni, and N. Isidoro, “A preliminary study of sounds emitted by honey bees in a beehive,” in Audio Engineering Society Convention 144, 2018.</p>
Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology
<p>The VocalImitationSet is a collection of crowd-sourced vocal imitations of a large set of diverse sounds collected from Freesound (<a href="https://freesound.org/">https://freesound.org/</a>), which were curated based on Google's AudioSet ontology (<a href="https://research.google.com/audioset/">https://research.google.com/audioset/</a>). We expect that this dataset will help research communities obtain a better understanding of human's vocal imitation and build a machine understand the imitations as humans do.</p> <p>See <a href="https://github.com/interactiveaudiolab/VocalImitationSet">https://github.com/interactiveaudiolab/VocalImitationSet</a> for more information about this dataset and its latest updates.</p> <p>For citations, please use this reference:</p> <p>Bongjun Kim, Madhav Ghei, Bryan Pardo, and Zhiyao Duan, "Vocal Imitation Set: a dataset of vocally imitated sound events using the AudioSet ontology," <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018)</em>, Nov. 2018.</p> <p>Contact Info:</p> <p>- Interactive Audio Lab: <a href="http://music.eecs.northwestern.edu/">http://music.eecs.northwestern.edu</a></p> <p>- Bongjun Kim <a href="mailto:bongjun@u.northwestern.edu">bongjun@u.northwestern.edu</a> | <a href="http://www.bongjunkim.com/">http://www.bongjunkim.com</a></p> <p>- Bryan Pardo <a href="mailto:pardo@northwestern.edu">pardo@northwestern.edu</a> | <a href="http://www.bryanpardo.com/">http://www.bryanpardo.com</a></p>
Sound source localization with varying amount of visual information in virtual reality [dataset]
<p>This is the dataset that belongs to the publication "Sound source localization with varying amount of visual information in virtual reality".</p> <p>-The file "AVIL_lab.fbx" contains the visual model of the loudspeaker environment that was used throughout the experiment.</p> <p>-The file "localization_task_instruction_final.docx" contains the information sheet that was handed to the subjects before the experiment.</p> <p>-The file "localizationData_public.xlsx" contains the responses from the subjects (column B). The responses in azimuth and elevation (columns F&G) are corrected for the pointing bias (columns H&I).</p>
DESRA sound files
<p>The .wav files listed in the desraJSON2.xls file.</p>
Artificial sound mixes with event insertions
<p><strong>Contains artificial sound mixes and meta data that were created for the task of<br> sound event detection.</strong></p> <p>The mixes were created using background and event audio recordings<br> from Tampere University's Detection and Classification of Acoustic Scenes and<br> Events (DCASE) Community. More information on the source data can be found at<br> http://www.cs.tut.fi/sgn/arg/dcase2017/challenge/task-rare-sound-event-detection#audio-dataset.</p> <p>Source data credits: Diment, Aleksandr et al (2017, 2018)</p> <p><strong>Created artificial sound mixes are 10 seconds long and contain:</strong></p> <p>- background audio from diverse scenes from start to end<br> - 0 to 4 event insertions</p> <p>With this formula, two datasets were created separately: a training, and an<br> evaluation dataset. These were created separately so that the backgrounds and<br> event recordings used for the evaluation dataset were not used in any of the<br> training audio mixes. Thus, keeping them unseen by the system during development.</p> <p><strong>Audio data:</strong><br> <strong>mixes_train.zip</strong>: Contains the audio mixes created for training.<br> Audio format: .wav<br> Count: 1000 tracks</p> <p><strong>mixes_eval.zip</strong>: Contains the audio mixes created for evaluation.<br> Audio format: .wav<br> Count: 500 tracks</p> <p><strong>Meta-data:</strong><br> Meta-data was maintained documenting the source background, the overlaid source<br> events, and the time onset and offset of each sound event or confusing sound.</p> <p><strong>meta_track_info_train.csv and meta_track_info_eval.csv columns:</strong></p> <p>- trackID: the unique ID of a mix<br> - class_dummy: whether a mix contain a glas break event (other events can be<br> determined using meta_clip_insertions_train and meta_clip_insertions_eval)<br> - background_file: the unique file reference used for background sound to the<br> source data (i.e. DCASE original audio data set)<br> - background_t0: the second in the original background sound recording in which<br> the 10 second background starts.</p> <p><strong>meta_clip_insertions_train.csv and meta_clip_insertions_eval.csv columns:</strong></p> <p>- trackID: the unique ID of a mix<br> - event: the type of event insertion<br> - event_start: at which time in the mix the event start<br> - event_end: at which time in the mix the event ends<br> - event_file: the unique file_number of event reference to the<br> source data (i.e. DCASE original audio data set)<br> an event_file with value 345584_4.wav means the event comes from file<br> 345584.wav in the DCASE audio set and is the 4th event in that audio file.</p> <p>To see the project for which this data set was created visit the github<br> repository at https://github.com/reyvaz/sound-event-detection</p> <p><strong>References:</strong><br> Diment, Aleksandr, Mesaros, Annamaria, Heittola, Toni, & Virtanen, Tuomas. (2017).<br> TUT Rare sound events, Development dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.401395</p> <p>Aleksandr Diment, Annamaria Mesaros, Toni Heittola, & Tuomas Virtanen. (2018).<br> TUT Rare sound events, Evaluation dataset [Data set]. Zenodo.<br> http://doi.org/10.5281/zenodo.1160455<br> </p>
Data of Listening Experiments for Azimuthal Localisation in (Local) Sound Field Synthesis
<p>Data of two listening experiments conducted at University of Rostock, Germany. The study investigated the four (Local) Sound Field Synthesis techniques</p> <ul> <li>Wave Field Synthesis</li> <li>Near-Field-Compensated Higher-Order Ambisonics</li> <li>Local Wave Field Synthesis using Spatial Bandwidth Limitation</li> <li>Local Wave Field Synthesis using Virtual Secondary Sources</li> </ul> <p>The corresponding binaural room scanning (BRS) files for the binaural simulation can be found in the directory `brs`. The employed noise stimulus is contained in `stimuli`. The localisation results are stored in `results`. The `analysis` directory includes scripts for parsing the data.</p>
The Acoustic Sounds for Wellbeing Dataset
<p>The field of sound healing includes ancient practices coming from a broad range of cultures. Across such practices there is a variety of instrumentation utilised. Practitioners suggest the ability of sound to target both mental and even physical health issues, e.g., chronic-stress, or joint-pain. Instruments including the Tibetan singing bowl and vocal chanting, are methods which are still widely encouraged today. With the noise-floor of modern urban soundscapes continually increasing and known to impact wellbeing, methods to approve daily soundscapes are needed. With this in mind, this study presents the Acoustic Sounds for Wellbeing (ASW) dataset. The ASW dataset is a dataset gathered from YouTube including 88+ hrs of audio from 5-classes of acoustic instrumentation (Chimes, Chanting, Drumming, Gongs, and Singing Bowl). We additionally present initial baseline classification results on the dataset, finding that conventional Mel-Frequency Cepstra coefficient features achieve at best an unweighted average recalled of 57.4 % for a 5-class support vector machine classification task. </p> <p>All research papers, presentations, or documents that report the use of the ASW Dataset will cite the following paper:</p> <p>@misc{baird2019acoustic,<br> title={Acoustic Sounds for Wellbeing: A Novel Dataset and Baseline Results},<br> author={Alice Baird and Bjoern Schuller},<br> year={2019},<br> eprint={1908.01671},<br> archivePrefix={arXiv},<br> primaryClass={cs.SD}<br> }</p> <p>For further information please contact:<br> alice.baird@ieee.org</p>
TAU Spatial Sound Events 2019 - Ambisonic and Microphone Array, Evaluation Datasets
<p>This package consists of two evaluation datasets, <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong> and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong>. These datasets contain recordings from an identical scene, with <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong> providing four-channel First-Order Ambisonic (FOA) recordings while <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong> provides four-channel directional microphone recordings from a tetrahedral array configuration. Both formats are extracted from the same microphone array. The recordings in the two datasets consist of stationary point sources from multiple sound classes each associated with a temporal onset and offset time, and DOA coordinate represented using azimuth and elevation angle. These evaluation datasets are part of the <a href="https://github.com/sharathadavanne/seld-dcase2019">DCASE 2019 Sound Event Localization and Detection Task</a>. The corresponding development datasets can be downloaded <a href="https://doi.org/10.5281/zenodo.2599196">here</a>.</p> <p>The IRs were collected in Finland by Tampere University between 12/2017 - 06/2018. The data collection received funding from the European Research Council, grant agreement 637422 EVERYSOUND.</p> <ul> <li>The <strong>foa_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong> evaluation dataset.</li> <li>The <strong>mic_eval.zip</strong>, correspond to audio data of <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong> evaluation dataset.</li> </ul> <p>-- Version 2 updates --</p> <p>The<a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection-results"> DCASE 2019 sound event localization and detection task has now ended</a>. Hence we are releasing the reference labels for the evaluation dataset in this version.</p> <ul> <li>The <strong><em>metadata_eval.zip</em></strong> is the common metadata for both <strong>TAU Spatial Sound Events 2019 - Ambisonic</strong> and <strong>TAU Spatial Sound Events 2019 - Microphone Array</strong> evaluation datasets. </li> <li>The <strong>short2longnames.txt</strong> file consists of the corresponding names for each recording in the dataset in the <a href="http://dcase.community/challenge2019/task-sound-event-localization-and-detection#development-dataset">development-set format</a>, i.e., including the information of the impulse response location and the maximum number of overlapping sound events in the recording.</li> </ul> <p>Download the zip files corresponding to the dataset of interest and use your favorite compression tool to unzip these split zip files.<br> </p> <p> </p> <p> </p>
MTG-Audio Problems Detection on Sound Collections
<p>Manual annotation for the Audio Tagging Competition 2019 (<a href="https://www.kaggle.com/c/freesound-audio-tagging-2019">https://www.kaggle.com/c/freesound-audio-tagging-2019</a>) for audio problems described in Victor Badenas' Thesis from <a href="https://github.com/pirulok02/MTG-Audio-Problems-Detection">https://github.com/pirulok02/MTG-Audio-Problems-Detection</a></p>
Underwater-sound records in glacier fjords (Inglefield Bredning, Baffin Bay, NW Greenland, Denmark), 19-28 July 2019
<p>Acoustic data (.wav) recorded by 2 hydrophones suspended from boats in Inglefield Bredning and Bowdoin fjords (Baffin Bay, NW Greenland, Denmark) in July 2019 for underwater soundscape documentation (narwhal vocalizations and environmental sources). </p> <p>*******************</p> <p>First set-up had a hydrophone AQH-020 by AquaSound Inc. (20Hz – 20kHz) connected to Amplifier Aquafeeler III (SQE-1001B, 50dB gain) by AquaSound Inc. and a recorder PCM-M10 by Sony (44.1 kHz, 16 bit, auto-mode). Recording depth was about 6.6 m, except a record collected on July 27, 2019 at 15:56:33 (depth was about 0.5 m).</p> <p>Channels: 2 (but records are only at “Left”/1 channel; the “Right”/2 is electric noise).</p> <p>File name: sony.YYMMDDhhmmss.wav</p> <p>Note that the strongest regularly-spaced impulsive sounds in two files (sony.190719130533.WAV, sony.190719132214.WAV) are seemingly not due to a whale nearby, but due to repetitive impacts of the hydrophone with a ballast-rope in strong current. This issue was fixed by adjusting the rope length, after which the sound was gone.</p> <p>*******************</p> <p>Second set-up had a hydrophone SoundTrap SD3000 by Ocean Instruments NZ (20Hz – 60kHz), integrated with amplifier and recorder, sampling at 96 kHz, 16 bit. Signal-to-pressure conversion constant was 176.2 dB for this particular device (ID number 5146, at High-Gain mode). Recording depth was about 10.8 m, except records collected on July 20, 2019 between 00:44:47 and 08:44:47 (depth was about 1 m). </p> <p>Channels: 1</p> <p>File name format: 5146.YYMMDDhhmmss.wav</p> <p>*******************</p> <p>Coordinates for each record by each set-up are shown below.</p> <p> </p> <p>Geographic position of each measurement with <strong>SoundTrap</strong> is as the following:</p> <p>Date, Record Start Time(UTC), lon, lat,</p> <p> </p> <p>19 July 2019, 12:55:17, 77.474752, -68.660610 </p> <p>19 July 2019, 13:19:53, 77.488215, -68.597227 </p> <p>19 July 2019, 14:19:53, 77.485103, -68.574985 </p> <p>19 July 2019, 16:19:41, 77.523033, -68.403958 </p> <p>19 July 2019, 19:06:57, 77.618543, -68.564536 </p> <p> </p> <p>20 July 2019, 00:44:47, 77.548649, -68.550752 </p> <p>20 July 2019, 13:02:59, 77.527026, -68.534285</p> <p>20 July 2019, 14:30:26, 77.487211, -68.483834</p> <p>20 July 2019, 15:10:13, 77.495954, -68.658084 </p> <p> </p> <p>*******************</p> <p>Geographic position of each measurement with <strong>Sony-AquaSound</strong> is as the following:</p> <p>Date, Record Start Time(UTC), lon, lat,</p> <p> </p> <p>19 July 2019, 13:05:33, 77.474752, -68.660610 </p> <p>19 July 2019, 13:22:14, 77.488215, -68.597227 </p> <p>19 July 2019, 16:11:54, 77.523033, -68.403958 </p> <p>19 July 2019, 19:09:11, 77.618543, -68.564536</p> <p> </p> <p>20 July 2019, 13:04:42, 77.527026, -68.534285</p> <p>20 July 2019, 14:01:42, 77.505033, -68.553648</p> <p>20 July 2019, 15:11:10, 77.495954, -68.658084 </p> <p> </p> <p>21 July 2019, 22:17:04, 77.675334, -68.636040 </p> <p>21 July 2019, 22:19:53, 77.671904, -68.639090</p> <p> </p> <p>22 July 2019, 00:12:34, 77.525553, -68.442136 </p> <p> </p> <p>27 July 2019, 13:28:58, 77.617588, -68.597946</p> <p>27 July 2019, 14:13:59, 77.667788, -68.643976 </p> <p>27 July 2019, 15:56:33, 77.665487, -68.778195 </p> <p>27 July 2019, 17:03:04, 77.672426, -68.658150 </p> <p>27 July 2019, 17:22:16, 77.669874, -68.657587 </p> <p>27 July 2019, 17:59:39, 77.676941, -68.664817 </p> <p>27 July 2019, 19:44:32, 77.668669, -68.656365 </p> <p>27 July 2019, 21:57:07, 77.628218, -68.637347</p> <p>27 July 2019, 23:07:12, 77.625571, -68.616875</p> <p> </p> <p>28 July 2019, 00:02:41, 77.619299, -68.595757</p>
Dataset for supervised learning with a deep neural network to assess azimuthal localisation in sound field synthesis
<p>Dataset for supervised learning with a deep neural network to assess azimuthal localisation in sound field synthesis.<br> Released as part of the Master Thesis 'An Auditory Model for Azimuthal Localisation in Sound Field Synthesis'.</p> <p>This database is calculated from the data of listening experiments.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.