Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.7.1
Dataset results
103 results for “soundscapes”
The International Soundscape Database: An integrated multimedia database of urban soundscape surveys -- questionnaires with acoustical and contextual information
<h1>Introduction</h1> <p>The International Soundscape Database contains the results of a series of soundscape assessment campaigns carried out across Europe and China. The data collection process was conducted according to the <a href="https://www.mdpi.com/2076-3417/10/7/2397">SSID Protocol [1]</a> which integrates in situ questionnaires about users' soundscape experience, with binaural recordings, sound level meter readings, and 360 degree video. The core of this database are individual soundscape questionnaires collected for 3,500+ participants completed in situ in cities across Europe and China, and the psychoacoustic analysis of 30s binaural recordings which can be matched up to each questionnaire.</p> <p>The SSID Protocol was based on the ISO 12913 standard for soundscape data collection [2]. For more information on the specifics of how this data is collected, please see [1].</p> <p>It is the intention that this dataset be added to and augmented with new locations, cities, and contexts in the future. This will be done both by the SSID team at University College London, but we also strongly welcome contributions from other researchers and practicioners. If a soundscape assessment is collected according to the SSID Protocol, it can be integrated with the rest of the database to form a large, cohesive, and ever-growing database of soundscape assessments. </p> <h2>Analysis</h2> <p>Code for exploring and analysing this dataset is included as part of the <a href="https://soundscapy.readthedocs.io/en/latest/">Soundscapy package</a>.</p> <h2>Included Files</h2> <p>This dataset incorporates surveys taken in multiple urban public spaces across several cities in Europe and China. These urban spaces include places like parks, urban squares, green spaces, and market streets. At each location, up to 100 questionnaires were collected over a series of multi-hour long sessions. Therefore the data is organised by LocationID, then SessionID, then GroupID.</p> <p>The basic directory structure and contents can be found below. </p> <h3>Survey Data (.csv)</h3> <p>'ISD v1.0 Data.csv' organises the data according to the labels given above.</p> <h3>Survey Metadata (.xlsx)</h3> <p>In addition a metadata file ('ISD v1.0 Metadata.xlsx') with photos and descriptions of each of the locations is provided. This metadata file also includes Data Dictionaries for each of the survey instrument versions included. These data dictionaries document precisely the questions asked and the available reponse labels and coding, along with the relevant translations.</p> <h3>Psychoacoustic Analysis (.csv)</h3> <p>The compiled csv file is formatted with a row for each individual participant's questionnaire response, then includes the psychoacoustic analysis of the 30s binaural recording taken while the participant was completing the questionnaire. Details about the psychoacoustic analyses is given in the 'Acoustic Settings' tab in the metadata file.</p> <p>The compiled survey and psychoacoustic analysis data is contained in 'ISD v1.0 Data.csv'. This is compiled from raw survey data files contained in 'Survey_Data', with individual cleaned survey and psychoacoustic data files included in 'Survey_Data/Interim_<date>'. The scripts for compiling this data are included in 'Scripts/'.</p> <h3>Sound Level Meter logs (.xlsx)</h3> <p>'SLM_<city>/' folders include session-long (i.e. ~3hrs) sound level meter log data in.xlsx files for each SessionID.</p> <h3>Binaural Recordings (32-bit floating point .wav)</h3> <p>'WAV_<city>/' folders include the ~30s binaural recordings in 32 bit floating point .wav format. Within each city folder are a set of LocationID folders containing their associated recordings. The wav files are titled with its GroupID, which is matched to the corresponding survey GroupIDs. </p> <h3>Cleaning and Compilation Scripts (.py)</h3> <p>Python code for cleaning and compiling the data from the raw survey data (within Survey_Data/source_data) are provided. These can be run within the provided demo notebook, or from the terminal by calling 'python -m ISDv1_main' with the relevant arguments. See the README.md file in this directory for more information.</p> <pre><code><br>├── ISD v1.0 Data.csv ├── ISD v1.0 Metadata.xlsx ├── SLM_Granada │ ├── CampoPrincipe1_SLM.xlsx │ ├── ... ├── SLM_Groningen │ └── Noorderplantsoen1_SLM.xlsx ├── SLM_etc ├── Scripts │ ├── ISDcleanDemo.ipynb │ ├── ISDcleaning.py │ ├── ISDpsycho.py │ ├── ISDv1_main.py │ ├── README.md │ └── pyproject.toml ├── Survey_Data │ ├── Interim_2024-02-08_cleaned │ └── source_data ├── WAV_Granada_1 │ ├── CampoPrincipe │ ├── ... ├── WAV_etc</code></pre> <p><strong>Citation</strong>: If you use the ISD or part of it, please cite our paper describing the data collection protocol [1] and this dataset itself.</p> <p><strong>License and reuse</strong>: All ISD recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SSID protocol and contribute new locations to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute new data or for any other collaborations.</p> <p> </p> <p>[1] Mitchell A, Oberman T, Aletta F, Erfanian M, Kachlicka M, Lionello M, Kang J. The Soundscape Indices (SSID) Protocol: A Method for Urban Soundscape Surveys—Questionnaires with Acoustical and Contextual Information. <em>Applied Sciences</em>. 2020; 10(7):2397. <a href="https://www.mdpi.com/2076-3417/10/7/2397">https://doi.org/10.3390/app10072397 </a></p> <p>[2] ISO/TS 12913-2:2018 (2018). “Acoustics – Soundscape – Part 2: Data collection and reporting requirements” International Organization for Standardization, Geneva, Switzerland, 2018</p> <p>[3] Mitchell A, Oberman T, Aletta F, Kachlicka M, Lionello M, Erfanian M, Kang J. Investigating Urban Soundscapes of the COVID-19 Lockdown: A predictive soundscape modeling approach.<em> Journal of the Acoustical Society of America</em>. 2021.</p>
Soundscape Attributes Translation Project (SATP) Dataset
<p>The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP). First introduced in Aletta et. al. (<a href="https://biblio.ugent.be/publication/8695720/file/8695735.pdf">2020</a>), the SATP is an attempt to provide validated translations of soundscape attributes in languages other than English. The recordings were used for headphones - based listening experiments.</p> <p>The data are provided to accompany publications resulting from this project and to provide a unique dataset of 1000s of perceptual responses to a standardised set of urban soundscape recordings. This dataset is the result of efforts from hundreds of researchers, students, assistants, PIs, and participants from institutions around the world. We have made an attempt to list every contributor to this Zenodo repo; if you feel you should be included, please get in touch.</p> <p><strong>Citation</strong>: If you use the SATP dataset or part of it, please cite our paper describing the data collection and this dataset itself.</p> <p><strong>Overview</strong>: The SATP dataset consists of 27 30-sec binaural audio recordings made in urban public spaces in London and one 60 sec stereo calibration signal.</p> <p>The recordings were made at locations as reported in Table 1 of the README.md (<strong>Recording locations</strong>), at various times of day by an operator wearing a binaural kit consisting of BHS II microphones and a SQobold (HEAD acoustics) device. Recordings were then exported to WAV via the ArtemiS SUITE software, using the original dynamic range from HDF. The listening experiment and the calibration procedure were intended for a headphone playback system (Sennheiser HD650 or similar open-back headphones recommended). </p> <p>The recordings were selected from an initial set of 80 recordings through a pilot study to ensure the test set had an even coverage of the soundscape circumplex space. These recordings were sent to the partner institutions (see Table 2 of the README.md) and assessed by approximately 30 participants in the institution's target language. The questionnaire used in each assessment is a translation of Method A Questionnaire, ISO 12913-2:2018. Each institution carried out their own lab experiment to collect data, then submitted their data to the team at UCL to compile into a single dataset. Some institutions included additional questions or translation options; the combined dataset (`SATP Dataset v1.x.xlsx`) includes only the base set of questions, the extended set of questions from each institution is included in the `Institution Datasets` folder.</p> <p>In all, SATP Dataset v1.4 contains 19,089 samples, including 707 participants, for 27 recordings, in 18 languages with contributions from 29 institutions.</p> <p><strong>Descriptions of the recordings, including GPS coordinates and sound sources, can be found in the README.md file.</strong></p> <p><strong>Format</strong>: The audio recordings are provided as 24 bit, 48 kHz, stereo WAV files. The combined dataset and Institutional datasets are provided as long tidy data tables in .xlsx files.</p> <p><strong>Calibration: </strong>The recommended calibration approach was based on the open-circuit voltage (OCV) procedure which was considered most accessible but other calibration procedures are also possible (Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>)). The provided calibration file is a computer generated sine wave at 1kHz, matching a sine wave recorded using the exact same setup at SPL of 94 dB. In case of the calibration signal playback level set to match SPL of 94 dB at the eardrum, all the 27 samples should be reproduced at realistic loudness. More details on OCV calibration procedure and other options you can find in Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>) and the attached documentation. PLEASE DO NOT EXPOSE YOURSELF NOR THE PARTICIPANTS TO THE CALIBRATION SIGNAL SET AT THE REALISTIC LEVEL AS IT CAN CAUSE HARM.</p> <p><strong>License and reuse</strong>: All SATP recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SATP protocol and contribute new languages to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute a new translation or for any other collaborations.</p>
Visualization and quantification of coral reef soundscapes using CoralSoundExplorer software
<p>Support material for the research paper "Visualization and quantification of coral reef soundscapes using CoralSoundExplorer software"</p>
SOUNDSCAPE North Adriatic Underwater Noise Sound Pressure Levels
<p>Within the Interreg Italy-Croatia <strong>SOUNDSCAPE</strong> <strong>project</strong> a basin-scale, cross-national, long-term underwater monitoring in the Northern Adriatic Sea was carried out (https://www.italy-croatia.eu/web/soundscape). A broad network of nine monitoring stations, characterized by different natural conditions and anthropogenic pressures, ensured acoustic data collection from March 2020 to June 2021, including the first full lockdown period related to the COVID-19 pandemic (March–April 2020). Calibrated stationary recorders featured with an omnidirectional Neptune Sonar D60 Hydrophone recorded continuously 24h a day (48 kHz, 16 bit).</p> <p>Data were analysed to Sound Pressure Levels (SPLs, dB re 1 uPa) that are here released as a dataset composed of 20 and 60 seconds averaged SPL output files for each station. Data are archived using structured hdf5 files, each one containing metadata and SPL data, according to ICES (International Council for the Exploration of the Sea) continuous noise data specification (https://www.ices.dk/data/data-portals/Pages/Continuous-Noise.aspx).</p> <p>A research object with a jupyter notebook developed to post process SPL data is available at https://doi.org/10.24424/hrhm-8849</p> <p>If data are used, please cite also Petrizzo, A., Barbanti, A., Barfucci, G. <em>et al.</em> Author Correction: First assessment of underwater sound levels in the Northern Adriatic Sea at the basin scale. <em>Sci Data</em> <strong>10</strong>, 179 (2023). https://doi.org/10.1038/s41597-023-02099-x</p>
Human auditory ecology : Extending hearing research to the perception of natural soundscapes by humans in rapidly-changing environments
<p>The audiomaterial corresponding to boreal, tropical and temperate forests, desert, savannah, sub-alpine meadow, and the construction site in New York is copyrighted (license from Wild Sanctuary) and cannot be used without explicit agreement of Bernie Krause. Additional audiomaterial (urban park and street traffic in Paris, France; fast street traffic in Marseille, France; English and French speech material) may only be used with the explicit agreement of the following authors: Jérôme Sueur and Sylvain Haupert (Museum National d'Histoire Naturelle in Paris, France); Sabine Meunier (LMA/CNRS in Marseille, France); Franck Ramus (CNRS in Paris, France) (see Figure legends).</p>
A collection of annotated soundscape recordings from western Kenya
<p>This collection contains 35 soundscape recordings of 32 hours total duration, which have been annotated with 10,294 labels for 176 different bird species from western Kenya. The data were recorded in 2021 and 2022 west and southwest of Lake Baringo in Baringo County, Kenya. This collection has partially been featured as test data in the 2023 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>For this collection, AudioMoths and SWIFT recording units were deployed at multiple locations west and southwest of Lake Baringo, Baringo County, Kenya between Dezember 2021 and February 2022. Recording locations cover a variety of habitats from open grasslands to semi-arid scrubland and mountain forests. Recordings were originally sampled at 48 kHz and converted to MP3 for faster file transfer. For publication, all files were resampled to 32 kHz and converted to FLAC.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>A total of 32 hours of audio from various sites west and southwest of Lake Baringo were selected for annotation. Annotators were tasked with identifying and labeling each bird call they could discern, excluding any calls that were too weak or indiscernible. The annotation process was carried out using Audacity. Provided labels mark the center of each bird call. In this collection, we use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list). Parts of this dataset have previously been used in the 2023 BirdCLEF competition. </p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in EAT (UTC+3). As an example, the file “KEN_001_20211207_153852.flac” has sequential ID 001 and was recorded on December 7th 2021 at 15:38:52 EAT. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements</strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection. In particular, our thanks go to Francis Cherutich for setting up recording units, collecting and annotating data, and to Alain Jacot for assisting in programming the units and transporting the recorders to Kenya.</p>
Data and supplementary material used for Soundscapes to Landscapes soundscape mapping
<p>This repository contains supporting data products to enable the soundscape mapping outlined in the associated publication (DOI forthcoming). Data were used to extract acoustic recording location environmental data for training random forest models to spatially predict 2021 ecoacoustic metrics. The accompanying code will be linked to the GitHub repository. Files include:</p> <p>Data:</p> <ul> <li>clustered_fold_k10.rsd: indices of the model data used if geoCV approach</li> <li>extracted_predictors_vif3.csv: site-specific predictor values extracted from predictors_annual_20230223.tif</li> <li>final_predictors_vif3.csv: a two column table summarizing the VIF selected predictors</li> <li>final_sites_2017-2021.csv: the list of 1,195 potential sites</li> <li>predictor_sprmn_corr.csv: correlation matrix for predictors in model data</li> <li>predictors_annual_20230223.tif: all predictors </li> <li>response_df_200623.csv: site level ecoacoustic metrics</li> </ul> <p>Results:</p> <ul> <li>map_correlations.tar: pairwise response map correlations</li> <li>pdps.tar: partial dependence plot data</li> <li>performance.tar: model performance summaries</li> <li>predictions_maps.tar: final median and IQR model prediction surfaces</li> <li>variable_importance.tar: summaries for variable importance analyses</li> </ul> <p>Contact Colin Quinn at cq73@nau.edu for questions related to this repository or the underlying work. Original wav recordings are expected to be made publicly available on the NASA DAACs in the near future. </p>
Worldwide Soundscapes project metadata and analysis scripts
<p>The Worldwide Soundscapes project is a global, open inventory of spatio-temporally replicated passive acoustic monitoring meta-datasets (i.e. meta-data collections). This Zenodo entry comprises the data tables that constitute its (meta-)database, as well as their description. Additionally, R scripts are provided to replicate the analysis published in [placeholder].</p> <p>The overview of all sampling sites and timelines can be found on the corresponding project on <a href="https://ecosound-web.de/ecosound_web/collection/index/106">ecoSound-web</a>, as well as a <a href="https://ecosound-web.de/ecosound_web/collection/show/49">demonstration collection</a> containing selected recordings. The recordings of this collection were annotated and analysed to explore macro-ecological trends.</p> <p>The audio recording criteria justifying inclusion into the meta-database are:</p> <ul> <li>Stationary (no transects, towed sensors or microphones mounted on cars)</li> <li>Passive (unattended, no human disturbance by the recordist)</li> <li>Ambient (no directional microphone or triggered recordings, non-experimental conditions)</li> <li>Spatially and/or temporally replicated (i.e. multiple sites sampled at the same time and/or multiple days - covering the same daytime - sampled at the same site)</li> </ul> <p>The individual columns of the provided data tables are described in the following. Data tables are linked through primary keys; joining them will result in a database. The data shared here only includes validated collections.</p> <p><strong>Changes from version 4.0.0</strong></p> <p>Added link to the published synthesis.</p> <p><strong>Meta-database CSV files</strong></p> <p><strong>collections</strong></p> <ul> <li>collection_id: unique integer, primary key</li> <li>name: name of the dataset. if it is repeated, incremental integers should be used in the "subset" column to differentiate them.</li> <li>ecoSound-web_link: link of validated meta-collection on ecoSound-web</li> <li>primary_contributors: full names of people deemed corresponding contributors who are responsible for the dataset</li> <li>secondary_contributors: full names of people who are not primary contributors but who have significantly contributed to the dataset, and who could be contacted for in-depth analyses</li> <li>date_added: when the datased was added (YYYY-MM-DD)</li> <li>URL_open_recordings: internet link for openly-available recordings from this collection</li> <li>URL_project: internet link for further information about the corresponding project</li> <li>DOI_publication: Digital Object Identifiers of corresponding publications</li> <li>core_realm_IUCN: The main, core realm of the dataset according to IUCN Global Ecosystem Typology (v2.0): https://global-ecosystems.org/</li> <li>medium: the physical medium the microphone is situated in</li> <li>locality: optional free text about the locality</li> <li>contributor_comments: free-text field for comments by the primary contributors</li> </ul> <p><strong>collections-sites</strong></p> <ul> <li>dataset_ID: primary key of collections table</li> <li>site_ID: primary key of sites table</li> </ul> <p><strong>sites</strong></p> <ul> <li>site_ID: unique integer, primary key</li> <li>site_name: internal name or code of sampling site as used in respective projects</li> <li>latitude_numeric: site's numeric degrees of latitude</li> <li>longitude_numeric: site's numeric degrees of longitude</li> <li>blurred_coordinates: whether latitude and longitude coordinates are inaccurate, boolean. Coordinates may be blurred with random offsets, rounding, snapping, etc. Indicate the blurring method inside the comments field</li> <li>topography_m: vertical position of the microphone relative to the sea level. for sites on land: elevation. For marine sites: depth (negative). in meters. Only indicate if the values were measured by the collaborator.</li> <li>freshwater_depth_m: microphone depth, only used for sites inside freshwater bodies that also have an elevation value above the sea level</li> <li>realm: Ecosystem type: main realm according to IUCN GET https://global-ecosystems.org/</li> <li>biome: Ecosystem type: main biome according to IUCN GET https://global-ecosystems.org/</li> <li>functional_group: Ecosystem type: main functional group according to IUCN GET https://global-ecosystems.org/</li> <li>contributor_comments: free text field for contributor comments</li> <li>GADM_0: Global ADMinistrative Database level 0 classification of terrestrial site or marine site that is within territorial waters. Source: https://gadm.org/download_world.html</li> <li>IHO: International Hydrographic Organization classification of marine site. Source: https://marineregions.org/downloads.php</li> <li>WDPA: World Database on Protected Areas classification of the site. Source: https://www.protectedplanet.net/en/thematic-areas/wdpa?tab=WDPA</li> </ul> <p><strong>deployments</strong></p> <ul> <li>dataset_ID: primary key of datasets table</li> <li>deployment: identical subscript letters to denote rows that belong to the same deployment. For instance, you may use different operation times and schedules for different target taxa within one deployment.</li> <li>subset_site_ID: If the deployment was not done in all the sites of the corresponding collection, site IDs where the deployment was conducted</li> <li>start_date: date of deployment start</li> <li>start_time_mixed: deployment start local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset). Corresponds to the recording start time for continuous recording deployments. If multiple start times were used, you should mention the latest start time (corresponds to the earliest daytime from which all recorders are active). If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>permanent: whether the deployment is permanent, boolean</li> <li>end_date: date of deployment end (date when last scheduled operation starts)</li> <li>end_time_mixed: deployment end local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight). Corresponds to the recording end time for continuous recording deployments.</li> <li>operation_mode: continuous: recording takes place from the deployment start date-time to deployment end date-time.<br>periodical: recording takes place periodically (i.e., with duty cycle) from the deployment start date-time to deployment end date-time.<br>scheduled: recording takes place during scheduled daily time intervals (optionally with duty cycle)</li> <li>duty_cycle_minutes: duty cycle of the recording (i.e. the fraction of minutes when it is recording), written as "recording(minutes)/period(minutes)". empty if no duty cycle is used. For example: "1/6" if the recorder is active for 1 minute and standing by for 5 minutes</li> <li>operation_start_time_mixed: only for scheduled recordings: start local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight). If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>operation_duration_minutes: only for scheduled recordings: duration of operation in minutes, if constant</li> <li>operation_end_time_mixed: only for scheduled recordings: end local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight). Only required if durations are variable. Do not use when end times are ambiguous (for instance, if a recording could be 1 hour or 25 hours long because the end is on the next day). If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>high_pass_filter_Hz: frequency of the high-pass filter of the recorder if applied, in Hz. Otherwise, write "none". This may be called a "low-cut" filter too.</li> <li>bit_depth: sampling bit depth of the recordings. Often constant for a particular recorder</li> <li>channels: number of recorded audio channels</li> <li>sampling_frequency_kHz: frequency at which the microphone signal was sampled by the recorder (sounds of half that frequency will be recorded)</li> <li>recorder: recorder used for deployment</li> <li>microphone: microphone used for deployment</li> <li>target_taxa: main IUCN animal taxa that were studied with this deployment, using the exact IUCN Red list names (http://www.iucnredlist.org/), separated by commas. Only genera, families, orders, and classes are accepted. Empty if there was no taxonomic focus (i.e., general soundscapes were the study focus).</li> <li>contributor_comments: free text field for contributor comments</li> <li>exact_recordings: whether the deployment data here have been superseded by inserting more exact recording date-time ranges into the meta-collection on ecoSound-web</li> </ul> <p><strong>recordings (partial download from <a href="https://ecosound-web.de/">ecoSound-web</a>)</strong></p> <ul> <li>recording_id: primary key of the recordings table</li> <li>collection_id: ID of the collection the recording belongs to</li> <li>name: name of the recording</li> <li>site_id: site ID the recording belongs to:</li> <li>recorder_id: ID of the recorder used for the recording (internal ecoSound-web code)</li> <li>microphone_id: ID of the microphone used for the recording (internal ecoSound-web code)</li> <li>recording_gain:recording gain applied for amplifying the audio signal, in decibels</li> <li>duty_cycle_recording: fraction of the recording periode when the recorder is actively recording audio</li> <li>duty_cycle_period: period of the duty cycle, i.e., time between the starts of two subsequent recordings</li> <li>note: comments (contains the target taxon)</li> <li>file_date: date of the recording start</li> <li>file_time: local time of the recording start</li> <li>sampling_rate: audio sampling rate in Hz</li> <li>bitdepth: depth in bits for each audio sample</li> <li>channel_num: number of channels</li> <li>duration: duration of the recording in seconds. Note: duty-cycled recordings cover only a proportion of this duration<strong><br></strong></li> </ul> <p><strong>affiliations</strong></p> <ul> <li>affiliation_id: primary key of affiliations table</li> <li>lab_research_group: Laboratory or research group name</li> <li>department_school_institute: department, school, or institute name</li> <li>university_institution: University or institution name</li> <li>street_address: street address</li> <li>region_state_province_city: region, state, province, or city name</li> <li>postal_code: postal code</li> <li>country: country name</li> </ul> <p><strong>primary_contributors</strong></p> <ul> <li>First_name: First, given name, anonymised when contributor is technically accepted but has not yet given publication authorisation</li> <li>Last_name: Last, family name, anonymised when contributor is technically accepted but has not yet given publication authorisation</li> <li>ORCiD</li> <li>affiliation_IDs: primary keys of the affiliations' table corresponding affiliations, separated by comma</li> <li>first_tier_position: Author position in first-tier</li> <li>publication_agreement: Has contributor explicitly agreed to share her/his meta-data in the collaboration agreement?</li> <li>co_author_first_synthesis: Has contributor confirmed co-authorship intention in the collaboration agreement?</li> </ul> <p>The following columns describe the contributor's role in the project accordint to <a href="https://credit.niso.org/">CRediT</a> taxonomy.</p> <p><strong>Auxiliary files for reproducing analysis</strong></p> <p><strong>R scripts</strong></p> <ul> <li><strong>acoustic analysis.R: </strong>reproduces the result of the soundscape case studies</li> <li><strong>metadata analysis.R:</strong> reproduces the metadata analysis results in the publication</li> </ul> <p><strong>Data from the demonstration collection (download from ecoSound-web)</strong></p> <ul> <li><strong>demo_recordings.csv:</strong> metadata of the recordings, see recordings table</li> <li><strong>demo_sites.csv: </strong>metadata of the sampling locations, see sites table</li> <li><strong>demo_tags.csv: </strong>data describing annotations made in demonstration recordings for the biophony, anthropophony, geophony, and unknown sound sources</li> <li><strong>spectrograms.zip:</strong> contains PNG format spectrograms used in generating Figure 5</li> </ul> <p><strong>Externally sourced data</strong></p> <ul> <li><strong>GET_areas_2.1.1.csv: </strong>raw data obtained from Keith et al. 2023 (https://doi.org/10.5281/zenodo.10081251), then summarized in QGIS to obtain areas per functional group</li> <li><strong>Havlik_sites.csv:</strong> data obtained from Havlik et al. 2022 supplementary material (https://www.frontiersin.org/articles/10.3389/fmars.2022.919418), originally named "Data Sheet 1.CSV"</li> <li><strong>Sugai_sites_updated.csv:</strong> data obtained from Sugai et al. 2019 (https://doi.org/10.1093/biosci/biy147), personal communication with permission</li> <li><strong>taxonomy.csv:</strong> raw data obtained from IUCN Red List for all animal taxa (https://www.iucnredlist.org/)</li> <li><strong>topography_range_latitude.csv:</strong> raw topography from GEBCO sub-ice data (https://www.gebco.net/data_and_products/gridded_bathymetry_data/), summarised by bins of 10 latitudinal rows</li> </ul>
STARSS22: Sony-TAu Realistic Spatial Soundscapes 2022 dataset
<p><strong>DESCRIPTION:</strong></p> <p>The **<strong>Sony-TAu Realistic Spatial Soundscapes 2022 (STARSS22)</strong>** dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of target classes. The dataset is collected in two different countries, in Tampere, Finland by the Audio Researh Group (ARG) of **<strong>Tampere University (TAU)</strong>**, and in Tokyo, Japan by **<strong>SONY</strong>**, using a similar setup and annotation procedure. The dataset is delivered in two 4-channel spatial recording formats, a microphone array one (**<strong>MIC</strong>**), and first-order Ambisonics one (**<strong>FOA</strong>**). These recordings serve as the development dataset for the <a href="https://dcase.community/challenge2022/task-sound-event-localization-and-detection">DCASE 2022 Sound Event Localization and Detection Task</a> of the <a href="https://dcase.community/challenge2022/">DCASE 2022 Challenge</a>.</p> <p>Contrary to the three previous datasets of synthetic spatial sound scenes of TAU Spatial Sound Events 2019 (<a href="https://zenodo.org/record/2599196">development</a>/<a href="https://zenodo.org/record/3377088">evaluation</a>), <a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a>, and <a href="https://zenodo.org/record/5476980">TAU-NIGENS Spatial Sound Events 2021</a> associated with the previous iterations of the DCASE Challenge, the STARS22 dataset contains recordings of real sound scenes and hence it avoids some of the pitfalls of synthetic generation of scenes. Some such key properties are:</p> <ul> <li>annotations are based on a combination of human annotators for sound event activity and optical tracking for spatial positions,</li> <li>the annotated target event classes are determined by the composition of the real scenes,</li> <li>the density, polyphony, occurences and co-occurences of events and sound classes is not random, and it follows actions and interactions of participants in the real scenes.</li> </ul> <p>The recordings were collected between September 2021 and January 2022. Collection of data from the TAU side has received funding from Google.</p> <p><strong>REPORT & REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the related DCASE2022 task setup:</p> <p>Archontis Politis, Kazuki Shimada, Parthasaarathy Sudarsanam, Sharath Adavanne, Daniel Krause, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Yuki Mitsufuji, Tuomas Virtanen (2022). <strong>STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events</strong>. In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022)</em>, Nancy, France.</p> <p>found <a href="https://dcase.community/documents/workshop2022/proceedings/DCASE2022Workshop_Politis_51.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>70 recording clips of 30 sec ~ 5 min durations, with a total time of ~2hrs, contributed by SONY (development dataset).</li> <li>51 recording clips of 1 min ~ 5 min durations, with a total time of ~3hrs, contributed by TAU (development dataset).</li> <li>52 recording clips with a total time of ~2hrs, contributed by SONY&TAU (evaluation dataset).</li> <li>A training-test split is provided for reporting results using the development dataset.</li> <li>40 recordings contributed by SONY for the training split, captured in 2 rooms (dev-train-sony).</li> <li>30 recordings contributed by SONY for the testing split, captured in 2 rooms (dev-test-sony).</li> <li>27 recordings contributed by TAU for the training split, captured in 4 rooms (dev-train-tau).</li> <li>24 recordings contributed by TAU for the testing split, captured in 3 rooms (dev-test-tau).</li> <li>A total of 11 unique rooms captured in the recordings, 4 from SONY and 7 from TAU (development set).</li> <li>Sampling rate 24kHz.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array (MIC).</li> <li>Recordings are taken in two different countries and two different sites.</li> <li>Each recording clip is part of a recording session happening in a unique room.</li> <li>Groups of participants, sound making props, and scene scenarios are unique for each session (with a few exceptions).</li> <li>To achieve good variability and efficiency in the data, in terms of presence, density, movement, and/or spatial distribution of the sounds events, the scenes are loosely scripted.</li> <li>13 target classes are identified in the recordings and strongly annotated by humans.</li> <li>Spatial annotations for those active events are captured by an optical tracking system.</li> <li>Sound events out of the target classes are considered as interference.</li> <li>Occurences of up to 3 simultaneous events are fairly common, while higher numbers of overlapping events (up to 5) can occur but are rare.</li> </ul> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>SOUND CLASSES:</strong></p> <p>13 target sound event classes are annotated. The classes follow loosely the <a href="https://research.google.com/audioset/ontology/index.html">Audioset ontology</a>.</p> <p> 0. <strong>Female speech, woman speaking</strong><br> 1. <strong>Male speech, man speaking</strong><br> 2. <strong>Clapping</strong><br> 3. <strong>Telephone</strong><br> 4. <strong>Laughter</strong><br> 5. <strong>Domestic sounds</strong><br> 6. <strong>Walk, footsteps</strong><br> 7. <strong>Door, open or close</strong><br> 8. <strong>Music</strong><br> 9. <strong>Musical instrument</strong><br> 10. <strong>Water tap, faucet</strong><br> 11. <strong>Bell</strong><br> 12. <strong>Knock</strong></p> <p>The content of some of these classes corresponds to events of a limited range of Audioset-related subclasses. For more information see the README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2022">here</a>. This implementation will serve as the baseline method in the DCASE 2022 Sound Event Localization and Detection Task.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current version (Version 1.1) of the dataset includes the 121 development audio recordings and labels, used by the participants of Task 3 of DCASE2022 Challenge to train and validate their submitted systems, and the 52 evaluation audio recordings without labels, for the evaluation phase of DCASE2022.</p> <p>If researchers wish to compare their system against the submissions of DCASE2022 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The file <strong><em>foa_dev.zip</em></strong>, correspond to audio data of the <strong>FOA </strong>recording format.<br> The file <strong><em>mic_dev.zip</em></strong>, correspond to audio data of the <strong>MIC</strong> recording format.<br> The <strong><em>metadata_dev.zip</em></strong> is the common metadata for both formats.</p> <p>The file <strong><em>foa_eval.zip</em></strong>, corresponds to audio data of the <strong>FOA</strong> recording format for the evaluation dataset.<br> The file <strong><em>mic_eval.zip</em></strong>, corresponds to audio data of the <strong>MIC</strong> recording format for the evaluation dataset.</p> <p>Download the zip files corresponding to the format of interest and use your favourite compression tool to unzip these zip files.</p>
A collection of fully-annotated soundscape recordings from the Western United States
<p>This collection contains 33 hour-long soundscape recordings, which have been annotated with 20,147 bounding box labels for 56 different bird species from the Western United States. The data were recorded in 2018 in the Sierra Nevada, California, USA. This collection has partially been featured as test data in the 2021 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>Measuring the effects of forest management activities in the Sierra Nevada, California, USA can reveal a potential correlation with avian population density and diversity. For this dataset, passive acoustic surveys were conducted in the Lassen and Plumas National Forests in May-August 2018. Survey grid cells (4 km<sup>2</sup>) were randomly selected from a 6,000-km<sup>2</sup> area, and SWIFT recording units were deployed at locations conducive to sound propagation (e.g., ridges rather than gullies) within those cells. The sensitivity of the used microphones was -44 (+/-3) dB re 1 V/Pa. The microphone's frequency response was not measured, but is assumed to be flat (+/- 2 dB) in the frequency range 100 Hz to 7.5 kHz. The analog signal was amplified by 38 dB and digitized (16-bit resolution) using an analog-to-digital converter (ADC) with a clipping level of -/+ 0.9 V. Recording units recorded continuously 17:00 - 23:59, 0:00 - 10:00, one-hour files were stored as uncompressed WAVE sampled at 32 kHz and later converted to FLAC. Parts of this dataset have previously been used in the 2021 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>We subsampled data for this collection by selecting locations that spanned the full elevational and latitudinal gradients of our study area (~840 – 1700 m asl and 39.41 – 40.71°N), and thus represent a broad range of plant communities. A single annotator boxed every bird call he could recognize, ignoring those that are too faint or unidentifiable. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. The annotator was allowed to combine multiple consecutive calls of one species into one bounding box label if pauses between calls were shorter than five seconds. We use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in PDT. As an example, the file “SNE_001_20180509_050002.flac” has sequential ID 001 and was recorded on May 9th 2018 at 05:00:02 PDT. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>The collection and annotation of this dataset was funded by the U.S. Forest Service Region 5 and the California Department of Fish and Wildlife.</p>
A collection of fully-annotated soundscape recordings from the Island of Hawai'i
<p>This collection contains 635 soundscape recordings with a total duration of almost 51 hours, which have been annotated by expert ornithologists who provided 59,583 bounding box labels for 27 different bird species from the Hawaiian Islands, including 6 threatened or endangered native birds. The data were recorded between 2016 and 2022 at four sites across Hawai‘i Island. This collection has partially been featured as test data in the 2022 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>Soundscapes for this collection were recorded for various research projects by the Listening Observatory for Hawaiian Ecosystems (LOHE) at the University of Hawai‘i at Hilo. The recordings were collected using Wildlife Acoustics Inc. Song Meters (models 2, 4, or Mini), as 16-bit wav files at a sampling rate of 44.1 kHz, using the default gain settings of each model. Further specifics for each recording, such as recording location and habitat type, can be found in the metadata provided. Soundscapes in this collection vary in length, ranging from just under a minute to 9 minutes in duration. All audio was unified, converted to FLAC, and resampled to 32 kHz for this collection. Parts of this dataset have previously been used in the 2022 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>This collection is a subset of the files recorded over the course of the LOHE lab’s respective studies. The data were subsampled for annotation by aurally scanning the recordings and visually scanning spectrograms generated using Raven Pro software for target species of interest to the individual research project for which each recording was collected. Recordings that did not contain vocalizations of the species of interest were excluded from full annotation and thus this collection. </p> <p>Using Raven Pro, annotators were asked to create a selection box around every bird call they could recognize, ignoring those that were too faint or unidentifiable at a spectrogram window size of 700 points. Provided labels contain full bird calls that are boxed in time and frequency. Annotators were allowed to combine multiple consecutive calls of the same species into one bounding box label if pauses between calls were shorter than 0.5 seconds. We converted labels to eBird species codes, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, site ID, recording date, and timestamp in HST. As an example, the file “UHH_001_S01_20161121_150000.flac” has sequential ID 001 and was recorded at site S01 on Nov 21st, 2016 at 15:00:00 HST. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz, and an eBird species code. These species codes can be assigned to the scientific and common name of a species with the “species.csv” file. The approximate recording location with Universal Transverse Mercator (UTM) coordinates and other metadata can be found in the “recording_location.csv” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection. Specifically, we want to thank Charlotte Forbes-Perry with the Pacific Cooperative Studies Unit, University of Hawai'i at Hawai‘i Volcanoes National Park as well as the following current and past members of the LOHE lab (in alphabetical order): Keith Burnett, Saxony Charlot, Noah Hunt, Caleb Kow, Elizabeth Lough, and Bret Mossman.</p> <p>Access and permits to record soundscapes were provided by (in alphabetical order): Hakalau Forest National Wildlife Refuge, the State of Hawai‘i Department of Land and Natural Resources Division of Forestry and Wildlife, and the U.S. Fish and Wildlife Service.</p> <p>We would also like to acknowledge our funding sources (in alphabetical order): The National Park Service Inventory and Monitoring Division, the National Science Foundation, and the U.S. Army Engineer Research and Development Center.</p>
A collection of fully-annotated soundscape recordings from the Northeastern United States
<p>This collection contains 285 hour-long soundscape recordings, which have been annotated by expert ornithologists who provided 50,760 bounding box labels for 81 different bird species from the Northeastern USA. The data were recorded in 2017 in the Sapsucker Woods bird sanctuary in Ithaca, NY, USA. This collection has (partially) been featured as test data in the 2019, 2020 and 2021 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>As part of the Sapsucker Woods Acoustic Monitoring Project (SWAMP), the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology deployed 30 first-generation SWIFT recorders in the surrounding bird sanctuary area in Ithaca, NY, USA. The sensitivity of the used microphones was -44 (+/-3) dB re 1 V/Pa. The microphone's frequency response was not measured, but is assumed to be flat (+/- 2 dB) in the frequency range 100 Hz to 7.5 kHz. The analog signal was amplified by 33 dB and digitized (16-bit resolution) using an analog-to-digital converter (ADC) with a clipping level of -/+ 0.9 V. This ongoing study aims to investigate the vocal activity patterns and seasonally changing diversity of local bird species. The data are also used to assess the impact of noise pollution on the behavior of birds. Recordings were recorded 24 h/day in 1-hour uncompressed WAVE files at 48 kHz, converted to FLAC and resampled to 32 kHz for this collection. Parts of this dataset have previously been used in the 2019, 2020 and 2021 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>We subsampled data for this collection by randomly selecting one 1-hour file from one of the 30 different recording units for each hour of one day per week between Feb and Aug 2017. For this collection, we excluded recordings that were shorter than one hour or did not contain a bird vocalization. Annotators were asked to box every bird call they could recognize, ignoring those that are too faint. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. Annotators were allowed to combine multiple consecutive calls of one species into one bounding box label if pauses between calls were shorter than five seconds. We use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date, and timestamp in UTC. As an example, the file “SSW_001_20170225_010000Z.flac” has sequential ID 001 and was recorded on Feb 25th, 2017 at 01:00:00 UTC. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz, and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. Unidentifiable calls have been marked with “????” and are included in the ground truth annotations. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection (individual contributors in alphabetic order): Jessie Barry, Sarah Dzielski, Cullen Hanks, W. Alexander Hopping, Robert Koch, Jim Lowe, Jay McGowan, Ashik Rahaman, Yu Shiu, Laurel Symes, and Matt Young. </p> <p><strong>Version history</strong></p> <p>Version 2: Unidentifiable calls have been marked with “????” and added as bounding box labels to the ground truth annotations.<br> Version 1: Initial release.</p>
A collection of fully-annotated soundscape recordings from the southern Sierra Nevada mountain range
<p>This collection contains 100 soundscape recordings of 10 minutes duration, which have been annotated with 10,296 bounding box labels for 21 different bird species from the Western United States. The data were recorded in 2015 in the southern end of the Sierra Nevada mountain range in California, USA. This collection has been featured as test data in the 2020 BirdCLEF and Kaggle Birdcall Identification competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>The recordings were made in Sequoia and Kings Canyon National Parks, two contiguous national parks in the southern Sierra Nevada mountain range in California, USA. The focus of the acoustic study was the high-elevation region of the Parks; specifically, the headwater lake basins above 3,000 km in elevation. The original intent of the study was to monitor seasonal activity of birds and bats at lakes containing trout and lakes without trout, because the cascading impacts of trout on the adjacent terrestrial zone remain poorly understood. Soundscapes were recorded for 24 h continuously at 10 lakes (5 fishless, 5 fish-containing) throughout Sequoia and Kings Canyon National Parks during June-September 2015. Song Meter SM2+ units (Wildlife Acoustics, USA) powered by custom-made solar panels were used to obviate the need to swap batteries, due to the recording locations being extremely difficult to access. Song Meters continuously recorded mono-channel, 16-bits uncompressed WAVE files at 48 kHz sampling rate. For this collection, recordings were resampled at 32 kHz and converted to FLAC.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>A total of 100 10-minute segments of audio between July 9 and 12, 2015 from morning hours (06:10-09:10 PDT) from all 10 sites were selected at random. Annotators were asked to box every bird call they could recognize, ignoring those that are too faint or unidentifiable. Every sound that could not be confidently assigned an identity was reviewed with 1-2 other experts in bird identification. To minimize observer bias, all identifying information about the location, date and time of the recordings was hidden from the annotator. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. In this collection, we use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list). Unidentifiable calls have been marked with “????” and were added as bounding box labels to the ground truth annotations. Parts of this dataset have previously been used in the 2020 BirdCLEF and Kaggle Birdcall Identification competition.</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in PDT (UTC-7). As an example, the file “HSN_001_20150708_061805.flac” has sequential ID 001 and was recorded on July 8th 2015 at 06:18:05 PDT. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection (individual contributors in alphabetic order): Anna Calderón, Thomas Hahn, Ruoshi Huang, Angelly Tovar</p>
The Effect of Soundscape Composition on Bird Vocalization Classification in a Citizen Science Biodiversity Monitoring Project
<p>This archive includes sound clips (.wav files) and associated mel-scale spectrograms of bird vocalizations for 54 species in Sonoma County, California, USA. These data were used for training and validating convolutional neural network (CNN) models for bird species detection. We also include xeno-canto training and validation mel spectrograms used to pretrain CNNs. Details on these data are explained in the paper by Clark et al. (2023) titled "The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project". These data are available for use without restrictions, with no warranty on data quality or utility for a given application. We request that any work that does use these data cite the Clark et al. (2023) paper.<br> <br> Clark, M.L., Salas, L., Baligar, S., Quinn, C., Snyder, R.L., Leland, D., Schackwitz, W., Goetz, S.J., Newsam, S. (2023). The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project. <em>Ecological Informatics</em>. <a href="https://doi.org/10.1016/j.ecoinf.2023.102065">https://doi.org/10.1016/j.ecoinf.2023.102065</a></p> <p>Associated code for training CNN models, performing inference, and applying post-classification corrections can be found in the GitHub archive <a href="https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species">https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species</a></p> <p>Raw sound data from the Soundscapes to Landscapes project are available upon request: Dr. Matthew Clark, matthew.clark@sonoma.edu</p> <p>These data were collected as part of the Soundscapes to Landscapes project (<a href="https://soundscapes2landscapes.org/">soundscapes2landscapes.org</a>), funded by NASA’s Citizen Science for Earth Systems Program (CSESP) 16-CSESP 2016-0009 under cooperative agreement 80NSSC18M0107.<br> <br> ----------------------------<br> This depository includes the following archives:</p> <ul> <li> <p>mel_specs.zip: contains 2-sec mel spectrograms split into training (“tr”), validation (“val”), testing (“test”) data for each target bird species (n = 54) used to fine-tune the CNNs. Select spectrogram files are appended with “aug” if they are augmented versions for the training data.</p> </li> <li> <p>wav.zip: contains the associated wav-format sound recordings used to generate the training, validation, testing mel spectrograms found in mel_specs.zip.</p> </li> <li> <p>Xeno-canto_pretrain.tar: contains 2-sec mel spectrograms split into training and validation data for 40 bird species used for CNN pre-training that were generated using a warbleR segmentation methodology described in the paper. The sound files used to generate these mel spectrograms came from the Kaggle competition, <a href="https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset">https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset</a><br> Mel spectrogram naming reflects the XC number used for cataloging on Xeno-canto in the format XC123456_2.png. The six numbers following the XC characters can be used to search for unique recordings on Xeno-canto (<a href="https://xeno-canto.org/">https://xeno-canto.org/</a>) using the search query “nr:123456” in the search tool or queried using the Xeno-canto API (<a href="https://xeno-canto.org/explore/api">https://xeno-canto.org/explore/api</a>). Unique recording names can be extracted from the mel spectrogram filenames.</p> </li> <li> <p>soundscape_test_wavs.zip: the wav-format sound recordings used to perform soundscape testing.</p> </li> </ul>
Extensive crowdsourced dataset of in-situ evaluated binaural soundscapes of private dwellings containing subjective sound-related and situational ratings along with person factors to study time-varying influences on sound perception — research data
<p><strong>Abstract:</strong></p> <p>The soundscape approach highlights the role of situational factors in sound evaluations; however, only a few studies have applied a multi‐domain approach including sound‐related, person‐related, and time‐varying situational variables. Therefore, we conducted a study based on the Experience Sampling Method to measure the relative contribution of a broad range of potentially relevant acoustic and non‐auditory variables in predicting indoor soundscape evaluations. Here we present the comprehensive dataset for which 105 participants reported temporally (rather) stable trait variables such as noise sensitivity, trait affect, and quality of life. They rated 6.594 situations regarding the soundscape standard dimensions, perceived loudness, and the saliency of its sound components and evaluated situational variables such as state affect, perceived control, activity, and location. To complement these subject‐centered data, we additionally crowdsourced object‐centered data by having participants make binaural measurements of each indoor soundscape at their homes using a low‐(self‐)noise recorder. These recordings were used to compute (psycho‐)acoustical indices such as the energetically averaged loudness level, the A‐weighted energetically averaged equivalent continuous sound pressure level, and the A‐weighted five‐percent exceedance level. This complex hierarchical data can be used to investigate time‐varying non‐auditory influences on sound perception and to develop soundscape indicators based on the binaural recordings to predict soundscape evaluations.</p> <p><strong>Content:</strong></p> <ul> <li><a href="https://zenodo.org/record/7858848/files/01%20StudyDescription.pdf">01 StudyDescription.pdf </a> <ul> <li>Description of the field study.</li> <li>Information about the methods and materials used.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/02%20Dataset.csv">02 Dataset.csv</a> <ul> <li>The dataset, consisting of 93 variables describing 6594 observations taken by 105 participants.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/03%20VariableDescriptions_EnglishPersonQuestionnaire.pdf">03 VariableDescriptions_EnglishPersonQuestionnaire.pdf</a> <ul> <li>Descriptions of all variables, their measurement scale, scale ranges and levels.</li> <li>Questions and task descriptions of the Experience Sampling Method questionnaire in German language with an English translation.</li> <li>English translations of questions asked in the person questionnaire.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/04%20ESM-Questionnaire.pdf">04 ESM-Questionnaire.pdf</a> <ul> <li>Screenshots of the original Experience Sampling Method questionnaire with English translations.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/05%20PersonQuestionnaire_OriginalGermanVersion.pdf">05 PersonQuestionnaire_OriginalGermanVersion.pdf</a> <ul> <li>Original version of the person questionnaire in German language.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/06%20HelpTexts.pdf">06 HelpTexts.pdf</a> <ul> <li>Descriptions of the study task.</li> <li>Explanations of the scales used in the questionnaire.</li> <li>Explanations of the sound categories and the soundscape composition.</li> <li>Explanation of the operation of the recording device.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_README.md">AcousticFeatures_README.md</a> <a href="https://zenodo.org/api/files/3d784540-c0f4-412f-8742-df1db6f5401d/TimeSeries_and_Spectrograms_README.md?versionId=9291496c-d2c6-4151-96f1-a2ad99e1a540"> </a> <ul> <li>Descriptions of the structure of the AcousticFeatures_xxx.csv and .zip files.</li> <li>Analyis settings used in Artemis Suite to generate the acoustic features.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_SingleValues.csv">AcousticFeatures_SingleValues.csv</a> <ul> <li>All acoustic features, aggregated to single values per feature, recording, and channel.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectra.csv">AcousticFeatures_Spectra.csv</a> <ul> <li>Time-averaged 1/3 octave spectra of each channel of each recording, A-weichted and un-weighted.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectrograms.zip">AcousticFeatures_Spectrograms.zip</a> <ul> <li>13188 .csv files with un-weighted spetrograms of each channel of each recording.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_TimeSeries.zip">AcousticFeatures_TimeSeries.zip</a> <ul> <li>A .csv file containing LAeq and LZeq time series of each channel of each recording.</li> </ul> </li> </ul> <p><strong>Publications refering to this dataset:</strong></p> <p>Versümer, Siegbert; Steffens, Jochen; Weinzierl, Stefan (currently under review): "The role of loudness predictions, personal and situational factors in day-to-day loudness assessments of indoor soundscapes."</p> <p><strong>Funding:</strong></p> <p>This study was sponsored by the German Federal Ministry of Education and Research. “FHprofUnt” funding code: 13FH729IX6. </p> <p><strong>License: </strong></p> <p>CC 4.0 BY, <a href="https://creativecommons.org/licenses/by/4.0/legalcode">https://creativecommons.org/licenses/by/4.0/legalcode</a></p> <p><strong>Version history:</strong></p> <p>Details can be found in the <a href="https://zenodo.org/api/files/a15d6a91-1a35-4b5e-a7ec-da8a9bcbee2b/Changelog.md">Changelog.md</a> file.</p> <ul> <li> V.01.0. March 7, 2023: Initial publication. <a href="https://doi.org/10.5281/zenodo.7193938">https://doi.org/10.5281/zenodo.7193938</a></li> <li> V.01.1. April 25, 2023. <a href="https://doi.org/10.5281/zenodo.7858848">https://doi.org/10.5281/zenodo.7858848</a></li> </ul>
STARSS23: Sony-TAu Realistic Spatial Soundscapes 2023
<p><strong>DESCRIPTION:</strong></p> <p>The <strong>Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23)</strong> dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of target classes. The dataset is collected in two different countries, in Tampere, Finland by the Audio Researh Group (ARG) of <strong>Tampere University (TAU)</strong>, and in Tokyo, Japan by <strong>SONY</strong>, using a similar setup and annotation procedure. The dataset is delivered in two 4-channel spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). These recordings serve as the development dataset for the <a href="https://dcase.community/challenge2023/task-sound-event-localization-and-detection-evaluated-in-real-spatial-sound-scenes">DCASE 2023 Sound Event Localization and Detection Task</a> of the <a href="https://dcase.community/challenge2023/">DCASE 2023 Challenge</a>.<br> <br> The STARSS23 dataset is a continuation of the <a href="https://zenodo.org/record/6600531">STARSS22 dataset</a>. It extends the previous version with the following:</p> <ul> <li>An additional <strong>additional 2hrs 30mins </strong>of recordings in the development set, from <strong>5 new rooms</strong> distributed in 47 new recording clips.</li> <li>An <strong>additional 1hr 40mins</strong> of recordings added in the evaluation set of the dataset.</li> <li><strong>360° videos</strong> spatially and temporally aligned to the audio recordings of the dataset (apart from 12 audio-only clips).</li> <li><strong>Distance labels</strong> (in cm) for the spatially annotated sound events, instead of the previous azimuth and elevation only labels.</li> </ul> <p>Contrary to the three previous datasets of synthetic spatial sound scenes of TAU Spatial Sound Events 2019 (<a href="https://zenodo.org/record/2599196">development</a>/<a href="https://zenodo.org/record/3377088">evaluation</a>), <a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a>, and <a href="https://zenodo.org/record/5476980">TAU-NIGENS Spatial Sound Events 2021</a> associated with previous iterations of the DCASE Challenge, the STARS22-23 dataset contains recordings of real sound scenes and hence it avoids some of the pitfalls of synthetic generation of scenes. Some such key properties are:</p> <ul> <li>annotations are based on a combination of human annotators for sound event activity and optical tracking for spatial positions,</li> <li>the annotated target event classes are determined by the composition of the real scenes,</li> <li>the density, polyphony, occurences and co-occurences of events and sound classes is not random, and it follows actions and interactions of participants in the real scenes.</li> </ul> <p>The first round of recordings was collected between September 2021 and January 2022. A second round of recordings was collected between November 2022 and February 2023.<br> <br> Collection of data from the TAU side has received funding from Google.</p> <p>A demo video combining the different modalities and spatial annotations can be found <a href="https://www.youtube.com/watch?v=ZtL-8wBYPow">here</a>.</p> <p><strong>REPORT & REFERENCE:</strong></p> <p>If you use this dataset you could cite this report on its design, capturing, and annotation process:</p> <p>Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji (2023). <strong>STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events</strong>,<br> <br> found <a href="https://arxiv.org/abs/2306.09126">here</a>, and</p> <p>Archontis Politis, Kazuki Shimada, Parthasaarathy Sudarsanam, Sharath Adavanne, Daniel Krause, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Yuki Mitsufuji, Tuomas Virtanen (2022). <strong>STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events</strong>. In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022)</em>, Nancy, France.</p> <p>found <a href="https://dcase.community/documents/workshop2022/proceedings/DCASE2022Workshop_Politis_51.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The STARSS22-23 dataset is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p>Specifically the STARSS23 allows evaluation of audiovisual processing methods with a spatial dimension, such as audiovisual source localization or audiovisual object recognition.</p> <p><strong>SPECIFICATIONS:</strong></p> <p>General:</p> <ul> <li>Recordings are taken in two different sites.</li> <li>Each recording clip is part of a recording session happening in a unique room.</li> <li>Groups of participants, sound making props, and scene scenarios are unique for each session (with a few exceptions).</li> <li>To achieve good variability and efficiency in the data, in terms of presence, density, movement, and/or spatial distribution of the sounds events, the scenes are loosely scripted.</li> <li>13 target classes are identified in the recordings and strongly annotated by humans.</li> <li>Spatial annotations for those active events are captured by an optical tracking system.</li> <li>Sound events out of the target classes are considered as interference.</li> <li>Occurrences of up to 3 simultaneous events are fairly common, while higher numbers of overlapping events (up to 5) can occur but are rare.</li> </ul> <p>Volume, duration, and data split:</p> <ul> <li>A total of 16 unique rooms captured in the recordings, 4 in Tokyo and 12 in Tampere (development set).</li> <li>70 recording clips of 30 sec ~ 5 min durations, with a total time of ~2hrs, captured in Tokyo (development dataset).</li> <li>98 recording clips of 40 sec ~ 9 min durations, with a total time of ~5.5hrs, captured in Tampere (development dataset).</li> <li>79 recordings clips of 40 sec ~ 7 min durations, with a total time of ~3.5hrs, captured in both sites (evaluation dataset).</li> <li>A training-testing split is provided for reporting results using the development dataset.</li> <li>40 recordings contributed by Sony for the training split, captured in 2 rooms (dev-train-sony).</li> <li>30 recordings contributed by Sony for the testing split, captured in 2 rooms (dev-test-sony).</li> <li>50 recordings contributed by TAU for the training split, captured in 7 rooms (dev-train-tau).</li> <li>48 recordings contributed by TAU for the testing split, captured in 5 rooms (dev-test-tau).</li> </ul> <p>Audio:</p> <ul> <li>Sampling rate: 24kHz.</li> <li>Bit depth: 16 bits.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array (MIC).</li> </ul> <p>Video:</p> <ul> <li>Video 360° format: equirectangular</li> <li>Video resolution: 1920x960</li> <li>Video frames per second (fps): 29.97</li> <li>All audio recordings are accompanied by synchronised video recordings, apart from 12 audio recordings with missing videos (<em>fold3_room21_mix001.wav - fold3_room21_mix012.wav</em>)</li> </ul> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>SOUND CLASSES:</strong></p> <p>13 target sound event classes are annotated. The classes follow loosely the <a href="https://research.google.com/audioset/ontology/index.html">Audioset ontology</a>.</p> <p> 0. <strong>Female speech, woman speaking</strong><br> 1. <strong>Male speech, man speaking</strong><br> 2. <strong>Clapping</strong><br> 3. <strong>Telephone</strong><br> 4. <strong>Laughter</strong><br> 5. <strong>Domestic sounds</strong><br> 6. <strong>Walk, footsteps</strong><br> 7. <strong>Door, open or close</strong><br> 8. <strong>Music</strong><br> 9. <strong>Musical instrument</strong><br> 10. <strong>Water tap, faucet</strong><br> 11. <strong>Bell</strong><br> 12. <strong>Knock</strong></p> <p>The content of some of these classes corresponds to events of a limited range of Audioset-related subclasses. For more information see the README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model performing <strong>audio-only</strong> joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2023">here</a>. This implementation will serve as the baseline method in the DCASE 2023 Sound Event Localization and Detection Task, under the audio-only inference track.</p> <p>Additionally, an implementation of a trainable model performing <strong>audiovisual</strong> SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sony/audio-visual-seld-dcase2023">here</a>. This implementation will serve as the baseline method in the DCASE 2023 Sound Event Localization and Detection Task, under the audiovisual inference track.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current version (Version 1.1) of the dataset includes development audio/video recordings and labels and the evaluation recordings without labels, used by the participants of Task 3 of DCASE2023 Challenge to train and validate their submitted systems (development), and produce system outputs for the challenge evaluation phase.</p> <p>If researchers wish to compare their system against the submissions of DCASE2023 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The file <strong><em>foa_dev.zip</em></strong>, correspond to audio data of the <strong>FOA </strong>recording format.<br> The file <strong><em>mic_dev.zip</em></strong>, correspond to audio data of the <strong>MIC</strong> recording format.</p> <p>The file <strong><em>video_dev.zip </em></strong>contains the common videos for both audio formats.<br> The file <strong><em>metadata_dev.zip</em></strong> contains the common metadata for both audio formats.</p> <p>The file <em><strong>foa_eval.zip</strong></em> corresponds to audio data of the <strong>FOA</strong> recording format for the evaluation dataset.<br> The file <em><strong>mic_eval.zip</strong></em> corresponds to audio data of the <strong>MIC</strong> recording format for the evaluation dataset.<br> The file <em><strong>video_eval.zip</strong></em> contains the common videos for both audio formats of the evaluation dataset.</p> <p>Download the zip files corresponding to the format of interest and use your favourite compression tool to unzip these zip files.</p>
Data and code used in analyses for Simulated soundscapes and transfer learning boost the performance of acoustic classifiers under data scarcity
<p>Evaluation datasets, Python scripts, and computation environments used to conduct analyses for Simulated soundscapes and transfer learning boost the performance of acoustic classifiers under data scarcity. <br><br>transfer_learning_project.zip also contains a vignette describing the use of a generalized script for adapting these methods to novel acoustic classification tasks. </p> <p> </p>
Rana sierrae annotated aquatic soundscapes (2022)
<p>This dataset is associated with the following manuscript, which contains details in the methodology of data collection and annotation: </p> <p>Lapp, S., Smith, T. C., Wilhelm, A, Knapp, R., Kitzes, J. In press. Aquatic soundscape recordings reveal diverse vocalizations and nocturnal activity of an endangered frog. The American Naturalist.</p> <p><em>Rana</em> <em>sierrae</em> (the Sierra Nevada yellow-legged frog) is an endangered species residing in high-elevation lakes in the Sierra Nevada mountains. The species is highly aquatic and, unlike most amphibians, primarily vocalizes while underwater. As a result, its vocalizations have rarely been recorded and its vocal repertoire is not well studied.</p> <p>This dataset contains an annotated set of underwater soundscape recordings containing <span>1236</span> annotations of <em>R. sierrae</em> vocalizations. We annotated five distinct vocalization types of<em> R. sierrae</em>, only two of which have been previously documented for this species. Besides the calls of <em>R.</em> <em>sierrae</em>, these audio recordings also contain stridulation sounds (not annotated), which were most likely produced by members of the family Corixidae or other aquatic invertebrates that stridulate underwater. </p>
Urban Soundscapes of the World
<p>The Urban Soundscapes of the World database currently contains about 130 high-quality audiovisual recordings performed within 9 cities worldwide. The csv and json files contain the recording locations. </p><p>Each recording consists of a 360-degree video file (4096 x 2048 resolution, 30 fps), a 4-channel first-order ambisonics (ACN/SN3D) audio file and/or a binaural audio file. All audio files have a sample rate of 48 kHz and are 24-bit PCM encoded. All audio and video files are time-synchronized.</p><p>Recordings are made during the day, in favorable weather conditions with little to no wind. Note that the recordings always present a snapshot in time. Combined and simultaneous audio and video recordings are performed using a portable, stationary recording setup as shown on the picture. The setup consists of the following components (from top to bottom):</p><ul><li>First order ambisonics: Core Sound TetraMic with windshield and Tascam DR-680 MkII 4-channel recording device;</li><li>360-degree video camera: GoPro Omni spherical camera system (only available upon request).</li><li>Binaural audio: HEAD acoustics HSU III.2 artificial head with windshield and SQobold 2-channel recording device;</li></ul><p>The ears of the artificial head, the video camera system and the ambisonics microphone are located at heights of about 1.5m, 1.7m and 1.9m, respectively. At each location, the recording system is oriented towards the most important sound source and/or the most prominent visual scene—this orientation defines the initial frontal viewing direction for the 360-degree video and ambisonics recordings, and the fixed orientation for the binaural recordings.</p><p>All audio files are calibrated to the same reference, so once you have your playback setup calibrated, it can be used to play all files. The csv and json file contains the one-minute LAeq values of the binaural recordings (average of left and right channel and left and right channels separately). These values are the most representative for the LAeq at the location. The second column presents the LAeq of the mono mix (superposition) of both left and right channels of the binaural recording. Note that this is not necessarily the same as the (energetic) average of the LAeq's of both left and right channels separately, because both channels are to some degree correlated (depending on the diffuseness of the sound field). This explains why the (energetic) average of the third and fourth column will not always exactly correspond to the value in the second column, but the difference is usually small. Roughly speaking, the larger the difference, the more the sound at both ears is correlated.</p><p>More details on the recording setup and protocol can be found in our <a href="http://urban-soundscapes.org/publications/">publications</a>. Note that some publications contain LAeq values that were calculated from the ambisonics recordings (W channel). There is not really a standard way of calculating LAeq values from ambisonics recordings, so these are maybe less suitable to use in most cases.</p>
Acoustic features as a tool to visualize and explore marine soundscapes: Applications illustrated using marine mammal Passive Acoustic Monitoring datasets
<p>Passive Acoustic Monitoring (PAM) is emerging as a solution for monitoring species and environmental change over large spatial and temporal scales. However, drawing rigorous conclusions based on acoustic recordings is challenging, as there is no consensus over which approaches, and indices are best suited for characterizing marine and terrestrial acoustic environments.</p> <p>Here, we describe the application of multiple machine-learning techniques to the analysis of a large PAM dataset. We combine pre-trained acoustic classification models (VGGish, NOAA & Google Humpback Whale Detector), dimensionality reduction (UMAP), and balanced random forest algorithms to demonstrate how machine-learned acoustic features capture different aspects of the marine environment.</p> <p>The UMAP dimensions derived from VGGish acoustic features exhibited good performance in separating marine mammal vocalizations according to species and locations. RF models trained on the acoustic features performed well for labelled sounds in the 8 kHz range, however, low and high-frequency sounds could not be classified using this approach.</p> <p>The workflow presented here shows how acoustic feature extraction, visualization, and analysis allow for establishing a link between ecologically relevant information and PAM recordings at multiple scales.</p> <p>The datasets and scripts provided in this repository allow replicating the results presented in the publication. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.