Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “Passive acoustics”
Acoustic features as a tool to visualize and explore marine soundscapes: Applications illustrated using marine mammal Passive Acoustic Monitoring datasets
<p>Passive Acoustic Monitoring (PAM) is emerging as a solution for monitoring species and environmental change over large spatial and temporal scales. However, drawing rigorous conclusions based on acoustic recordings is challenging, as there is no consensus over which approaches, and indices are best suited for characterizing marine and terrestrial acoustic environments.</p> <p>Here, we describe the application of multiple machine-learning techniques to the analysis of a large PAM dataset. We combine pre-trained acoustic classification models (VGGish, NOAA & Google Humpback Whale Detector), dimensionality reduction (UMAP), and balanced random forest algorithms to demonstrate how machine-learned acoustic features capture different aspects of the marine environment.</p> <p>The UMAP dimensions derived from VGGish acoustic features exhibited good performance in separating marine mammal vocalizations according to species and locations. RF models trained on the acoustic features performed well for labelled sounds in the 8 kHz range, however, low and high-frequency sounds could not be classified using this approach.</p> <p>The workflow presented here shows how acoustic feature extraction, visualization, and analysis allow for establishing a link between ecologically relevant information and PAM recordings at multiple scales.</p> <p>The datasets and scripts provided in this repository allow replicating the results presented in the publication. </p>
Estimating the abundance of the critically endangered Baltic Proper harbour porpoise (Phocoena phocoena) population using passive acoustic monitoring
<p>Knowing the abundance of a population is a crucial component to assess its conservation status and develop effective conservation plans. For most cetaceans, abundance estimation is difficult given their cryptic and mobile nature, especially when the population is small and has a transnational distribution. In the Baltic Sea, the number of harbour porpoises (<i>Phocoena phocoena</i>) has collapsed since the mid-20<sup>th</sup> century and the Baltic Proper harbour porpoise is listed as Critically Endangered by the IUCN and HELCOM; however, its abundance remains unknown. Here, one of the largest ever passive acoustic monitoring studies was carried out by eight Baltic Sea nations to estimate the abundance of the Baltic Proper harbour porpoise for the first time. By logging porpoise echolocation signals at 298 stations during May 2011-April 2013, calibrating the loggers' spatial detection performance at sea, and measuring the click rate of tagged individuals, we estimated an abundance of 71-1,105 individuals (95% CI, point estimate 491) during May-October within the population's proposed management border. The small abundance estimate strongly supports that the Baltic Proper harbour porpoise is facing an extremely high risk of extinction, and highlights the need for immediate and efficient conservation actions through international cooperation. It also provides a starting point in monitoring the trend of the population abundance to evaluate the effectiveness of management measures and determine its interactions with the larger neighbouring Belt Sea population. Further, we offer evidence that design-based passive acoustic monitoring can generate reliable estimates of the abundance of rare and cryptic animal populations across large spatial scales.</p>
Unlabeled AnuraSet: A dataset for leveraging unlabeled data in machine learning models for passive acoustic monitoring
<p>The Unlabeled AnuraSet (U-AnuraSet) is an extension of the original AnuraSet dataset. It consists of soundscape recordings from passive acoustic monitoring conducted in Brazil. The recording sites are identical to those in the original AnuraSet. Each site comprises 2,666 one-minute raw audio files of unlabeled data. The U-AnuraSet is publicly available to encourage machine learning researchers to explore innovative methods for leveraging unlabeled data in the training of models aimed at solving problems such as anuran call identification.</p> <p>If you find the Unlabeled AnuraSet useful for your research, please consider citing it as follows:</p> <p>Cañas, J.S., Toro-Gómez, M.P., Sugai, L.S.M., et al. A dataset for benchmarking Neotropical anuran calls identification in passive acoustic monitoring. Sci Data 10, 771 (2023). https://doi.org/10.1038/s41597-023-02666-2</p>
Figure 2 in The Distribution Of The Northern Bat Eptesicus Nilssonii (Keyserling & Blasius, 1839) In Latvia Assessed By Passive Acoustic Survey
Figure 2. Dot plot showing the activity of E. nilssonii in each observation site. The outer dot in each region represents the mean and the whiskers the standard error of the mean. Different letters indicate statistically significant differences (p<0.05) (A). Data visualization depicting the gradient of the activity of E. nilssonii in four parts of Latvia. The darker color represents the higher activity (B).
Figure 1 in The Distribution Of The Northern Bat Eptesicus Nilssonii (Keyserling & Blasius, 1839) In Latvia Assessed By Passive Acoustic Survey
Figure 1. The map of Latvia divided in four regions under LKS92 25x25 km square network. Bat activity was studied in randomly selected squares (visited squares marked grey). In total, 60 squares were surveyed with six survey sites chosen in each square (n=360).
Data from: Performance of unmarked abundance models with data from machine-learning classification of passive acoustic recordings
<p>The ability to conduct cost-effective wildlife monitoring at scale is rapidly increasing due to availability of inexpensive autonomous recording units (ARUs) and automated species recognition, presenting a variety of advantages over human-based surveys. However, estimating abundance with such data collection techniques remains challenging because most abundance models require data that are difficult for low-cost monoaural ARUs to gather (e.g., counts of individuals, distance to individuals), especially when using the output of automated species recognition. Statistical models that do not require counting or measuring distances to target individuals in combination with low-cost ARUs provide a promising way of obtaining abundance estimates for large-scale wildlife monitoring projects but remain untested. We present a case study using avian field data collected in forests of Pennsylvania during the Spring of 2020 and 2021 using both traditional point counts and passive acoustic monitoring at the same locations. We tested the ability of the Royle-Nichols and time-to-detection models to estimate abundance of two species from detection histories generated by applying a machine-learning classifier to ARU-gathered data. We compared abundance estimates from these models to estimates from the same models fit using point-count data and to two additional models appropriate for point counts, the N-mixture model and distance models. We found that the Royle-Nichols and time-to-detection models can be used with ARU data to produce abundance estimates similar to those generated by a point-count based study but with greater precision. ARU-based models produced confidence or credible intervals that were on average 31.9% ( 11.9 SE) smaller than their point-count counterpart. Our findings were consistent across two species with differing relative abundance and habitat use patterns. The higher precision of models fit using ARU data is likely due to higher cumulative detection probability, which itself may be the result of greater survey effort using ARUs and machine-learning classifiers to sample significantly more time for focal species at any given point. Our results provide preliminary support the use of ARUs in abundance-based study applications, and thus may afford researchers a better understanding of habitat quality and population trends, while allowing them to make more informed conservation actions and recommendations.</p>
Data for: Large-scale long-term passive-acoustic monitoring reveals spatiotemporal activity patterns of boreal bats
<p class="MsoNormal"><span>The distribution ranges and spatio-temporal patterns in the occurrence and activity of boreal bats are yet largely unknown due to their cryptic lifestyle and lack of suitable and efficient study methods. We approached the issue by establishing a permanent passive-acoustic sampling setup spanning the area of Finland to gain an understanding on how latitude affects bat species composition and activity patterns in northern Europe. The recorded bat calls were semi-automatically identified for three target taxa; <em>Myotis</em> spp., <em>Eptesicus nilssonii</em> or <em>Pipistrellus nathusii</em> and the seasonal activity patterns were modeled for each taxa across the seven sampling years (2015–2021). We found an increase in activity since 2015 for <em>E. nilssonii</em> and <em>Myotis </em>spp. For <em>E. nilssonii</em> and <em>Myotis</em> spp. we found significant latitude -dependent seasonal activity patterns, where seasonal variation in patterns appeared stronger in the north. Over the years, activity of <em>P. nathusii</em> increased during activity peak in June and late season but decreased in mid season. We found the passive-acoustic monitoring </span><span>network to be an effective and cost-efficient method for gathering b</span><span>at activity data to analyze spatio-temporal patterns. Long-term data on the composition and dynamics of bat communities facilitates better estimates of abundances and population trend directions for conservation purposes and predicting the effects of cli</span><span>mate change.</span></p>
Passive acoustic monitoring applied to black-and-white ruffed lemurs (Varecia variegata) in Ranomafana National Park, Madagascar
<p>Data accompanying the paper: <strong>"An integrated passive acoustic monitoring and deep learning pipeline applied to black-and-white ruffed lemurs (\textit{Varecia variegata}) in Ranomafana National Park, Madagascar"</strong></p> <p>Fieldwork was conducted at Mangevo (21.3833S, 47.4667E), an isolated and undisturbed forest location within Ranomafana National Park (RNP), located in southeastern Madagascar, during the period of May to July 2019. To facilitate passive acoustic monitoring, we deployed a total of two SongMeter SM4 devices (manufactured by Wildlife Acoustics) and two Swift units (provided by the Cornell Yang Center for Conservation Bioacoustics). The placement of these recorders was strategically chosen within the central regions of known subgroups, ensuring a minimum distance of 300 meters between each device. The SongMeter devices operated at a sampling rate of 48 kHz, while the Swift units operated at 32 kHz, respectively, enabling comprehensive audio data collection throughout the study period.</p> <p>We provide the audio data (.wav) used to train and test our neural network classifier along with the corresponding labelled text files (.data).</p> <p><strong>Files provided</strong></p> <ul> <li><strong>Test_Audio.zip </strong>-- contains (.wav) testing audio files</li> <li><strong>Test_Annotations.zip </strong>-- contains (.svl) manually annotated testing files which can be read in using Sonic Visualiser or by parsing the XML file in Python or another programming language. Load in the audio file into Sonic Visualiser and then drag-and-drop the corresponding .svl file.</li> <li><strong>Training_Audio_batch_x.zip -</strong>- several .zip files were created to simplify downloading. There are 10 batches, each is roughly 4GB. Each batch contains (.wav) training audio files</li> <li><strong>Training_Annotations.zip</strong> -- contains (.svl) manually annotated training files which can be read in using Sonic Visualiser or by parsing the XML file in Python or another programming language. Load in the audio file into Sonic Visualiser and then drag-and-drop the corresponding .svl file.</li> <li><strong>model_weights_tensorflow.hdf5 </strong>-- the Tensorflow model. Load the model using: model = tf.keras.models.load_model(model_filepath) note that the model expects a three channel input as explained in the research article.</li> </ul>
Data from: Narwhal (Monodon monoceros) echolocation click rates to support cue counting passive acoustic density estimation
<p class="MsoNormal"><span>The datasets correspond to the data used to obtain the results shown in the manuscript "Narwhal (<em>Monodon monoceros</em>) echolocation click rates to support cue counting passive acoustic density estimation".</span></p> <p class="MsoNormal"><span>When the manuscript is accepted we will also edit and add here the full reference including the DOI of the publication.</span></p>
Data from: Optimizing passive acoustic monitoring (PAM) for Biodiversity Studies: using species-area relationship (SAR) to predict species richness
Open the record for dataset details and reuse information.
A novel method for estimating avian roost sizes using passive acoustic recordings
Open the record for dataset details and reuse information.
Acoustic features as a tool to visualize and explore marine soundscapes: Applications illustrated using marine mammal Passive Acoustic Monitoring datasets
Open the record for dataset details and reuse information.
Data from: Performance of unmarked abundance models with data from machine-learning classification of passive acoustic recordings
Open the record for dataset details and reuse information.
Data for: Large-scale long-term passive-acoustic monitoring reveals spatiotemporal activity patterns of boreal bats
Open the record for dataset details and reuse information.
Estimating the abundance of the critically endangered Baltic Proper harbour porpoise (Phocoena phocoena) population using passive acoustic monitoring
Open the record for dataset details and reuse information.
Data from: Narwhal (Monodon monoceros) echolocation click rates to support cue counting passive acoustic density estimation
Open the record for dataset details and reuse information.
Comparing distribution of harbour porpoises (Phocoena phocoena) derived from satellite telemetry and passive acoustic monitoring
<p>Data used for publication in Plos One. Two excel files. The satellite_filtered_data is the filtered satellite positions used for MaxEnt modelling in R. The CPOD_data_PPH is the raw C-POD data expressed here as porpoises positive hours (PPH) and can easily be converted to porpoise positive days (PPD).</p>
A labelled dataset of the loud calls of four vertebrates collected using passive acoustic monitoring in Malaysian Borneo
<p><em>Passive acoustic monitoring data collection</em></p> <p>We collected data using first generation Swift autonomous recording units (ARUs) (Koch et al. 2016) with a microphone sensitivity of −44 (+/−3) dB re 1 V/Pa. The microphone frequency response was not measured but is assumed to be flat (+/− 2 dB) in the frequency range 100 Hz to 7.5 kHz. The analog signal was amplified by 40 dB and digitized (16-bit resolution) using an analog-to-digital converter (ADC) with a clipping level of −/+ 0.9 V. We collected acoustic data from one primary conservation area in Sabah, Malaysia: Danum Valley Conservation Area (with 11 recording units from March to July 2018). Danum Valley covers an area of roughly 440 km², and is characterized by lowland dipterocarp forest. Unlike many tropical forest regions, this area is considered 'aseasonal' due to its lack of clearly differentiated wet and dry seasons (Walsh and Newbery 1999). In Danum Valley, the ARUs recorded at a sampling rate of 16 kHz. All recordings were saved in waveform audio (.wav) format, with files of 2-hr duration. We affixed each recording unit to trees approximately 2-m above the ground and recorded continuously over 24 hours. We set the units on a 750 m grid structure, and preliminary field tests indicate that with these recording settings the detection range of gibbon vocalizations is ~ 400 m.</p> <p> </p> <p><em>Acoustic data processing </em></p> <p>We randomly chose approximately 500 h of recordings from Danum Valley Conservation Area to use to create a training dataset. We used a band-limited energy detector (BLED) to identify potential sounds of interest in the gibbon frequency range. For the BLED detector, we convert the 2-hr recordings into a spectrogram using a 1,600-point (100 ms) Hamming window (3 dB bandwidth = 13 Hz) with 0% overlap and a 2,048-point DFT, with the "seewave" package (Sueur et al. 2008). We then filtered the spectrogram to focus on the desired frequency range, specifically 0.5–1.6 kHz for Northern grey gibbons. For each unique time window in the recording, we determined the total energy across frequency bins which gave a single value for every 100 ms interval. Utilizing the "quantile" function in base R, we established the threshold to delineate signal from noise. Preliminary tests with varied quantile values revealed that the 15th quantile led to optimized recall for our target signal. This approach resulted in 1,439 unique sound events. The sound events were then annotated by a single observer (DJC) using a custom-written function in R to visualize the spectrograms into the following categories: great argus pheasant (<em>Argusianus argus</em>) long and short calls (Clink et al. 2021), helmeted hornbills (<em>Rhinoplax vigil</em>), rhinoceros hornbills (<em>Buceros rhinoceros</em>), female gibbons (<em>Hylobates funereus</em>) and a catch-all “noise” category. </p> <p><em>Update Version 5 and later</em></p> <p>Includes labeled test clips from a second conservation area, Maliau Basin Conservation Area, Sabah, Malaysia recorded during August 2019. The ARUs recorded at a sampling rate of 16 kHz. All recordings were saved in waveform audio (.wav) format, with files of 2-hr duration. We affixed each recording unit to trees approximately 2-m above the ground and recorded continuously over 24 hours.</p> <p> </p> <p><strong>References</strong></p> <p>Clink, D. J., Groves, T., Ahmad, A. H., & Klinck, H. (2021). Not by the light of the moon: Investigating circadian rhythms and environmental predictors of calling in Bornean great argus. <em>PloS one</em>, <em>16</em>(2), e0246564.</p> <p>Koch, R., Raymond, M., Wrege, P., & Klinck, H. (2016). SWIFT: A small, low-cost acoustic recorder for terrestrial wildlife monitoring applications. In <em>North American Ornithological Conference</em> (p. 619). Washington, D.C.</p> <p>Sueur, J., Aubin, T., & Simonis, C. (2008). Seewave: a free modular tool for sound analysis and synthesis. <em>Bioacoustics</em>, <em>18</em>, 213–226.</p> <p>Walsh, R. P., & Newbery, D. M. (1999). The ecoclimatology of Danum, Sabah, in the context of the world’s rainforest regions, with particular reference to dry periods and their impact. <em>Philosophical transactions of the Royal Society of London. Series B, Biological sciences</em>, <em>354</em>(1391), 1869–83. https://doi.org/10.1098/rstb.1999.0528</p> <p>Webb, C. O., & Ali, S. (2002). Plants and vegetation of the Maliau Basin Conservation Area, Sabah, East Malaysia. <em>Final Report to Maliau Basin Management Committee</em>.</p> <p> </p>
Thyolo alethe (Chamaetylas choloensis) calls for passive acoustic monitoring
<p>Data accompanying the paper: "Passive Acoustic Monitoring and Transfer Learning"</p> <p><strong>Please cite this dataset as:</strong></p> <blockquote> <p>Dufourq, Emmanuel and Batist, Carly and Foquet, Ruben and Durbach, Ian. (2022). Passive Acoustic Monitoring and Transfer Learning. BioRxiv doi: </p> </blockquote> <p>This dataset contains approximately 10 hours of audio that contained calls of the vulnerable Thyolo Alethe (Chamaetylas choloensis). The audio data was collected in the Mount Mulanje Biosphere Reserve, Malawi using 10 Audiomoths. The sampling rate was set to 32,000Hz and the recordings were obtained over five days in November 2020. A larger dataset exists.</p> <p>The annotations files are in (.svl) format which is compatible with SonicVisualiser (https://www.sonicvisualiser.org/). Each audio file has a corresponding .svl file. Each .svl has segments of audio that were manually annotated as either ''thyolo-alethe" (presence class) or "noise" (absence class) -- this dataset can be used to train a binary classification model.</p> <p>The audio files are provided in "Audio.zip" and the manually verified annotation in "Annotations.zip".</p>
Pin-tailed whydah (Vidua macroura) calls for passive acoustic monitoring
<p>Data accompanying the paper: "Passive Acoustic Monitoring and Transfer Learning"</p> <p><strong>Please cite this dataset as:</strong></p> <blockquote> <p>Dufourq, Emmanuel and Batist, Carly and Foquet, Ruben and Durbach, Ian. (2022). Passive Acoustic Monitoring and Transfer Learning. BioRxiv doi: </p> </blockquote> <p>This dataset contains approximately 6 hours of audio that contained calls of the pin-tailed whydah (Vidua macroura). The audio data was collected in the Intaka Island Nature Reserve in Cape Town, South Africa using 1 Audiomoth. The sampling rate was set to 48,000Hz and the recordings were obtained over four days in January 2021. A larger dataset exists.</p> <p>The annotations files are in (.svl) format which is compatible with SonicVisualiser (https://www.sonicvisualiser.org/). Each audio file has a corresponding .svl file. Each .svl has segments of audio that were manually annotated as either ''thyolo-alethe" (presence class) or "noise" (absence class) -- this dataset can be used to train a binary classification model.</p> <p>The audio files are provided in "Audio.zip" and the manually verified annotation in "Annotations.zip".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.