Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “audio tagging”
Audio tagging of avian dawn chorus recordings in California, Oregon, and Washington
<p><strong>General Summary</strong></p> <p>This acoustic data collection includes 1,575 5-minute soundscape recordings randomly selected from passive acoustic recordings made at 525 sites during 2022 on federally managed lands in western California, Oregon, and Washington, USA. We fully labeled 141 recordings (11.75 hrs) with 39,717 annotations for 118 sound types, including 58 avian species, two mammalian species, six aggregated biotic sounds, and eight non-biotic sound types. An additional 215 recordings were partially annotated with 1,466 annotations. The remaining unlabeled recordings have been included to facilitate novel research applications and methodological evaluations. Beyond the labeled soundscape recordings, we have included township and range identifications and 38 environmental covariates for each recording location.</p> <p><strong>Data Collection</strong></p> <p>Lesmeister et al. (2021) collected passive acoustic recordings during 2022 in support of long-term monitoring of federally threatened northern spotted owl (<em>Strix occidentalis caurina) </em>populations under the Northwest Forest Plan Effective Monitoring Program (U. S. Fish and Wildlife Service 1990, U. S. Department of Agriculture and U. S. Department of the Interior 1994). These data were collected at 643 hexagons that were randomly selected from a tessellation of 5 km2 hexagons covering the entire range of the northern spotted owl (Northern California, Oregon, Washington) under a selective constraint that hexagons contain ≥ 50 % forest-capable lands (<em>def.</em> forested lands or lands capable of developing closed-canopy forests) and be ≥ 25% federal ownership (Davis et al., 2011).</p> <p>Each hexagon was sampled by four Song Meter 4 (SM4) acoustic recording units (Wildlife Acoustics, Maynard, MA) deployed in a standardized spatial arrangement, such that recorders on a site were placed ≥ 500 m apart and were ≥ 200 m from the edge of the sampling hexagon boundary. Recorders were mounted to small trees (15 – 20 cm diameter at breast height) approximately 1.5 m above the ground and were placed on mid-to-upper slopes and ≥ 50 m from roads, trails, and streams. The SM4 devices each have two built-in omnidirectional microphones with a signal-to-noise ratio of 80 dB, typical at 1 kHz, and a recording bandwidth of 20 Hz – 48 kHz. Each device recorded ~11 hours of audio daily for six weeks from March to August at a sampling rate of 32 kHz. The daily recording schedule included a 4-hour window from two hours before sunrise to two hours after sunrise, a 4-hour window from one hour before sunset to 3 hours after sunset, and 10-minute recordings outside the two longer recording blocks at the start of every hour.</p> <p><strong>Data Sampling</strong></p> <p>The goal of this project was to develop a tagged audio dataset (hereafter project dataset) focused on the avian dawn chorus, which is an ecologically important period for the study of avian behavior (McNamara et al. 1987, Staicer et al. 1996, Zhang et al. 2015) and monitoring avian biodiversity (Bibby et al. 2000), but remains a challenging problem for acoustic classification systems (Duan et al. 2013, Stowell 2022). Passive acoustic monitoring on our sites occurs throughout the day. We filtered the full dataset to recordings collected between May and August during the hour immediately after sunrise. From the recordings meeting our filtering criteria, we randomly selected three 5-minute files from each site, which were assigned ordinal labels 'A, 'B,' or 'C.' The final project dataset comprised 131.25 hours of acoustic data.</p> <p><strong>Annotation Protocol</strong></p> <p>We randomly selected 141 sites from the project dataset and fully annotated each recording at a 2-second resolution. We applied labels to each 2-second window of the selected recordings following a predefined sound phonology library (available in the 'metadata.tsv' file), which concatenated the 2021 eBird taxonomy codes (Clements list; Clements et al. 2022) with standardized sonotype codes that incremented depending on the species repertoire (i.e., 'call_1,' 'song_1,' 'drum_1'). For example, 'herthr_song_1' is the label for Hermit Thrush, song_1. Unknown signals were labeled 'unknown,' and clips with no biotic signals (or noise classes of interest documented in metadata.tsv) were labeled 'empty.' Windows were labeled 'complete' and considered fully annotated when every signal was assigned an annotation. Files were deemed fully annotated when every 2-second window contained the 'complete' label.</p> <p><strong>Environmental Covariates</strong></p> <p>Sampling locations will not be published to afford protections for Federally Threatened or Endangered species which may occur on our sites. However, we provide the State, Township, and Range for each sampling location along with the site-specific values for 38 forest structure, topographic, and climatic environmental covariates developed by the Landscape Ecology, Modeling, Mapping, and Analysis group in the Pacific Northwest (<a href="https://lemma.forestry.oregonstate.edu/data">https://lemma.forestry.oregonstate.edu/data</a>; Ohmann and Gregory 2002). State, Township, and Range values are sufficient to explore geographic variation in species- or community-specific call and song phenology and the extracted environmental covariates may provide useful contextual information for novel machine-learning developments (Liu et al. 2018). </p> <p><strong>Description of Data Format</strong></p> <p>The fully annotated audio files can be accessed by downloading and extracting "annotated_recordings.zip." Partially annotated and non-annotated audio files can be accessed by downloading and extracting "additional_recordings_part_1.zip" or "additional_recordings_part_2.zip." Acoustic file names contain site and replicate indicators, such that file "Site_001_Rep_A.wav' was recorded on site 1 and is the A replicate random draw from the available set of dawn chorus recordings. The site and replicate numbers link to additional recording information in "files.tsv," annotations in "annotations.tsv" and "partial_annotations.tsv," as well as site and replicate specific environmental characteristics in "environmental_characteristics.tsv."</p> <p>Metadata describing sound classes and environmental characteristics can be found in "metadata.tsv," and "environmental_characteristics_metadata.tsv."</p> <p><strong>Acknowledgments</strong></p> <p>Acoustic data collection was funded and collected by the US Forest Service and the US Bureau of Land Management. Annotation work was funded by Google. We would also like to thank the many biologists that collected and processed the data compiled here. The use of trade or firm names in this publication is for reader information and does not imply endorsement by the U.S. Government of any product or service.</p>
LOLA. Flamenco tagged audios Dataset
<p>This dataset holds ~1500 audio samples of 3 different flamenco styles. Those styles are:</p> <ul> <li>Bulerías</li> <li>Alegrías</li> <li>Sevillanas</li> </ul> <p>Each audio is in a folder with the same name as the style to which it belongs. Each audio has a duration between 10 and 15 seconds and is saved in mp3 format.</p>
Text to audio grounding (TAG) dataset: AudioGrounding
<p>AudioGrounding dataset, including audio files and timestamp annotations.</p><p>Changes in version 2: The train/validation/test sets are re-split. The validation and test annotations are refined.</p><p> </p><p>----------------------------------------------------------</p><p><strong>References</strong></p><p>[1] Xuenan Xu, Heinrich Dinkel, Mengyue Wu and Kai Yu. "Text-to-audio grounding: Building correspondence between captions and sound events." In <i>Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</i>. IEEE, 2021, pp. 606-610.</p><p>[2] Xuenan Xu, Mengyue Wu, and Kai Yu. "Investigating Pooling Strategies and Loss Functions for Weakly-Supervised Text-to-Audio Grounding via Contrastive Learning." In <i>Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW)</i>. IEEE, 2023, pp. 1-5.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.