Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
639
datasets available to search
ShareScore release 0.7.1
Dataset results
639 results for “audio”
Audio Cartography
Open the record for dataset details and reuse information.
BiVib - Audio-Tactile Piano Sample Library
<p><strong>BiVib</strong> is an extensive piano sample library consisting of <strong>bi</strong>naural sounds and keyboard <strong>vib</strong>ration signals.<br>Samples were acquired with high-quality audio and vibration measurement equipment on two <a href="https://en.wikipedia.org/wiki/Disklavier">Yamaha Disklavier pianos</a> (one grand and one upright model) by means of computer-controlled playback of each key at ten different MIDI velocity values.<br>Project files (<em>instruments</em> and <em>multis</em>) are provided for use with the software sampler <a href="https://www.native-instruments.com/en/products/komplete/samplers/kontakt-6/">Native Instruments Kontakt</a> (version 5 and above, available for Windows and Mac OS).<br>The nominal specifications of the equipment used in the acquisition chain are reported in a companion document, allowing researchers to calculate physical quantities (e.g. acoustic pressure, vibration acceleration) from the recordings.<br>The library is especially suited for acoustic and vibration research on the piano, as well as for research on multimodal interaction with musical instruments.</p>
Audio and Water Movement Data for Oyster Reef and Mudflat Sites on the Coast of Virginia, 2018
Paired marine audio soundscape recordings with concurrent acoustic Doppler velocimeter (ADV) turbulence measurements at three intertidal sites: a natural oyster reef, a restored oyster reef, and a bare mudflat. The objective was to establish the link between oyster reef soundscapes and hydrodynamics, specifically the role that turbulence plays in generating near field pressure waves that may be used as physical cues by oyster larvae when initiating settlement behaviors. Data were collected at three different locations, with water turbulance measurements at rates up to 25Hz concurrent with audio recording.
Audio tagging of avian dawn chorus recordings in California, Oregon, and Washington
<p><strong>General Summary</strong></p> <p>This acoustic data collection includes 1,575 5-minute soundscape recordings randomly selected from passive acoustic recordings made at 525 sites during 2022 on federally managed lands in western California, Oregon, and Washington, USA. We fully labeled 141 recordings (11.75 hrs) with 39,717 annotations for 118 sound types, including 58 avian species, two mammalian species, six aggregated biotic sounds, and eight non-biotic sound types. An additional 215 recordings were partially annotated with 1,466 annotations. The remaining unlabeled recordings have been included to facilitate novel research applications and methodological evaluations. Beyond the labeled soundscape recordings, we have included township and range identifications and 38 environmental covariates for each recording location.</p> <p><strong>Data Collection</strong></p> <p>Lesmeister et al. (2021) collected passive acoustic recordings during 2022 in support of long-term monitoring of federally threatened northern spotted owl (<em>Strix occidentalis caurina) </em>populations under the Northwest Forest Plan Effective Monitoring Program (U. S. Fish and Wildlife Service 1990, U. S. Department of Agriculture and U. S. Department of the Interior 1994). These data were collected at 643 hexagons that were randomly selected from a tessellation of 5 km2 hexagons covering the entire range of the northern spotted owl (Northern California, Oregon, Washington) under a selective constraint that hexagons contain ≥ 50 % forest-capable lands (<em>def.</em> forested lands or lands capable of developing closed-canopy forests) and be ≥ 25% federal ownership (Davis et al., 2011).</p> <p>Each hexagon was sampled by four Song Meter 4 (SM4) acoustic recording units (Wildlife Acoustics, Maynard, MA) deployed in a standardized spatial arrangement, such that recorders on a site were placed ≥ 500 m apart and were ≥ 200 m from the edge of the sampling hexagon boundary. Recorders were mounted to small trees (15 – 20 cm diameter at breast height) approximately 1.5 m above the ground and were placed on mid-to-upper slopes and ≥ 50 m from roads, trails, and streams. The SM4 devices each have two built-in omnidirectional microphones with a signal-to-noise ratio of 80 dB, typical at 1 kHz, and a recording bandwidth of 20 Hz – 48 kHz. Each device recorded ~11 hours of audio daily for six weeks from March to August at a sampling rate of 32 kHz. The daily recording schedule included a 4-hour window from two hours before sunrise to two hours after sunrise, a 4-hour window from one hour before sunset to 3 hours after sunset, and 10-minute recordings outside the two longer recording blocks at the start of every hour.</p> <p><strong>Data Sampling</strong></p> <p>The goal of this project was to develop a tagged audio dataset (hereafter project dataset) focused on the avian dawn chorus, which is an ecologically important period for the study of avian behavior (McNamara et al. 1987, Staicer et al. 1996, Zhang et al. 2015) and monitoring avian biodiversity (Bibby et al. 2000), but remains a challenging problem for acoustic classification systems (Duan et al. 2013, Stowell 2022). Passive acoustic monitoring on our sites occurs throughout the day. We filtered the full dataset to recordings collected between May and August during the hour immediately after sunrise. From the recordings meeting our filtering criteria, we randomly selected three 5-minute files from each site, which were assigned ordinal labels 'A, 'B,' or 'C.' The final project dataset comprised 131.25 hours of acoustic data.</p> <p><strong>Annotation Protocol</strong></p> <p>We randomly selected 141 sites from the project dataset and fully annotated each recording at a 2-second resolution. We applied labels to each 2-second window of the selected recordings following a predefined sound phonology library (available in the 'metadata.tsv' file), which concatenated the 2021 eBird taxonomy codes (Clements list; Clements et al. 2022) with standardized sonotype codes that incremented depending on the species repertoire (i.e., 'call_1,' 'song_1,' 'drum_1'). For example, 'herthr_song_1' is the label for Hermit Thrush, song_1. Unknown signals were labeled 'unknown,' and clips with no biotic signals (or noise classes of interest documented in metadata.tsv) were labeled 'empty.' Windows were labeled 'complete' and considered fully annotated when every signal was assigned an annotation. Files were deemed fully annotated when every 2-second window contained the 'complete' label.</p> <p><strong>Environmental Covariates</strong></p> <p>Sampling locations will not be published to afford protections for Federally Threatened or Endangered species which may occur on our sites. However, we provide the State, Township, and Range for each sampling location along with the site-specific values for 38 forest structure, topographic, and climatic environmental covariates developed by the Landscape Ecology, Modeling, Mapping, and Analysis group in the Pacific Northwest (<a href="https://lemma.forestry.oregonstate.edu/data">https://lemma.forestry.oregonstate.edu/data</a>; Ohmann and Gregory 2002). State, Township, and Range values are sufficient to explore geographic variation in species- or community-specific call and song phenology and the extracted environmental covariates may provide useful contextual information for novel machine-learning developments (Liu et al. 2018). </p> <p><strong>Description of Data Format</strong></p> <p>The fully annotated audio files can be accessed by downloading and extracting "annotated_recordings.zip." Partially annotated and non-annotated audio files can be accessed by downloading and extracting "additional_recordings_part_1.zip" or "additional_recordings_part_2.zip." Acoustic file names contain site and replicate indicators, such that file "Site_001_Rep_A.wav' was recorded on site 1 and is the A replicate random draw from the available set of dawn chorus recordings. The site and replicate numbers link to additional recording information in "files.tsv," annotations in "annotations.tsv" and "partial_annotations.tsv," as well as site and replicate specific environmental characteristics in "environmental_characteristics.tsv."</p> <p>Metadata describing sound classes and environmental characteristics can be found in "metadata.tsv," and "environmental_characteristics_metadata.tsv."</p> <p><strong>Acknowledgments</strong></p> <p>Acoustic data collection was funded and collected by the US Forest Service and the US Bureau of Land Management. Annotation work was funded by Google. We would also like to thank the many biologists that collected and processed the data compiled here. The use of trade or firm names in this publication is for reader information and does not imply endorsement by the U.S. Government of any product or service.</p>
Dataset for the study Late development of audio-visual integration in the vertical plane
<p>It is not clear how multisensory skills develop and how visual experience impacts on multisensory spatial development. Conflicting results show that visual calibration precedes multisensory integration for the audio-visual spatial bisection task (Gori et al., 2012a, 2012b) while in other tasks such as spatial localization, visual calibration occurs after multisensory development (Rohlf et al., 2020). Results in blind individuals can say something about the role of vision on perceptual development. Scientific evidences show that blind individuals have impairments in bisecting the auditory space (Gori et al., 2014) but not in localizing auditory sources (Lessard et al., 1998). Such results suggest that sensory calibration and impairment are linked. We studied the development of audio-visual multisensory localization in the vertical plane in sighted individuals from 5 years to adulthood to address this hypothesis. We hypothesize that typical children would show late audio-visual integration for the vertical plane, preceded by visual dominance. Unimodal and bimodal audio-visual thresholds and PSEs were measured and compared with the Bayesian optimal-integration model (maximum likelihood estimation). Results show that the development of multisensory integration in the vertical plane is not evident at 5 years, suggesting visual dominance for vertical audio-visual localization. These results support the idea that multisensory perception in the vertical domain depends on sensory calibration. We discuss these scientific results proposing that the process of cross-sensory calibration is task-specific and highlighting the importance of linking the impairment and development to better determine how our brain works.</p> <p>Data are in textual tab delimited format. Columns report for each subject: age, age_bin, condition, jnd.</p> <p> </p>
PartialSpoof Database - Partially Spoofed Audio Dataset for Anti-spoofing
<p>All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utterances that are only partially spoofed. By definition, partially-spoofed utterances contain a mix of both spoofed and bona fide segments, which will likely degrade the performance of countermeasures trained with entirely spoofed utterances. This hypothesis raises the obvious question: ‘Can we detect partially spoofed audio?’ This paper introduces a new database of partially-spoofed data, named <strong>PartialSpoof</strong>, to help address this question. This new database enables us to investigate and compare the performance of countermeasures on both utterance- and segmental- level labels. Experimental results using the utterance-level labels reveal that the reliability of countermeasures trained to detect fully-spoofed data is found to degrade substantially when tested with partially-spoofed data, whereas training on partially-spoofed data performs reliably in the case of both fully- and partially- spoofed utterances. Additional experiments using segmental-level labels show that spotting injected spoofed segments included in an utterance is a much more challenging task even if the latest countermeasure models are used.</p> <p> </p> <ul> <li><strong>!!!NEW!!! For detailed (bonafide/spoofing methods/nonspeech/concatenated parts) timestamps of PartialSpoof v1.3</strong> <ul> <li><a href="https://drive.google.com/drive/folders/1kKW3GBuooPkAl64Zyv6WPgICH5LZtnQR">Google Drive</a> </li> <li>The official version is under preparation. Please download this one if you urgently need it.</li> </ul> </li> <li>For fine-grained labels of PartialSpoof v1.2 <ul> <li>Arxiv: http://arxiv.org/abs/2204.05177</li> <li>PartialSpoof Database v1.2<strong> </strong>(including segmental-level labels in different temporal resolutions and timestamp labels)<strong>: This one</strong></li> </ul> </li> <li>For the multi-task version of PartialSpoof <strong>v1.1</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2107.14132</li> <li>PartialSpoof Database v1.1 (including 0.16s segmental level labels): https://zenodo.org/record/5112031</li> </ul> </li> <li>For the initial version of PartialSpoof <strong>v1.0</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2104.02518</li> <li>Samples: https://nii-yamagishilab.github.io/zlin-demo/IS2021/index.html</li> <li>PartialSpoof Database v1.0: https://zenodo.org/record/4817532</li> </ul> </li> </ul> <p>P.S.</p> <p>1. Compared to the <a href="../record/4817532#.YLO07S2l1hE">PartialSpoof_v1.0</a> and <a href="../record/5112031">PartialSpoof_v1.1</a>, only <strong>database_segment_labels_v1.2.tar.gz, database_vad.tar.gz, </strong> and<strong> README_v1.2</strong> are updated for version 1.2, you don't need to download other files if you already downloaded version1.0 or 1.1.</p> <p>2. File database_eval.tar.gz is a little large, if you cannot download it smoothly, you can download the split database_eval.tar.gz from <a href="../record/4817532#.YLO07S2l1hE">PartialSpoof_v1.0</a> </p>
Transcribing audio data: overview and transcripts of several automatic transcription tools
<p>Throughout institutions, audio recordings are being made regularly. To be able to further process these recordings, the audio often needs to be transcribed. In order to avoid having to transcribe the audio manually, there is a wealth of tools available for doing so automatically. In this record, we present an overview of several often-used tools to automatically transcribe pre-recorded audio data, including their features, costs, and security.</p> <p>To check the quality of the tool, we also recorded an audio fragment in Dutch that we ran through all tools in this overview in March of 2022. This original audio fragment (Test_interview_20220203.mp3), the cleaned-up transcription (Test_interview_cleaned_transcript.odt) and each tool’s raw transcript of the audio fragment (Test_interview_[name-tool]_raw_[date-run]) are included in this record as well. The raw transcripts were downloaded as .docx or .txt files and the .docx files saved as .odt. No edits to the transcripts were made before saving them, except an incidental removal of a personal email address or hyperlink.</p> <p>The overview contains information and transcripts of following transcription tools:</p> <ul> <li>Amberscript</li> <li>HappyScribe</li> <li>Kaldi</li> <li>NVIVO transcription</li> <li>Sonix</li> <li>SpokenOnline</li> <li>Transcribe</li> <li>Trint</li> <li>Microsoft Word 365 Online</li> </ul> <p><strong>About</strong></p> <p>This overview was created through a collaboration between Utrecht University’s Research Data Management (RDM) Support and the <a href="https://datahub.sites.uu.nl/">DataHub SSH</a> programme situated at the faculty of Humanities.</p> <p>The details in the overview have last been updated April 19, 2022. Please note that at the time you are downloading these files, the quality of the (Dutch) speech-to-text conversion may have been improved by the respective supplier.</p>
Supplementary materials to the paper: Automatic Parameters Tuning of Late Reverberation Algorithms for Audio Augmented Reality
<p>Supplementary materials to the paper:</p> <blockquote> <p>Riccardo Bona, Davide Fantini, Giorgio Presti, Marco Tiraboschi, Isaac Engel and Federico Avanzini. 2022. Automatic Parameters Tuning of Late Reverberation Algorithms for Audio Augmented Reality. In <em>Proceedings of International Conference on Audio Mostly</em>.</p> </blockquote> <p>The supplementary materials include the reverberated audio stimuli employed in the MUSHRA listening test reported in the paper. For each type of audio stimuli (Drums, Sax and Speech) the version reverberated with each of the six target Room Impulse Responses (RIRs) is provided along with the versions reverberated using the reverb matching method proposed in the paper (two different artificial reverberators have been considered: FDN and Freeverb).</p> <p>Further, the reverberation times (<span class="math-tex">\(T_{20}\)</span>) per octave band for each considered RIR are provided.</p>
The Audio Database of Hatoma Example Sentences
<p>This is a set of sound files of Hatoma Language, Southern Ryukyuan, spoken on Hatoma island, Okinawa.</p> <p>The database has 37611 sentences in Hatoma, included in Hatoma-Japanese Dictionary.</p> <p>See the audio database of Hatoma lexicon (Southern Ryukyuan) for a set of lexicons. https://doi.org/10.5281/zenodo.4560935</p>
CVoiceFake (crafted by SafeEar: Content Privacy-Preserving Audio Deepfake Detection)
<h1><strong>Introduction:</strong></h1> <p>CVoiceFake (small) is a dataset that features a random selection of 10% of samples from the entire collection. This dataset encompasses <strong>five common languages (English, Chinese, German, French, and Italian)</strong> and utilizes <strong>multi-advanced and classical voice cloning techniques</strong> (Parallel WaveGAN, Multi-band MelGAN, Style MelGAN, Griffin-Lim, WORLD, and DiffWave) to produce audio samples that bear a high resemblance to authentic audio.</p> <ol> <li><strong>Parallel WaveGAN</strong>: As a non-autoregressive vocoder-based model, Parallel WaveGAN produces high-fidelity audio rapidly, ideal for efficient and quality deepfake generation.</li> <li><strong>Multi-band MelGAN</strong>: Multi-band MelGAN is a variant of MelGAN that divides the frequency spectrum into sub-bands for faster and more stable multi-lingual vocoder training, enhancing the robustness and scalability of the dataset.</li> <li><strong>Style MelGAN</strong>: Style MelGAN is designed to capture fine prosodic and stylistic nuances of speech, making it particularly compelling for deepfake applications that require high levels of expressivity and variation in speech synthesis.</li> <li><strong>Griffin-Lim</strong>: This algorithm reconstructs waveforms from spectrograms using an iterative phase estimation method. Though less high-fidelity than neural vocoders, it serves as a traditional baseline for comparing deepfake generation.</li> <li><strong>WORLD</strong>: WORLD is a statistical parameter-based voice synthesis system that offers fine control over the spectral and prosodic features of the synthesized audio. Its fine manipulation is useful for crafting the nuanced variations needed in deepfake datasets.</li> <li>We have also built the SOTA diffusion-based deepfake audio (DiffWave); please contact the author at <code>xinfengli@zju.edu.cn</code> if you are interested in the dataset, particularly the DiffWave portion. Furthermore, any additional discussions are welcomed.<br><strong>DiffWave</strong>: DiffWave is a diffusion probability model for waveform generation. It converts the white noise signal into structured waveform through a Markov chain, capable of both conditional and unconditional generation tasks. DiffWave represents the advanced synthesis method for its fast synthesis speed and high synthesis quality.</li> </ol> <h1><strong>🔥</strong><strong>News:</strong></h1> <p>Please note that we recently released our DiffWave subset in Version 2 in comparison to Version 1, which is available on <a href="../records/14062964" target="_blank" rel="noopener">CVoiceFake Full</a>. You can download the file named CVoiceFake_Large_diffwave_update.tar.gz.xx, and after unzipping it, you will find it retains the same file structure as before.<br> <strong>| CVoiceFake_Large_diffwave_update.tar.gz.00 |<br> | CVoiceFake_Large_diffwave_update.tar.gz.01 |</strong></p> <p> </p> <h1><strong>Full Dataset & Project Page:</strong></h1> <p>The whole dataset is available on <a href="../records/14062964" target="_blank" rel="noopener">CVoiceFake Full</a> as well. Please kindly also refer to the project page: <a title="SafeEar Website" href="https://safeearweb.github.io/Project/" target="_blank" rel="noopener">SafeEar Website</a>.</p> <p> </p> <h1><strong>Citation:</strong></h1> <p>If you find our paper/code/benchmark helpful, please kindly consider citing this work with the following reference:</p> <pre><code>@inproceedings{li2024safeear,<br> author = {Li, Xinfeng and Li, Kai and Zheng, Yifan and Yan, Chen and Ji, Xiaoyu, and Xu, Wenyuan},<br> title = {{SafeEar: Content Privacy-Preserving Audio Deepfake Detection}},<br> booktitle = {Proceedings of the 2024 {ACM} {SIGSAC} Conference on Computer and Communications Security (CCS)}<br> year = {2024},<br>} </code></pre> <div> <div> </div> </div>
Path Following in Non-Visual Conditions - Screen shots and audio sample
<p>Screen shots and audio file in addition to the publication "Path Following in Non-Visual Conditions" by Alan Del Piccolo, Davide Rocchesso, and Stefano Papetti. Under revision for IEEE Transaction on Haptics (1 June 2018).</p> <p>Developed in Max (https://cycling74.com).</p> <p>Short description:</p> <ul> <li><em>interface.png</em>: the interface for managing the experiment. Output levels, trial repetitions and feedback conditions can be adjusted from here. The underlying Max patch is depicted in "main patch.png"</li> <li><em>main patch.png</em>: the main patch controlling the experiment. It receives the data from the Soundplane's controller (top left), invokes finger position detection and feedback generation (bottom left), manages the trial repetition and feedback modes (bottom center), and enables the adjustment of the feedback levels (right). </li> <li><em>mapToImage.png</em>: the patch that maps the participant's finger position on the Soundplane to the relative position over the loaded path shape. The position is shown by the white circle on the bottom left of the image.</li> <li><em>rolling_feedback.png</em>: the SDT "rolling model" configured for the use in the experiment. Note that the input levels are generated in the track_detection_SP2 patch.</li> <li><em>sdt_rolling.png</em>: a configuration of the SDT "rolling model" adjusted to output a signal similar to the one used in the experiment (see record.wav) as a stand-alone, namely without using the experiment's patches.</li> <li><em>track_detection_SP2.png</em>: the patch that manages the feedback generation (top left), the recording of execution time (bottom left), and the recording of position and force (center).</li> <li><em>record.wav</em>: a recording of the signal used in the experiment for both audio and vibrotactile feedback.</li> </ul>
Audio Commons Ground Truth Data for deliverables D4.4, D4.10 and D4.12
<p>This dataset contains the ground truth data used to evaluate the musical <strong>pitch</strong>, <strong>tempo</strong> and <strong>key </strong>estimation algorithms developed during the AudioCommons H2020 EU project and which are part of the <a href="https://www.audiocommons.org/2018/07/15/audio-commons-audio-extractor.html">Audio Commons Audio Extractor tool</a>. It also includes ground truth information for the <strong>single-event<em>ness</em> </strong>audio descriptor also developed for the same tool.</p> <p>This ground truth data has been used to generate the following documents:</p> <ul> <li><strong>Deliverable D4.4</strong>: Evaluation report on the first prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.10</strong>: Evaluation report on the second prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.12</strong>: Release of tool for the automatic semantic description of music samples</li> </ul> <p>All these documents are available in the <a href="https://www.audiocommons.org/materials/">materials section </a>of the AudioCommons website.</p> <p>All ground truth data in this repository is provided in the form of CSV files. Each CSV file corresponds to one of the individual datasets used in one or more evaluation tasks of the aforementioned deliverables. This repository <strong>does not include the audio files</strong> of each individual dataset, but includes references to the audio files. The following paragraphs describe the structure of the CSV files and give some notes about how to obtain the audio files in case these would be needed.</p> <p><br> <strong>Structure of the CSV files</strong></p> <p>All CSV files in this repository (with the sole exception of <em>SINGLE EVENT - Ground Truth.csv</em>) feature the following 5 columns:</p> <ol> <li><strong>Audio reference</strong>: reference to the corresponding audio file. This will either be a string withe the <strong>filename</strong>, or the <strong>Freesound ID </strong>(for one dataset based on Freesound content). See below for details about how to obtain those files. </li> <li><strong>Audio reference type</strong>: will be one of <em>Filename</em> or <em>Freesound ID</em>, and specifies how the previous column should be interpreted. </li> <li><strong>Key annotation</strong>: tonality information as a string with the form "RootNote minor/major". Audio files with no ground truth annotation for tonality are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> <li><strong>Tempo annotation</strong>: tempo information as an integer representing beats per minute. Audio files with no ground truth annotation for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10. Note that integer values are used here because we only have tempo annotations for <em>music loops</em> which typically only feature integer tempo values.</li> <li><strong>Pitch annotation</strong>: pitch information as an integer representing the MIDI note number corresponding to annotated pitch's frequency. Audio files with no ground truth pitch for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> </ol> <p>The remaining CSV file, <em>SINGLE EVENT - Ground Truth.csv</em>, has only the following 2 columns:</p> <ul> <li><strong>Freesound ID</strong>: sound ID used in Freesound to identify the audio clip.</li> <li><strong>Single Event: </strong>boolean indicating whether the corresponding sound is considered to be a single event or not. Single event annotations were collected by the authors of the deliverables as described in deliverable D4.10.</li> </ul> <p> </p> <p><strong>How to get the audio data</strong></p> <p>In this section we provide some notes about how to obtain the audio files corresponding to the ground truth annotations provided here. Note that due to licensing restrictions we are not allowed to re-distribute the audio data corresponding to most of these ground truth annotations.</p> <ul> <li><strong>Apple Loops (APPL)</strong>: This dataset includes some of the music loops included in Apple's music software such as Logic or GarageBand. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#appl">provided here</a>. </li> <li><strong>Carlos Vaquero Instruments Dataset (CVAQ)</strong>: This dataset includes single instrument recordings carried out by <a href="https://www.linkedin.com/in/carlosvaquero/">Carlos Vaquero</a> as part of this <a href="http://mtg.upf.edu/node/2609">master thesis</a>. Sounds are available as Freesound packs and can be downloaded at this page: https://freesound.org/people/Carlos_Vaquero/packs</li> <li><strong>Freesound Loops 4k (FSL4)</strong>: This dataset set includes a selection of music loops taken from Freesound. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#instructions-for-setting-up-datasets">provided here</a>.</li> <li><strong>Giant Steps Key Dataset (GSKY)</strong>: This dataset includes a selection of previews from Beatport annotated by key. Audio and original annotations <a href="https://github.com/GiantSteps/giantsteps-key-dataset">available here</a>.</li> <li><strong>Good-sounds Dataset (GSND)</strong>: This dataset contains monophonic recordings of instrument samples. Full description, original annotations and audio are <a href="https://zenodo.org/record/820937#.XEYMiy2ZN25">available here</a>.</li> <li><strong>University of IOWA Musical Instrument Samples (IOWA)</strong>: This dataset was created by the Electronic Music Studios of the University of IOWA and contains recordings of instrument samples. The dataset is available upon request by <a href="http://theremin.music.uiowa.edu/MIS.html">visiting this website</a>.</li> <li><strong>Mixcraft Loops (MIXL)</strong>: This dataset includes some of the music loops included in Acoustica's Mixcraft music software. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#mixl">provided here</a>.</li> <li><strong>NSynth Dataset Test and Validation sets (NSYT and NSYV)</strong>: NSynth is a large-scale and high-quality dataset of annotated musical notes built with synthesized sounds by Google's Magenta team. Full dataset description including original annotations and audio files is <a href="https://magenta.tensorflow.org/datasets/nsynth">available here</a>.</li> <li><strong>Philarmonia Orchestra Sound Samples Dataset (PHIL)</strong>: This includes thousands of free, downloadable sound samples specially recorded by Philharmonia Orchestra players. Audio files are freely downloadable from the <a href="http://www.philharmonia.co.uk/explore/sound_samples">philarmonia orchestra website</a>.</li> <li><strong>Freesound Single Events Dataset (SINGLE EVENT)</strong>: This includes a selection of Freesound audio clips representing audio signals containing either a single audio <em>event</em> or multiple ones. Original audio files can be retrieved by downloading individual audio clips from Freesound using the ID identifier provided in the CSV file. A similar procedure to that described <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#getting-fsl4-by-downloading-content-from-freesound">here</a> could be followed.</li> </ul>
Audio Commons Estimation Results Data for deliverables D4.4, D4.10 and D4.12
<p>This dataset contains the results of running the automatic audio annotation algorithms for <strong>pitch</strong>, <strong>tempo</strong> and <strong>key </strong>used for the evaluation of algorithms developed during the AudioCommons H2020 EU project and which are part of the <a href="https://www.audiocommons.org/2018/07/15/audio-commons-audio-extractor.html">Audio Commons Audio Extractor tool</a>. It also includes estimation results information for the <strong>single-event<em>ness</em> </strong>audio descriptor also developed for the same tool.</p> <p>These estimation results data has been used to generate the following documents:</p> <ul> <li><strong>Deliverable D4.4</strong>: Evaluation report on the first prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.10</strong>: Evaluation report on the second prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.12</strong>: Release of tool for the automatic semantic description of music samples</li> </ul> <p>All these documents are available in the <a href="https://www.audiocommons.org/materials/">materials section </a>of the AudioCommons website.</p> <p>All data in this repository is provided in the form of CSV files. Each CSV file corresponds to the analysis results of one musical task and one of the individual datasets used in the aforementioned deliverables. This repository <strong>does not include the audio files </strong>of each individual dataset, but includes references to the audio files. The following paragraphs describe the structure of the CSV files and give some notes about how to obtain the audio files in case these would be needed.</p> <p><br> <strong>Structure of the CSV files</strong></p> <p>All the CSV files in this repository (with the sole exception of <em>SINGLE EVENT - Estimation Results Truth.csv</em>) are named according to the following convention: "<em>DATASET_NAME</em> - <em>ESTIMATION_TASK</em> Estimation Results.csv". Therefore, estimation results for pitch, tempo and tonality music tasks are separated in different files. All these files share the same structure for the first 2 CSV columns:</p> <ol> <li><strong>Audio reference</strong>: reference to the corresponding audio file. This will either be a string withe the <strong>filename</strong>, or the <strong>Freesound ID </strong>(for one dataset based on Freesound content). See below for details about how to obtain those files. </li> <li><strong>Audio reference type</strong>: will be one of <em>Filename</em> or <em>Freesound ID</em>, and specifies how the previous column should be interpreted. </li> </ol> <p>The rest of the columns include the estimation results for each one of the algorithms included in the evaluation of each music facet. For <strong>each algorithms two columns</strong> are reserved, the first one containing the actual <strong>estimation</strong> and the second one the <strong>confidence</strong> of this estimation (see CSV file previews below). The format of actual estimations depends on the musical task, check the description of the <a href="https://zenodo.org/deposit/2545728">corresponding ground truth dataset</a> for more information on that. The confidence value is a float number, typically in the range from 0.0 to 1.0. It can happen that one or both columns are empty for a given analysis algorithm and CSV row. This will be the case if the algorithm could not successfully produce an estimation for the audio file row corresponding to the CSV row.</p> <p>The remaining CSV file, <em>SINGLE EVENT - Estimation Results.csv</em>, has the following 4 columns:</p> <ul> <li><strong>Freesound ID</strong>: sound ID used in Freesound to identify the audio clip.</li> <li><strong>ACExtractorV2</strong>: single-event<em>ness</em> estimation of the algorithm included in the second version of the Audio Commons Audio Extractor tool (bool).</li> <li><strong>ACExtractorV2-opt</strong>: single-event<em>ness</em> estimation of the algorithm included in the second version of the Audio Commons Audio Extractor tool with optimized parameters (bool).</li> <li><strong>ACExtractorV3</strong>: single-event<em>ness</em> estimation of the algorithm included in the third version of the Audio Commons Audio Extractor tool (bool).</li> </ul> <p> </p> <p><strong>How to get the audio data</strong></p> <p>In this section we provide some notes about how to obtain the audio files corresponding to the estimation results provided here. Note that due to licensing restrictions we are not allowed to re-distribute the audio data corresponding to most of these automatic annotations.</p> <ul> <li><strong>Apple Loops (APPL)</strong>: This dataset includes some of the music loops included in Apple's music software such as Logic or GarageBand. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#appl">provided here</a>.</li> <li><strong>Carlos Vaquero Instruments Dataset (CVAQ)</strong>: This dataset includes single instrument recordings carried out by <a href="https://www.linkedin.com/in/carlosvaquero/">Carlos Vaquero</a>as part of this <a href="http://mtg.upf.edu/node/2609">master thesis</a>. Sounds are available as Freesound packs and can be downloaded at this page: https://freesound.org/people/Carlos_Vaquero/packs</li> <li><strong>Freesound Loops 4k (FSL4)</strong>: This dataset set includes a selection of music loops taken from Freesound. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#instructions-for-setting-up-datasets">provided here</a>.</li> <li><strong>Giant Steps Key Dataset (GSKY)</strong>: This dataset includes a selection of previews from Beatport annotated by key. Audio and original annotations <a href="https://github.com/GiantSteps/giantsteps-key-dataset">available here</a>.</li> <li><strong>Good-sounds Dataset (GSND)</strong>: This dataset contains monophonic recordings of instrument samples. Full description, original annotations and audio are <a href="https://zenodo.org/record/820937#.XEYMiy2ZN25">available here</a>.</li> <li><strong>University of IOWA Musical Instrument Samples (IOWA)</strong>: This dataset was created by the Electronic Music Studios of the University of IOWA and contains recordings of instrument samples. The dataset is available upon request by <a href="http://theremin.music.uiowa.edu/MIS.html">visiting this website</a>.</li> <li><strong>Mixcraft Loops (MIXL)</strong>: This dataset includes some of the music loops included in Acoustica's Mixcraft music software. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#mixl">provided here</a>.</li> <li><strong>NSynth Dataset Test and Validation sets (NSYT and NSYV)</strong>: NSynth is a large-scale and high-quality dataset of annotated musical notes built with synthesized sounds by Google's Magenta team. Full dataset description including original annotations and audio files is <a href="https://magenta.tensorflow.org/datasets/nsynth">available here</a>.</li> <li><strong>Philarmonia Orchestra Sound Samples Dataset (PHIL)</strong>: This includes thousands of free, downloadable sound samples specially recorded by Philharmonia Orchestra players. Audio files are freely downloadable from the <a href="http://www.philharmonia.co.uk/explore/sound_samples">philarmonia orchestra website</a>.</li> <li><strong>Freesound Single Events Dataset (SINGLE EVENT)</strong>: This includes a selection of Freesound audio clips representing audio signals containing either a single audio <em>event</em>or multiple ones. Original audio files can be retrieved by downloading individual audio clips from Freesound using the ID identifier provided in the CSV file. A similar procedure to that described <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#getting-fsl4-by-downloading-content-from-freesound">here</a> could be followed.</li> </ul>
Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024
<h1><strong>Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024</strong></h1> <h3><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2"><em>Navigating the World of Wikidata for Research, Science and Cultural Heritage</em></a></h3> <h3><em><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht" target="_blank" rel="noopener">Wikidata's Twelfth Birthday: Workshop in Maastricht </a></em></h3> <p>Wikidata is a free, collaborative, multilingual database, collecting structured open data for anyone in the world to use. It also plays a crucial role in supporting Wikimedia projects, such as Wikipedia and Wikimedia Commons. Over the last 12 years it has strongly increased in popularity among the scientific and cultural heritage communities.</p> <p>In this 2,5 hours workshop you will learn the basics of working with Wikidata, both in theory and practice. You will learn</p> <ol> <li>The basics of Wikidata: A first look at what Wikidata is and how it works, both technically and socially (Wikidata community)</li> <li>How Wikidata can be relevant for research, science and cultural heritage (GLAM), and</li> <li>First steps in contributing to Wikidata yourself, with a focus on the topic of UM professors from past and present.</li> </ol> <p>As part of the <a title="Wikidata:Twelfth Birthday" href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday">Wikidata 12th Birthday celebrations</a> this workshop is open to academics, researchers, students, and professionals interested in working with Wikidata in the intersection of open data, research, and science. Whether you are new to Wikidata or looking to deepen your understanding, this session will provide valuable insights for improving your work.</p> <h2><strong>Workshop outline</strong></h2> <h3><strong>Part 1: Theory, Wikidata basics (45-60 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><em>1) Wikidata basics</em></p> <ul> <li>What is Wikidata?</li> <li>What are the principles of Wikidata?</li> <li>How are things described in Wikidata?</li> <li>Who builds Wikidata? - The Wikidata community</li> </ul> <p><em>2) Wikidata for research, science and cultural heritage</em></p> <ul> <li>To what extent is Wikidata used throughout science, research and GLAM?</li> <li>Six anecd<em>a</em>tic cases <ol> <li>Scientometrics - Scholia</li> <li>Life and biomedical sciences</li> <li>Astronomy</li> <li>Language technology / AI / LLMs</li> <li>GLAM – KB collection highlights</li> <li>Representation of (female) scientists</li> </ol> </li> </ul> <h3><strong>Break (15 minutes)</strong></h3> <h3><strong>Part 2: Practice, contributing to Wikidata (75-90 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm?download=1" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> <li><strong>Handout for participants in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.docx?download=1" rel="nofollow">Word</a> or <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.pdf?download=1" rel="nofollow">PDF</a></strong> <strong>are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><strong> </strong>The goals of this hands-on part are:</p> <ul> <li>Get familiar with basic data editing via the Wikidata interface</li> <li>Understand WD data models and structures related to professors (of Maastricht University)</li> <li>Extend existing <a title="Wikidata:Wiki-wetenschappers/Universiteit Maastricht/hoogleraren" href="https://www.wikidata.org/wiki/Wikidata:Wiki-wetenschappers/Universiteit_Maastricht/hoogleraren">Wikidata items about UM professors</a>, based on information in public sources.</li> <li>If time allows: Create new Wikidata items about UM professors</li> </ul> <p>The visual slides and the textual handout explain the same content, blocks and exercises, albeit in a slightly different order.<strong> </strong></p> <h2><strong>Required preparation</strong></h2> <p>To make optimal use of our time, participants must create a Wikidata account in the weeks before the workshop. See <a href="https://www.wikidata.org/w/index.php?title=Special:CreateAccount" target="_blank" rel="noopener">https://www.wikidata.org/w/index.php?title=Special:CreateAccount</a>.</p> <p>This is important because very fresh accounts may have limited editing rights. Furthermore only 6 Wikidata accounts can be created per day from UM IP addresses, so creating a lot of new accounts during the workshop might overstretch this limit.</p> <h2><strong>Workshop leader</strong></h2> <p>This workshop was given by <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Olaf Janssen</a>, the Wikimedia coordinator of the <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Koninklijke Bibliotheek</a>, the national library of the Netherlands.</p> <p>In this role he stimulates and facilitates collaboration between the collections, knowledge, open data and staff of the KB on the one hand, and the projects of the Wikimedia movement, such as Wikipedia, Wikimedia Commons and Wikidata on the other. He is also active as a volunteer within the community. Feel free to contact Olaf via olaf.janssen(at)<a href="http://kb.nl">kb.nl</a></p> <h2><strong>Materials on Wikimedia Commons</strong></h2> <p>Photos , videos and presentations related to this event can be found on Wikimedia Commons: <a title="c:Category:Wikidata Workshop at Maastricht University, 15 October 2024" href="https://commons.wikimedia.org/wiki/Category:Wikidata_Workshop_at_Maastricht_University,_15_October_2024">Category:Wikidata Workshop at Maastricht University, 15 October 2024</a></p> <h2>Relevant URLs </h2> <ul> <li><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht">https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht </a></li> <li><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2" target="_blank" rel="noopener">https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2</a> + <a href="https://web.archive.org/web/20240926154021/https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2/">archived version</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/">https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/</a></li> </ul> <p>Earlier LinkedIn posts (April-May 2024, before rescheduling the worlshop to October)</p> <ul> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/</a></li> </ul> <h3> </h3>
Pre-training Audio Embeddings
<p>Pre-trained audio embeddings (VGGish, OpenL3, YAMNet) of <a href="https://www.upf.edu/web/mtg/irmas">IRMAS</a> and <a href="https://zenodo.org/record/1432913">OpenMIC-2018</a> datasets, released in the following paper:</p> <p>Changhong Wang, Brian McFee, and Gaël Richard. "<strong>Transfer Learning and Bias Correction with Pre-trained Audio Embeddings</strong>". <em>Proceedings of the <a href="https://ismir2023.ismir.net/">International Society for Music Information Retrieval (ISMIR) Conference</a></em>, 2023.</p>
FOAMS: Processed Audio Files
<p>The processed audio files included in the Free Open-Access Misophonia Stimuli (FOAMS) project to curate a freely available database of sound stimuli intended for misophonia research.</p> <p>If you use this database, please credit it as follows:</p> <p>Orloff, D. M., Benesch, D., & Hansen, H. A. (2023). Curation of FOAMS: a Free Open-Access Misophonia Stimuli Database. <em>Journal of Open Psychology Data</em>, <em>11</em>(1).</p>
Guinea baboon vocalizations dataset automatically extracted with a deep neural network from natural audio recordings
<p><strong>Abstract</strong></p> <p>The data collection process consisted of continuously recording during one month a group of Guinea baboons living in semi-liberty at the CNRS primatology center in Rousset-sur-Arc (France). Two microphones we placed nearby their enclosure to continuously record the sounds produced by the group. A convolutional neural network (CNN) was used on these large and noisy audio recordings to automatically extract segments of sound containing a baboon vocal production by following the method of <a href="https://arxiv.org/abs/2302.07640">Bonafos et al. (2023)</a>. The resulting dataset consists of one-second to several-minute wav files of automatically detected vocalizations segments. The dataset thus provides a wide range of baboon vocalizations produced at all times of the day. It can be used to study vocal productions of non-human primates, their repertoire, their distribution over the day, their frequency, and their heterogeneity. In addition to the analysis of animal communication, the dataset can also be used as a learning base for sound classification models.</p> <p> </p> <p><strong>Data acquisition</strong></p> <p>The data are audio recordings of baboons. The recordings were made with a H6 Zoom recorder, using the included XYH-6 stereo microphone. The sample size is 44100 Hertz, 16 bits. The microphones were placed in the vicinity of the enclosure for one month and recorded continuously on a PC computer. A CNN passed over the data with a sliding window of 1 second and an overlap of 80% to detect the vocal productions of the baboons. The dataset consists of the segments predicted by the CNN to contain a baboon vocalization. Windows containing signal less than one second apart were merged into a single vocalization.</p> <p> </p> <p><strong>Data source location</strong></p> <ul> <li>Institution: CNRS, Primate Facility</li> <li> <p>City/Town/Region: Rousset-sur-Arc</p> </li> <li> <p>Country: France</p> </li> <li> <p>Latitude and longitude for collected samples/data: 43.47033535251509, 5.6514732876668905</p> </li> </ul> <p> </p> <p><strong>Value of the data</strong></p> <ul> <li> <p>This dataset is relatively unique in terms of the quantity of vocalizations available.</p> </li> <li> <p>This massive dataset can be very useful to two types of scientific communities: experts in primatology who study the vocal productions of non-human primates, and experts in data science and audio signal processing.</p> </li> <li> <p>The machine learning research community has at its disposal a database of several dozen hours of animal vocalizations, which will make it possible to build up a large learning base, very useful for Environemental Sound Recognition tasks, for example.</p> </li> </ul> <p> </p> <p><strong>Objective</strong></p> <p>This dataset is a follow-up of two studies on the vocal productions of Guinea baboons (Papio papio) in which we carried out analyses of their vocal productions on the basis of a relatively large vocalization sample containing around 1300 vocalizations (<a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0169321">Boë, Berthommier, Legou, Captier, Kemp, Sawallis, Becker, Rey, & Fagot, 2017</a>; <a href="https://hal.science/hal-01649539">Kemp, Rey, Legou, Boë, Berthommier, Becker, & Fagot, 2017</a>). The aim was to collect a larger database using the technique of deep convolutional neural networks in order to 1) automatically detect vocal productions in a large continuous audio recording and 2) perform a categorization of these vocalizations on a more massive sample. A description of the pipeline that enabled these automatic detections and categorizations is given in <a href="https://arxiv.org/abs/2302.07640">Bonafos, Pudlo, Freyermuth, Legou, Fagot, Tronçon, & Rey (2023)</a>.</p> <p> </p> <p><strong>Data description</strong></p> <p>The data is a set of audio files in wav format. They are at least one second long (the size of the window), up to several minutes, if several windows are consecutively predicted as containing signal. Moreover, we add the labeled data we used to train the CNN which did the prediction. We also provide two hours of the continuous recordings to have an idea of the continuous recordings and test the code of the paper provided on <a href="https://gitlab.com/papers4375727/detection-and-classification-of-vocal-productions">gitlab</a>.</p> <p>In addition, there is a database in csv format listing all the vocalizations, the day and time of their production, and the prediction probabilities of the model.</p> <p> </p> <p><strong>Experimental design, materials and methods</strong></p> <p>The original recordings represent one month of continuous audio recording. Seven hours of this month were manually labelled. They were segmented and labelled according to whether or not there was a monkey vocalization (i.e., noise or vocalization) and, if there was a vocalization, according to the type of vocalization (6 possible classes: bark, copulation grunt, grunt, scream, yak, wahoo). These manually labelled data were used as a training set for a CNN, which was automatically trained following the pipeline of Bonafos et al. (2023). This model was then used to automatically detect and classify vocalization during the whole month of audio recording. It processes the data in the same way when predicting new data as it does when training. It uses a sliding window of one second with an overlap of 80%. It does not take into account information from previous predictions, but calculates the probability of a vocalization in each one-second window independently. It then iterates through the month. For each window, the model predicts two outputs: the probability that there is a vocalization and the probability of each class of vocalization.</p> <p>For the purpose of generating the wav files, if a window has a probability of a vocalization greater than 0.5, it is considered to contain a vocalization. If it is the first one, a vocalization is started at that moment. If the time windows that follow a vocalization also contain a vocalization, then the signal they contain is added to the first segment for which a vocalization has been detected. As soon as a one-second segment no longer contains a signal corresponding to a vocalization, the wav file is closed. If windows are predicted to contain no vocalizations, but are between two windows that contain vocalizations within 1 second of each other, then all windows are merged.</p>
An Open-set Recognition and Few-Shot Learning Dataset for Audio Event Classification in Domestic Environments
<p>The problem of training a deep neural network with a small set of positive samples is known as few-shot learning (FSL). It is widely known that traditional deep learning (DL) algorithms usually show very good performance when trained with large datasets. However, in many applications, it is not possible to obtain such a high number of samples. In the image domain, typical FSL applications are those related to face recognition. In the audio domain, music fraud or speaker recognition can be clearly benefited from FSL methods. This paper deals with the application of FSL to the detection of specific and intentional acoustic events given by different types of sound alarms, such as door bells or fire alarms, using a limited number of samples. These sounds typically occur in domestic environments where many events corresponding to a wide variety of sound classes take place. Therefore, the detection of such alarms in a practical scenario can be considered an open-set recognition (OSR) problem. To address the lack of a dedicated public dataset for audio FSL, researchers usually make modifications on other available datasets. This paper is aimed at providing the audio recognition community with a carefully annotated dataset for FSL and OSR comprised of 1360 clips from 34 classes divided into pattern sounds and unwanted sounds. To facilitate and promote research in this area, results with two baseline systems (one trained from scratch and another based on transfer learning), are presented.</p> <p> </p>
OrchideaSOL: an audio dataset of isolated musical notes, including mutes and extended playing techniques
<p>OrchideaSOL<br> ==========<br> Version 2.0, April 2020.<br> </p> <p> </p> <p>Created By<br> --------------</p> <p>Carmine-Emanuele Cella (1), Daniele Ghisi (1), Vincent Lostanlen (2), Fabien Lévy (3), Joshua Fineberg (4), Yan Maresz (5)<br> <br> (1): UC Berkeley<br> (2): New York University<br> (3): Columbia University<br> (4): Boston University<br> (5): Conservatoire de Paris</p> <p> </p> <p>Description<br> ---------------</p> <p><br> OrchideaSOL is a dataset of 13265 samples, each containing a single musical note from one of 14 different instruments:</p> <ol> <li>Bass Tuba</li> <li>French Horn</li> <li>Trombone</li> <li>Trumpet in C</li> <li>Accordion</li> <li>Contrabass</li> <li>Violin</li> <li>Viola</li> <li>Violoncello</li> <li>Bassoon</li> <li>Clarinet in B-flat</li> <li>Flute</li> <li>Oboe</li> <li>Alto Saxophone</li> </ol> <p> </p> <p>These sounds were originally recorded at Ircam in Paris (France) between 1996 and 1999, as part of a larger project named Studio On Line (SOL). One asset of OrchideaSOL is that it contains many combinations of mutes and extended playing techniques.<br> <br> The OrchideaSOL audio data can be used for creative purposes insofar at the use complies with the Ircam Forum License. Please visit: https://forum.ircam.fr/legal/contrat-de-licence-forum-ircam/</p> <p><br> The OrchideaSOL metadata can be used for creative purposes insofar at the use complies with the Creative Commons Attribution 4.0 International license (see below).<br> <br> OrchideaSOL can be used for education and research purposes. In particular, it can be employed as a dataset for training and/or evaluating music information retrieval (MIR) systems, for tasks such as instrument recognition, playing technique recognition, or fundamental frequency estimation. For this purpose, we provide an official 5-fold split of OrchideaSOL. This split has been carefully balanced in terms of instrumentation, pitch range, and dynamics. For the sake of research reproducibility, we encourage users of OrchideaSOL to adopt this split and report their results in terms of average performance across folds.</p> <p> </p> <p>Data Files<br> --------------</p> <p>OrchideaSOL contains 13265 audio clips as WAV files, sampled at 44.1 kHz, with a single channel (mono), at a bit depth of 16. This is equivalent to the audio quality of a compact disc. Audio clips vary in duration between two and ten seconds.</p> <p>Every audio file has a file path of the form:<br> <FAMILY>/<INSTRUMENT><+MUTE>/<TECHNIQUE>/<INSTR><+M>-<TECH>-<PITCH>-<DYN>-<INSTANCE>-<MISC>.wav</p> <p><br> where:</p> <ul> <li><FAMILY> corresponds to the instrument family: "Brass", "Keyboards" (includes accordion), "Strings", and "Winds" (i.e., woodwinds).</li> <li><INSTRUMENT> is the full name of the instrument.</li> <li><+MUTE> is the type of mute being used, such as "wah", "harmon", "piombo", or "sordina". If there is no mute, this field is absent.</li> <li><TECHNIQUE> is the type of playing technique.</li> <li><INSTR> is the abbreviation of the instrument.</li> <li><+M> is the abbreviation of the type of mute, if applicable.</li> <li><TECH> is the abbreviation of playing technique.</li> <li><PITCH> denotes the pitch of the musical note. This pitch is encoded in the American standard pitch notation: pitch class (C means "do") followed by pitch octave. According to this convention, A4 has a fundamental frequency of 440 Hz.</li> <li><DYN> denotes the intensity dynamics, ranked from pp (pianissimo) to ff (fortissimo).</li> <li><INSTANCE> contains additional information, when applicable. For example, for bowed string instruments, the same pitch may sometimes be achieved on different positions and different strings, resulting in small timbre differences. In this case the label "1c", "2c", "3c", or "4c" denotes the string which is being bowed. (The letter c originates from the word "corde", which means string in French.) By convention, the first string is the one with the highest pitch when played as an open string. Furthermore, on some wind instruments, the same note was played multiple times, e.g. at multiple durations. In this case, we use the label "alt1", "alt2", etc. to denote alternative instances of the note. If none of these tags apply, the <INSTANCE> field becomes "N", which stands for "Not Applicable".</li> <li><MISC> contains additional information, if applicable. In OrchideaSOL, some pitches were never recorded, and thus missing from the chromatic scale. In this case, the <MISC> tag contains a letter "R", to denote the fact that the corresponding WAV file has been obtained by transforming a different audio clip via some digital frequency transposition (similar to Auto-Tune). The letter "R" stands for "resampled". Furthermore, some pitches were slightly out of tune in comparison with the A440 tuning standard. Again, we applied some digital frequency transposition to correct them and put them exactly in tune. The amount of frequency transposition is measured in "cents" of an equal-tempered semitone. The letter "T" stands for "tuned". Because we employed a high-fidelity algorithm for frequency transposition, and because the amount of digital frequency transposition is small, the timbre of pitch-corrected notes remains faithful to the instrument. If none of these tags apply, the <MISC> field becomes "N", which stands for "natural"; in this case, the note is distributed exactly as it was recorded in the studio.</li> </ul> <p>For example, "Strings/Violin+sordina/tremolo/Vn+S-trem-A4-mf-4c-T13d_R200d.wav" corresponds to:</p> <ul> <li>a violin sound;</li> <li>equipped with a sordina mute;</li> <li>played in the tremolo playing technique;</li> <li>at pitch A4 (440 Hz);</li> <li>with mezzoforte dynamics;</li> <li>on the fourth string (i.e. the lowest);</li> <li>resampled from a B4 by lowering pitch by a semitone, i.e. 100 cents (R100d)</li> <li>lowered by 13 cents (T22d) to match the A440 tuning standard.</li> </ul> <p> </p> <p>The audio data for OrchideaSOL is not directly downloadable on Zenodo. Rather, it can be downloaded for free after registering to the Ircam forum. Please visit: https://forum.ircam.fr/</p> <p> </p> <p>Metadata File<br> -------------------</p> <p>The OrchideaSOL_metadata.csv file contains 13265 rows, one for each audio clip. It can be opened by a text editor or by a spreadsheet software application. It contains 13 columns:</p> <ol> <li>Path to the WAV file, in UNIX filesystem format. For Windows compatibility, replace the slashes ("/") by backslashes ("\"). Ex: "Strings/Violin+sordina/tremolo/Vn+S-trem-A4-mf-4c-T13d_R200d.wav"</li> <li>Fold ID. Either equal to 0, 1, 2, 3, or 4.</li> <li>Family. Ex: "Brass"</li> <li>Instrument abbreviation. Ex: "BTb"</li> <li>Instrument name in full. Ex: "Bass Tuba"</li> <li>Technique abbreviation.</li> <li>Technique name in full.</li> <li>Pitch. Ex: "A#1"</li> <li>Pitch ID in MIDI format. Ex: 34. Integer in the range 0-127.</li> <li>Dynamics. Ex: "ff".</li> <li>Dynamics ID. Integer. pp maps to 0 and ff maps to 4. The higher, the louder.</li> <li>Instance ID. Integer in the range 0-4</li> <li>String ID. Equal to 1, 2, 3, 4, or empty if not applicable.</li> <li>"Needed digital retuning". TRUE if the file has been pitch-shifted with digital audio effects; FALSE otherwise.</li> </ol> <p> </p> <p>Conditions of Use<br> ------------------------</p> <p>OrchideaSOL was created in 2020 by Carmine-Emanuele Cella, Daniele Ghisi, Vincent Lostanlen, Fabien Lévy, Joshua Fineberg, and Yan Maresz.</p> <p>OrchideaSOL is a derivative of SOL. We wish to thank Hugues Vinet, Greg Beller, and all coordinators of the Ircam Forum for their authorization to upload the metadata of OrchideaSOL to Zenodo.</p> <p>The audio samples in OrchideaSOL are offered free of charge under the Ircam Forum License. Please visit: https://forum.ircam.fr/legal/contrat-de-licence-forum-ircam/</p> <p>The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, the authors are not liable for, and expressly exclude all liability for, loss or damage however and whenever caused to anyone by any use of the OrchideaSOL dataset or any part of it.</p> <p> </p> <p>Versions<br> -----------<br> 1.0 was released on February 24th, 2020.<br> 2.0 was released on April 4th, 2020. It fixes a bug in the instance IDs of oboe sounds in the "blow without reed" technique.</p> <p> </p> <p>Feedback<br> -------------</p> <p>Please help us improve OrchideaSOL by sending your feedback to:<br> carmine.cella@berkeley.edu</p> <p>For issues regarding the metadata encoding, the five-fold split, or the OrchideaSOL module in mirdata, please write to:<br> vincent.lostanlen@nyu.edu</p> <p>In case of a problem, please include as many details as possible.</p>
A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)
<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie "Forrest Gump" (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.