Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “audio data”
Audio and Water Movement Data for Oyster Reef and Mudflat Sites on the Coast of Virginia, 2018
Paired marine audio soundscape recordings with concurrent acoustic Doppler velocimeter (ADV) turbulence measurements at three intertidal sites: a natural oyster reef, a restored oyster reef, and a bare mudflat. The objective was to establish the link between oyster reef soundscapes and hydrodynamics, specifically the role that turbulence plays in generating near field pressure waves that may be used as physical cues by oyster larvae when initiating settlement behaviors. Data were collected at three different locations, with water turbulance measurements at rates up to 25Hz concurrent with audio recording.
Transcribing audio data: overview and transcripts of several automatic transcription tools
<p>Throughout institutions, audio recordings are being made regularly. To be able to further process these recordings, the audio often needs to be transcribed. In order to avoid having to transcribe the audio manually, there is a wealth of tools available for doing so automatically. In this record, we present an overview of several often-used tools to automatically transcribe pre-recorded audio data, including their features, costs, and security.</p> <p>To check the quality of the tool, we also recorded an audio fragment in Dutch that we ran through all tools in this overview in March of 2022. This original audio fragment (Test_interview_20220203.mp3), the cleaned-up transcription (Test_interview_cleaned_transcript.odt) and each tool’s raw transcript of the audio fragment (Test_interview_[name-tool]_raw_[date-run]) are included in this record as well. The raw transcripts were downloaded as .docx or .txt files and the .docx files saved as .odt. No edits to the transcripts were made before saving them, except an incidental removal of a personal email address or hyperlink.</p> <p>The overview contains information and transcripts of following transcription tools:</p> <ul> <li>Amberscript</li> <li>HappyScribe</li> <li>Kaldi</li> <li>NVIVO transcription</li> <li>Sonix</li> <li>SpokenOnline</li> <li>Transcribe</li> <li>Trint</li> <li>Microsoft Word 365 Online</li> </ul> <p><strong>About</strong></p> <p>This overview was created through a collaboration between Utrecht University’s Research Data Management (RDM) Support and the <a href="https://datahub.sites.uu.nl/">DataHub SSH</a> programme situated at the faculty of Humanities.</p> <p>The details in the overview have last been updated April 19, 2022. Please note that at the time you are downloading these files, the quality of the (Dutch) speech-to-text conversion may have been improved by the respective supplier.</p>
Audio Commons Ground Truth Data for deliverables D4.4, D4.10 and D4.12
<p>This dataset contains the ground truth data used to evaluate the musical <strong>pitch</strong>, <strong>tempo</strong> and <strong>key </strong>estimation algorithms developed during the AudioCommons H2020 EU project and which are part of the <a href="https://www.audiocommons.org/2018/07/15/audio-commons-audio-extractor.html">Audio Commons Audio Extractor tool</a>. It also includes ground truth information for the <strong>single-event<em>ness</em> </strong>audio descriptor also developed for the same tool.</p> <p>This ground truth data has been used to generate the following documents:</p> <ul> <li><strong>Deliverable D4.4</strong>: Evaluation report on the first prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.10</strong>: Evaluation report on the second prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.12</strong>: Release of tool for the automatic semantic description of music samples</li> </ul> <p>All these documents are available in the <a href="https://www.audiocommons.org/materials/">materials section </a>of the AudioCommons website.</p> <p>All ground truth data in this repository is provided in the form of CSV files. Each CSV file corresponds to one of the individual datasets used in one or more evaluation tasks of the aforementioned deliverables. This repository <strong>does not include the audio files</strong> of each individual dataset, but includes references to the audio files. The following paragraphs describe the structure of the CSV files and give some notes about how to obtain the audio files in case these would be needed.</p> <p><br> <strong>Structure of the CSV files</strong></p> <p>All CSV files in this repository (with the sole exception of <em>SINGLE EVENT - Ground Truth.csv</em>) feature the following 5 columns:</p> <ol> <li><strong>Audio reference</strong>: reference to the corresponding audio file. This will either be a string withe the <strong>filename</strong>, or the <strong>Freesound ID </strong>(for one dataset based on Freesound content). See below for details about how to obtain those files. </li> <li><strong>Audio reference type</strong>: will be one of <em>Filename</em> or <em>Freesound ID</em>, and specifies how the previous column should be interpreted. </li> <li><strong>Key annotation</strong>: tonality information as a string with the form "RootNote minor/major". Audio files with no ground truth annotation for tonality are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> <li><strong>Tempo annotation</strong>: tempo information as an integer representing beats per minute. Audio files with no ground truth annotation for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10. Note that integer values are used here because we only have tempo annotations for <em>music loops</em> which typically only feature integer tempo values.</li> <li><strong>Pitch annotation</strong>: pitch information as an integer representing the MIDI note number corresponding to annotated pitch's frequency. Audio files with no ground truth pitch for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> </ol> <p>The remaining CSV file, <em>SINGLE EVENT - Ground Truth.csv</em>, has only the following 2 columns:</p> <ul> <li><strong>Freesound ID</strong>: sound ID used in Freesound to identify the audio clip.</li> <li><strong>Single Event: </strong>boolean indicating whether the corresponding sound is considered to be a single event or not. Single event annotations were collected by the authors of the deliverables as described in deliverable D4.10.</li> </ul> <p> </p> <p><strong>How to get the audio data</strong></p> <p>In this section we provide some notes about how to obtain the audio files corresponding to the ground truth annotations provided here. Note that due to licensing restrictions we are not allowed to re-distribute the audio data corresponding to most of these ground truth annotations.</p> <ul> <li><strong>Apple Loops (APPL)</strong>: This dataset includes some of the music loops included in Apple's music software such as Logic or GarageBand. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#appl">provided here</a>. </li> <li><strong>Carlos Vaquero Instruments Dataset (CVAQ)</strong>: This dataset includes single instrument recordings carried out by <a href="https://www.linkedin.com/in/carlosvaquero/">Carlos Vaquero</a> as part of this <a href="http://mtg.upf.edu/node/2609">master thesis</a>. Sounds are available as Freesound packs and can be downloaded at this page: https://freesound.org/people/Carlos_Vaquero/packs</li> <li><strong>Freesound Loops 4k (FSL4)</strong>: This dataset set includes a selection of music loops taken from Freesound. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#instructions-for-setting-up-datasets">provided here</a>.</li> <li><strong>Giant Steps Key Dataset (GSKY)</strong>: This dataset includes a selection of previews from Beatport annotated by key. Audio and original annotations <a href="https://github.com/GiantSteps/giantsteps-key-dataset">available here</a>.</li> <li><strong>Good-sounds Dataset (GSND)</strong>: This dataset contains monophonic recordings of instrument samples. Full description, original annotations and audio are <a href="https://zenodo.org/record/820937#.XEYMiy2ZN25">available here</a>.</li> <li><strong>University of IOWA Musical Instrument Samples (IOWA)</strong>: This dataset was created by the Electronic Music Studios of the University of IOWA and contains recordings of instrument samples. The dataset is available upon request by <a href="http://theremin.music.uiowa.edu/MIS.html">visiting this website</a>.</li> <li><strong>Mixcraft Loops (MIXL)</strong>: This dataset includes some of the music loops included in Acoustica's Mixcraft music software. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#mixl">provided here</a>.</li> <li><strong>NSynth Dataset Test and Validation sets (NSYT and NSYV)</strong>: NSynth is a large-scale and high-quality dataset of annotated musical notes built with synthesized sounds by Google's Magenta team. Full dataset description including original annotations and audio files is <a href="https://magenta.tensorflow.org/datasets/nsynth">available here</a>.</li> <li><strong>Philarmonia Orchestra Sound Samples Dataset (PHIL)</strong>: This includes thousands of free, downloadable sound samples specially recorded by Philharmonia Orchestra players. Audio files are freely downloadable from the <a href="http://www.philharmonia.co.uk/explore/sound_samples">philarmonia orchestra website</a>.</li> <li><strong>Freesound Single Events Dataset (SINGLE EVENT)</strong>: This includes a selection of Freesound audio clips representing audio signals containing either a single audio <em>event</em> or multiple ones. Original audio files can be retrieved by downloading individual audio clips from Freesound using the ID identifier provided in the CSV file. A similar procedure to that described <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#getting-fsl4-by-downloading-content-from-freesound">here</a> could be followed.</li> </ul>
Audio Commons Estimation Results Data for deliverables D4.4, D4.10 and D4.12
<p>This dataset contains the results of running the automatic audio annotation algorithms for <strong>pitch</strong>, <strong>tempo</strong> and <strong>key </strong>used for the evaluation of algorithms developed during the AudioCommons H2020 EU project and which are part of the <a href="https://www.audiocommons.org/2018/07/15/audio-commons-audio-extractor.html">Audio Commons Audio Extractor tool</a>. It also includes estimation results information for the <strong>single-event<em>ness</em> </strong>audio descriptor also developed for the same tool.</p> <p>These estimation results data has been used to generate the following documents:</p> <ul> <li><strong>Deliverable D4.4</strong>: Evaluation report on the first prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.10</strong>: Evaluation report on the second prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.12</strong>: Release of tool for the automatic semantic description of music samples</li> </ul> <p>All these documents are available in the <a href="https://www.audiocommons.org/materials/">materials section </a>of the AudioCommons website.</p> <p>All data in this repository is provided in the form of CSV files. Each CSV file corresponds to the analysis results of one musical task and one of the individual datasets used in the aforementioned deliverables. This repository <strong>does not include the audio files </strong>of each individual dataset, but includes references to the audio files. The following paragraphs describe the structure of the CSV files and give some notes about how to obtain the audio files in case these would be needed.</p> <p><br> <strong>Structure of the CSV files</strong></p> <p>All the CSV files in this repository (with the sole exception of <em>SINGLE EVENT - Estimation Results Truth.csv</em>) are named according to the following convention: "<em>DATASET_NAME</em> - <em>ESTIMATION_TASK</em> Estimation Results.csv". Therefore, estimation results for pitch, tempo and tonality music tasks are separated in different files. All these files share the same structure for the first 2 CSV columns:</p> <ol> <li><strong>Audio reference</strong>: reference to the corresponding audio file. This will either be a string withe the <strong>filename</strong>, or the <strong>Freesound ID </strong>(for one dataset based on Freesound content). See below for details about how to obtain those files. </li> <li><strong>Audio reference type</strong>: will be one of <em>Filename</em> or <em>Freesound ID</em>, and specifies how the previous column should be interpreted. </li> </ol> <p>The rest of the columns include the estimation results for each one of the algorithms included in the evaluation of each music facet. For <strong>each algorithms two columns</strong> are reserved, the first one containing the actual <strong>estimation</strong> and the second one the <strong>confidence</strong> of this estimation (see CSV file previews below). The format of actual estimations depends on the musical task, check the description of the <a href="https://zenodo.org/deposit/2545728">corresponding ground truth dataset</a> for more information on that. The confidence value is a float number, typically in the range from 0.0 to 1.0. It can happen that one or both columns are empty for a given analysis algorithm and CSV row. This will be the case if the algorithm could not successfully produce an estimation for the audio file row corresponding to the CSV row.</p> <p>The remaining CSV file, <em>SINGLE EVENT - Estimation Results.csv</em>, has the following 4 columns:</p> <ul> <li><strong>Freesound ID</strong>: sound ID used in Freesound to identify the audio clip.</li> <li><strong>ACExtractorV2</strong>: single-event<em>ness</em> estimation of the algorithm included in the second version of the Audio Commons Audio Extractor tool (bool).</li> <li><strong>ACExtractorV2-opt</strong>: single-event<em>ness</em> estimation of the algorithm included in the second version of the Audio Commons Audio Extractor tool with optimized parameters (bool).</li> <li><strong>ACExtractorV3</strong>: single-event<em>ness</em> estimation of the algorithm included in the third version of the Audio Commons Audio Extractor tool (bool).</li> </ul> <p> </p> <p><strong>How to get the audio data</strong></p> <p>In this section we provide some notes about how to obtain the audio files corresponding to the estimation results provided here. Note that due to licensing restrictions we are not allowed to re-distribute the audio data corresponding to most of these automatic annotations.</p> <ul> <li><strong>Apple Loops (APPL)</strong>: This dataset includes some of the music loops included in Apple's music software such as Logic or GarageBand. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#appl">provided here</a>.</li> <li><strong>Carlos Vaquero Instruments Dataset (CVAQ)</strong>: This dataset includes single instrument recordings carried out by <a href="https://www.linkedin.com/in/carlosvaquero/">Carlos Vaquero</a>as part of this <a href="http://mtg.upf.edu/node/2609">master thesis</a>. Sounds are available as Freesound packs and can be downloaded at this page: https://freesound.org/people/Carlos_Vaquero/packs</li> <li><strong>Freesound Loops 4k (FSL4)</strong>: This dataset set includes a selection of music loops taken from Freesound. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#instructions-for-setting-up-datasets">provided here</a>.</li> <li><strong>Giant Steps Key Dataset (GSKY)</strong>: This dataset includes a selection of previews from Beatport annotated by key. Audio and original annotations <a href="https://github.com/GiantSteps/giantsteps-key-dataset">available here</a>.</li> <li><strong>Good-sounds Dataset (GSND)</strong>: This dataset contains monophonic recordings of instrument samples. Full description, original annotations and audio are <a href="https://zenodo.org/record/820937#.XEYMiy2ZN25">available here</a>.</li> <li><strong>University of IOWA Musical Instrument Samples (IOWA)</strong>: This dataset was created by the Electronic Music Studios of the University of IOWA and contains recordings of instrument samples. The dataset is available upon request by <a href="http://theremin.music.uiowa.edu/MIS.html">visiting this website</a>.</li> <li><strong>Mixcraft Loops (MIXL)</strong>: This dataset includes some of the music loops included in Acoustica's Mixcraft music software. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#mixl">provided here</a>.</li> <li><strong>NSynth Dataset Test and Validation sets (NSYT and NSYV)</strong>: NSynth is a large-scale and high-quality dataset of annotated musical notes built with synthesized sounds by Google's Magenta team. Full dataset description including original annotations and audio files is <a href="https://magenta.tensorflow.org/datasets/nsynth">available here</a>.</li> <li><strong>Philarmonia Orchestra Sound Samples Dataset (PHIL)</strong>: This includes thousands of free, downloadable sound samples specially recorded by Philharmonia Orchestra players. Audio files are freely downloadable from the <a href="http://www.philharmonia.co.uk/explore/sound_samples">philarmonia orchestra website</a>.</li> <li><strong>Freesound Single Events Dataset (SINGLE EVENT)</strong>: This includes a selection of Freesound audio clips representing audio signals containing either a single audio <em>event</em>or multiple ones. Original audio files can be retrieved by downloading individual audio clips from Freesound using the ID identifier provided in the CSV file. A similar procedure to that described <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#getting-fsl4-by-downloading-content-from-freesound">here</a> could be followed.</li> </ul>
WaveFake: A data set to facilitate audio DeepFake detection
<p>The main purpose of this data set is to facilitate research into audio DeepFakes. We hope that this work helps in finding new detection methods to prevent such attempts. These generated media files have been increasingly used to commit <a href="https://www.vice.com/en/article/pkyqvb/deepfake-audio-impersonating-ceo-fraud-attempt">impersonation attempts</a> or <a href="https://www.wired.com/story/telegram-still-hasnt-removed-an-ai-bot-thats-abusing-women/">online harassment</a>. You can find the accompanying code repository on <a href="https://github.com/RUB-SysSec/WaveFake">GitHub</a>.</p> <p>The data set consists of 104,885 generated audio clips (16-bit PCM wav). We examine multiple networks trained on two reference data sets. First, the <a href="https://keithito.com/LJ-Speech-Dataset/">LJSpeech</a> data set consisting of 13,100 short audio clips (on average 6 seconds each; roughly 24 hours total) read by a female speaker. It features passages from 7 non-fiction books and the audio was recorded on a MacBook Pro microphone. Second, we include samples based on the <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a> data set, specifically, basic5000 corpus. This corpus consists of 5,000 sentences covering all basic kanji of the Japanese language (4.8 seconds on average; roughly 6.7 hours total). The recordings were performed by a female native Japanese speaker in an anechoic room. Finally, we include samples from a full text-to-speech pipeline (16,283 phrases; 3.8s on average; roughly 17.5 hours total). Thus, our data set consists of approximately 175 hours of generated audio files in total. Note that we do not redistribute the reference data.</p> <p>We included a range of architectures in our data set:</p> <ul> <li><a href="https://arxiv.org/abs/1910.06711">MelGAN</a></li> <li><a href="https://arxiv.org/abs/1910.11480">Parallel WaveGAN</a></li> <li><a href="https://arxiv.org/abs/2005.05106">Multi-Band MelGAN</a></li> <li><a href="http://arxiv.org/abs/2005.05106">Full-Band MelGAN</a></li> <li><a href="https://arxiv.org/abs/2010.05646">HiFi-GAN</a></li> <li><a href="https://arxiv.org/abs/1811.00002">WaveGlow</a></li> </ul> <p>Additionally, we examined a bigger version of MelGAN and include samples from a full TTS-pipeline consisting of a conformer and parallel WaveGAN model.</p> <p><strong>Collection Process</strong></p> <p>For WaveGlow, we utilize the <a href="https://github.com/NVIDIA/waveglow">official implementation</a> (commit 8afb643) in conjunction with the official pre-trained network on <a href="https://pytorch.org/hub/nvidia_deeplearningexamples_waveglow/">PyTorch Hub</a>. We use a popular implementation available on <a href="https://github.com/kan-bayashi/ParallelWaveGAN">GitHub</a> (commit 12c677e) for the remaining networks. The repository also offers pre-trained models. We used the pre-trained networks to generate samples that are similar to their respective training distributions, <a href="https://keithito.com/LJ-Speech-Dataset/">LJ Speech</a> and <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a>. When sampling the data set, we first extract Mel spectrograms from the original audio files, using the pre-processing scripts of the corresponding repositories. We then feed these Mel spectrograms to the respective models to obtain the data set. For sampling the full TTS results, we use the <a href="https://github.com/espnet/espnet">ESPnet</a> project. To make sure the generated phrases do not overlap with the training set, we downloaded the <a href="https://commonvoice.mozilla.org/en/datasets">common voices data set</a> and extracted 16.285 phrases from it.</p> <p>This data set is licensed with a CC-BY-SA 4.0 license.</p> <p>This work was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany's Excellence Strategy -- EXC-2092 CaSa -- 390781972.</p>
Saarland Music Data: MIDI-Audio Piano Music
<p>This is an improved version of the dataset originally referred to as <strong>SMD MIDI-Audio Piano Music.</strong> For more details, please visit the website: <a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi">https://www.audiolabs-erlangen.de/resources/MIR/SMD/midi</a></p> <p>Saarland Music Data provides audio recordings along with perfectly synchronized MIDI files for various piano pieces. The pieces were performed by students of the <a href="http://www.hfm.saarland.de">Hochschule für Musik Saar</a> on a hybrid acoustic/digital piano <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a>. The Disklavier allows for capturing key and pedal movements of the piano while playing. This information, which can be stored in a MIDI file, yields an accurate annotation of the corresponding audio recording in form of a symbolic description of all played musical note events. The SMD MIDI-Audio pairs constitute a valuable dataset for various music analysis tasks such as music transcription, performance analysis, music synchronization, audio alignment, or source separation. </p> <p>All performances were recorded in the studios of the <a href="http://www.hfm.saarland.de">Hochschule für Musik Saar</a>, played by students of piano classes of different levels, on a <a href="http://www.yamaha.com/Products/Disklavier.html">Yamaha Disklavier</a> model <a href="http://www.yamaha.com/yamahavgn/CDA/ContentDetail/ModelSeriesDetail.html?CNTID=556850&CNTYP=PRODUCT">DCFIIISM4PRO</a>. Using two cardioid-condenser microphones fixed over the resonating body of the piano, all performances were directly recorded into Steinberg Cubase 4. Except for trimming the beginnings and ends of the recordings, no further post-processing (filters, effects) was applied to the musical material. From each Cubase project, an audio file (44.1 kHz, stereo) as well as a synchronized standard MIDI file (SMF) were exported. Besides these files, we also provide the audio files as WAV (22.05 kHz, mono) and the MIDI files encoded as CSV files and as WAV files (22.05 kHz, mono) rendered using the Software synthesizer <a href="https://www.fluidsynth.org/">FluidSynth</a>.</p> <p>SMD MIDI-Audio Piano Music (V1) contains the following data:</p> <ul> <li>wav_44100_stereo: Audio file (44.1 kHz, stereo)</li> <li>wav_22050_mono: Audio file (22.05 kHz, mono)</li> <li>midi: MIDI file</li> <li>csv: Export of note events from MIDI file into CSV format</li> <li>midi_wav_22050_mono: MIDI file rendered as audio file (22.05 kHz, mono)</li> </ul> <p>If you publish results obtained using this dataset, please cite:</p> <p>Meinard Müller, Verena Konz, Wolfgang Bogler, Vlora Arifi-Müller: Saarland Music Data (SMD). In Late-Breaking and Demo Session of the 12th International Conference on Music Information Retrieval (ISMIR), 2011. [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/2011_MuellerKonzBoglerArifi_SaarlandMusicData_ISMIR-LateBreaking.pdf">pdf</a>] [<a href="https://www.audiolabs-erlangen.de/resources/MIR/SMD/bibtex.html">bib</a>]</p>
Audio data from thesis Perception and Production of Nanning Mandarin Fourth Tone
<p>Recordings of 4 female speakers (S1, S2, S3, and S4) of Nanning-accented Mandarin Chinese reading preconstructed sentences. Recordings of those 4 female speakers telling a story based on 5 pages from Mercer Mayer's wordless picture book <em>Frog on His Own</em> (FOHO). Recording of perception test and perception test warm-up given to 26 Chinese living in Nanning, Guangxi. List of corpus sentences read, test prompts, and test sheet.</p>
Data Sets ''Kombucha--Proteinoid Biosynthetic Classifiers of Audio Signals''
<p>These voltage signals represent the electrical response generated by the Kombucha-Proteinoid biosynthetic system when exposed to various audio signals. The analysis of this voltage signal data contributes to the understanding and development of novel audio classification methods based on organic systems.</p>
Audio samples from generative models trained on the TIMIT speech data.
<p>This is a posting of audio snippets to accompany the paper "Benchmarking Generative Latent Variable Models for Speech".</p> <p>The snippets include samples and reconstructions. All samples are completely unconditional and utilise only the prior internal representations learned by the model. Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>
Data to accompany "Automatic text clustering for audio attribute elicitation experiment responses", AES 143rd Convention, New York, NY, USA, 2017
<p>This work was supported by the EPSRC Programme Grant S3A: Future Spatial Audio for an Immersive Listener Experience at Home (EP/L000539/1) and the BBC as part of the BBC Audio Research Partnership. Details about the data underlying this work, along with the terms for data access, are available from http://dx.doi.org/10.15126/surreydata.00841589.</p> <p>If you use the data, please cite the following paper:</p> <p>J. Francombe, T. Brookes, and R. Mason, “Automatic text clustering for audio attribute elicitation experiment responses”, AES 143rd Convention, New York, NY, USA, 2017</p>
Assessment of Simulations in Faust and Tascar for the Development of Audio Algorithms in Acoustic Environments - Code and Data
<p>Developing and testing audio algorithms with hard real-time constraints can be a complex task, requiring certain programming skills and/or<br>specialized equipment. However, many things can be tested in simulations on an ordinary computer, using <a href="https://tascar.org/" target="_blank" rel="noopener">TASCAR</a> for acoustic scene creation and<br><a href="https://faust.grame.fr/" target="_blank" rel="noopener">FAUST</a> for signal processing. Their capability are evaluated and compared to measurements using an FxLMS algorithm for active noise control as<br>example. This repository contains code and measured data of the publication “Assessment of simulations in FAUST and TASCAR for the development of<br>audio algorithms in acoustic environments”, presented at the International Faust Conference 2024 in Turin, Italy.</p>
Atmosphéries and the poetics of the in situ: the role and impact of sensors in data-to-sound transposition installations - Supplementary audio material
<p>This audio file wishes to give an insight into NXI Gestatio Design Lab's <em>Atmosphéries</em> research program, and to accompany the publication "Atmosphéries and the poetics of the in situ: the role and impact of sensors in data-to-sound transposition installations". The file consists in a 7-minutes recording of the <em>Meridian Probe</em>, the last instrument designed within the <em>Atmosphéries</em> program, which allows to generate sound from on-site atmospheric data. The recorded extract then acts as a sonic representation of the atmospheric conditions found at the place of exhibition, at the time of recording - in this case, gardens of the Bussy-Rabutin's Castle (France), respectively July 31st, 2021, 17h17. For further information about the <em>Meridian Probe</em>'s design and functioning, we invite you to read the aforementioned publication.</p>
Sample audio data for JSEALS article "Tonal variation in Pyen"
<p>Audio data samples corresponding to sections 1.1, 5.1 and 6 in the JSEALS article "Tonal variation in Pyen." </p>
Data for: The buzzOmeter system: In situ audio recordings of pollinators in flight
<ol> <li><span>The role of sounds produced by free-flying insects is challenging to research due to technical difficulties in obtaining audio recordings suitable for playback experiments. Experimental studies using flight sounds are needed to understand if buzzes carry information and by whom it is perceived.</span></li> <li><span>We developed the 'buzzOmeter system' for recording untethered, flying insects in their habitat, followed by file processing that allows precise measurements of acoustic parameters, including those dependent on the distance of the sound source from the microphone, i.e. signal magnitude measurements. The system consists of commercially available elements and open-source software.</span></li> <li><span>We provide a practical guide for the assembly and use of two alternative setups of the buzzOmeter system, followed by a video tutorial on file processing and an R script for the assignment of audio recordings to the corresponding species based on mixture discriminant analysis. Recordings of nine insect species (bees, wasps and lepidopterans) obtained with the use of our system in various habitats demonstrate its feasibility for field studies. </span></li> <li><span>Diverse species interactions are based on sound, and our new tool can aid researchers studying acoustical signalling in predator-prey, pollinator-plant and mimic-model complexes, among others.</span></li> </ol>
Data from: Audio-visual crossmodal correspondences in domestic dogs (Canis familiaris)
Open the record for dataset details and reuse information.
Data for: The buzzOmeter system: In situ audio recordings of pollinators in flight
Open the record for dataset details and reuse information.
Accompanying dataset for "Nappe oscillations on free-overfall structures, data from laboratory experiments (audio and video)"
<p>This dataset accompanies the manuscript "Nappe Oscillations on Free-Overfall Structures: Data from Laboratory Experiments" submitted to Scientific Data.</p> <p>This dataset contains raw audio and video data, which complement the dataset uploaded at: <a href="https://zenodo.org/record/3381078#.XlUGVSFKiUk">https://zenodo.org/record/3381078#.XlUGVSFKiUk</a></p> <p>The names of the folders describe each one of the 52 experiments, with respect to the submitted paper in Scientific Data:</p> <p>- M1 and M2 denote Model 1 and Model 2, respectively.</p> <p>- C and UC denote confined and unconfined nappe, respectively.</p> <p>- QR, THR, HR, R, and RR denote the crest type of the weir as explained in the paper.</p> <p>- W is the width of the crest and L is the falling height.</p>
December 2021 raw audio data at 9999 kHz using Grape 1 for radio station JJ1BDX
<p>This dataset contains all audio data (WAV, 8kHz S16_LE monaural) for the analysis of 10MHz standard frequency stations (mostly BPM, including WWVH and WWV) for HamSCI December 2021 Eclipse Festival recorded at radio station JJ1BDX in Setagaya City, Tokyo, Japan. </p> <p>Other references:</p> <ul> <li>See <a href="https://github.com/jj1bdx/hamsci-202112-freqdata">https://github.com/jj1bdx/hamsci-202112-freqdata</a> for the analysis results.</li> <li>See <a href="https://github.com/jj1bdx/dcrx-10MHz-design/">https://github.com/jj1bdx/dcrx-10MHz-design/</a> for the technical specification of the direct conversion receiver HamSCI Grape 1.</li> </ul> <p> </p>
Audio, Data and Codes
<p>This folder contains all the audio files, data and codes for analyses of the paper by Ponsot et al. (in press, 2018)</p>
Cane Toad Acoustic Classifier Audio Training Data
<h3><strong>Cane Toad Audio Dataset for Machine Learning Classifier Development</strong></h3> <h3><strong>Description:</strong></h3> <p>This dataset was created as part of a study aimed at developing a machine learning classifier to detect the advertisement calls of the cane toad (<em>Rhinella marina</em>) using BirdNET. The dataset comprises 3-second audio snippets that capture a variety of sounds, including cane toad vocalizations, calls from spectrally overlapping species, environmental noises, and unidentified sounds. These labelled sound data were collected from various Australian Acoustic Observatory' recording sites, covering a broad range of geographic locations and environmental conditions in Australia.</p> <h3><strong>Sound Classes:</strong></h3> <ul> <li>Background</li> <li>Canis lupus dingo (Dingo)</li> <li>Centropus phasianinus (Pheasant Coucal)</li> <li>Cyclorana australis (Water Holding Frog)</li> <li>Cyclorana cryptotis (Hidden Ear Frog)</li> <li>Cyclorana novaehollandiae (New Holland Frog)</li> <li>Dacelo novaeguineae (Kookaburra)</li> <li>Ninox boobook (Southern Boobook)</li> <li>Notaden melanoscaphus (Northern Spadefoot Toad)</li> <li>Rhinella marina (Cane Toad) – Processed with a low-pass filter to remove frequencies above 1300 Hz, minimizing interference from co-occurring sounds.</li> <li>Unidentified Sounds</li> </ul> <h3><strong>Use and Applications:</strong></h3> <p>This dataset is valuable for training machine learning models focused on the acoustic detection of cane toads, could be useful for researchers and professionals working in bioacoustics, machine learning, ecological monitoring and invasive species management.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.