Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

13

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

13 results for “audio analysis”

Learn how ShareScore rates datasets ↗
zenodo44/100

A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)

<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie &quot;Forrest Gump&quot; (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation)&nbsp;and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

ESSENTIA analysis of audio snippets from the Million Song Dataset Taste Profile subset

<p>This upload includes the ESSENTIA analysis output of (a subset of) song snippets from the Million Song Dataset, namely those included in the Taste Profile subset. The audio snippets were collected from 7digital.com and were subsequently analyzed with ESSENTIA 2.1-beta3. Pre-trained SVM models provided by the ESSENTIA authors on their website were applied.</p> <p>The file <strong>msd_song_jsons.rar </strong>contains the ESSENTIA analysis output after applying the SVM models for highlevel feature extraction. Please note that these are 204317 files.</p> <p>The file <strong>msd_played_songs_essentia.csv.gz </strong>contains all one-dimensional real-valued fields of the jsons merged into one csv file with 204317 rows.</p> <p>The full procedure and subsequent analysis is described in</p> <p>Fricke, K. R., Greenberg, D. M., Rentfrow, P. J., &amp; Herzberg, P. Y. (2019). Measuring musical preferences from listening behavior: Data from one million people and 200,000 songs. <em>Psychology of Music</em>, 0305735619868280.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Audio files for spectrum analysis demonstration

<div>This data set consists of 6 real-world audio files (in .wav format, 48000Hz, mono) that are carefully crafted as examples for spectral analysis (i.e for teaching or as sample test data for algorithms).</div> <div>&nbsp;</div> <div>The files have clear discernible sound with an added true random background ambient noise (composed of: distant fan noise + distant street traffic + close harddisk clicking noise). The sound is clearly discernible by a human, despite the noise.&nbsp;</div> <div>&nbsp;</div> <div>- There are some musical sounds (the note G3 on several instruments; fundamental frequency 196Hz) on a tubular bell, classical piano, trumpet, violin. The sounds of the instrument was generated from MIDI banks with FluidSynth software, played on a loudspeaker and re-recorded with an analogical microphone (with the ambient noises). Audio processing was performed with Tenacity software;</div> <div>- Sample from human speech (wovel &ldquo;o&rdquo;), with the same processing as above;&nbsp;</div> <div>- The &ldquo;Noise&rdquo; file is purely digitally generated (white noise).</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>Each audio set is composed of three files:&nbsp;</div> <div>- The audio file (*.wav), each sampled at 48000 Hz, Mono.&nbsp;</div> <div>- an amplitude file (*_amplitude.csv, corresponding linear amplitudes recorded by the microphone of the .wav file). Numeric format in simple text format (.csv) with labeled column names.&nbsp;</div> <div>-a spectrum file (*_spectrum.csv, frequency/amplitude(dB) ) with the results of a FFT (Fast Fourier Transform). Numeric format in simple text format (.csv) with labeled column names.&nbsp;</div> <div>&nbsp;</div> <div>Detailed description of each set is provided below.</div> <div>&nbsp;</div> <div> <ul> <li><strong>Bell_G3.wav</strong></li> <li>Bell_G3_amplitude.csv:<br>Length processed: 113851 samples 2.37190 seconds.<br>Sample Rate: 48000 Hz. <br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 113851 samples, 2.37190 seconds.<br>Peak amplitude: 0.59001 (linear) -4.58276 dB.&nbsp; <br>Unweighted RMS: -19.87796 dB.<br>DC offset: 0.00069 linear, -63.18732 dB.</li> <li>Bell_G3_spectrum.csv:<br>FFT transform (Hz / dB)</li> </ul> </div> <div> <ul> <li><strong>Noise.wav</strong></li> <li>Noise_amplitude.csv:<br>Length processed: 31765 samples 0.66177 seconds.<br>Sample Rate: 48000 Hz. <br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 31765 samples, 0.66177 seconds.<br>Peak amplitude: 0.52797 (linear) -5.54788 dB.<br>Unweighted RMS: -16.61153 dB.<br>DC offset: -0.00013 linear, -77.53051 dB</li> <li>Noise_spectrum.csv)<br>FFT transform (Hz / dB)</li> </ul> </div> <div> <ul> <li><strong>Piano_G3.wav</strong></li> <li>Piano_G3_amplitude.csv:<br>Sample Rate: 48000 Hz.<br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 133063 samples, 2.77215 seconds.<br>Peak amplitude: 0.41070 (linear) -7.72946 dB.<br>Unweighted RMS: -23.95101 dB.<br>DC offset: 0.00030 linear, -70.58058 dB.</li> <li>Piano_G3_spectrum.csv:<br>FFT transform (Hz / dB)</li> </ul> </div> <div> <ul> <li><strong>Trumpet_G3.wav</strong></li> <li>Trumpet_G3_amplitude.csv:<br>Sample Rate: 48000 Hz.<br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 66931 samples, 1.39440 seconds.<br>Peak amplitude: 0.29551 (linear) -10.58853 dB.&nbsp; <br>Unweighted RMS: -22.00611 dB.<br>DC offset: 0.00076 linear, -62.39740 dB.</li> <li>Trumpet_G3_spectrum.csv:<br>FFT transform (Hz / dB)</li> </ul> </div> <div> <ul> <li><strong>Violin_G3.wav</strong></li> <li>Violin_G3_amplitude.csv:<br>Sample Rate: 48000 Hz.<br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 59252 samples, 1.23442 seconds.<br>Peak amplitude: 0.41633 (linear) -7.61119 dB.<br>Unweighted RMS: -18.36738 dB.<br>DC offset: 0.00021 linear, -73.46784 dB.</li> <li>Violin_G3_spectrum.csv:<br>FFT transform (Hz / dB)</li> </ul> </div> <div> <ul> <li><strong>Wovel_O.wav</strong></li> <li>Wovel_O_amplitude.csv<br>Sample Rate: 48000 Hz.<br>Sample values on linear scale. 1 channel (mono).<br>Length processed: 4975 samples, 0.10365 seconds.<br>Peak amplitude: 0.46174 (linear) -6.71214 dB. <br>Unweighted RMS: -14.88086 dB.<br>DC offset: 0.00009 linear, -80.54917 dB.</li> <li>Wovel_O_spectrum.csv:<br>FFT transform (Hz / dB)</li> </ul> </div> <div>&nbsp;</div> <div>These files are created by A. Iftime and released under Creative Commons Licence, 2024.&nbsp;</div> <div>&nbsp;</div> <div>You might cite the dataset as:&nbsp;</div> <div>&ldquo;Audio files for spectrum analysis demonstration&rdquo; [dataset] (2024), in &ldquo;Medical Biophysics for 1st year medical students&rdquo;, by Călinescu O., Babeș R., Iftime A., Băran I., Ionescu D., Ganea C., in publishing&nbsp;</div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Jamendo content analysed with the Audio Commons Analysis Service

<p>This dataset contains partial output of running the final release of the <a href="https://github.com/AudioCommons/faas-ac-analysis">Audio Commons Analysis Service</a> on 99960 music pieces of the <a href="https://licensing.jamendo.com">Jamendo Licensing</a> catalogue. The service is described in <a href="https://www.audiocommons.org/assets/files/AC-WP4-QMUL-D4.13%20Release%20of%20tool%20for%20the%20automatic%20semantic%20description%20of%20music%20pieces.pdf">Deliverable D4.13</a> of the Audio Commons project. The analysis results are comprised of chord output including confidence for all pieces (using the software described in the deliverable) and of the output of the Essentia music extractor (v2.1_beta4-447-gc5ea2738) for 77190 of the pieces (of which tempo, beats, tuning and global-key are exposed through the API of the analysis service).</p> <p>The dataset is formatted as a single JSON file containing an array of documents. Each document in the array has two or three top-level keys. The first key is &quot;_id&quot; with values of the form &quot;jamendo-tracks:&lt;jamendo-id&gt;&quot;. This id can be used to request metadata and audio through the <a href="https://developer.jamendo.com">Jamendo API</a>. The second key of the document is &quot;chords&quot;, containing a subdocument with the output of the chord extraction algorithm. The third key is &quot;essentia-music&quot;, which contains a subdocument with the output of the Essentia music extractor.</p> <div>&nbsp;</div>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Vocalization Patterns in Laying Hens - An Analysis of Stress-Induced Audio Responses

<p>This repository houses a comprehensive collection of data and resources from the study "Vocalization Patterns in Laying Hens - An Analysis of Stress-Induced Audio Responses." Led by Dr. Suresh Neethirajan at Mooanalytica, Department of Agriculture &amp; Aquaculture, Faculty of Agriculture &amp; Computer Science, Dalhousie University, this research represents a significant foray into the field of poultry ethology and welfare monitoring using advanced machine learning techniques.</p> <p><strong>Key Components of the Repository</strong></p> <ol> <li> <p><strong>Experimental Audio Data</strong></p> <ul> <li><strong>Control and Treatment Vocalizations</strong>: Audio recordings of laying hens under two different stress conditions &ndash; sudden umbrella opening (Treatment 1) and simulated dog barking sounds (Treatment 2), along with control groups. The processed dataset is approximately 460 MB for the control group experimental data and about 2 GB for the 2 treatment group experimental data, capturing the nuanced responses of hens to these stressors.</li> <li><strong>Original Raw Data</strong>: The original, unprocessed audio data is around 9 GB in size. Though not included in the repository, it can be made available upon reasonable request.</li> </ul> </li> <li> <p><strong>Algorithm and Code Files</strong></p> <ul> <li><strong>CNN Feature Extraction and Classification Algorithms</strong>&nbsp;Python scripts used for the extraction of features from the audio data using Convolutional Neural Networks (CNN) and subsequent classification.</li> <li><strong>Supplementary Algorithms</strong>&nbsp;Additional code files that support the processing and analysis of the audio data.</li> </ul> </li> <li> <p><strong>MFCC Feature Dataset</strong></p> <ul> <li>An Excel file containing the 40 Mel Frequency Cepstral Coefficients (MFCC) features extracted from the vocalization data. This dataset provides a detailed spectral analysis of the hen's vocalizations, crucial for understanding their response to stress.</li> </ul> </li> </ol> <p><strong>Study Overview</strong></p> <p>This study aimed to classify and analyze the vocalization patterns of laying hens subjected to different stressors. Using a CNN model, the research identified distinct vocal patterns between control and treated groups, indicating unique vocal responses to different types of stressors. This study is pivotal in understanding the impact of environmental stressors on poultry welfare and behavior. The age of the chickens and the timing of stressor application were also critical factors influencing vocalization patterns.</p> <p><strong>Implications and Applications:</strong></p> <p>The findings from this study have significant implications for poultry welfare monitoring and management. By providing a non-invasive method to assess the well-being of chickens, this research contributes valuable insights into enhancing poultry management practices and welfare standards.</p> <p>The resources in this repository are intended for researchers, academicians, and professionals in animal behavior, veterinary science, and poultry management. We encourage the use of these data and tools for further research and practical applications in the field of precision (Digital) livestock farming and animal welfare.</p> <p>For any queries or requests related to the raw dataset, please contact Dr. Suresh Neethirajan.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

IDMT Audio Provenance Analysis Dataset

<p>This dataset contains two distinct collections tailored for evaulating audio provenance analysis solutions within specified scenarios: Singular Composition and Multi-Source Composition. For a comprehensive understanding of these scenarios and the process behind generating the test files, please consult the referenced publication.</p> <p>This dataset is accompanied to publication and in case you use it please cite:</p> <p>M. Gerhardt, L. Cuccovillo, P. Aichroth, &ldquo;Audio Provenance Analysis in Heterogeneous Media Sets&rdquo;, &nbsp;<em>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW),</em> in press.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Auditory Scene Analysis dataset (Multichannel universal sound separation & polyphonic audio classification)

<p>We constructed a new dataset for <strong>multichannel universal sound separation</strong> and <strong>polyphonic audio classification</strong> tasks.</p> <p>We constructed a new dataset for multichannel USS and polyphonic audio classification tasks. The proposed dataset is designed to reflect various conditions, including moving sources with temporal onsets and offsets. For foreground sound sources, signals from 13 audio classes were selected from open-source databases (Pixabay and FSD50K, Librispeech, MUSDB18, Vocalsound). These signals were resampled to 16 kHz and pre-processed by either padding zeros or cropping to 4 seconds. Each sound source has a 75% probability of being a moving source, with speeds ranging from 0 to 3 m/s. The dataset features between 2 to 4 foreground sound sources, along with one background noise from the diffused TAU-SNoise dataset with a signal-to-noise ratio (SNR) ranging from 6 to 30 dB. The simulations were conducted using gpuRIR. Room dimensions were set to a width and length between 5 and 8 meters, and a height between 3 and 4 meters, with reverberation times ranging from 0.2 to 0.6 seconds. These parameters were sampled from uniform distributions. We simulated spatialized sound sources using a 4-channel tetrahedral microphone array with a radius of 4.2 cm. The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p> <p>The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

SARdB: A Dataset for Audio Scene Source Counting and Analysis

<p>This dataset contains audio and textual data for audio source counting. A balanced&nbsp;audio dataset is contained under &#39;SAR&#39; containing speech and environmental audio mixtures. The corresponding speech transcripts are contained in &#39;transcript&#39; as .txt files. More information can be found <a href="https://github.com/mnigro9/SARdB/tree/v1.0.0">here</a>&nbsp;&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo28/100

Audio-Polygraphy Dataset for Sleep Apnea Analysis (APSAA)

<p>The&nbsp;<strong>Audio-Polygraphy Dataset for Sleep Apnea Analysis (APSAA)</strong> provides synchronized, full-night audio and polygraph recordings from 32 subjects, along with manual annotations for labeled events in the polygraph studies. &nbsp;All subjects were provided with detailed information about the study, and written informed consent was obtained from those&nbsp;who agreed to take part. The recordings were collected between September 2021 and April 2022 at the Sleep Unit of Dr. Sagaz Hospital in Ja&eacute;n (Spain). The study was approved by the Provincial Research Ethics Committee of Ja&eacute;n (Spain).</p> <p>Each subject's data is organized in a designated folder named according to the subject's unique identification code. Inside each folder, users will find: (1) the audio recording in WAV format, (2) separate CSV files for each polygraph signal, and (3) a CSV file containing manual annotations of polygraph events. The polygraph signals included are as follows:</p> <ul> <li><strong>Abdomen_EG</strong>: Abdominal respiratory effort.</li> <li><strong>Flow_EG</strong>: Nasal airflow from the nasal cannula.</li> <li><strong>Pulse_EG</strong>: Pulse rate, measured in beats per minute.</li> <li><strong>Snore_EG</strong>: Respiratory snore pressure envelope from the nasal cannula.</li> <li><strong>SpO2_EG</strong>: Peripheral oxygen saturation percentage.</li> <li><strong>Thermistor_EG</strong>: Oronasal thermal airflow.</li> <li><strong>Thorax_EG</strong>: Thoracic respiratory effort.</li> </ul> <p>Additionally, an automated algorithm for synchronizing audio and polygraph signals in sleep studies is provided, accessible via the following <a href="https://github.com/fdgonzal/Polygraph-Audio-Sync" target="_blank" rel="noopener">Github repository</a></p> <p><em>Funding: This work was supported in part under grant 1257914 funded by Programa Operativo FEDER Andalucia 2014&ndash;2020, grant P18-RT-1994 funded by the Ministry of Economy, Knowledge and University (Junta de Andaluc&iacute;a, Spain), by MCIN/AEI/10.13039/501100011033 under the project grants PID2020-119082RB-{C21,C22} and by the Ministerio de Ciencia, Innovaci&oacute;n y Universidades (Gobierno de Espa&ntilde;a) under the grants PID2023-146520OB-{C21,C22}.</em></p>

restrictedcc-by-4.0Nov 2024View details →
zenodo28/100

Data to accompany "Evaluation of Spatial Audio Reproduction Methods (Part 2): Analysis of Listener Preference", J. AES, 2016

<p>This work was supported by the EPSRC Programme Grant S3A: Future Spatial Audio for an Immersive Listener Experience at Home (EP/L000539/1). Details about the data underlying this work, along with the terms for data access, are available from http://dx.doi.org/10.15126/surreydata.00809533</p> <p>If you use the data, please cite the following paper:</p> <p>J. Francombe, T. Brookes, and R. Mason, 2016: Evaluation of Spatial Audio Reproduction Methods (Part 2): Analysis of Listener Preference. Journal of the Audio Engineering Society</p>

opencc-by-nc-4.0Nov 2020View details →
ClinicalTrials.gov20/100

Audio Health Engagement Analysis in Diabetes: The AHEAD Study

ClinicalTrials.gov study NCT01938807. IPD Sharing: Not stated. Countries: 0. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo12/100

Supplementary audio files for lecture slides "Machine Listening for Music and Sound Analysis (MLMSA)"

<p>Supplementary material / audio examples for lecture slides provided at https://machinelistening.github.io/</p> <p>Files need to be placed in a separated folder &quot;audio&quot;</p>

restrictedNov 2021View details →
zenodo8/100

ISMIR2018 paper: comparative analysis audio

<p>This is the audio used to train and test the algorithms that we have used in the comparative analysis of the paper submitted to ISMIR 2018: <em>Music detecion in broadcast media recordings: a non-binary approach with relative loudness annotations.</em></p> <p>It contains, separately, the training and testing splits.</p>

restrictedApr 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record