Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6 results for “Audio Quality”

Learn how ShareScore rates datasets ↗
zenodo36/100

Binaural detection thresholds and audio quality of speech and music signals in complex acoustic environments

<p>Every-day acoustical environments are often complex, typically comprising one attended target sound in the presence of interfering sounds (e.g., disturbing conversations) and reverberation. Here we assessed binaural detection thresholds and (supra-threshold) binaural audio quality ratings of four distortions types: spectral ripples, non-linear saturation, intensity and spatial modifications applied to speech, guitar, and noise targets in such complex acoustic environments (CAEs). The target and (up to) two masker sounds were either co-located as if contained in a common audio stream, or were spatially separated as if originating from different sound sources. The amount of reverberation was systematically varied. Masker and reverberation had a significant effect on the distortion-detection thresholds of speech signals. Quality ratings were affected by reverberation, whereas the effect of maskers depended on the distortion. The results suggest that detection thresholds and quality ratings for distorted speech in anechoic conditions are also valid for rooms with mild reverberation, but not for moderate reverberation. Furthermore, for spectral ripples, a significant relationship between the listeners&rsquo; individual detection thresholds and quality ratings was found. The current results provide baseline data for detection thresholds and audio quality ratings of different distortions of a target sound in CAEs, supporting the future development of binaural auditory models.</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Dataset: "Increasing loudness in audio signals: a perceptually motivated approach to preserve audio quality"

<p>This archive contains the audio files that were used to test the algorithms presented in the paper entitled: &quot;Increasing loudness in audio signals: a perceptually motivated approach to preserve audio quality&quot;.</p> <p>Listening is to the samples is interesting, however, as the samples are quite short, the listening experience might strongly differ depending on the software and hardware used. We recommend to also look at the waveforms (using Audacity or others) and spectrum to see how the algorithm perform masking in order to increase the headroom of the samples.</p> <p>&nbsp;</p> <p>The archive is composed as follow:</p> <p>In the folder &quot;Original&quot; are the original kick drums used numbered from 1 to 16.</p> <p>The folder &quot;perceptual_alt&quot; contains the samples modified according to Eq(3). Each subfolder number corresponds to the number of the original sample used. The naming is &quot;x_a_c&quot;+lambda+&quot;.wav&quot;. Lambda ranges from 0 to 0.95.</p> <p>The folder &quot;perceptual&quot; contains the samples modified according to Eq(2). The naming is &quot;x_pc_c&quot;+c+&quot;.wav&quot;. c ranges from 0 to 40.</p> <p>The folder &quot;compressor&quot;, &quot;hardclip&quot;, &quot;softclip&quot; contains samples modified according to, respectively, the compressor, hardclipper and softclipper detailled in the paper. The naming is &quot;x_&quot;+{&quot;c_c&quot;,&quot;hc_c&quot;,&quot;sc_c&quot;}+c+&quot;.wav&quot;, where in this case c ranges from 0 to 40 and correspond to the threshold needed to get the same peak value as the perceptual algorithm (2) for a given perceptual constant c.</p> <p>&nbsp;</p> <p>The samples are taken from https:/ /www.musicradar.com/news/drums/1000-free-drum-samples, the names of each individual sample is:</p> <p>1: CYCdh_AcouKick-19.wav</p> <p>2: CyCdh_K3Snr-10.wav</p> <p>3: CyCdh_K3Tom-05.wav</p> <p>4: CYCdh_K1close_Flam-05.wav</p> <p>5: CYCdh_K1close_Kick-08.wav</p> <p>6: CYCdh_K1close_Rim-07.wav</p> <p>7: CYCdh_K1close_SdSt-07.wav</p> <p>8: CYCdh_K1close_SnrOff-08.wav</p> <p>9: CYCdh_K4-Kick05.wav</p> <p>10: CYCdh_K5-Kick93.wav</p> <p>11: CYCdh_KesKick-08.wav</p> <p>12: CYCdh_LooseKick-08.wav</p> <p>13: CYCdh_VinylK3-Kick01.wav</p> <p>14: CYCdh_VinylK3-Kick02.wav</p> <p>15: CYCdh_VinylK3-Snr02.wav</p> <p>16: CYCdh_VinylK4-Tom03.wav</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Personalized Audio Quality Preference Prediction Dataset

<p>Dataset for the following paper. Please cite this paper if our dataset is used in your research.</p> <p>Chung-Che Wang, Yu-Chun Lin, Yu-Teng Hsu, and Jyh-Shing Roger Jang, &quot;Personalized Audio Quality Preference Prediction&quot;, APSIPA ASC 2023.</p> <p>Here is a brief description of our dataset. For more details, please see our paper.</p> <p>This dataset is designed for personalized audio quality preference prediction. It includes recordings from 5 different mobile phones playing 7 distinct song segments at 2 volume settings. The played audio is recorded by using a binaural microphone and a computer interface. For each volume type, 70 pairs of recorded audio files are formed, where each of the two audio files in one pair corresponds to same song segment played by different mobile phones. For each volume type, each subject is asked to compare at least 14 of the 70 pairs. Subject information, which includes age, gender, and headphone/earphone specifications such as impedance, frequency response range, and sensitivity, are also collected.</p>

opencc-by-4.0Aug 2023View details →
ClinicalTrials.gov32/100

Audio-guided Home Physical Activity to Support People With Low Vision: Physical Performance and Quality of Life

ClinicalTrials.gov study NCT05956730. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

WebRTC-QoE: A Dataset of Quality of Experience in Audio-Video Communications

<p>In the realm of real-time communications, WebRTC-based multimedia applications are increasingly prevalent as these can be smoothly integrated within Web browsing sessions. The browsing experience is then significantly improved concerning scenarios where browser add-ons and/or plug-ins are used; still, the end user's Quality of Experience (QoE) in WebRTC sessions may be affected by network impairments, such as delays and losses. Due to the variability in user perceptions under different communications scenarios, comprehending and enhancing the resulting service quality is a complex endeavour. To address this, we present a dataset that provides a comprehensive perspective on the conversational quality of a two-party WebRTC-based audiovisual telemeeting service. This dataset was gathered through subjective evaluations involving 20 subjects across 15 different test conditions (TCs). A specialized system was developed to induce controlled network disruptions such as delay, jitter, and packet loss rate, which adversely affected the communication between the parties. This methodology offered insight into user perceptions under various network impairments. The dataset encompasses a blend of objective and subjective data, including ACR (Absolute Category Rating) subjective scores, <em>webrtc-internals</em>&nbsp;parameters, facial expressions features, and speech features. Consequently, it serves as a substantial contribution to the improvement of WebRTC-based video call systems, offering practical and real-world data that can drive the development of more robust and efficient multimedia communication systems, thereby enhancing the user&rsquo;s experience.</p> <div> <p>&nbsp;</p> <p>&nbsp;In the following, we discuss the details of the provided datasets.</p> <p><strong>Subjective_results_dataset.csv:</strong>&nbsp;This dataset encompasses subjective evaluation results from 20 subjects (users) who assessed the quality of WebRTC-based video calls under 15 distinct test conditions (TCs), which included combinations of 3 network impairments (delay, jitter, packet loss) to disturb the communication. The single discrete Absolute Category Rating (ACR) scale with five category labels (1-Bad, 2-Poor, 3-Fair, 4-Good, and 5-Excellent) was used by the users to rate the perceived QoE. A total of 300 ACR scores were obtained (20 participants x 15 TCs). The size of this dataset is 4.00 KB.</p> <p>The significance of each column is explained as follows:</p> <ul> <li>Test Condition (TC): It enumerates the TC numbers, which span from 1 to 15.</li> <li>Delay [ms]: It refers to the time it takes for a signal to travel from one point to another and is represented at three levels: 0 ms (no delay), 500 ms (moderate delay), and 1000 ms (significant delay).</li> <li>Jitter [ms]: It refers to the variability in delay and is represented at two levels: 0 ms (no jitter) and 500 ms (moderate jitter).</li> <li>Packet Loss Rate [%]: It refers to the loss of data packets during transmission and is represented at three levels: 0% (no packet loss), 15% (moderate packet loss), and 30% (significant packet loss).</li> <li>User: The users, identified from 1 to 20, participated in 15 video calls and evaluated the quality of these calls on a scale from 1 to 5.</li> </ul> <p><strong>Webrtc_internals_dataset.zip:</strong>&nbsp;This dataset contains text files collected using the webrtc-internals tool during the video calls. The zip is organized into 20 distinct folders, each labeled from &lsquo;User1&rsquo; to &lsquo;User20&rsquo;. Each user folder contains 15 text files (.txt), each named following the pattern &lsquo;webrtc_internals_dump-TCx_y-z-t.txt&rsquo;. In this naming convention, &lsquo;x&rsquo; denotes the TC number, ranging from 1 to 15. The &lsquo;y-z-t&rsquo; segment varies with each TC, where &lsquo;y&rsquo; signifies the delay value (0, 500, or 1000), &lsquo;z&rsquo; indicates the jitter value (0 or 500), and &lsquo;t&rsquo; represents the packet loss rate (0, 15, or 30). Each text file includes application-level data concerning WebRTC sessions' statistics in a JSON format. A total of 300 webrtc-internals dump text files were obtained (20 participants x 15 TCs). The size of this zip dataset is 28.7 MB (396 MB uncompressed).&nbsp;</p> <p><strong>Facial_expression_features_dataset.zip:</strong>&nbsp;This dataset contains facial expression features extracted from the recorded videos with face images using the OpenFace toolkit. The zip is organized into 20 distinct folders, each labelled from &lsquo;User1&rsquo; to &lsquo;User20&rsquo;. Each user folder contains 15 .csv files, which are named following the pattern &lsquo;TCx_y-z-t.csv&rsquo;. In this naming convention, &lsquo;x&rsquo; denotes the TC, ranging from 1 to 15. The &lsquo;y-z-t&rsquo; segment varies with each TC, where &lsquo;y&rsquo; signifies the delay value (0, 500, or 1000), &lsquo;z&rsquo; indicates the jitter value (0 or 500), and &lsquo;t&rsquo; represents the packet loss rate (0, 15, or 30). A total of 300 facial expression feature files in .csv format were obtained (20 participants x 15 TCs).&nbsp;For each face image of each TC, the OpenFace outputs 6 gaze direction features, 280 eye region landmarks and 35 Action Units (AUs). The size of this zip dataset is 547 MB (2.14 GB uncompressed).&nbsp;</p> <p>Each column carries a specific significance, which is elaborated as follows:</p> <ul> <li>frame: the frame number in the context of sequences.</li> <li>face_id: the identifier assigned to each face when multiple faces are present.</li> <li>timestamp: the elapsed time in seconds during the processing of a video sequence.</li> <li>confidence: the level of confidence the tracker has in the current landmark detection estimate.</li> <li>success: a face has been detected in the frame and it has been tracked accurately.</li> <li>gaze_0_x, gaze_0_y, gaze_0_z: the normalized eye gaze direction vector in world coordinates for eye 0, which is the eye on the left in the image.</li> <li>gaze_1_x, gaze_1_y, gaze_1_z: the normalized eye gaze direction vector in world coordinates for eye 1, which is the eye on the right in the image.</li> <li>gaze_angle_x, gaze_angle_y: the eye gaze direction, averaged for both eyes and expressed in world coordinates in radians, is converted into a format that is easier to use than gaze vectors.</li> <li>eye_lmk_x_0, eye_lmk_x_1, ..., eye_lmk_x55, eye_lmk_y_1, ... eye_lmk_y_55: the pixel coordinates of 2D landmarks in the eye region.</li> <li>eye_lmk_X_0, eye_lmk_X_1, ..., eye_lmk_X55, eye_lmk_Y_0, ..., eye_lmk_Z_55: the position of landmarks in the eye region in 3D space, measured in millimeters.</li> <li>17 AUr: detect the activation intensity (from 1 to 5) of a particular facial muscle. These are: AU01_r, AU02_r, AU04_r, AU05_r, AU06_r, AU07_r, AU09_r, AU10_r, AU12_r, AU14_r, AU15_r, AU17_r, AU20_r, AU23_r, AU25_r, AU26_r, AU45_r.</li> <li>18 AUc: Identify the activation of a particular muscle and note its presence (0 for absent, 1 for present). These are: AU01_c, AU02_c, AU04_c, AU05_c, AU06_c, AU07_c, AU09_c, AU10_c, AU12_c, AU14_c, AU15_c, AU17_c, AU20_c, AU23_c, AU25_c, AU26_c, AU28_c, AU45_c.</li> </ul> <p><strong>Speech_features_dataset.csv:</strong>&nbsp;This dataset is a robust assembly of speech features extracted from the recorded audio files using the OpenSMILE toolkit. The dataset is further enhanced with the integration of Absolute Category Rating (ACR) scores, which were assigned by each subject for every TC. These scores are embedded within the speech features of each subject. The dataset is exhaustive, encompassing a total of 1,911,900 speech features (calculated as 15 TCs x 20 subjects x 6373 speech features per subject). The total size of this dataset is 18.9 MB.&nbsp;&nbsp;</p> <p>Each column carries a specific significance, which is elaborated as follows:</p> <ul> <li>file: It presents the audio files from which the speech features were extracted.</li> <li>Users: The users, identified from 1 to 20, participated in 15 video calls and evaluated the quality of these calls on a scale from 1 to 5.</li> <li>Test Condition (TC): It enumerates the Test Condition numbers, which span from 1 to 15.</li> <li>Delay [ms]: It refers to the time it takes for a signal to travel from one point to another and is represented at three levels: 0 ms (no delay), 500 ms (moderate delay), and 1000 ms (significant delay).</li> <li>Jitter [ms]: It refers to the variability in delay and is represented at two levels: 0 ms (no jitter) and 500 ms (moderate jitter).</li> <li>Packet Loss Rate [%]: It refers to the loss of data packets during transmission and is represented at three levels: 0% (no packet loss), 15% (moderate packet loss), and 30% (significant packet loss).</li> <li>OpenSmile Speech Features Columns: This presents a detailed view of the speech features. Specifically, from column G to column IKI, it enumerates 6373 distinct speech features for each user under each TC of each processed audio file in .wav.</li> <li>ACR Score: The Absolute Category Rating (ACR) scale ranges from 1, representing the lowest quality, to 5, indicating the highest quality evaluated by the subjects.</li> </ul> <p>&nbsp;</p> <p><strong>If you make use of this dataset, please consider citing the following publication:</strong></p> <p>Bingol G., Porcu, S., Floris, A., &amp; Atzori, L. (2024). WebRTC-QoE: A dataset of QoE assessment of subjective scores, network impairments, and facial &amp; speech features. Computer Networks, 244, 110356, doi: 10.1016/j.comnet.2024.110356.</p> <p>BibTex format:</p> <p>@article{bingol2024datasetwebrtc, title={WebRTC-QoE: A dataset of QoE assessment of subjective scores, network impairments, and facial &amp; speech features}, author={Bingol, Gulnaziye and Porcu, Simone and Floris, Alessandro and Atzori, Luigi}, journal={Computer Networks}, volume={244}, pages={110356}, year={2024}, publisher={Elsevier}, doi = {https://doi.org/10.1016/j.comnet.2024.110356} }</p> </div>

opencc-by-4.0Dec 2022View details →
ClinicalTrials.gov28/100

Study of Effectiveness of Audio Guided Deep Breathing on Improving the Quality of Life of Physically Disabled Group

ClinicalTrials.gov study NCT05396027. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record