Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.7.1
Dataset results
13 results for “urban sound”
Dataset-AOB: urban sounds events classification
<p>The dataset Dataset-AOB is an audio dataset collected and manually edited for urban sounds events classification using Convolutional Neural Networks for the Master Thesis: </p> <p>Ospina, A. "Audio Event Classification using Deep Learning. Use case: Urban Sounds Events classification with Convolutional Neural Networks," M.Eng. thesis, Beuth University of Applied Sciences, Berlin, 2020.</p> <p>- 10 audio events: alarm-siren, children playing, dog bark, engine, footsteps, glass breaking, gun shot, metro train, rain and screams.</p> <p>- duration: < 4 seconds</p> <p>- format: (.wav)</p> <p>- sampling rate: 22KHz - 44KHz</p> <p>- files: Dataset-AOB: development dataset (4831 samples), DatasetEVAL-AOB: evaluation (218 samples)</p> <p>- metadata: (.csv)</p> <p>- sources per class: (.png)</p> <p>Contact: aospinab@gmail.com</p>
USM Dataset - A Dataset for Polyphonic Sound Event Tagging in Urban Sound Monitoring Scenarios
<p>This dataset includes 24,000 5-seconds-long polyphonic stereo soundscapes composed of sounds taken from the FSD50k dataset:</p> <p>- Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra. FSD50K: an Open Dataset of Human-Labeled Sound Events (<a href="https://arxiv.org/abs/2010.00475">https://arxiv.org/abs/2010.00475</a>)</p> <p>FSD50k samples used in the USM dataset were selected to allow for commercial usage.</p> <p>Find more details about the USM dataset at <a href="https://github.com/jakobabesser/USM">https://github.com/jakobabesser/USM</a></p>
Urban Sound & Sight (Urbansas) - Labeled set
<p><strong>Urban Sound & Sight (Urbansas): </strong></p> <p>Version 1.0, May 2022</p> <p><strong>Created by</strong><br> Magdalena Fuentes (1, 2), Bea Steers (1, 2), Pablo Zinemanas (3), Martín Rocamora (4), Luca Bondi (5), Julia Wilkins (1, 2), Qianyi Shi (2), Yao Hou (2), Samarjit Das (5), Xavier Serra (3), Juan Pablo Bello (1, 2)<br> 1. Music and Audio Research Lab, New York University<br> 2. Center for Urban Science and Progress, New York University<br> 3. Universitat Pompeu Fabra, Barcelona, Spain<br> 4. Universidad de la República, Montevideo, Uruguay<br> 5. Bosch Research, Pittsburgh, PA, USA</p> <p><strong>Publication</strong></p> <p>If using this data in academic work, please cite the following paper, which presented this dataset:<br> M. Fuentes, B. Steers, P. Zinemanas, M. Rocamora, L. Bondi, J. Wilkins, Q. Shi, Y. Hou, S. Das, X. Serra, J. Bello. “Urban Sound & Sight: Dataset and Benchmark for Audio-Visual Urban Scene Understanding”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022.</p> <p><strong>Description</strong></p> <p>Urbansas is a dataset for the development and evaluation of machine listening systems for audiovisual spatial urban understanding. One of the main challenges to this field of study is a lack of realistic, labeled data to train and evaluate models on their ability to localize using a combination of audio and video.<br> We set four main goals for creating this dataset: <br> 1. To compile a set of real-field audio-visual recordings;<br> 2. The recordings should be stereo to allow exploring sound localization in the wild;<br> 3. The compilation should be varied in terms of scenes and recording conditions to be meaningful for training and evaluation of machine learning models;<br> 4. The labeled collection should be accompanied by a bigger unlabeled collection with similar characteristics to allow exploring self-supervised learning in urban contexts.<br> Audiovisual data<br> We have compiled and manually annotated Urbansas from two publicly available datasets, plus the addition of unreleased material. The public datasets are the TAU Urban Audio-Visual Scenes 2021 Development dataset (street-traffic subset) and the Montevideo Audio-Visual Dataset (MAVD):</p> <p><br> Wang, Shanshan, et al. "A curated dataset of urban scenes for audio-visual scene analysis." ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021.</p> <p>Zinemanas, Pablo, Pablo Cancela, and Martín Rocamora. "MAVD: A dataset for sound event detection in urban environments." Detection and Classification of Acoustic Scenes and Events, DCASE 2019, New York, NY, USA, 25–26 oct, page 263--267 (2019).</p> <p><br> The TAU dataset consists of 10-second segments of audio and video from different scenes across European cities, traffic being one of the scenes. Only the scenes labeled as traffic were included in Urbansas. MAVD is an audio-visual traffic dataset curated in different locations of Montevideo, Uruguay, with annotations of vehicles and vehicle components sounds (e.g. engine, brakes) for sound event detection. Besides the published datasets, we include a total of 9.5 hours of unpublished material recorded in Montevideo, with the same recording devices of MAVD but including new locations and scenes.</p> <p>Recordings for TAU were acquired using a GoPro Hero 5 (30fps, 1280x720) and a Soundman OKM II Klassik/studio A3 electret binaural in-ear microphone with a Zoom F8 audio recorder (48kHz, 24 bits, stereo). Recordings for MAVD were collected using a GoPro Hero 3 (24fps, 1920x1080) and a SONY PCM-D50 recorder (48kHz, 24 bits, stereo). </p> <p>When compiled in Urbansas, it includes 15 hours of stereo audio and video, stored in separate 10 second MPEG4 (1280x720, 24fps) and WAV (48kHz, 24 bit, 2 channel) files. Both released video datasets are already anonymized to obscure people and license plates, the unpublished MAVD data was anonymized similarly using this anonymizer. We also distribute the 2fps video used for producing the annotations.</p> <p>The audio and video files both share the same filename stem, meaning that they can be associated after removing the parent directory and extension.</p> <p>MAVD:<br> video/<location_id>_<mavd_clip_id>_<clip_split_id>.mp4<br> audio/<location_id>_<mavd_clip_id>_<clip_split_id>.wav</p> <p>TAU:<br> video/<location_id>_<tau_clip_id>.mp4<br> audio/<location_id>_<tau_clip_id>.wav</p> <p><br> where location_id in both cases includes the city and an ID number.</p> <p><br> city & places & clips & mins & frames & labeled mins \\<br> Montevideo & 8 & 4085 & 681 & 980400 & 92 \\<br> Stockholm & 3 & 91 & 15 & 21840 & 2 \\<br> Barcelona & 4 & 144 & 24 & 34560 & 24 \\<br> Helsinki & 4 & 144 & 24 & 34560 & 16 \\<br> Lisbon & 4 & 144 & 24 & 34560 & 19 \\<br> Lyon & 4 & 144 & 24 & 34560 & 6 \\<br> Paris & 4 & 144 & 24 & 34560 & 2 \\<br> Prague & 4 & 144 & 24 & 34560 & 2 \\<br> Vienna & 4 & 144 & 24 & 34560 & 6 \\<br> London & 5 & 144 & 24 & 34560 & 4 \\<br> Milan & 6 & 144 & 24 & 34560 & 6 \\<br> \midrule<br> Total & 50 & 5472 & 912 & 1.3M & 180 \\</p> <p><br> <strong>Annotations</strong></p> <p><br> Of the 15 hours of audio and video, 3 hours of data (1.5 hours TAU, 1.5 hours MAVD) are manually annotated by our team both in audio and image, along with 12 hours of unlabeled data (2.5 hours TAU, 9.5 hours of unpublished material) for the benefit of unsupervised models. The distribution of clips across locations was selected to maximize variance across different scenes. The annotations were collected at 2 frames per second (FPS) as it provided a balance between temporal granularity and clip coverage.</p> <p>The annotation data is contained in video_annotations.csv and audio_annotations.csv. </p> <p><strong>Video Annotations</strong></p> <p>Each row in the video annotations represents a single object in a single frame of the video. The annotation schema is as follows:</p> <ul> <li>frame_id: The index of the frame within the clip the annotation is associated with. This index is 0-based and goes up to 19 (assuming 10-second clips with annotations at 2 FPS)</li> <li>track_id: The ID of the detected instance that identifies the same object across different frames. These IDs are guaranteed to be unique within a clip.</li> <li>x, y, w, h: The top-left corner and width and height of the object’s bounding box in the video. The values are given in absolute coordinates with respect to the image size (1280x720). </li> <li>class_id: The index of the class corresponding to: [0, 1, 2, 3, -1] — see label for the index mapping. The -1 value corresponds to the case where there are no events, but still clip-level annotations, like night and city. When operating on bounding boxes, class_id of -1 should be filtered.</li> <li>label: The label text. This is equivalent to LABELS[class_id], where LABELS=[car, bus, motorbike, truck, -1]. The label -1 has the same role as above.</li> <li>visibility: The visibility of the object. This is 1 unless the object becomes obstructed, where it changes to 0.</li> <li>filename: The file ID of the associated file. This is the file’s path minus the parent directory and extension.</li> <li>city: The city where the clip was collected in.</li> <li>location_id: The specific name of the location. This may include an integer ID following the city name for cases where there are multiple collection points.</li> <li>time: The time (in seconds) of the annotation, relative to the start of the file. Equivalent to frame_id / fps .</li> <li>night: Whether the clip takes place during the day or at night. This value is singular per clip.</li> <li>subset: Which data source the data originally belongs to (TAU or MAVD).</li> </ul> <p><strong>Audio Annotations</strong></p> <p>Each row represents a single object instance, along with the time range that it exists within the clip. The annotation schema is as follows:</p> <ul> <li>filename: The file ID odd the associated audio file. See filename above. </li> <li>class_id, label: See above. Audio has an additional class_id of 4 (label=offscreen) which indicates an off-screen vehicle - meaning a vehicle that is heard but not seen. A class_id of -1 indicates a clip-level annotation for a clip that has no object annotations (an empty scene).</li> <li>non_identifiable_vehicle_sound: True if the region contains the sound of vehicles where individual instances cannot be uniquely identified. </li> <li>start, end: The start and end times (in seconds) of the annotation relative to the file. </li> </ul> <p><strong>Conditions of use</strong></p> <p>Dataset created by Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, and Juan Pablo Bello.</p> <p>The Urbansas dataset is offered free of charge under the following terms:</p> <ul> <li>Urbansas annotations are release under the CC BY 4.0 license</li> <li>Urbansas video and audio replicates the original sources licenses: <ul> <li> MAVD subset is released under CC BY 4.0 </li> <li> TAU subset is released under a Non-Commercial license</li> </ul> </li> </ul> <p><strong>Feedback</strong></p> <p>Please help us improve Urbansas by sending your feedback to:</p> <ul> <li>Magdalena Fuentes: mfuentes@nyu.edu</li> <li>Bea Steers: bsteers@nyu.edu </li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p><strong>Acknowledgments</strong></p> <p>This work was partially supported by the National Science Foundation award 1955357 and Bosch RTC.</p>
FuSA: recording of urban sounds in Valdivia, Chile, between 23 May and 6 June 2022
<p>This dataset includes 115 audio recordings and corresponding metadata. The audio recordings were obtained through two monitoring stations installed in different parts of Valdivia (city ubicated at the south of Chile).</p> <p>The two monitoring stations listened to a grand total of 30,240 minutes during two-week and one-week operation periods, respectively. From this dataset only 115 minutes were flagged as surpassing the established sound pressure level.</p> <p>The FuSA system (<a href="https://www.acusticauach.cl/fusa/">https://www.acusticauach.cl/fusa/</a>) was then used to obtain event predictions for the 115 minutes subset.<br> The file metadata.csv includes predictions for each audio recordings and their corresponding probabilities.</p>
"I was the class teacher at that time. It was a class trip, usually organized near the end of the schoolterm in summer. The pupils went there by bike to have a barbecue at the sandy banks of the river Rhine near Dusseldorf. The landscape around is mostly dominated by agriculture and glasshouse cultures. You find a mixture of former villages nowadays completely suburbanized. The population finds jobs in the nearby urban centers like Dusseldorf, Neuss and other big cities. The reason why Irecorded the scene is simply because Iam interested in collecting sounds in general by doing recordings in different surroundings like nature, cities and everything between. My memories about the event are that it was a relaxing and funny atmosphere, which is not always the case while teaching in a classroom" [Reinhard/reinsamba]15 in Collecting Sounds. Online Sharing of Field Recordings as Cultural Practice
"I was the class teacher at that time. It was a class trip, usually organized near the end of the schoolterm in summer. The pupils went there by bike to have a barbecue at the sandy banks of the river Rhine near Dusseldorf. The landscape around is mostly dominated by agriculture and glasshouse cultures. You find a mixture of former villages nowadays completely suburbanized. The population finds jobs in the nearby urban centers like Dusseldorf, Neuss and other big cities. The reason why Irecorded the scene is simply because Iam interested in collecting sounds in general by doing recordings in different surroundings like nature, cities and everything between. My memories about the event are that it was a relaxing and funny atmosphere, which is not always the case while teaching in a classroom" [Reinhard/reinsamba]15
SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network
<p><strong>SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network</strong></p> <p>Version 2.3, September 2020</p> <p> </p> <p><strong>Created by</strong></p> <p>Mark Cartwright (1,2,3), Jason Cramer (1), Ana Elisa Mendez Mendez (1), Yu Wang (1), Ho-Hsiang Wu (1), Vincent Lostanlen (1,2,4), Magdalena Fuentes (1), Graham Dove (2), Charlie Mydlarz (1,2), Justin Salamon (5), Oded Nov (6), Juan Pablo Bello (1,2,3)</p> <ol> <li>Music and Audio Research Lab, New York University</li> <li>Center for Urban Science and Progress, New York University</li> <li>Department of Computer Science and Engineering, New York University</li> <li>Cornell Lab of Ornithology</li> <li>Adobe Research</li> <li>Department of Technology Management and Innovation, New York University</li> </ol> <p> </p> <p><strong>Publication</strong></p> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Cartwright, M., Cramer, J., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In <em>Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE)</em>, 2020.<br> <a href="https://arxiv.org/abs/2009.05188">[pdf]</a></p> <p> </p> <p><strong>Description</strong></p> <p>SONYC Urban Sound Tagging (SONYC-UST) is a dataset for the development and evaluation of machine listening systems for realistic urban noise monitoring. The audio was recorded from the <a href="https://wp.nyu.edu/sonyc">SONYC</a> acoustic sensor network. Volunteers on the <a href="https://zooniverse.org">Zooniverse</a> citizen science platform tagged the presence of 23 classes that were chosen in consultation with the New York City Department of Environmental Protection. These 23 fine-grained classes can be grouped into 8 coarse-grained classes. The recordings are split into three sets: training, validation, and test. The training and validation sets are disjoint with respect to the sensor from which each recording came, and the test set is displaced in time. For increased reliability, three volunteers annotated each recording. In addition, members of the SONYC team subsequently created a subset of verified, ground-truth tags using a two-stage annotation procedure in which two annotators independently tagged and then collectively resolved any disagreements. This subset of recordings with verified annotations intersects with all three recording splits. All of the recordings in the test set have these verified annotations. In v2 version of this dataset, we have also included coarse spatiotemporal context information to aid in tag prediction when time and location is known. For more details on the motivation and creation of this dataset see the <a href="http://dcase.community/challenge2020/task-urban-sound-tagging-with-spatiotemporal-context">DCASE 2020 Urban Sound Tagging with Spatiotemporal Context Task website</a>.</p> <p> </p> <p><strong>Audio data</strong></p> <p>The provided audio has been acquired using the SONYC acoustic sensor network for urban noise pollution monitoring. Over 60 different sensors have been deployed in New York City, and these sensors have collectively gathered the equivalent of over 50 years of audio data, of which we provide a small subset. The data was sampled by selecting the nearest neighbors on VGGish features of recordings known to have classes of interest. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings. To maintain privacy, we quantized the spatial information to the level of a city block, and we quantized the temporal information to the level of an hour. We also limited the occurrence of recordings with positive human voice annotations to one per hour per sensor.</p> <p> </p> <p><strong>Label taxonomy</strong></p> <p>The label taxonomy is as follows:</p> <ol> <li>engine<br> 1: small-sounding-engine<br> 2: medium-sounding-engine<br> 3: large-sounding-engine<br> X: engine-of-uncertain-size</li> <li>machinery-impact<br> 1: rock-drill<br> 2: jackhammer<br> 3: hoe-ram<br> 4: pile-driver<br> X: other-unknown-impact-machinery</li> <li>non-machinery-impact<br> 1: non-machinery-impact</li> <li>powered-saw<br> 1: chainsaw<br> 2: small-medium-rotating-saw<br> 3: large-rotating-saw<br> X: other-unknown-powered-saw</li> <li>alert-signal<br> 1: car-horn<br> 2: car-alarm<br> 3: siren<br> 4: reverse-beeper<br> X: other-unknown-alert-signal</li> <li>music<br> 1: stationary-music<br> 2: mobile-music<br> 3: ice-cream-truck<br> X: music-from-uncertain-source</li> <li>human-voice<br> 1: person-or-small-group-talking<br> 2: person-or-small-group-shouting<br> 3: large-crowd<br> 4: amplified-speech<br> X: other-unknown-human-voice</li> <li>dog<br> 1: dog-barking-whining</li> </ol> <p>The classes preceded by an <code>X</code> code indicate when an annotator was able to identify the coarse class, but couldn’t identify the fine class because either they were uncertain which fine class it was or the fine class was not included in the taxonomy. <code>dcase-ust-taxonomy.yaml</code> contains this taxonomy in an easily machine-readable form.</p> <p> </p> <p><strong>Data splits</strong></p> <p>This release contains a training subset (13538 recordings from 35 sensors), and validation subset (4308 recordings from 9 sensors), and a test subset (669 recordings from 48 sensors). The training and validation subsets are disjoint with respect to the sensor from which each recording came. The sensors in the test set will not disjoint from the training and validation subsets, but the test recordings are displaced in time, occurring after any of the recordings in the training and validation subset. The subset of recordings with verified annotations (1380 recordings) intersects with all three recording splits. All of the recordings in the test set have these verified annotations.</p> <p> </p> <p><strong>Annotation data</strong></p> <p>The annotation data are contained in <code>annotations.csv</code>, and encompass the training, validation, and test subsets. Each row in the file represents one multi-label annotation of a recording—it could be the annotation of a single citizen science volunteer, a single SONYC team member, or the agreed-upon ground truth by the SONYC team (see the <em>annotator_id</em> column description for more information). Note that since the SONYC team members annotated each class group separately, there may be multiple annotation rows by a single SONYC team annotator for a particular audio recording.</p> <p> </p> <p> </p> <p><strong>Columns</strong></p> <p><em>split</em></p> <p>The data split. (<em>train</em>, <em>validate, test</em>)</p> <p><em>sensor_id</em></p> <p>The ID of the sensor the recording is from.</p> <p><em>audio_filename</em></p> <p>The filename of the audio recording</p> <p><em>annotator_id</em></p> <p>The anonymous ID of the annotator. If this value is positive, it is a citizen science volunteer from the Zooniverse platform. If it is negative, it is a SONYC team member. If it is <code>0</code>, then it is the ground truth agreed-upon by the SONYC team.</p> <p><em>year</em></p> <p>The year the recording is from.</p> <p><em>week</em></p> <p>The week of the year the recording is from.</p> <p><em>day</em></p> <p>The day of the week the recording is from, with Monday as the start (i.e. <code>0</code>=Monday).</p> <p><em>hour</em></p> <p>The hour of the day the recording is from</p> <p><em>borough</em><br> The NYC borough in which the sensor is located (<code>1</code>=Manhattan, <code>3</code>=Brooklyn, <code>4</code>=Queens). This corresponds to the first digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>block</em></p> <p>The NYC block in which the sensor is located. This corresponds to digits 2—6 digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>latitude</em></p> <p>The latitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em>longitude</em></p> <p>The longitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_presence</em></p> <p>Columns of this form indicate the presence of fine-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset.</p> <p><em><coarse_id>_<coarse_name>_presence</em></p> <p>Columns of this form indicate the presence of a coarse-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset. These columns are computed from the fine-level class presence columns and are presented here for convenience when training on only coarse-level classes.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_proximity</em></p> <p>Columns of this form indicate the proximity of a fine-level class. After indicating the presence of a fine-level class, citizen science annotators were asked to indicate the proximity of the sound event to the sensor. Only the citizen science volunteers performed this task, and therefore this data is not included in the verified annotations. This column may take on one of the following four values: (<code>near</code>, <code>far</code>, <code>notsure</code>, <code>-1</code>). If <code>-1</code>, then the proximity was not annotated because either the annotation was not performed by a citizen science volunteer, or the citizen science volunteer did not indicate the presence of the class.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>Dataset created by Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez, Yu Wang, Ho-Hsiang Wu, Vincent Lostanlen, Magdalena Fuentes, Graham Dove, Charlie Mydlarz, Justin Salamon, Oded Nov, and Juan Pablo Bello</p> <p>The SONYC-UST dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-UST dataset or any part of it.</p> <p> </p> <p><strong>Feedback</strong></p> <p>Please help us improve SONYC-UST by sending your feedback to:</p> <ul> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>We would like to thank all the Zooniverse volunteers who continue to contribute to our project. This work is supported by <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">National Science Foundation award 1544753</a>.</p> <p> </p> <p><strong>Change log</strong></p> <ul> <li>2.3 Added the ground truth annotations for the test set, and regrouped the audio files for upload to Zenodo.</li> <li>2.2 Added the audio for the test set (audio-eval.tar.gz).</li> <li>2.1 The DCASE 2020 development dataset. 14778 new recordings added along with coarse spatiotemporal context information.</li> <li>1.0 Data is the same as v0.4. Publication added to README.</li> <li>0.4 Fixed error in annotations. Previously, the coarse class "machinery-impact" was accidentally indicated as present whenever "non-machinery-impact" was present regardless of the presence of "machinery-impact". This error has been fixed.</li> <li>0.3 Test set annotations added</li> <li>0.2 Test set audio files added</li> </ul>
Urban background sound recordings for virtual acoustics under various weather conditions at IHTApark
<p>This is an open database of calibrated background sound recordings with metadata on meteorological and acoustical parameters ready to be used in virtual reality applications, for soundsacpe studies.</p> <p>The soundscapes have been recorded at the IHTApark (green space next to the Institute for Hearing Technology and Acoustics) in Aachen (Germany) in winter and spring of 2022.</p> <p>They have been segemented into 30-seconds auralizable snippets.</p> <p>For an<strong> <a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener">interactive exploration</a></strong> of the soundscapes, please visit the follwing website:</p> <p><a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener"></a></p> <p><a href="https://paad-group.github.io/IHTApark-ambient-recordings/" target="_blank" rel="noopener">https://paad-group.github.io/IHTApark-ambient-recordings/</a></p> <p>The dabase is available as binaural files, first-order ambisonics files, and as omni files.</p> <p>The metadata is included as a .xlsx file.</p> <p> </p>
Realistic urban sound mixture dataset
<p>This dataset resumes an urban sound corpus whose the realism has been proved through a perceptual test [1]. This corpus has been used in order to estimate the traffic sound level with the Non-negative Matrix Factorization formula [2].</p> <p>This dataset presents 4 folders :</p> <ul> <li><em>dictionary</em> where the traffic audio samples dedicated to the dictionary design of NMF are,</li> <li><em>recordings</em> which contains the 74 original recordings,</li> <li><em>annotation</em> which contains the annotations text files of the 74 audio files,</li> <li><em>transcribed scenes </em>which contains the 74 transcribed audio files generated with <a href="https://bitbucket.org/mlagrange/simscene"><em>SimScene</em></a> software. 4 folders composed it, according to the sound environment of the audio files (<em>park, quiet street, noisy street, very noisy street). </em> In each folder, one can find the global sound mixtures, the audio of each sound class as well as the files that include all the elements associated with the traffic and the interfering class (which contains all the other sound sources).</li> </ul> <p>[1] Gloaguen, J. R., Can, A., Lagrange, M., & Petiot, J. F. (2017, June). Creation of a corpus of realistic urban sound scenes with controlled acoustic properties. In <em>173rd Meeting of the Acoustical Society of America and the 8th Forum Acusticum (Acoustics' 17)</em>.</p> <p>[2] Gloaguen, J. R., Can, A., Lagrange, M., & Petiot, J. F. (2018), Road traffic sound level from realistic urban sound mixtures by Non-negative Matrix Factorization, submitted for publication</p>
Isolated urban sound database
<p>The Isolated urban sound database contains the audio samples used to design urban sound mixtures using <a href="https://bitbucket.org/mlagrange/simscene">SimScene</a> software. </p> <p>This database has already been used to design urban sound mixtures that can be found in <a href="https://zenodo.org/record/1145855">Estimation of the road traffic sound levels based on Non-Negative Matrix Factorization dataset</a> [1] and in <a href="https://zenodo.org/record/1184443">Realistic urban sound mixture dataset</a> [2]</p> <p>The dataset contains two folders :</p> <p>- 'event' which includes includes 231 brief sound samples considered as salient, with a 1 to 20 seconds duration and classified among 21 sound classes (ringing bell, whistling bird, car horn, passing car, hammer, barking dog, siren, footstep, metallic noise, voice...)</p> <p>- 'background' which includes 162 long duration sounds (~1mn30), whose acoustic properties do not vary in time. This category includes among others, whistling bird, crowd noise, rain, children playing in schoolyard, constant traffic noise ...</p> <p>More details on this sound database can be found in [3]</p> <p> </p> <p>[1] J.-R. Gloaguen, M. Lagrange, A. Can, J.-F. Petiot, Estimation of the road traffic sound levels in urban areas based on non-negative matrix factorization techniques, submitted for publication</p> <p>[2] J.-R. Gloaguen, A. Can, M. Lagrange, J.-F. Petiot, Road traffic sound level estimation from realistic urban sound mixtures by Non-negative Matrix Factorization, submitted for publication</p> <p>[3] J.-R. Gloaguen, A. Can, M. Lagrange, J.-F. Petiot, Creation of a corpus of realistic urban sound scenes with controlled acoustic properties, in: Acoustics ’17 Boston, Vol. 141 of The Journal of the Acoustical Society of America, Acoustical Society of America and the European Acoustics Association, Boston, United States, 2017, pp. 4044–4044.</p>
Plane-wave propagation path data from wideband MIMO channel sounding in an urban microcellular scenario
<p>We provide plane-wave propagation path data from a wideband MIMO radio channel sounding measurement in an urban microcell scenario. The binary Matlab file includes: direction of departure (DOD: variables "par.PhiTx" and "par.ThetaTx" in [rad]), direction of arrival (DOA: "par.PhiRx" and "par.ThetaRx" in [rad]), delay ("par.Tau" to be multiplied with 8.3ns, the tab length), and complex polarimetric path gain ("par.Alpha" is a 2x2 matrix, where element [1,1]=TXtheta ->RXtheta, [1,2]=TXphi ->RXtheta, [2,1]=TXtheta->RXphi, and [2,2]=TXphi->RXphi), for the 30 strongest signal paths (from TX to RX) at each of the 4574 RX locations along the route described below. In the element names above, the term "theta" refers to the vertically polarised component, and accordingly the term "phi" referes to the horizontally polarised component.<br> Note #1: The exact RX location for each individual measured radio channel was NOT recorded (see route description below). <br> Note #2: The complex path gain (par.Alpha) is NOT calibrated, but depends on the initially fixed AGC level in the receiver, which was chosen to provide the best dynamic range for the given mobile (RX) route.<br> Both these limitations are seen reasonable since this dataset is meant for the realistic *statistical comparison* of the performance of different RX antennas in a microcell environment (and not to determine the actual received power at each exact location of the measured route).<br> Background information: The provided dataset is processed and is based on a radio channel sounder measurement at 5.3 GHz, carried out in downtown Helsinki, Finland, in April 2004. The uniform rectangular transmit (TX) array was placed at 10 m height in Aleksanterinkatu-street (an approx. 15-m wide street canyon), in front of the Nordea building, broadside pointing westwards (towards Stockmann building). The semishperical receive (RX) array was moved at 1.6-m height and for about 50 m along Aleksanterinkatu-street in line-of-sight (LOS), i.e. from in front of Kluuvi shopping centre westwards just across the crossing of Kluuvikatu-street. The TX and RX arrays cover the relevant azimuth and elevation ranges, so that this plane wave propagation path data can directly be combined with the polarimetric directional radiation pattern(s) of an antenna (array).</p>
Urban junco flight initiation distances correlate with approach velocities of anthropogenic sounds
Open the record for dataset details and reuse information.
BCN Dataset: an Annotated Night Urban Sounds dataset
<p>This dataset contains 363 minutes and 53 seconds of real-world audio recordings made at the city center of Barcelona. The dataset has been carefully labelled and is organized in three different audio files (.wav files) with their corresponding labels files. </p>
Data from: Chatty maps: constructing sound maps of urban areas from social media data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.