Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “urban sensors”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data from multi-sensor devices and reference station to monitoring urban air quality

<p>Data from electrochemical and optical sensors.</p> <h3>Files names</h3> <ul> <li>ECT01, ECT02, ECT06, ECT07 = device name</li> <li>ISSEP = reference station <ul> <li>"c" = calibration data</li> <li>"v" = validation data</li> </ul> </li> </ul> <h3>Variable description</h3> <table> <tbody> <tr> <td><strong>Electrochemical sensor</strong></td> <td><strong>Optical sensor</strong></td> <td><strong>Probe</strong></td> <td><strong>Reference</strong></td> </tr> <tr> <td> <p>AE = auxiliary electrode (mV)</p> <p>WE = working electrode (mV)</p> <p>N = temperature correction&nbsp;</p> <ul> <li>ch0 = CO sensor</li> <li>ch1 = OX sensor</li> <li>ch2 = NO2 sensor</li> <li>ch3 = NO sensor</li> </ul> </td> <td> <p>PM1, PM2.5 and PM10 in &micro;g/m&sup3;</p> </td> <td> <p>Prs_mbar = pressure (mbar)</p> <p>Temp = temperature (&deg;C)</p> <p>RH = relative humidity (%)</p> </td> <td> <p>DV30 = wind direction @ 30m (&deg;)</p> <p>HR = relative humidity (%)</p> <p>NO, NO2, O3, PM10 and PM2.5 (&micro;g/m&sup3;)</p> <p>Precipita = precipitation (mm)</p> <p>TC3 = temperature @ 3m (&deg;C)</p> <p>VV30 = wind speed @ 30m (m/s)</p> </td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

PoqueiraOccupancy: Dataset and metrics from occupancy sensors of urban areas and establishments in the region of Barranco del Poqueira in the Alpujarra Granadina

<p>This dataset is linked to the analysis of different aspects related to the conservation of the Sierra Nevada National Park through advanced digital systems. The devices have been deployed in the municipalities of Pampaneira and Capileira in the Alpujarra region of the province of Granada. The data is collected by 11 BOSCH fixed cameras 11.00 387.4900 Interior IR 5.3 MP and 4 TURRET type cameras Interior IR Lens 2.8 mm 5.3 MP 100&ordm; H.265 multi-streaming (H.265; H.264; M-JPEG).</p> <p>The devices have been installed as follows: 3 devices have been placed in establishments in Capileira, 8 in establishments in Pampaneira, and 4 in urban passage areas in the municipality of Pampaneira. All devices are capable of measuring the entry and exit to the establishment or area they are designated for. In some cases, there is also a metric which measures the number of people present within that area. The information related to the establishments has been anonymized to ensure the privacy of the collaborating companies in the project and the flow of customers during the studied period.</p> <p>The data attached in the CSV files DATA_OCCUPANCY_2022 and DATA_OCCUPANCY_2023 contain information about individuals detected by the cameras in the years 2022 (from February to December) and 2023 (from January to August). The calculation of the number of people is done cumulatively in hourly intervals. The collected variables include:</p> <ul> <li> <p>device_ID: The name of the device recording the value.</p> </li> <li> <p>type: The metric measuring the recording, which can be ENTRADA (entry), SALIDA (exit), or AFORO (occupancy).</p> </li> <li> <p>date: The date and time at which the cumulative people count is recorded for the specific metric.</p> </li> <li> <p>counter: The number of people counted for a specific metric in that time period.</p> </li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo36/100

IoT Sensor Deployment in the Wildland Urban Interface: Leveraging Fire Risk Analysis

<p>Included here are individual burn maps used for evaluating algorithm results in the paper: IoT Sensor Deployment in the Wildland Urban Interface: Leveraging Fire Risk Analysis. This paper will be presented at the IEEE World Forum on Internet of Things in November 2024 and available on IEEE Xplore after that.</p> <p>Also included are maps of fuel load and elevation (geotifs) and the daily weather (in .csv format) for the region of interest used in the burn probability simulator Burn-P3+ to generate the individual burn maps.</p> <p>This paper investigates various algorithms for distributing Internet of Things sensors within the Wildland-Urban Interface to enhance early wildland fire detection. Utilizing geospatial data analysis and a validated wildland fire growth model burn maps were generated to guide sensor placement strategies across a defined region of interest. The algorithms evaluated include an even grid distribution, random distributions, and genetic algorithm-based methods. Each algorithm was tested against 50,000 selected burn maps to assess detection rates, with sensor counts ranging from 50 to 800 across 500 experimental runs. Results indicate that while the even grid distribution yielded the highest detection rates, the practicality of such a method in real-world applications is limited. Genetic algorithms showed promise, but require further exploration to more accurately simulate random distribution used in field deployment. Surprisingly, weighting sensor placement based on wildland fire growth risk did not significantly impact detection effectiveness, suggesting the need for additional research into the representativeness of selected burn maps.</p> <p>Partial code for the sensor deployment algorithms discussed in the above mentioned paper is <a href="https://github.com/richardjpurcell/sensor-deployment-algorithms">available on GitHub</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Dataset supporting publication: "Data collected by coupling fix and wearable sensors for addressing urban microclimate variability in an historical Italian city"

<p>Dataset supporting publication: &ldquo;Data collected by coupling fix and wearable sensors for addressing urban microclimate variability in an historical Italian city&rdquo;&nbsp;(publication available for download:&nbsp;<a href="https://zenodo.org/record/3901556">GEOFIT Zenodo</a>)</p> <p>Datasets resulting from monitoring activities of Sant&#39;Apollinare systems and climatic parameters inside and outside the building (post-intervention monitoring).</p> <p>The&nbsp;article presents the data collected through an extensive research work conducted in a historic hilly town in central Italy during the period 2016-2017. Data concern two different datasets: long-term hygrothermal histories collected in two specific positions of the town object of the research, and three environmental transects collected following on foot the same designed path at three different time of the same day, i.e. during a heat wave event in summer. The short-term monitoring campaign is carried out by means of an innovative wearable weather station specifically developed by the authors and settled upon a bike helmet. Data provided within the short-term monitoring campaign are analysed by computing the apparent temperature, a direct indicator of human thermal comfort in the outdoors. All provided environmental data are geo-referenced. These data are used in order to examine the intra-urban microclimate variability. Outcomes from both long- and short-term monitoring campaigns allow to confirm the existing correlation between the urban forms and functionalities and the corresponding local microclimate conditions, also generated by anthropogenic actions. In detail, higher fractions of built surfaces are associated to generally higher temperatures as emerges by comparing the two long-term air temperature data series, i.e. temperature collected at point 1 is higher than temperature collated at point 2 for the 75% of the monitored period with an average of &thorn;2.8 [1]C. Furthermore, gathered environmental transects demonstrate the high variability of the main environmental parameters below the Urban Canopy. Diversification of the urban thermal behaviour leads to a computed apparent temperature range in between 33.2 [1]C and 46.7 [1]C at 2 p.m. along the monitoring path. Reuse of these data may be helpful for further investigating interesting correlations among urban configuration, anthropogenic actions and microclimate variables affecting outdoor comfort. Additionally, the proposed dataset may be compared to other similar datasets collected in other urban contexts around the world. Finally, it can be compared to other monitoring methodologies such as weather stations and satellite measurements available in the location at the same time.</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Seasonally optimized calibrations improve low-cost sensor performance: Long-term field evaluation of PurpleAir sensors in urban and rural India

Open the record for dataset details and reuse information.

publicAug 2023View details →
zenodo32/100

SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network

<p><strong>SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network</strong></p> <p>Version 2.3, September 2020</p> <p>&nbsp;</p> <p><strong>Created by</strong></p> <p>Mark Cartwright (1,2,3), Jason Cramer (1), Ana Elisa Mendez Mendez (1), Yu Wang (1), Ho-Hsiang Wu (1), Vincent Lostanlen (1,2,4), Magdalena Fuentes (1), Graham Dove (2), Charlie Mydlarz (1,2), Justin Salamon (5), Oded Nov (6), Juan Pablo Bello (1,2,3)</p> <ol> <li>Music and Audio Research Lab, New York University</li> <li>Center for Urban Science and Progress, New York University</li> <li>Department of Computer Science and Engineering, New York University</li> <li>Cornell Lab of Ornithology</li> <li>Adobe Research</li> <li>Department of Technology Management and Innovation, New York University</li> </ol> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Cartwright, M., Cramer, J., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In <em>Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE)</em>, 2020.<br> <a href="https://arxiv.org/abs/2009.05188">[pdf]</a></p> <p>&nbsp;</p> <p><strong>Description</strong></p> <p>SONYC Urban Sound Tagging (SONYC-UST) is a dataset for the development and evaluation of machine listening systems for realistic urban noise monitoring. The audio was recorded from the <a href="https://wp.nyu.edu/sonyc">SONYC</a>&nbsp;acoustic sensor network. Volunteers on the &nbsp;<a href="https://zooniverse.org">Zooniverse</a>&nbsp;citizen science platform tagged the presence of 23 classes that were chosen in consultation with the New York City Department of Environmental Protection. These 23 fine-grained classes can be grouped into 8 coarse-grained classes. The recordings are split into three sets: training, validation, and test. The training and validation sets are disjoint with respect to the sensor from which each recording came, and the test set is displaced in time. For increased reliability, three volunteers annotated each recording. In addition, members of the SONYC team subsequently created a subset of verified, ground-truth tags using a two-stage annotation procedure in which two annotators independently tagged and then collectively resolved any disagreements. This subset of recordings with verified annotations intersects with all three recording splits. All of the recordings in the test set have these verified annotations.&nbsp; In v2 version of this dataset, we have also included coarse spatiotemporal context information to aid in tag prediction when time and location is known. For more details on the motivation and creation of this dataset see the <a href="http://dcase.community/challenge2020/task-urban-sound-tagging-with-spatiotemporal-context">DCASE 2020 Urban Sound Tagging with Spatiotemporal Context Task website</a>.</p> <p>&nbsp;</p> <p><strong>Audio data</strong></p> <p>The provided audio has been acquired using the SONYC acoustic sensor network for urban noise pollution monitoring. Over 60 different sensors have been deployed in New York City, and these sensors have collectively gathered the equivalent of over 50 years of audio data, of which we provide a small subset. The data was sampled by selecting the nearest neighbors on VGGish features of recordings known to have classes of interest. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings. To maintain privacy, we quantized the spatial information to the level of a city block, and we quantized the temporal information to the level of an hour. We also limited the occurrence of recordings with positive human voice annotations to one per hour per sensor.</p> <p>&nbsp;</p> <p><strong>Label taxonomy</strong></p> <p>The label taxonomy is as follows:</p> <ol> <li>engine<br> 1: small-sounding-engine<br> 2: medium-sounding-engine<br> 3: large-sounding-engine<br> X: engine-of-uncertain-size</li> <li>machinery-impact<br> 1: rock-drill<br> 2: jackhammer<br> 3: hoe-ram<br> 4: pile-driver<br> X: other-unknown-impact-machinery</li> <li>non-machinery-impact<br> 1: non-machinery-impact</li> <li>powered-saw<br> 1: chainsaw<br> 2: small-medium-rotating-saw<br> 3: large-rotating-saw<br> X: other-unknown-powered-saw</li> <li>alert-signal<br> 1: car-horn<br> 2: car-alarm<br> 3: siren<br> 4: reverse-beeper<br> X: other-unknown-alert-signal</li> <li>music<br> 1: stationary-music<br> 2: mobile-music<br> 3: ice-cream-truck<br> X: music-from-uncertain-source</li> <li>human-voice<br> 1: person-or-small-group-talking<br> 2: person-or-small-group-shouting<br> 3: large-crowd<br> 4: amplified-speech<br> X: other-unknown-human-voice</li> <li>dog<br> 1: dog-barking-whining</li> </ol> <p>The classes preceded by an <code>X</code> code indicate when an annotator was able to identify the coarse class, but couldn&rsquo;t identify the fine class because either they were uncertain which fine class it was or the fine class was not included in the taxonomy. <code>dcase-ust-taxonomy.yaml</code> contains this taxonomy in an easily machine-readable form.</p> <p>&nbsp;</p> <p><strong>Data splits</strong></p> <p>This release contains a training subset (13538 recordings from 35 sensors), and validation subset (4308 recordings from 9 sensors), and a test subset (669 recordings from 48 sensors). The training and validation subsets are disjoint with respect to the sensor from which each recording came. The sensors in the test set will not disjoint from the training and validation subsets, but the test recordings are displaced in time, occurring after any of the recordings in the training and validation subset. The subset of recordings with verified annotations (1380 recordings) intersects with all three recording splits.&nbsp; All of the recordings in the test set have these verified annotations.</p> <p>&nbsp;</p> <p><strong>Annotation data</strong></p> <p>The annotation data are&nbsp;contained in <code>annotations.csv</code>, and&nbsp;encompass the training, validation, and test subsets. Each row in the file represents one multi-label annotation of a recording&mdash;it could be the annotation of a single citizen science volunteer, a single SONYC team member, or the agreed-upon ground truth by the SONYC team (see the <em>annotator_id</em> column description for more information).&nbsp; Note that since the SONYC team members annotated each class group separately, there may be multiple annotation rows by a single SONYC team annotator for a particular audio recording.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Columns</strong></p> <p><em>split</em></p> <p>The data split. (<em>train</em>, <em>validate, test</em>)</p> <p><em>sensor_id</em></p> <p>The ID of the sensor the recording is from.</p> <p><em>audio_filename</em></p> <p>The filename of the audio recording</p> <p><em>annotator_id</em></p> <p>The anonymous ID of the annotator. If this value is positive, it is a citizen science volunteer from the Zooniverse platform. If it is negative, it is a SONYC team member. If it is <code>0</code>, then it is the ground truth agreed-upon by the SONYC team.</p> <p><em>year</em></p> <p>The year the recording is from.</p> <p><em>week</em></p> <p>The week of the year the recording is from.</p> <p><em>day</em></p> <p>The day of the week the recording is from, with Monday as the start (i.e. <code>0</code>=Monday).</p> <p><em>hour</em></p> <p>The hour of the day the recording is from</p> <p><em>borough</em><br> The NYC borough in which the sensor is located (<code>1</code>=Manhattan, <code>3</code>=Brooklyn, <code>4</code>=Queens). This corresponds to the first digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>block</em></p> <p>The NYC block in which the sensor is located. This corresponds to digits 2&mdash;6 digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>latitude</em></p> <p>The latitude coordinate of the <strong>block</strong>&nbsp;in which the sensor is located.</p> <p><em>longitude</em></p> <p>The longitude coordinate of the <strong>block</strong>&nbsp;in which the sensor is located.</p> <p><em>&lt;coarse_id&gt;-&lt;fine_id&gt;_&lt;fine_name&gt;_presence</em></p> <p>Columns of this form indicate the presence of fine-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset.</p> <p><em>&lt;coarse_id&gt;_&lt;coarse_name&gt;_presence</em></p> <p>Columns of this form indicate the presence of a coarse-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset. These columns are computed from the fine-level class presence columns and are presented here for convenience when training on only coarse-level classes.</p> <p><em>&lt;coarse_id&gt;-&lt;fine_id&gt;_&lt;fine_name&gt;_proximity</em></p> <p>Columns of this form indicate the proximity of a fine-level class. After indicating the presence of a fine-level class, citizen science annotators were asked to indicate the proximity of the sound event to the sensor. Only the citizen science volunteers performed this task, and therefore this data is not included in the verified annotations. This column may take on one of the following four values: (<code>near</code>, <code>far</code>, <code>notsure</code>, <code>-1</code>). If <code>-1</code>, then the proximity was not annotated because either the annotation was not performed by a citizen science volunteer, or the citizen science volunteer did not indicate the presence of the class.</p> <p>&nbsp;</p> <p><strong>Conditions of use</strong></p> <p>Dataset created by Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez, Yu Wang, Ho-Hsiang Wu, Vincent Lostanlen, Magdalena Fuentes, Graham Dove, Charlie Mydlarz, Justin Salamon, Oded Nov, and Juan Pablo Bello</p> <p>The SONYC-UST dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an &ldquo;as is&rdquo; basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-UST dataset or any part of it.</p> <p>&nbsp;</p> <p><strong>Feedback</strong></p> <p>Please help us improve SONYC-UST&nbsp;by sending your feedback to:</p> <ul> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p>&nbsp;</p> <p><strong>Acknowledgments</strong></p> <p>We would like to thank all the Zooniverse volunteers who continue to contribute to our project. This work is supported by <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">National Science Foundation award 1544753</a>.</p> <p>&nbsp;</p> <p><strong>Change log</strong></p> <ul> <li>2.3 Added the ground truth annotations for the test set, and regrouped the audio files for upload to Zenodo.</li> <li>2.2&nbsp;Added the audio for the test set (audio-eval.tar.gz).</li> <li>2.1 The DCASE 2020 development dataset. 14778 new recordings added along with coarse spatiotemporal context information.</li> <li>1.0 Data is the same as v0.4. Publication added to README.</li> <li>0.4 Fixed error in annotations. Previously, the coarse class &quot;machinery-impact&quot; was accidentally indicated as present whenever &quot;non-machinery-impact&quot; was present regardless of the presence of &quot;machinery-impact&quot;. This error has been fixed.</li> <li>0.3 Test set annotations added</li> <li>0.2 Test set audio files added</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo32/100

SONYC-Backgrounds: a collection of urban background recordings from an acoustic sensor network

<p><strong>Created by</strong></p> <p>Aurora Cramer <sup>(1, 2)</sup>, Mark Cartwright <sup>(3)</sup>, Fatemeh Pishdadian <sup>(4)</sup>, Juan Pablo Bello <sup>(1,2,5,6)</sup></p> <p>&nbsp;&nbsp;&nbsp; 1. Music and Audio Research Lab, New York University<br> &nbsp;&nbsp;&nbsp; 2. Department of Electrical and Computer Engineering, New York University<br> &nbsp;&nbsp;&nbsp; 3. Department of Informatics, New Jersey Institute of Technology<br> &nbsp;&nbsp;&nbsp; 4. Interactive Audio Lab, Northwestern University<br> &nbsp;&nbsp;&nbsp; 5. Center for Urban Science and Progress, New York University<br> &nbsp;&nbsp;&nbsp; 6. Department of Computer Science and Engineering, New York University</p> <p><br> <strong>Publication</strong></p> <p>If you use this data in your work, please cite the following paper, which introduced this dataset:</p> <p>[1] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021. [<a href="https://arxiv.org/pdf/2105.02911">pdf</a>]</p> <p><br> <strong>Description</strong></p> <p>SONYC-Backgrounds is an open dataset of recordings of urban background noise obtained from the <a href="https://wp.nyu.edu/sonyc/">SONYC</a> acoustic sensor network [2]. This dataset was developed with the goal of synthesizing soundscapes with a diverse set of realistic sounding background activity, for use in developing and evaluating machine listening systems in urban settings.</p> <p><br> <strong>Data acquisition</strong></p> <p>The provided audio has been acquired using the <a href="https://wp.nyu.edu/sonyc/">SONYC</a> acoustic sensor network for urban noise pollution monitoring [2]. Over 50 different sensors have been deployed in New York City. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings.</p> <p><br> <strong>Recording selection</strong></p> <p>From the large collection of audio recordings acquired in 2017, we obtain a much smaller subset of likely background recordings. We first process the dataset using a sensor fault detector to filter out recordings with artifacts caused by hardware failures in the sensors. The sensor fault detector is a random forest, trained with a small collection of audio examples using active learning [3].</p> <p>We then determine if a recording is background or not using an urban sound classifier trained to detect the presence of sources of interest to urban noise pollution monitoring [4, 5]. We use the classifier to find recordings that *do not* contain the sound classes of interest. The classifier model is a multi-layer perception with two hidden layers, which takes as input an OpenL3 embedding [6]&nbsp; for a 1 s clip of audio and produces multi-label prediction probabilities for each class. This model is nearly identical to the one used for the <a href="http://dcase.community/challenge2019/task-urban-sound-tagging">DCASE 2019 Challenge Urban Sound Tagging Task</a> baseline model, aside from the addition of an extra hidden layer.</p> <p>Predictions for entire recordings are obtained by max-pooling the predictions for each class across time. A recording is considered background if the probabilities of the target classes fall below their respective detection thresholds, i.e. no target classes are detected. The classifier was trained on the SONYC-UST v1 dataset [4], and the detection thresholds for each class were tuned to correspond to 70% <em>negative</em> recall (true negative rate) on the test set to increase the likelihood that recordings are background.</p> <p>After this selection process, we obtain 441 background clips.</p> <p>&nbsp;</p> <p><strong>Metadata</strong></p> <p>To maintain privacy, the recordings in this release have been distributed in time and location, and recording times have been quantized to the hour. Sensor IDs are consistent with those SONYC-UST dataset [4]. The corresponding location of the sensors can be found in the SONYC-UST v2 dataset [5], though these locations have been mapped to the &quot;block&quot; level to maintain privacy. See the <a href="http://dcase.community/challenge2020/task-urban-sound-tagging-with-spatiotemporal-context">DCASE 2020 Challenge Urban Sound Tagging with Spatiotemporal Context Task page</a> for more information on the metadata.</p> <p><br> <strong>Data splits</strong></p> <p>The dataset is partitioned into a train/valid/test split of roughly 60/20/20, using a simple greedy method to assign sensors to subsets.</p> <p><br> <strong>Files</strong></p> <p>The dataset directory contains the directories `train`, `valid`, and `test` for each of the respective data subsets. Each directory contains recordings, with the file format: `&lt;sensor-id&gt;_&lt;year&gt;-&lt;month&gt;-&lt;day&gt;_&lt;hour&gt;_&lt;instance-num&gt;.wav`, where `&lt;instance-num&gt;` is used to distinguish recordings from the same sensor occurring during the same hour. Aside from `&lt;year&gt;`, each of these fields in the format are lead zero padded to two places (i.e. `printf` format `&quot;%02d&quot;`).</p> <p><br> <strong>Conditions of use</strong></p> <p>Dataset created by Aurora Cramer, Mark Cartwright, Fatemeh Pishdadian, and Juan Pablo Bello.</p> <p>The SONYC-Backgrounds dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license: <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an &ldquo;as is&rdquo; basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-Backgrounds dataset or any part of it.</p> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If you have any questions, comments, or concerns, please direct correspondence to Aurora Cramer (aurora (dot) linh (dot) cramer (at) gmail (dot) com).</p> <p><br> <strong>References and Links</strong></p> <p>[1] Cramer, A., Cartwright, M., Pishdadian, F., and Bello, J.P. Weakly Supervised Source-Specific Sound Level Estimation in Noisy Soundscapes. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021.</p> <p>[2] Bello, J. P., Silva, C., Nov, O., Dubois, R. L., Arora, A., Salamon, J., C. Mydlarz, and Doraiswamy, H. (2019). Sonyc: A system for monitoring, analyzing, and mitigating urban noise pollution. Communications of the ACM, 62(2), 68-77.</p> <p>[3] Wang, Y., Mendez, A.E.M., Cartwright, M., and Bello, J.P. Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.</p> <p>[4] Cartwright, M., Mendez, A.E.M., Cramer, A., Lostanlen, V., Dove, G., Wu, H., Salamon, J., Nov, O., and Bello, J.P. SONYC Urban Sound Tagging (SONYC-UST): A Multilabel Dataset from an Urban Acoustic Sensor Network. In Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2019.</p> <p>[5] Cartwright, M., Cramer, A., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., and Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2020.</p> <p>[6] Look, Listen and Learn More: Design Choices for Deep Audio Embeddings<br> Cramer, A., Wu, H.-H., Salamon J., and Bello. J.P. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019.</p> <p><br> <strong>Acknowledgements</strong></p> <p>We would like to thank <a href="https://wp.nyu.edu/sonyc/people/">all those involved in the SONYC project</a>. This work is partially supported by National Science Foundation <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1633259">award 1633259</a> and <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">award 1544753</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Dataset for "Improving data quality of low-cost light-scattering PM sensors: Towards automatic air quality monitoring in urban environments"

<p>The dataset contains the data used in the article "Improving data quality of low-cost light-scattering PM sensors: Towards automatic air quality monitoring in urban environments".</p> <p>A low-cost monitoring system composed of 14 monitoring stations was positioned at the official monitoring station of Torino Rubino in the city of Turin (Italy). The official station is managed by the environmental agency ARPA Piemonte.</p> <p>Each low-cost station contains four low-cost light-scattering PM sensors (Honeywell HPMA115S0-XXX), one temperature and relative humidity sensor (DHT22), and one atmospheric pressure sensor (BME/BMP280).<br>The sampling time of the PM sensors was set to one second, while the other sensors generated measurements every 3-4 seconds.</p> <p>The official monitoring station uses both a gravimetric and a beta attenuation instrument for measuring PM.</p> <p>The data contained in this dataset was collected from October 2020 to November 2021. It contains the PM2.5, relative humidity, and temperature measurements of the low-cost monitoring system and the official measurements of the beta attenuation device.</p> <p>Measurements of low-cost sensors are expressed in UTC, while official measurements are expressed in UTC+1.</p> <p>Official PM measurements can be also found at https://aria.ambiente.piemonte.it/qualita-aria/dati.</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

Hydrostatic pressure sensor logger data for floodX urban flooding experiments

<p>Film of pressure sensor loggers for floodX dataset</p>

opencc-by-4.0Dec 2016View details →
zenodo24/100

Evaluation of optical particulate matter sensors under realistic conditions of strong and mild urban pollution

<p>In this paper we evaluate characteristics of three optical particulate matter sensors/sizers (OPS): high-end spectrometer 11-D (Grimm, Germany), low-cost sensor OPC-N2 (Alphasense, United Kingdom) and in-house developed MAQS which is based on another low-cost sensor - PMS5003 (Plantower, China), under realistic conditions of strong and mild urban pollution. Results were compared against a reference gravimetric system, based on Gemini (Dadolab, Italy), 2.3 m3/h air sampler, with two channels (simultaneously measuring PM2.5 and PM10 concentrations). The measurements were performed in Sarajevo, the capital of Bosnia-Herzegovina, from December 2019 until May 2020. This interval is divided into period 1 - strong pollution (December 2019 - March 2020) and period 2 - mild pollution (March 2020 - May 2020). The city of Sarajevo is one of the most polluted cities in Europe in terms of aerosols: the average concentration of PM2.5 during the period 1 was 83 ug/m3, with daily aveverage values exceeding 500 ug/m3. During period 2, the average concentration of PM2.5 was 20 ug/m3. These conditions represent a good opportunity to<br> test optical devices against reference instrument in a wide range of ambient particulate matter (PM) concentrations. The effect of an in-house developed diffusion dryer for 11-D is discussed as well. In order to analyze the mass distribution of aerosols, a scanning mobility particle sizer (SMPS) spectrometer, which together with the 11-D spectrometer gives the full spectrum from nanoparticles of diameter 10 nm to coarse particles of diameter 35 um, was used. All tested devices showed excellent correlation with the reference instrument in period 1, with R^2 values between 0.90 and 0.99 for daily average PM concentrations. However, in period 2, where the range of concentrations was much narrower, R^2 values decreased significantly, to values from 0.28 to 0.92. We have also included results of a 13.5 month long-term comparison of our MAQS sensor with a nearby beta attenuation monitor (BAM) 1020 (Met One Instruments, USA) operated by the United States Environmental Protection Agency (US EPA), which showed similar correlation and no observable change of performance over time.</p> <p>This dataset contains:<br> 1. preliminary.csv - hourly average PM2.5 and PM10 values for 8 MAQS sensors during the preliminary test period<br> 2. meteo.csv - ambient air temperature and relative humidity, hourly average values<br> 3. PM_reference_timeseries.csv - timeseries of data from reference air sampler, including PM2.5, PM10, average air temperature, average air pressure, actual air volume and mass difference (loaded-blank) of each filter<br> 4. main.csv - comparisons of 11-D, OPC-N2 and MAQS sensor against the reference during the campaign<br> 5. dryer_effect.csv - hourly average data of ambient air vs 11-D internal, temperature and relative humidity. Dryer was installed on 03/12/2020 15:30<br> 6. SMPS.zip (4 files) - wide-range spectrometer (SMPS with 11-D) csv files for density plots: time, Log(Dp) and concentration values (both absolute and normalized), based on hourly average data<br> 7. histograms.zip (8 files) - selected wide-range histograms, based on hourly average data from SMPS and 11-D<br> 8. BAM.zip (5 files) - long-term comparisons of MAQS sensor with a nearby BAM. Hourly (BAM1.csv), daily (BAM2.csv) and monthly (BAM3.csv) average values and comaparison of two MAQS sensors 300 meters apart, hourly (BAM4.csv) and daily (BAM5.csv) values<br> &nbsp;</p>

opencc-by-4.0Jun 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record