Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

206

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

206 results for “air quality”

Learn how ShareScore rates datasets ↗
zenodo44/100

Britain Breathing 2016-2019 Air Quality and Meteorological Regional Estimates Dataset

<p>This data set is a collection of estimated daily mean and maximum values for a range of air quality and meterological measurements and model forecasts for the <em>UK and crown dependencies</em> postcode districts (e.g. &#39;AB&#39;) for the years 2016-2019, inclusive.</p> <p>The paper describing this dataset is available here:&nbsp;<a href="https://www.nature.com/articles/s41597-022-01135-6">https://www.nature.com/articles/s41597-022-01135-6</a></p> <p>The data uses a &#39;concentric regions&#39;&nbsp;method to estimate the measurement for all regions, as follows. If measurements exist within the region, the mean of those measurements is used, if not, then a ring of neighbouring postcode regions are selected, and the mean of their measurement values used. If no measurement sites/data are found in the first ring, the process continues, taking the next&nbsp;ring of postcode district regions, working outwards until one or more sensors are found in a ring.&nbsp; As well as the measurement estimations, the number of rings required to find site data and make the estimations is also published.&nbsp;<strong>As a result, please note that estimations with higher ring counts (&#39;rings&#39;) are likely to be calculated from more distant sensors. This distance depends upon the size of the postcode regions surrounding the location being estimated. Please use the ring count (&#39;rings&#39;) to limit/filter estimations based on your required level of confidence.</strong><br> <br> The meteorological, pollen and air quality measurement data used to make the regional estimations can be found at&nbsp;<a href="https://zenodo.org/record/4416028#.YABxNnX7RhF">this Zenodo archive</a>.&nbsp; The data&nbsp;there contains Temperature, Relative Humidity, and Pressure data, downloaded from the Met Office MIDAS archives via the MEDMI server (https://www.data-mashup.org.uk/). Also downloaded from the MEDMI server are daily pollen measurements for the UK. PM10, PM2.5, NO2, NOx (as NO2), O3, and SO2 measurements from the DEFRA AURN network, and also model forecasts of the same made using the EMEP model.</p> <p>The code used to make the&nbsp;estimations is&nbsp;available at <a href="https://zenodo.org/record/4518866">this Zenodo archive</a>.</p> <p>The postcode data in postcode_district_data.csv are collated from several sources:&nbsp;</p> <ul> <li><a href="https://www.doogal.co.uk/UKPostcodes.php">https://www.doogal.co.uk/UKPostcodes.php</a>&nbsp;(population figures for the UK (UK Census 2011))</li> <li><a href="https://www.freemaptools.com/download-uk-postcode-outcode-boundaries.htm">https://www.freemaptools.com/download-uk-postcode-outcode-boundaries.htm</a>&nbsp;(postcode boundary polygons for UK and crown dependancies)</li> <li><a href="https://www.gov.gg/population">https://www.gov.gg/population</a>&nbsp;(Guernsey (GY) population data for end June 2020)&nbsp;</li> <li><a href="https://www.gov.je/Government/JerseyInFigures/Population/Pages/Population.aspx">https://www.gov.je/Government/JerseyInFigures/Population/Pages/Population.aspx</a>&nbsp;(Jersey (JE) population data for end 2019)&nbsp;</li> <li><a href="https://www.gov.im/media/1369690/isle-of-man-in-numbers-july-2020.pdf">https://www.gov.im/media/1369690/isle-of-man-in-numbers-july-2020.pdf</a>&nbsp;(Isle of Man&nbsp;(IM) population data for April 2016)</li> </ul> <p>The data-set is presented in CSV format, as six files:</p> <ol> <li>postcode_district_data.csv: location metadata (region_id, geometry, description, population, country)</li> <li>regional_site_counts.csv: a table showing the number of sites for each measurement (columns), for each region_id (rows). region_id&#39;s match those in the postcode_district_data.csv file.</li> <li>turing_regional_estimates_aq_daily_met_pollen_pollution_imputed_data.csv: uses imputed site data (timestamp, region_id, ...[measurement name, rings]) (&#39;rings&#39; is the number of rings required to make the estimation)</li> <li>turing_regional_estimates_aq_daily_met_pollen_pollution_original_data.csv: uses original site data (timestamp, region_id, ...[measurement name, rings]) (&#39;rings&#39; is the number of rings required to make the estimation)</li> <li>turing_regional_estimates_aq_loc_type_daily_imputed_data.csv: uses imputed site data. Air quality regional estimates are calculated using specific AQ site location types* separately.&nbsp;(To prevent,&nbsp;for example, &#39;Traffic Urban&#39; type sites being used to estimate&nbsp;&#39;non-traffic&#39; or rural regions.)</li> <li>turing_regional_estimates_aq_loc_type_daily_original_data.csv: uses original data.&nbsp;Air quality regional estimates are calculated using specific AQ site location types* separately.&nbsp;(To prevent,&nbsp;for example, &#39;Traffic Urban&#39; type sites being used to estimate&nbsp;&#39;non-traffic&#39; or rural regions.)</li> </ol> <p>* Air quality site types:&nbsp;</p> <ul> <li>Industrial: comprises &#39;urban industrial&#39; (9 sites) and suburban industrial (2 sites)</li> <li>&#39;Rural background&#39; (14 sites)</li> <li>&#39;Urban background&#39; (48 sites)</li> <li>&#39;Urban traffic&#39; (47 sites)</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Britain Breathing 2020 Air Quality and Meteorological Regional Estimates Dataset

<p>This data set is a collection of estimated daily mean and maximum values for a range of air quality and meterological measurements and model forecasts for&nbsp;UK postcode districts (e.g. &#39;AB&#39;) for the year 2020.</p> <p>The data uses a &#39;concentric regions&#39;&nbsp;method to estimate the measurement for all regions, as follows. If measurements exist within the region, the mean of those measurements is used, if not, then a ring of neighbouring postcode regions are selected, and the mean of their measurement values used. If no measurement sites/data are found in the first ring, the process continues, taking the next&nbsp;ring of postcode district regions, working outwards until one or more sensors are found in a ring.&nbsp; As well as the measurement estimations, the number of rings required to find site data and make the estimations is also published.&nbsp;<strong>As a result, please note that estimations with higher ring counts (&#39;rings&#39;) are likely to be calculated from more distant sensors. This distance depends upon the size of the postcode regions surrounding the location being estimated. Please use the ring count (&#39;rings&#39;) to limit/filter estimations based on your required level of confidence.</strong></p> <p>The meteorological, pollen and air quality measurement data used to make the regional estimations can be found at&nbsp;<a href="https://zenodo.org/record/4740965#.YPWJf3VKhhF">this Zenodo archive</a>.&nbsp; The data&nbsp;there contains Temperature, Relative Humidity, and Pressure data, downloaded from the Met Office MIDAS archives via the MEDMI server (https://www.data-mashup.org.uk/). Also downloaded from the MEDMI server are daily pollen measurements for the UK. PM10, PM2.5, NO2, NOx (as NO2), O3, and SO2 measurements from the DEFRA AURN network, and also model forecasts of the same made using the EMEP model.</p> <p>The code used to make the&nbsp;estimations is&nbsp;available at&nbsp;<a href="https://zenodo.org/record/4518866">this Zenodo archive</a>.</p> <p>The data-set is presented in CSV format, as two files:</p> <ol> <li>turing_regional_estimates_aq_daily_met_pollen_pollution_original_data.csv: uses original site data (timestamp, region_id, ...[measurement name, rings]) (&#39;rings&#39; is the number of rings required to make the estimation)</li> <li>turing_regional_estimates_aq_loc_type_daily_original_data.csv: uses original data.&nbsp;Air quality regional estimates are calculated using specific AQ site location types* separately.&nbsp;(To prevent,&nbsp;for example, &#39;Traffic Urban&#39; type sites being used to estimate&nbsp;&#39;non-traffic&#39; or rural regions.)</li> </ol> <p>* Air quality site types:&nbsp;</p> <ul> <li>Industrial: comprises &#39;urban industrial&#39; (9 sites) and suburban industrial (2 sites)</li> <li>&#39;Rural background&#39; (14 sites)</li> <li>&#39;Urban background&#39; (48 sites)</li> <li>&#39;Urban traffic&#39; (47 sites)</li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

AgrImOnIA: Open Access dataset correlating livestock and air quality in the Lombardy region, Italy

<p>The AgrImOnIA dataset is a comprehensive dataset relating air quality and livestock (expressed as the&nbsp;density of bovines and swine bred) along with weather and other variables. The AgrImOnIA Dataset represents the first step of the <a href="http://www.agrimonia.net">AgrImOnIA project</a>. The purpose of this dataset is to give the opportunity to assess the impact of agriculture on air quality in Lombardy through statistical techniques capable of highlighting the relationship between the livestock sector and air pollutants concentrations.</p> <p>The building process of the dataset is detailed in the <strong>companion paper:</strong></p> <p>A. Fass&ograve;, J. Rodeschini, A. Fusta Moro, Q. Shaboviq, P. Maranzano, M. Cameletti, F. Finazzi, N. Golini, R. Ignaccolo, and P. Otto&nbsp;(2023). Agrimonia: a dataset on livestock, meteorology and air quality in the Lombardy region, Italy.&nbsp;<em>SCIENTIFIC DATA</em>, 1-19.</p> <p>available <a href="https://rdcu.be/c7T9H">here</a>.</p> <p>This dataset is a collection of estimated daily values for a range of measurements of different dimensions as: air quality, meteorology, emissions, livestock animals and land use. Data are related to Lombardy and the surrounding area for&nbsp;2016-2021, inclusive. The surrounding area is obtained by applying a 0.3&deg; buffer on Lombardy borders.</p> <p>The data uses several aggregation and interpolation methods to estimate the measurement for all days.</p> <p>The files in the record, renamed according to their version (es. .._v_3_0_0),&nbsp;are:</p> <ul> <li> <p>Agrimonia_Dataset.csv(.mat and .Rdata) which is built by joining the daily time series related to the AQ, WE, EM, LI and LA variables. In order to simplify access to variables in the Agrimonia dataset, the variable name starts with the dimension of the variable, i.e., the name of the variables related to the AQ dimension start with &#39;AQ_&#39;. This file is archived also in the&nbsp;format for MATLAB and R software.&nbsp;</p> </li> <li> <p>Metadata_Agrimonia.csv which provides further information about the Agrimonia variables: e.g. sources used, original names of the variables imported, transformations applied.</p> </li> <li> <p>Metadata_AQ_imputation_uncertainty.csv which contains the daily uncertainty estimate of the imputed observation for the AQ to mitigate missing data in the hourly time series.&nbsp;&nbsp;</p> </li> <li> <p>Metadata_LA_CORINE_labels.csv which contains the label and the description associated with the CLC class.&nbsp;&nbsp;</p> </li> <li> <p>Metadata_monitoring_network_registry.csv which contains all details about the AQ monitoring station used to build the dataset. Information about air quality monitoring stations include: station type, municipality code, environment type, altitude, pollutants sampled and other. Each row represents a single sensor.</p> </li> <li> <p>Metadata_LA_SIARL_labels.csv which contains the label and the description associated with the SIARL class.</p> </li> <li> <p>AGC_Dataset.csv(.mat and .Rdata)&nbsp;that&nbsp;includes daily data of almost all variables available in&nbsp;the Agrimonia&nbsp;Dataset (excluding AQ variables)&nbsp;on an&nbsp;equidistant grid covering the Lombardy region and its surrounding area.&nbsp;</p> </li> </ul> <p>The Agrimonia dataset can be reproduced using&nbsp;the code available at the GitHub page: <a href="https://github.com/AgrImOnIA-project/AgrImOnIA_Data">https://github.com/AgrImOnIA-project/AgrImOnIA_Data</a></p> <p><strong>UPDATE 31/05/2023</strong>&nbsp;<strong>- NEW RELEASE - V 3.0.0</strong></p> <p>A new version of the dataset is released:&nbsp;Agrimonia_Dataset_v_3_0_0.csv (.Rdata and .mat), where variable&nbsp;<em>WE_rh_min, WE_rh_mean and WE_rh_max&nbsp;</em>have been recomputed due to some bugs<em>.</em></p> <p>In addition, two new columns are added, they are&nbsp;<em>LI_pigs_v2 and LI_bovine_v2&nbsp;</em>and represents the density of the pigs and bovine (expressed as animals per kilometer squared) of a square of size ~ 10 x 10 km centered at the station localisation.</p> <p>A new dataset is released: the Agrimonia Grid Covariates (AGC) that includes daily information for the period from 2016 to 2020 of almost all variables within the Agrimonia Dataset on a equidistant grid containing the Lombardy region and its surrounding area. The AGC does not include AQ variables as they come from&nbsp;the monitoring stations that are irregularly spread over the area considered.</p> <p><strong>UPDATE 11/03/2023</strong>&nbsp;<strong>- NEW RELEASE - V 2.0.2</strong></p> <p>A new version of the dataset is released:&nbsp;Agrimonia_Dataset_v_2_0_2.csv (.Rdata), where variable&nbsp;<em>WE_tot_precipitation&nbsp;</em>have been recomputed due to some bugs<em>.</em></p> <p>A new version of the metadata is available:&nbsp;Metadata_Agrimonia_v_2_0_2.csv where the spatial resolution of the variable <em>WE_precipitation_t&nbsp;</em>is corrected.</p> <ul> </ul> <p><strong>UPDATE 24/01/2023</strong>&nbsp;<strong>- NEW RELEASE - V 2.0.1</strong></p> <p>minor bug fixed</p> <p><strong>UPDATE 16/01/2023</strong>&nbsp;<strong>- NEW RELEASE - V 2.0.0</strong></p> <p>A new version of the dataset is released, Agrimonia_Dataset_v_2_0_0.csv (.Rdata) and Metadata_monitoring_network_registry_v_2_0_0.csv.&nbsp;Some minor points have been addressed:</p> <ul> <li>Added&nbsp;values for <em>LA_land_use</em> variable for Switzerland stations (in Agrimonia Dataset_v_2_0_0.csv)</li> <li>Deleted&nbsp;incorrect values for <em>LA_soil_use</em> variable for stations outside Lombardy region during 2018 (in Agrimonia Dataset_v_2_0_0.csv)</li> <li>Fixed duplicate&nbsp;sensors corresponding&nbsp;to the same pollutant within the same&nbsp;station<em>&nbsp;</em>(in&nbsp;Metadata_monitoring_network_registry_v_2_0_0.csv)</li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo44/100

NO2, O3, PM10 and PM2.5 concentrations - Daily geographical aggregates at ZIP-code level from CAMS European Air Quality Re-analyses.

<p>This dataset offers daily aggregated measurements of air pollutants &ndash; NO2, O3, PM10, and PM2.5 &ndash; across distinct ZIP-code areas in Germany. The temporal coverage spans from January 1, 2013, to December 31, 2022, providing a comprehensive temporal context for analyzing long-term air quality dynamics.</p> <p>Each daily entry comprises key statistical descriptors, encompassing mean, maximum, minimum, and standard deviation values of pollutant concentrations specific to each ZIP-code area. Additionally, for O3, the dataset includes an eight-hour rolling mean daily maximum.</p> <p>Spatial reference is established via shapefiles provided by ESRI Deutschland (<a href="https://opendata-esri-de.opendata.arcgis.com/datasets/5b203df4357844c8a6715d7d411a8341_0">https://opendata-esri-de.opendata.arcgis.com/datasets/5b203df4357844c8a6715d7d411a8341_0</a>). These shapefiles link the air quality data to precise ZIP-code areas .</p> <p>The concentration data spanning from 2018 to 2022 originate from the European Air Quality Reanalyses dataset of the Atmosphere Data Store (ADS), an initiative by the Copernicus Atmosphere Monitoring Service (CAMS). Accessible via <a href="https://ads.atmosphere.copernicus.eu/cdsapp#!/dataset/cams-europe-air-quality-reanalyses?tab=doc">https://ads.atmosphere.copernicus.eu/cdsapp#!/dataset/cams-europe-air-quality-reanalyses?tab=doc</a>, this dataset offers a robust foundation for assessing air quality. For the years 2013 to 2017, data were previously obtained from a former download platform for the same dataset. Important: in future all data will be migrated to the Atmosphere Data Store (ADS) platform.</p> <p>The native resolution of the CAMS data is 0.1&deg; x 0.1&deg; spatially and hourly temporally. To enhance spatial accuracy, the spatial resolution was virtually increased by a factor of 5 using bilinear interpolation, resulting in a refined grid. The daily mean concentrations were subsequently computed for this augmented grid.</p> <p>Aggregated statistics were derived for each ZIP-code polygon, employing all grid cells intersecting with the polygons. The computation was based on the proportion of cell area included within the respective polygons.</p> <p>This dataset constitutes a valuable resource for conducting ecologically designed epidemiological studies, as it facilitates the exploration of potential associations between air quality and health trends across broad geographical areas.</p> <p>Generated using Copernicus Atmosphere Monitoring Service Information 2013-2022</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Low-energy Museum Storage Buildings: Climate, Energy Consumption and Air Quality. Data Set for Final Data Report

<p>The 43 txt-files included in this dataset relate to the report: Ryhl-Svendsen, Jensen, B&oslash;hm, and Klenz Larsen (2012): <em>Low-energy Museum Storage Buildings: Climate, Energy Consumption and Air Quality. UMTS Research Project 2007</em>&ndash;<em>2011: Final Data Report</em>, Kgs. Lyngby: National Museum of Denmark, 122&nbsp;pp.</p> <p>The document <a href="https://zenodo.org/api/files/145584b0-46b5-4341-8b02-7dfea90fa97c/00_List-of-data-files.pdf?versionId=a3e9691f-6e73-4a7b-aaab-c8ccee7c419b">00_List-of-data-files.pdf</a> contain a full list of the data files with&nbsp;a description of their structure and content, and&nbsp;is the key to how the individual data files relate to the report.&nbsp;</p> <p>The research project focussed on four modern museum storage facilities in Denmark, for which the indoor climate, air quality, and the energy consumption of the climate control systems was measured at several locations, typically for a period of between two and four years. The storage facilities were Museum of Southwest Jutland&rsquo;s storage building in Ribe (&lsquo;Ribe&rsquo;), The Shared Storage Facility at The Centre for Preservation of Cultural Heritage in Vejle (&lsquo;Vejle&rsquo;), The Joint Storage Facility for museums in East Jutland/ Museum &Oslash;stjylland (&lsquo;Randers&rsquo;), and from The National Museum of Denmark the storage building Hall P at the &Oslash;rholm Storage Facility (&lsquo;&Oslash;rholm&rsquo;). For description of the sites, monitoring campaigns, and graphed data, the report should be consulted.</p> <p>For completeness, the report is included with the dataset (<a href="https://zenodo.org/api/files/145584b0-46b5-4341-8b02-7dfea90fa97c/Report_low-energy-museum-storage-buildings.pdf?versionId=44097d39-775b-4031-9e07-6978c68912a9">Report_low-energy-museum-storage-buildings.pdf</a>).</p>

opencc-by-4.0Oct 2023View details →
edi44/100

High-frequency water quality and air parameters from large lake Võrtsjärv, Estonia: 2010-2019

This high frequency water-air dataset was collected from large shallow Lake Võrtsjärv (Estonia) with an automated lake monitoring buoy system and was used in the analyses described in the manuscript by Thayne, M.W., B. Kraemer, J.P. Mesman, A. Laas, E. de Eyto, B. W. Ibelings, R. Adrian. “Trophic state effects on antecedent lake characteristics shape the resistance and resilience of lakes following extreme storms”. Dataset contains every 10-minute measurements from the open water periods in years 2010-2019. Underwater measurements were conducted with Yellow Springs Instrument multiparameter sonde model YSI 6600 V2-4. Underwater measurements contain the collected values of pH, dissolved oxygen, dissolved oxygen saturation, water temperature, specific conductance, water turbidity, chlorophyll concentration (calculated by the sensor from fluorescence values), and the abundance of the bluegreen algae (calculated by the sensor from phycocyanin fluorescence values). Air parameters were collected with the Vaisala multiparameter weather station model WXT520 and contain measurements of wind direction, wind speed, air temperature, atmospheric pressure, and cumulative amount of rain between measurement periods. All air measurements were collected 2 m above from the water surface. Additionally, solar irradiance measurements were collected from the automated buoy system with the Licor pyranometer model LI-200, those measurements contain values in two different columns: average solar irradiance in W/m2 and average photon flux density values in micromoles per square meter per second.

openCC (other)Apr 2022View details →
zenodo40/100

The features of the selected papers in the field of air quality prediction

<p>The table is a part of a submitted manuscript (Iskandaryan, D., Ramos, F., &amp; Trilles, S. The Role of Datasets in Air Quality Prediction. &nbsp;Submitted to Atmosphere.)&nbsp;and includes the following features extracted from the selected papers: <em>Year, Case Study, Prediction Target, Dataset Type, Data Rate, Period (Days), Open Data, Algorithm, Time Granularity and Evaluation Metric</em>. The relevant papers&nbsp;were selected from a systematic review in <em>Air Quality Prediction Using Machine Learning Technologies. </em>The works were&nbsp;queried in Association for Computing Machinery, IEEE Xplore, Scopus and Web of Science databases using the following query: (&quot;machine learning&quot;) AND (&quot;prediction&quot;OR &quot;forecast&quot;) AND (&quot;air quality&quot; OR &quot;air pollution&quot;), which was being applied to title, abstract and keywords. After filtering the results guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses, &nbsp;ninety-three papers were selected. The goal of this review is to understand which features are used in the field, in particular to answer the following questions:&nbsp;&nbsp;1) What types of datasets are used to improve air quality predictions?; and 2) What characteristics of the dataset are important for efficient and effective air quality forecasting?&nbsp;<br> Twenty-six datasets were used by the authors as supplemental air quality data in order to predict air quality more accurately. Those datasets are: &quot;MET&quot;- meteorological data; &quot;Spatial&quot;- topographical characteristics, the locations of the stations; &quot;Temporal&quot;-includes the day of the month, day of the week, the hour of the day; &quot;AOD&quot;- aerosol optical depth; &quot;Social Media&quot;- microblog data; &quot;Traffic&quot;; &quot;PBL Height&quot;- planetary boundary layer height; &nbsp;&quot;Land Use&quot;; &quot;BEV&quot;- Built Environment Variables; &quot;UV Index&quot;; &quot;SP&quot;- Sound Pressure; &quot;PD&quot;-Population Density; &quot;Human Movements&quot;- floating population and estimated traffic volume; &nbsp;&quot;Altitude&quot;; &nbsp;&quot;OMI-SO2&quot;-Satellite-retrieved SO2 from Ozone Monitoring Instrument-SO2; &quot;PPS&quot;- Pollution Point Source; &quot;TS&quot;-Transportation Source; &quot;WFD&rsquo;&quot;- weather forecast data; &quot;POI Distribution&quot;; &quot;FAPE&quot;- factory air pollution emission; &quot;RND&quot;- Road Network Distribution; &quot;Elevation&quot;; &quot;AEI&quot;- Anthropogenic Emission Inventory; &quot;NDVI&quot;; &quot;Chemical&quot;- chemical component forecast data (organic carbon, black carbon, sea salt, etc.); &quot;Emission&quot;.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Unraveling a black box: An open-source methodology for the field calibration of small air quality sensors

<p>This repository contains&nbsp;data for the manuscript:&nbsp;&quot;Unraveling a black box: An open-source methodology for the field calibration of small air quality sensors.&quot;</p> <p>&nbsp;</p> <p>This includes:</p> <p>Raw data from the low-cost prototype EarthSense Zephyrs, as well as raw data from reference instrumentation.</p> <p>SC stands for &quot;Summer Campaign&quot; and WC stands for &quot;Winter Campaign&quot;, denoting the two different campaigns assessed in this study.</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>The last two decades have seen substantial technological advances in the development of low-cost air pollution instruments using small sensors. While their use continues to spread across the field of atmospheric chemistry, challenges remain in ensuring data quality and comparability of calibration methods. This study introduces a seven-step methodology for the field calibration of low-cost sensors using reference instrumentation with user-friendly guidelines, open access code, and a discussion of common barriers to such an approach. The methodology has been developed and is applicable for gas-phase pollutants, such as for the measurement of nitrogen dioxide (NO<sub>2</sub>) or ozone (O<sub>3</sub>). A full example of the application of this methodology to a case study in an urban environment using both Multiple Linear Regression (MLR) and the Random Forest (RF) machine-learning technique is presented with relevant R code provided, including error estimation. In this case, we have applied it to the calibration of metal oxide gas-phase sensors (MOS). Results reiterate previous findings that MLR and RF are similarly accurate, though with differing limitations. The methodology presented here goes a step further than most studies by including explicit, transparent steps for addressing model selection, validation, and tuning, as well as addressing the common issues of autocorrelation and multicollinearity. We also highlight the need for standardized reporting of methods for data cleaning and flagging, model selection and tuning, and model metrics. In the absence of a standardized methodology for the calibration of low-cost sensors, we suggest a number of best practices for future studies using low-cost sensors to ensure greater comparability of research.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

SAQI: An Ontology based Knowledge GraphPlatform for Social Air Quality Index

<p>This dataset consists of all contributions made by Social AQI (SAQI)&nbsp;project. The description of dataset is as below -</p> <p>Local Sensor Data (hyperlocal-air-quality-sensor-data) - contains all sensors values recorded through local neighbourhood sensors throught the length of the project&nbsp;<br> Locations for all these sensors are as below - In Najafgarh, Delhi, India : Jharoda Kalan, Nangli Dairy and&nbsp;DTC Bus terminal.<br> In Okhla : Sanjay Colony, Tekhand, Shaheen Bagh.</p> <p>Data from <a href="https://cpcb.nic.in/">Central Pollution Control Board</a>&nbsp;(central-air-quality-sensor-data) - Najafgarh_CPCB.csv,&nbsp;Okhla_CPCB.csv : Contains data provided by CPCB from Najafgarh,Delhi&nbsp;and Oklha, Delhi</p> <p><br> PollutionODP.owl : Ontology Design Pattern for pollution -&nbsp;http://ontologydesignpatterns.org/wiki/Submissions:Pollution.<br> <br> Ontology&nbsp;: SAQI ontology as triples (ttl), xml (rdf) and json-ld (json) serialization format<br> Ontology documentation : ontology/diagram contains figures describing ontology, ontology/documentation/saqi.html contains LODE documentation for the ontology</p> <p><br> ethnographic-survey-data - anonymized survey responses for initial pollution perception and literacy survey as well as SAQI app feedback survey.</p> <p>SHACL-shapes - for validating against SAQI ontology.</p> <p>&nbsp;sparql-queries - sample queries to run on our ontology.</p> <p>setup-rdf-store-script - script to setup rdf store with given data using rml mapper.</p> <p>&nbsp;</p> <p><br> &nbsp;&nbsp; &nbsp;</p>

openapache2.0Dec 2022View details →
zenodo40/100

TURDATA: a database of low-cost air quality and remote sensing measurements for the validation of micro-scale models in the real Prague urban environments

<p><strong>README</strong></p> <p>TURDATA is a supplementary data set for the TURBAN project Prague observation campaign described in the manuscript Bauerov&aacute; et al. 2024 (submitted for publication). The measurement campaign was focused on air pollution and meteorological measurement, including vertical profiles in selected part of Prague city centre called here as Legerova domain. Within this area, one professional meteorological station (MS) Prague Karlov and one reference traffic air quality monitoring (AQM) station Prague 2-Legerova (classified as traffic hotspot) are located. To gain high spatial and temporal resolution data, the supplementary measurement network was established, which consisted of:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 20 combined low-cost sensor (LCS) stations for monitoring of PM<sub>10</sub>, PM<sub>2.5</sub>, NO<sub>2</sub> and O<sub>3</sub> concentrations (using Plantower PMS7003 particle counters and Envea Cairsense electrochemical sensors) placed in different sites and different height levels AGL (higher = H, lower = L),</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1 mobile telescopic meteorological mast for measuring temperature, relative humidity, wind velocity and direction and air pressure (using 2D ultrasonic anemometer Gill WindSonic 60 and weather station Gill MetConnect THP),</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1 MTP-5-He microwave radiometer (MWR; Attex) for temperature vertical profile,</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1 StreamLine XR Doppler LIDAR (HALO Photonics) for wind vertical profile. &nbsp;</p> <p>The main Legerova campaign lasted from 30 May 2022 to 28 March 2023 with some exceptions (see <em>TURDATA_metadata.xlsx</em> with all details). Because LCSs are known for their highly variable measurement quality, before their deployment the Legerova campaign, a sufficiently long-term initial field comparative measurement of all LCSs at RM Prague 4-Libu&scaron; was carried out (lasting from 16/12/2021 to 30/5/2022). The results showed that most of the LCSs were in raw measurement differently zero-shifted against each other and against gaseous reference or aerosol optical equivalent monitors (RMs or EMs).&nbsp; Therefore, the Multivariate Adaptive Regression Splines (MARS) method was applied to calculate corrected LCS concentrations based on initial field comparative measurement complemented by meteorological data from MS Prague Libu&scaron;. To check the quality of raw and MARS corrected LCS concentrations at the end of the measurement campaign, the final comparative field measurement of all LCSs at Prague 4-Libu&scaron; RM station was performed.</p> <p>Therefore, in case of LCSs measurement (both raw and corrected) the important columns of location (measurement placement: RM_Prague_4-Libus and Legerova_domain) and measurement_program (Initial_comparative_measurement, Legerova_campaign and Final_comparative_measurement) were added.</p> <p>In case of PM<sub>10</sub> and PM<sub>2.5</sub> measurement the maximum raw and MARS-corrected concentrations were influenced by temporary pollution episode on 26 July 2022 around 4 a.m. and 9 p.m. (both UTC) caused by aerosol pollution transported from large forest fire in Hřensko&nbsp;(the northern part of the Czech Republic).&nbsp;</p> <p>&nbsp;</p> <p>TURDATA includes the following files:</p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp; <strong>TURDATA_metadata_and_photos.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>TURDATA_metadata.xlsx</em>" with the important list of metadata about devices placement, locations parameters and measurement periods</p> <p>-&nbsp; &nbsp; &nbsp; &nbsp; Folder "<em>Photos_from_Legerova_campaign</em>" with photos from Legerova measurement campaign</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp; <strong>AQ_LCSs_raw_measurement_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>NO2_RAW_LCSs_TURDATA.xlsx</em>" with complete data set of NO<sub>2</sub> raw measured concentrations by all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>O3_RAW_LCSs_TURDATA.xlsx</em>" with complete data set of O<sub>3</sub> raw measured concentrations by all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>PM10_RAW_LCSs_TURDATA.xlsx</em>" with complete data set of PM<sub>10</sub> raw measured concentrations by all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>PM2_5_RAW_LCSs_TURDATA.xlsx</em>" with complete data set of PM<sub>2.5</sub> raw measured concentrations by all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>AQ_LCSs_raw_measurement_TURDATA_readme.txt</em>" with all necessary information for correct data use</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp; <strong>AQ_data_RM_stations_Prague_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>AQ_data_Prague_RM_stations_TURDATA_12-2021_06-2023.xlsx</em>" with air quality data measured by reference AQM stations in Prague</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>AQ_data_RM_stations_Prague_TURDATA_readme.txt</em>" with all necessary information for correct data use</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp; <strong>Meteo_data_Prague_MS_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>Meteo_data_Prague_MS_TURDATA_12-2021_06-2023.xlsx</em>" with meteorological data measured by professional meteorological stations in Prague</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>Meteo_data_Prague_MS_TURDATA_readme.txt</em>" with all necessary information for correct data use</p> <p>5.&nbsp;&nbsp;&nbsp;&nbsp; <strong>AQ_LCSs_MARS-corrected_measurement_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>NO2_COR_LCSs_TURDATA.xlsx</em>" with complete data set of NO<sub>2</sub> MARS-corrected concentrations for all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>O3_COR_LCSs_TURDATA.xlsx</em>" with complete data set of O<sub>3</sub> MARS-corrected concentrations for all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>PM10_COR_LCSs_TURDATA.xlsx</em>" with complete data set of PM<sub>10</sub> MARS-corrected concentrations for all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>PM2_5_COR_LCSs_TURDATA.xlsx</em>" with complete data set of PM<sub>2.5</sub> MARS-corrected concentrations for all LCSs</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>AQ_LCSs_MARS-corrected_measurement_TURDATA_readme.txt</em>" with all necessary information for correct data use and brief description of MARS correction method</p> <p>6.&nbsp;&nbsp;&nbsp;&nbsp; <strong>Meteo-mast_PVK_measurement_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>Meteo-mast_PVK_TURDATA_06-2022_06_2023.xlsx</em>&ldquo; with non-referential meteorological data measured by mobile meteo-mast</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>Meteo-mast_data_PVK_TURDATA_readme.txt</em>" with all necessary information for correct data use</p> <p>7.&nbsp;&nbsp;&nbsp;&nbsp; <strong>MWR_temperature_profile_TURDATA.zip</strong> containing:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>MWR_5min_temperature_TURDATA_02-2022_03-2023.xlsx</em>" with raw temperature vertical profile measurement from microwave radiometer</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>MWR_1hour_temperature_TURDATA.xlsx</em>" with 1-hour averaged temperature vertical profile from microwave radiometer</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>MWR_1hour_TMP_gradient_TURDATA.xlsx</em>" with 1hour temperature gradient calculated from raw temperature profiles measured by microwave radiometer</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>MWR_temperature_profile_TURDATA_readme.txt</em>" with all necessary information for correct data use</p> <p>8.&nbsp;&nbsp;&nbsp;&nbsp; <strong>LIDAR_wind_profile_TURDATA.zip</strong> contains:</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Individual folders "yyyymm&ldquo; -&gt; "yyyymmdd"</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Each daily folder "yyyymmdd" contains files:</p> <p>a)&nbsp;&nbsp;&nbsp;&nbsp; "<em>Processed_Wind_Profile_188_yyyymmdd_hhmmss.hpl</em>" with processed WV and WS data</p> <p>b)&nbsp;&nbsp;&nbsp;&nbsp; "<em>Wind_Profile_188_yyyymmdd_hhmmss.hpl</em>" with non-processed Doppler wind profile data</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; "<em>LIDAR_wind_profile_TURADATA_readme.txt</em>" with all necessary information for correct data use</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Data for the submitted paper by Yasunari et al., "Comprehensive Impact of Changing Siberian Wildfire Severities on Air Quality, Climate, and Economy: MIROC5 Global Climate Model's Sensitivity Assessments"

<p>The dataset contains some of the outputs from the global climate model experiments by MIROC5 on changing Siberian wildfire severities, the other data used in the paper (see READ_ME files on the data sources), the analyzed data, and the scripts for analyses, which were used in the following submitted paper. Note that this dataset also includes unused data for the paper:</p> <p><br>Yasunari, T. J., D. Narita, T. Takemura, S. Wakabayashi, and A. Takeshima, Comprehensive Impact of Changing Siberian Wildfire Severities on Air Quality, Climate, and Economy: MIROC5 Global Climate Model's Sensitivity Assessments, submitted.</p> <p>Please read the READ_ME files for detailed information in each directory (especially see the "about_figures_and_tables/" directory first). Because of their large sizes, the data were separated into three zipped files.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Air Quality Stripes

<p>Timeseries of annual mean particulate matter (PM2.5) concentrations (in micrograms per meter cubed) from 1850 to 2021 in 177 cities around the globe.&nbsp;</p> <p>Updated 10/12/2024 to include data for more cites.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Fig. 5 in Landuse Patterns, Air Quality And Bird Diversity In Urban Landscapes Of Delhi

Fig. 5. CCA plot showing the relationships of landuse patterns, air pollutants and bird species abundance.

opencc-by-4.0Apr 2022View details →
zenodo40/100

Seasonal analysis comparison of three air-cooling systems in terms of thermal comfort, air quality and energy consumption for school buildings in Mediterranean climates

<p>Efficient air-cooling systems for hot climatic conditions, such as Southern Europe, are required in the context of nearly Zero Energy Buildings, nZEB. Innovative air-cooling systems such as regenerative indirect evaporative coolers, RIEC and desiccant regenerative indirect evaporative coolers, DRIEC, can be considered an interesting alternative to direct expansion air-cooling systems, DX. The main aim of the present work was to evaluate the seasonal performance of three air-cooling systems in terms of air quality, thermal comfort and energy consumption in a standard classroom. Several annual energy simulations were carried out to evaluate these indexes for four different climate zones in the Mediterranean area. The simulations were carried out with empirically validated models. The results showed that DRIEC and DX improved by 29.8% and 14.6% over RIEC regarding thermal comfort, for the warmest climatic conditions, Lampedusa and Seville. However, DX showed an energy consumption three and four times higher than DRIEC for these climatic conditions, respectively. RIEC provided the highest percentage of hours with favorable indoor air quality for all climate zones, between 46.3% and 67.5%. Therefore, the air-cooling systems DRIEC and RIEC have a significant potential to reduce energy consumption, achieving the user&rsquo;s thermal comfort and improving indoor air quality.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Data for "Measurement Report: A Multi-Year Study on the Impacts of Chinese New Year Celebrations on Air 1 Quality in Beijing, China."

<p>These are the datasets that have been used&nbsp;for the article &quot;Measurement Report: A Multi-Year Study on the Impacts of Chinese New Year Celebrations on Air 1 Quality in Beijing, China,&quot; which is published in the journal&nbsp;<em>Atmospheric Chemistry and Physic</em><em>s</em>, by Foreback et al. (2022).</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Air quality and noise data in Santo Domingo and Santiago de los Caballeros, Dominican Republic from May 18,2020 to January 26, 2021

<p>Air and noise pollution affect the quality of life of any community. The data presented observes the noise level, eight air quality parameters and three weather parameters. The air quality parameters are carbon monoxide (CO), sulfur dioxide (SO2), ozone (O3), nitrogen dioxide (NO2); three particle-matter variables: ultrafine particulate matter (PM1), fine particulate matter (PM2.5), and coarse particulate matter (PM10). The weather parameters are temperature, relative humidity and atmospheric pressure. The empirical measurements were collected every ten or twenty minutes in two cities of the Dominican Republic from May 18th, 2020 to January 26th, 2021. The data can provide insight on the changes in air quality and noise level and can be used to compare with other variables such as traffic conditions.&nbsp;&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Code and data used in "A Tool for Air Pollution Scenarios (TAPS v1.0) to enable global, long-term, and flexible study of climate and air quality policies"

<p>Data and code for Tool for Air Pollution Scenarios (TAPS v1.0) as submitted to Geoscientific Model Development for publication. See the enclosed README and full user manual (https://github.com/watkin-mit/TAPS/wiki) for more information.&nbsp;</p>

openmit-licenseApr 2022View details →
zenodo40/100

Results from Air Quality Monitoring and Surveys in UK Residences - Appendix (Survey Questions)

<p>Survey questions used in the paper titled &quot;Results from Air Quality Monitoring and Surveys in UK Residences&quot;. Abstract below.</p> <p>Air pollution is a persistent issue in dwellings worldwide, costing an estimated 10-25 billion US dollars per year to the United Kingdom&rsquo;s national health service alone. However, it is an &ldquo;invisible problem&rdquo; since background pollutants are often imperceptible except during acute pollution events such as wildfires. Although public awareness of ventilation has increased due to the COVID-19 pandemic, there are few tools available to assess its efficacy. Widely available sensor systems that can measure these pollutants tend to be single units with simple apps and little connection to mitigation, whereas different rooms in a house may have different pollution issues with different recommended actions. In this study, we present the results of a measurement study conducted using a multi-room sensor kit in twenty-nine dwellings across the UK. We also analyze the occupants&rsquo; reaction to the hardware, data, and a prototype alerting system. The study shows broad awareness of air quality in the participants. However, this awareness rarely corresponded to effective mitigation actions or ventilation provision. The concept of alerts was welcomed by participants if accompanied by actionable recommendations.</p> <p>The data showed significant pollution events, as measured by proxies such as total VOC and CO<sub>2</sub>, occurring almost daily, particularly in households with gas appliances. These incidents were concentrated around particular times of day and behaviors, indicating that the capacity of infiltration and extract ventilation to bring in adequate fresh air was overwhelmed. No significant outdoor pollution was detected in houses, which was expected given their sheltered peri-urban locations. The study highlights the need for comprehensive implementation of measurement, ventilation, and treatment measures in the UK housing stock to reduce the impact of indoor pollution on health.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Air Quality Index for Indian Cities 2015-2023

<p>This repository contains analysis of AQI Daily bulletins released by the Central Pollution Control Board (CPCB) since 2015. These daily bulletins can be currently obtained here: <a href="https://cpcb.nic.in/AQI_Bulletin.php" rel="nofollow">CPCB Daily AQI Bulletins (link as accessed May 8th, 2024)</a></p> <p>https://github.com/urbanemissions-info/AQI_bulletins contains all the codes for parsing the data from the daily PDFs, data repository, and codes to make the calendar plots.</p> <p>The data contains (available) day-wise&nbsp;<code>city-average AQI value</code>,&nbsp;<code>AQI category</code>, and&nbsp;<code>conditional pollutant</code>&nbsp;information. This data for all cities can be obtained in the&nbsp;<code>AllIndiaBulletins_master.csv</code>&nbsp;from&nbsp;<code>data/Processed</code> folder (on github) or as <code>India_AQI_Bulletins_Master.csv</code>&nbsp; here.</p> <p>Calendar plot of the AQI values is produced for each city, along with the average number of stations reporting AQI value in each year.&nbsp;City wise AQI bulletins CSV and calendar plots can be obtained on UrbanEmissions website:&nbsp;<a href="https://urbanemissions.info/india-air-quality/india-ncap-aqi-indian-cities-2015-2023/" rel="nofollow">Link</a></p>

opencc-by-4.0May 2024View details →
zenodo40/100

Soil Moisture, Soil NOx and Regional Air Quality in the Agricultural Central United States: Data

<p>The data in this repository are associated with the manuscript from Huber et al. (2024) titled "Soil Moisture, Soil NOx and Regional Air Quality in the Agricultural Central United States" in the Journal of Geophysical Research: Atmospheres. Additional information regarding these data can be found in the attached readme file.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record