Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
708
datasets available to search
ShareScore release 0.9.0
Dataset results
708 results for “global dataset”
Global dataset of nitrogen fixation rates across inland and coastal waters based on a coordinated synthesis effort
Biological nitrogen fixation converts inert di-nitrogen gas into bioavailable nitrogen and can be an important source of bioavailable nitrogen to organisms. This dataset synthesizes the aquatic nitrogen fixation rate measurements across inland and coastal waters. Data were derived from papers and datasets published by April 2022 and include rates measured using the acetylene reduction assay (ARA), 15N2 labeling, or the N2/Ar technique. The dataset is comprised of 4793 nitrogen fixation rates measurements from 267 studies, and is structured into four tables: 1) a reference table with sources from which data were extracted, 2) a rates table with nitrogen fixation rates that includes habitat, substrate, geographic coordinates, and method of measuring N2 fixation rates, 3) a table with supporting environmental and chemical data for a subset of the rate measurements when data were available, and 4) a data dictionary with definitions for each variable in each data table. This dataset was compiled and curated by the NSF-funded Aquatic Nitrogen Fixation Research Coordination Network (award number 2015825).
Global eutrophication and antibiotic resistance genes dataset for "Coupling mechanisms between cyanobacteria and antibiotic resistance genes in freshwater ecosystems"
This dataset compiles global records of cyanobacteria, antibiotic resistance genes (ARGs), and associated water quality parameters to support research on freshwater ecosystem dynamics. It includes 990 metagenomes, 16,648 chlorophyll-a (Chl-a) records, and over 90 documented cases of ARGs–cyanobacteria co-occurrence under comparable spatiotemporal conditions. The dataset covers the years 2000–2024 and provides both raw measurements and harmonized tables for cross-study comparisons. Data were extracted from previously published literature and public repositories, with references to source publications included. This archive is intended to facilitate reproducible analyses, enable large-scale meta-studies, and support further exploration of microbial interactions in freshwater systems.
Global Extra-tropical Circulation Database based on the Jenkinson-Collison Classification calculated with 6-hourly mean sea-level pressure fields from various reanalysis datasets
<h1>Dataset Description</h1> <p>Global Extra-tropical Circulation Database based on the Jenkinson-Collison Classification calculated with 6-hourly mean sea-level pressure fields from several reanalysis datasets. This dataset is the result of an extension of the Jenkinson-Collison circulation type classification to the entire globe, including a modification of its original formulation for the southern hemisphere.</p> <p>A modified version of the IPCC-AR6 Reference Regions that excludes the intertropical range where the method is not applicable is also included, as used in the reference paper for global assessment.</p> <p>Further details in <a href="https://doi.org/10.1007/s00382-022-06658-7" target="_blank" rel="noopener">https://doi.org/10.1007/s00382-022-06658-7 </a></p> <h2>Note for version 1.1.0</h2> <p>This version corrects an issue in the previous release, which was incorrectly labeled as <em>version 0.1</em>. That version was incomplete due to the omission of previously existing files, and should be considered <strong>incomplete</strong>. Version 1.1.0 restores all original files alongside the newly added one, ensuring the dataset is now complete and consistent. We apologize for any inconvenience this may have caused and appreciate your understanding.</p>
GlobalHighPM₂.₅: Global Daily Seamless 1 km Ground-Level PM₂.₅ Dataset over Land (2017–Present)
<p>GlobalHighPM<sub>2.5</sub> is part of a series of long-term, seamless, global, high-resolution, and high-quality datasets of air pollutants over land (i.e., GlobalHighAirPollutants, GHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>This dataset contains input data, analysis codes, and generated dataset used for the following article. If you use the GlobalHighPM<sub>2.5</sub> dataset in your scientific research, please cite the following reference (Wei et al., NC, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Lyapustin, A., Wang, J., Dubovik, O., Schwartz, J., Sun, L., Li, C., Liu, S., and Zhu, T. <a href="https://weijing-rs.github.io/publications/Wei_et_al-NC-2023.pdf" target="_blank" rel="noopener">First close insight into global daily gapless 1 km PM<sub>2.5</sub> pollution, variability, and health impact</a>. <em>Nature Communications</em>, 2023, 14, 8349. https://doi.org/10.1038/s41467-023-43862-3</p> </li> </ul> <p><strong>Input Data</strong></p> <p>Relevant raw data for each figure (compiled into a single sheet within an Excel document) in the manuscript.</p> <p><strong>Code</strong></p> <p>Relevant Python scripts for replicating and ploting the analysis results in the manuscript, as well as codes for converting data formats.</p> <p><strong>Generated Dataset</strong></p> <p>Here is the first big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) global ground-level PM<sub>2.5</sub> dataset over land from 2017 to the present. This dataset exhibits high quality, with cross-validation coefficients of determination (CV-R<sup>2</sup>) of 0.91, 0.97, and 0.98, and root-mean-square errors (RMSEs) of 9.20, 4.15, and 2.77 µg m<sup>-3</sup> on the daily, monthly, and annual bases, respectively.</p> <p><strong>Due to data volume limitations, </strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2022 </strong>is accessible at: <strong><a href="../records/10795661">GlobalHighPM2.5 (2022)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2021 </strong>is accessible at: <strong><a href="../records/10398385">GlobalHighPM2.5 (2021)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2020 </strong>is accessible at: <strong><a href="../records/10402639">GlobalHighPM2.5 (2020)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2019 </strong>is accessible at: <strong><a href="../records/10402723">GlobalHighPM2.5 (2019)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2018 </strong>is accessible at: <strong><a href="../records/10402824">GlobalHighPM2.5 (2018)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2017 </strong>is accessible at: <strong><a href="../records/10403497">GlobalHighPM2.5 (2017)</a></strong></p> <p> continuously updated...</p> <p><strong>More GHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
Metabolism dataset: one year of high-frequency temperature, dissolved oxygen, wind, photosynthetically active radiation observations and low-frequency nutrient data for 58 lakes in the Global Lake Ecological Observatory Network
Understanding controls on primary productivity is essential for describing ecosystems and their responses to environmental change. Lake primary production is strongly controlled by inputs of nutrients and colored dissolved organic matter. While past studies have developed mathematical models of this nutrient-color paradigm, broad empirical tests of these models are scarce. We compiled data from 58 diverse and globally distributed and mostly temperate lakes to test such a model and improve understanding and prediction of the controls on lake primary production. These lakes varied widely in size (0.02-2300 km2), pelagic gross primary production (20-8000 mg C m-2 d-1), and other characteristics. The data package includes high-frequency dissolved oxygen, water temperature, wind speed, and solar radiation data as well as daily estimates of GPP and ER derived from those data. In addition, the data package includes median in-lake and stream concentrations of dissolved organic carbon and total phosphorus for a subset of 18 of those lakes.
ML-TOMCAT V2.0: Machine-Learning-Based Satellite-Corrected Global Stratospheric Ozone Profile Dataset
<p>MLTOMCAT V2 is 46 years (1979-2024) of gap free ozone profile data sets that is created by correcting biases in a TOMCAT Chemical Transport Model (CTM) simulated ozone profiles. We use Random Forest regression model to correct model biases. </p> <p>Each file contain monthly mean zonal mean ozone profiles. There are 6 data files.</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_vmr_V2.nc</a> contains ozone profiles on geometric height levels (1 to 60 km) in mixing ratio units, whereas <a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_nd_V2.nc</a> contains ozone profile in number density units.</p> <p>Similarly, </p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_vmr_V2.nc</a> contains ozone profiles on 43 MLS pressure levels (1000 to 0.1 hPa) in mixing ratio units, whereas <a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_nd_V2.nc</a> contains ozone profile in number density units.</p> <p>Please note that data below 300 hPa (~8km) and 1 hPa (~50 km) should be used with caution.</p> <p>There are two straospheric column files</p> <p>ML-TOMCAT-SCO_120ppb_boundary_V2_197901-202412.nc and</p> <p>ML-TOMCAT-SCO_150ppb_boundary_V2_197901-202412.nc</p> <p>Stratospheric column files calculated using 120 ppb and 150 ppb as a chemical ozone boundaries.</p> <p>A manuscript describing MLTOMCAT would be published in EESD (Dhomse et al., 2021).</p>
DeepOWT v2.25.1: An updated and improved global offshore wind turbine dataset until 2025Q1
<p>DeepOWT (deep learning derived global offshore wind turbines) is an independent and openly accessible data set of offshore wind energy infrastructure locations and their temporal deployment dynamics on a global scale. It is derived by applying deep learning based object detection on ESA's spaceborne Sentinel-1 synthetic aperture radar (SAR) archive. DeepOWT provides OWT locations along with their quarterly deployment stages from 2016Q1 until 2025Q1. It differentiates between platforms under construction, OWTs which are readily deployed and offshore wind farm substations, such as transformer stations.<br><br>The dataset continues the work of <a href="https://essd.copernicus.org/articles/14/4251/2022/">10.5194/essd-14-4251-2022</a>.</p> <p>File metadata</p> <table> <tbody> <tr> <th>File</th> <th>Time</th> <th>Geometry</th> <th>Spatial extent</th> </tr> <tr> <td>DeepOWT.geojson (Dataset)</td> <td>2016Q1-2025Q1</td> <td>points</td> <td>Global</td> </tr> <tr> <td>gt_2021Q2_nsb.geojson (Ground Truth Location)</td> <td>2021Q2</td> <td>polygons</td> <td>North Sea Basin</td> </tr> <tr> <td>gt_2021Q2_ecs.geojson (Ground Truth Location)</td> <td>2021Q2</td> <td>polygons</td> <td>East China Sea</td> </tr> <tr> <td>gt_2021Q2_vtn.geojson (Ground Truth Location)</td> <td>2021Q2</td> <td>polygons</td> <td>Southeast Vietnamese Coast</td> </tr> <tr> <td>gt_nsb_gridded.geojson (Ground Truth Region)</td> <td>-</td> <td>polygon</td> <td>North Sea Basin</td> </tr> <tr> <td>gt_ecs_gridded.geojson (Ground Truth Region)</td> <td>-</td> <td>polygon</td> <td>East China Sea</td> </tr> <tr> <td>gt_ecs_gridded.geojson (Ground Truth Region)</td> <td>-</td> <td>polygon</td> <td>Southeast Vietnamese Coast</td> </tr> </tbody> </table> <p> </p> <table> <thead> <tr> <th>Used semantic label</th> </tr> </thead> <tbody> <tr> <td>open sea</td> </tr> <tr> <td>under construction</td> </tr> <tr> <td>offshore wind turbine</td> </tr> <tr> <td>offshore wind farm substation</td> </tr> </tbody> </table> <p> </p>
A global dataset gathering 37 field experiments involving cereal-legume intercrops and their corresponding sole crops.
<p>The overall description of the dataset is reported in the <strong>data_report.pdf</strong> file. The methodology for data curation and tidying is published in Peer Community Journal (<a href="https://doi.org/10.24072/pcjournal.389">Mahmoud2024</a>).</p> <p>This dataset gathers the results of 37 field experiments, which involved cereal-legume intercrops and their corresponding sole crops. The field experiments were carried in 5 European countries (France, Denmark, Italy, Germany and England) from 2001 to 2017. The dataset includes:</p> <ul> <li>5 legume species , <em>i.e.</em> chickpea (<em>Cicer arietinum</em> L.), faba bean (<em>Vicia faba</em> L.), lentil (<em>Lens culinaris</em> Med.), lupin (<em>Lupinus albus</em> L.) and pea (<em>Pisum sativum</em> L.),</li> <li>3 cereal species, <em>i.e.</em> barley (<em>Hordeum vulgare</em> L.), durum wheat (<em>Triticum turgidum</em> L.) and soft wheat (<em>Triticum aestivum</em> L.), </li> <li>8 resulting intercrops, <em>i.e.</em> i) barley associated with faba bean, lupin or pea, ii) durum wheat associated with chickpea, faba bean or pea, and iii) soft wheat associated with lentil or pea. </li> </ul> <p>In total, the dataset contains 299 sole crop and 308 intercrop experimental units, one given experimental unit being defined as the unique combination of {site, year, crop management}, with the crop management including species and cultivar choice as well as agricultural interventions (sowing conditions, inputs).</p> <p>The global dataset includes four tables, all sharing a common identifier (experiment_id):</p> <ul> <li>data_trials.csv: the global features describing the experimental sites,</li> <li>data_management.csv: the agricultural management actions carried out on each of the experimental sites,</li> <li>data_traits.csv: measured plant and crop characteristics,</li> <li>data_climate.csv: climate for the experimental sites, retrieved from NASA POWER API.</li> </ul> <p>Additionally, a metadata file is provided (<strong>metadata.xlsx</strong>), describing the table to which the variables belong (variable_type, i.e. trials, management, traits or climate), their name (variable_name), their significance (description) and their unit (unit). Finally, a table including the original references related to experimental files gathered (<strong>references.xlsx</strong>) is also provided.</p> <p>Data providers and field experiments: Laurent Bedoussac, Eric Justes, Etienne-Pascal Jour- net, Christophe Naudin, Henrik Hauggaard-Nielsen, Erik Steen Jensen, Elise Pelzer, Guénaëlle Corre-Hellou, Bochra Kammoun, Loic Viguier, Romain Barillot, Antoine Couëdel, Philippe Hinsinger</p> <p>Database and management: Noémie Gaudio, Rémi Mahmoud, Pierre Casadebaig</p>
A global dataset of specialty crop biomass and N2O emissions
<div> <p>We reviewed global field studies of vineyard, orchard, and vegetable cropping systems, which were also included in a meta-analysis (<a href="https://doi.org/10.1111/gcb.17233">https://doi.org/10.1111/gcb.17233</a>). We narrowed down the studies to those with field measurements of adequate variables (biomass C, N, and N<sub>2</sub>O) covering at least one growing season. As a result, cumulative N₂O emission measurements (per growing rotation, season, or year), along with biomass data of different plant organs from the same regions, were compiled for grape (<em>Vitis vinifera</em>), almond [<em>Prunus dulcis</em> (Mill.) D.A. Webb], peach (<em>Prunus persica</em> L.), walnut (<em>Juglans regia</em>), lettuce (<em>Lactuca sativa</em>), broccoli (<em>Brassica oleracea</em> var. <em>italica </em>P.), cauliflower (<em>Brassica oleracea</em> var. <em>botrytis </em>L.), and tomato (<em>Lycopersicon esculentum</em> L.) planting system. These observations were collected from fields spanning seven Koppen-Geiger climate types and five countries (the United States, Germany, Spain, France, and Australia). When only dry mass was measured, biomass C content for aboveground vegetable crops and berry fruit was assumed at 43%; nut fruit and woody organs of orchard tree at 48%. Area-weighted averages of N<sub>2</sub>O emissions were used (tree/vine row and interrow).</p> <p> </p> <p>Corresponding author: Mu Hong (mu.hong@colostate.edu)</p> </div> <p> </p>
Supplementary dataset for "Rising CO2 and warming reduce global canopy demand for nitrogen"
<p>This repository contains the dataset used for “<strong>Rising CO<sub>2</sub> and warming reduce global canopy demand for nitrogen” </strong></p> <p>The deposition consists of:</p> <ol> <li>An satellite-derived leaf chlorophyll vcmax25 database (Luo<em> et al.</em>, 2019)</li> <li>Simulated <em>V<sub>cmax</sub></em> with all the factors based on the coordination hypothesis</li> <li>Simulated <em>V<sub>cmax </sub></em>with CO<sub>2</sub> fixed at 340 ppm based on the coordination hypothesis</li> <li>Simulated <em>V<sub>cmax</sub> </em>with fixed climate based on the coordination hypothesis</li> <li>Simulated turnover time.</li> <li>Simulated leaf-level <em>N</em><sub>rubisco</sub> (g m<sup>–2</sup> leaf area), canopy-level <em>N<sub>rubisco</sub></em> (g m<sup>–2</sup> ground area), annual leaf-level <em>N<sub>rubisco</sub></em>demand (g m<sup>–2</sup> leaf area year<sup>–1</sup>), and annual canopy-level of <em>N<sub>rubisco</sub></em> demand (g m<sup>–2</sup> ground area year<sup>–1</sup>) in figure 4.</li> <li>Lifespan of evergreen</li> </ol> <p>Note. </p> <ol> <li>LAI products used in the paper , such as TCDR LAI during 1982­–2016; GLASS LAI during 1982–2014; and GLOBMAP LAI during 1982–2011 are public available, the details information see Jiang <em>et al </em>(2017).</li> <li>Evergreen, deciduous and herbaceous vegetation fractions data derived from ESA CCI land cover products is publicly available, the details information see Li <em>et al </em>(2018).</li> <li>The climate force for <em>V<sub>cmax </sub></em>simulation was used CRU TS4.3 (Harris <em>et al,</em> 2020) for 1982–2016 at 0.5° resolution, which is publicly available at </li> </ol> <p><a href="https://crudata.uea.ac.uk/cru/data/hrg/">https://crudata.uea.ac.uk/cru/data/hrg/</a>.</p> <p>The data files are all in netcdf format at 0.5 resolution </p> <p>Reference:</p> <ol> <li><strong>Luo X, Croft H, Chen JM, He L, Keenan TF. 2019.</strong> Improved estimates of global terrestrial photosynthesis using information on leaf chlorophyll content. <em>Global Change Biology</em> <strong>25</strong>(7): 2499-2514.</li> <li><strong>Jiang C, Ryu Y, Fang H, Myneni R, Claverie M, Zhu Z. 2017.</strong> Inconsistencies of interannual variability and trends in long-term satellite leaf area index products. <em>Global Change Biology</em> <strong>23</strong>(10): 4133-4146.</li> <li><strong>Li W, MacBean N, Ciais P, Defourny P, Lamarche C, Bontemps S, Houghton RA, Peng S. 2018.</strong> Gross and net land cover changes in the main plant functional types derived from the annual ESA CCI land cover maps (1992–2015). <em>Earth Syst. Sci. Data</em> <strong>10</strong>(1): 219-234.</li> <li><strong>Harris I, Osborn TJ, Jones P, Lister D. 2020.</strong> Version 4 of the CRU TS monthly high-resolution gridded multivariate climate dataset. <em>Scientific Data</em> <strong>7</strong>(1): 109.</li> </ol>
Country Compendium of the Global Register of Introduced and Invasive Species. Dataset.
<p>The Country Compendium of the Global Register of Introduced and Invasive Species (GRIIS) is a collation of data across 196 individual country checklists of alien species, along with a designation of those species associated with evidence of impact at a country level. </p>
Global Mangrove Watch (1996 - 2020) Version 3.0 Dataset
<p>This study has used L-band Synthetic Aperture Radar (SAR) global mosaic datasets from the Japan Aerospace Exploration Agency (JAXA) for 11 epochs from 1996 to 2020 to develop a long-term time-series of global mangrove extent and change. The study used a map-to-image approach to change detection where the baseline map (GMW v2.5) was updated using thresholding and a contextual mangrove change mask. This approach was applied between all image-date pairs producing 10 maps for each epoch, which were summarised to produce the global mangrove time-series. The resulting mangrove extent maps had an estimated accuracy of 87.4 % (95th conf. int.: 86.2 - 88.6 %), although the accuracies of the individual gain and loss change classes were lower at 58.1 % (52.4 - 63.9 %) and 60.6 % (56.1 - 64.8 %), respectively. Sources of error included a mis-registration in the SAR mosaic datasets, which could only be partially corrected for, but also confusion in fragmented areas of mangroves, such as around aquaculture ponds. Overall, 152,604 km<sup>2</sup> (133,996 - 176,910) of mangroves were identified for 1996, with this decreasing by -5,245 km<sup>2</sup> (-13,587 - 3686) resulting in a total extent of 147,359 km<sup>2</sup> (127,925 - 168,895) in 2020, and representing an estimated loss of 3.4 % over the 24-year time period. The Global Mangrove Watch Version 3.0 represents the most comprehensive record of global mangrove change achieved to date and is expected to support a wide range of activities, including the ongoing monitoring of the global coastal environment, defining and assessments of progress towards conservation targets, protected area planning and risk assessments of mangrove ecosystems worldwide.</p> <p>The paper which goes along with this dataset is available at the following reference:</p> <p>Bunting, P.; Rosenqvist, A.; Hilarides, L.; Lucas, R.M.; Thomas, T.; Tadono, T.; Worthington, T.A.; Spalding, M.; Murray, N.J.; Rebelo, L-M. Global Mangrove Extent Change 1996 – 2020: Global Mangrove Watch Version 3.0. Remote Sensing. 2022</p>
Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)
<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1 contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE). </p> <p> </p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called "Sentinel2LULC_GeoTiff.zip" contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called "Sentinel2LULC_JPEG.zip" contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called "Sentinel2LULC_CSV.zip" includes 29 zip-compressed CSV files with as many rows as provided images and with 12 columns containing the following metadata (this same metadata is provided in the image filenames): </p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class </li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products </li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing and exporting its corresponding image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class </li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with "including_non_downloaded_images".</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset, we included in the version v2.1, a compressed folder called "Geographic_Representativeness.zip". This zip-compressed folder contains a csv file for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>© <a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset </a>by Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujemâa Achchab, Francisco Herrera & Siham Tabik is marked with Attribution 4.0 International (CC-BY 4.0)</p>
Global Dataset of Cyber Incidents V.1.2
<p>The dataset contains data on 2889 cyber incidents between 01.01.2000 and 02.05.2024 using 60 variables, including the start date, names and categories of receivers along with names and categories of initiators. The database was compiled as part of the <strong><a href="https://eurepoc.eu">European Repository of Cyber Incidents (EuRepoC)</a> </strong>project.</p> <p><br>EuRepoC gathers, codes, and analyses publicly available information from over 200 sources and 600 Twitter accounts daily to report on dynamic trends in the global, and particularly the European, cyber threat environment.<br><br>For more information on the scope and data collection methodology see: <a href="https://eurepoc.eu/methodology">https://eurepoc.eu/methodology</a><br><br><strong>Codebook available <a href="https://eurepoc.eu/wp-content/uploads/2023/07/EuRepoC_Codebook_1_2.pdf">here</a><br><br>Information about each file:</strong></p> <p><strong>Global Database (csv or xlsx):<br></strong>This file includes all variables coded for each incident, organised such that one row corresponds to one incident - our main unit of investigation. Where multiple codes are present for a single variable for a single incident, these are separated with semi-colons within the same cell.</p> <p><strong>Receiver Dataset (csv):<br></strong>In this file, the data of affected entities and individuals (receivers) is restructured to facilitate analysis. Each cell contains only a single code, with the data "unpacked" across multiple rows. Thus, a single incident can span several rows, identifiable through the unique identifier assigned to each incident (incident_id). </p> <p><strong>Attribution Dataset (csv):</strong><br>This file follows a similar approach to the receiver dataset. The attribution data is "unpacked" over several rows, allowing each cell to contain only one code. Here too, a single incident may occupy several rows, with the unique identifier enabling easy tracking of each incident (incident_id). In addition, some attributions may also have multiple possible codes for one variable, these are also "unpacked" over several rows, with the attribution_id enabling to track each attribution.<br><br><strong>eurepoc_global_database_1.2 (json):</strong><br>This file contains the whole database in JSON format. </p>
S1S2-Water: A global dataset for semantic segmentation of water bodies from Sentinel-1 and Sentinel-2 satellite images
<p>The S1S2-Water dataset is a global reference dataset for training, validation and testing of convolutional neural networks for semantic segmentation of surface water bodies in publicly available Sentinel-1 and Sentinel-2 satellite images. The dataset consists of 65 triplets of Sentinel-1 and Sentinel-2 images with quality checked binary water mask. Samples are drawn globally on the basis of the Sentinel-2 tile-grid (100 x 100 km) under consideration of pre-dominant landcover and availability of water bodies. Each sample is complemented with metadata and Digital Elevation Model (DEM) raster from the Copernicus DEM.</p><p>This work was supported by the German Federal Ministry of Education and Research (BMBF) through the project "Künstliche Intelligenz zur Analyse von Erdbeobachtungs- und Internetdaten zur Entscheidungsunterstützung im Katastrophenfall" (AIFER) under Grant 13N15525, and by the Helmholtz Artificial Intelligence Cooperation Unit through the project "AI for Near Real Time Satellite-based Flood Response" (AI4FLOOD) under Grant ZT-IPF-5-39. </p>
SM2RAIN-CCI (1 Jan 1998 – 31 December 2015) global daily rainfall dataset
<p>A NEW GLOBAL SCALE RAINFALL PRODUCT obtained from satellite soil moisture data through the SM2RAIN algorithm (<em>Brocca et al., 2014</em>), at 0.25 degree/daily spatial-temporal resolution, has been delivered (Ciabatta et al., 2018). The SM2RAIN method was applied to the ESA CCI soil moisture Active and Passive products (<em>Liu et al., 2011, 2012; Wagner et al., 2012</em>) for the period from January 1998 to December 2015 (18 years).</p> <p>The CCI-derived rainfall datasets (in mm/day) is gridded over a 0.25-degree grid on a global scale. The number of dates is 6574 (1998/01/01 – 2015/12/31). The product represents the cumulated rainfall between the 00:00 and the 23:59 UTC of the indicated day. A climatological correction has been applied to the data at monthly scale.</p> <p>The rainfall dataset is provided in netCDF format. A total of 18 netCDF files, one per year, are provided.</p> <p>The rainfall dataset is obtained by applying the SM2RAIN algorithm to the ESA CCI soil moisture Active and Passive products at version 03.1 separately. Then, an integration procedure based on a weighted average is applied in order to obtain the rainfall estimate. The algorithm has been calibrated during three different periods (1998-2001, 2002-2006 and 2007-2013) against the Global Precipitation Climatology Centre Full-Data daily dataset (GPCC-FDD, Schamm et al., 2015). The quality flag provided within the raw soil moisture observations has been used to mask out low quality data, as well as the areas characterized by high topographic complexity, high frozen soil and snow probability and presence of tropical forests.</p> <p><strong>References</strong></p> <p>Brocca, L., Ciabatta, L., Massari, C., Moramarco, T., Hahn, S., Hasenauer, S., Kidd, R., Dorigo, W., Wagner, W., Levizzani, V. (2014). Soil as a natural rain gauge: estimating global rainfall from satellite soil moisture data. <em>Journal of Geophysical Research</em>, 119(9), 5128-5141, doi:10.1002/2014JD021489.</p> <p>Ciabatta, L., Massari, C., Brocca, L., Gruber, A., Reimer, C., Hahn, S., Paulik, C., Dorigo, W., Kidd, R., and Wagner, W.: SM2RAIN-CCI: a new global long-term rainfall data set derived from ESA CCI soil moisture, Earth Syst. Sci. Data, 10, 267-280, https://doi.org/10.5194/essd-10-267-2018, 2018.</p> <p>Liu, Y. Y., Parinussa, R. M., Dorigo, W. A., De Jeu, R. A. M., Wagner, W., van Dijk, A. I. J. M., McCabe, M. F., Evans, J. P. (2011). Developing an improved soil moisture dataset by blending passive and active microwave satellite-based retrievals. Hydrology and Earth System Sciences, 15, 425-436, doi:10.5194/hess-15-425-2011.</p> <p>Liu, Y.Y., Dorigo, W.A., Parinussa, R.M., de Jeu, R.A.M., Wagner, W., McCabe, M.F., Evans, J.P., van Dijk, A.I.J.M. (2012). Trend-preserving blending of passive and active microwave soil moisture retrievals, Remote Sensing of Environment, 123, 280-297, doi: 10.1016/j.rse.2012.03.014.</p> <p>Schamm, K., Ziese, M., Raykova, K., Becker, A., Finger, P., Meyer-Christoffer, A., Schneider, U. (2015). GPCC Full Data Daily Version 1.0 at 1.0°: Daily Land-Surface Precipitation from Rain-Gauges built on GTS-based and Historic Data. DOI: 10.5676/DWD_GPCC/FD_D_V1_100.</p> <p>Wagner, W., Dorigo, W., de Jeu, R., Fernandez, D., Benveniste, J., Haas, E., Ertl, M. (2012). Fusion of active and passive microwave observations to create an Essential Climate Variable data record on soil moisture, ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences (ISPRS Annals), Volume I-7, XXII ISPRS Congress, Melbourne, Australia, 25 August-1 September 2012, 315-321.</p>
RRING Global Survey Research Dataset (WP3)
<p>The RRING Work Package 3 (WP3) objective was to clarify how Research Funding Organisations (RFOs) and Research Performing Organisations (RPOs) operated within region-specific research and innovation environments. It explored how they navigated the governance and regulatory frameworks for Responsible Research and Innovation (RRI), as well as offering their perspectives on the entities responsible for RRI-related policy and action in their locales.</p> <p>This data set covers the global survey research part, which was designed to contextualise how RPOs and RFOs interacted within the research environment and with non-academic stakeholders. Countries were grouped according to the UNESCO regions of the world and key results per region are listed below. For a detailed analysis and further findings of the work completed under WP3 of the RRING project, please refer to the full deliverable document "State of the Art of RRI in the Five UNESCO World Regions" [link to be inserted].</p> <p> </p> <p><strong>European and North American States</strong></p> <ul> <li>‘Diverse and inclusive': Respondents were most attitudinally supportive of the importance of ensuring ethical principles were applied in R&I (92%), followed by diverse perspectives (88%), and gender equality (79%). Including ethnic minorities was the area which garnered the least attitudinal support (71%). Respondents took the most practical steps towards engaging with diverse perspectives (63%), and the least towards inclusion of ethnic minorities (24%).</li> <li>‘Anticipative and reflective’: Respondents widely agreed (82%) with the importance of ensuring R&I work does not cause concerns for society, but only 37% confirmed they had taken practical steps to ensure this.</li> <li>‘Open and transparent’: Vast majorities of respondents agreed on the importance of keeping R&I methods open and transparent (94%), with 65% also confirming they take practical steps to do this. An equally high number agreed on the importance of making the results of R&I work accessible to as wide a public as possible (94%), and 68% confirmed this through their reported actions. This indicated the smallest value-action gap of all RRI measures for respondents from European and North American countries. Attitudinal agreement on the importance of making data freely available to the public was lower (83%), as was the practical action aspect for this measure (45%).</li> <li>‘Responsive and adaptive to change’: Most respondents agreed (89%) that it was important to ensure their work addresses societal needs, and 62% confirmed that they take practical steps towards this aim.</li> </ul> <p> </p> <p><strong>Latin American and Caribbean States</strong></p> <ul> <li>‘Diverse and inclusive': Respondents were most attitudinally supportive of the importance of gender equality in R&I (86%), followed by ensuring ethical principles are applied (85%), and diverse perspectives incorporated (83%). Including ethnic minorities was the area which garnered the least attitudinal support (77%). Respondents took the most practical steps towards ensuring ethical principles guide their work (50%), and the least towards including ethnic minorities (25%), but the smallest value action gap was found for gender equality.</li> <li>‘Anticipative and reflective’: Respondents agreed (79%) that it is important to ensure R&I work does not cause concerns for society, but only 29% confirmed they had taken practical steps to ensure this.</li> <li>‘Open and transparent’: The majority of respondents agreed on the importance of keeping R&I methods open and transparent (89%), with 45% indicating they had taken practical action. A majority also agreed on the importance of making the results of R&I work accessible to as wide a public as possible (88%), and 44% backed this up with practical action. Attitudinal agreement on the importance of making data freely available to the public was slightly lower (81%), as was the practical action aspect for this measure (35%).</li> <li>‘Responsive and adaptive to change’: Most respondents agreed (84%) that it was important to ensure their work addresses societal needs, and 49% confirmed that they take practical steps towards this aim.</li> </ul> <p> </p> <p><strong>Asian and Pacific States</strong></p> <ul> <li>‘Diverse and inclusive': Respondents were most attitudinally supportive of the importance of ensuring ethical principles were applied in R&I (90%), followed by diverse perspectives (89%), and gender equality (86%). Including ethnic minorities was the area which garnered the least attitudinal support (76%). Respondents took the most practical steps towards engaging with diverse perspectives (65%), and the least towards including ethnic minorities (30%).</li> <li>‘Anticipative and reflective’: Respondents widely agreed (78%) with the importance of ensuring R&I work does not cause concerns for society, and 42% confirmed they had taken practical steps to ensure this.</li> <li>‘Open and transparent’: The majority of respondents agreed on the importance of keeping R&I methods open and transparent (91%), with 58% indicating they take practical steps to do this. A majority also agreed on the importance of making the results of R&I work accessible to as wide a public as possible (89%), and 64% backed this up with practical action. Attitudinal agreement on the importance of making data freely available to the public was lower (79%), as was the practical action aspect for this measure (40%).</li> <li>‘Responsive and adaptive to change’: Most respondents agreed (92%) that it was important to ensure their work addresses societal needs, and 69% confirmed that they take practical steps towards this aim. This was the RRI measure with the smallest valueaction gap for respondents from the Asian and Pacific region.</li> </ul> <p> </p> <p><strong>Arab States</strong></p> <ul> <li>‘Diverse and inclusive': Respondents were most attitudinally supportive of the importance of ensuring ethical principles were applied in R&I (93%), followed by diverse perspectives (81%), and gender equality (85%). Including ethnic minorities was the area which garnered the least attitudinal support (74%). Respondents took the most practical steps towards engaging with diverse perspectives (66%), which equated to one of two equally small value-action gaps for respondents from Arab states, and the least practical steps towards inclusion of ethnic minorities (22%).</li> <li>‘Anticipative and reflective’: A high proportion of respondents (85%) agreed that it is important to ensure R&I work does not cause concerns for society. However, only 38% confirmed they had taken practical steps to ensure this.</li> <li>‘Open and transparent’: The majority of respondents agreed on the importance of keeping R&I methods open and transparent (89%), with 59% also confirming they take practical steps to do this. A majority also agreed on the importance of making the results of R&I work accessible to as wide a public as possible (90%), and 66% backed this up with practical action. Ensuring public accessibility of research results was the second of two measures with equally small value-action gaps. Attitudinal agreement on the importance of making data freely available to the public was much lower (78%), which also reflected the practical action aspect for this measure (49%).</li> <li>‘Responsive and adaptive to change’: Most respondents agreed (96%) that it was important to ensure their work addresses societal needs, and 68% confirmed that they take practical steps to achieve this.</li> </ul> <p><strong>African States</strong></p> <ul> <li>‘Diverse and inclusive': Respondents were most attitudinally supportive of the importance of ensuring engagement with diverse perspectives and expertise in R&I (91%), followed by ensuring ethical principles are applied (90%), and gender equality (89%). Including ethnic minorities was the area which garnered the least attitudinal support (74%). Respondents took the most practical steps towards ensuring ethical principles guide their work (57%), and the least towards including ethnic minorities (32%).</li> <li>‘Anticipative and reflective’: The majority of respondents (85%) agreed that it is important to ensure R&I work does not cause concerns for society, with 59% confirming that they take practical steps to ensure this.</li> <li>‘Open and transparent’: A high proportion of respondents agreed on the importance of keeping R&I methods open and transparent (90%), with 54% also confirming they take practical steps to do this. A majority also agreed on the importance of making the results of R&I work accessible to as wide a public as possible (86%), and 56% backed this up with practical action. Attitudinal agreement on the importance of making data freely available to the public was significantly lower (73%), as was the practical action aspect for this measure (38%).</li> <li>‘Responsive and adaptive to change’: Respondents mostly agreed (92%) that it was important to ensure their work addresses societal needs, and 64% confirmed that they take practical steps towards this aim. This was the RRI measure with the smallest valueaction gap for respondents from African states.</li> </ul> <p> </p> <p><em>Note: Please refer to the "RRING WP3 - Survey Data Documentation" document for detailed instructions on how to use this dataset.</em></p>
SInAS: A global dataset of native and alien distributions of alien species
<p>The SInAS dataset represents a collection of regional lists of alien (also called non-native or non-indigenous) species and includes information about their native ranges, alien ranges, invasion status for alien ranges, habitats and year of first record. This dataset has been generated by standardising and integrating large global databases of alien species occurrences using the SInAS workflow version 2.0. </p> <p>The SInAS dataset is described in more detail in the following scientific article, which need to be cited when using this dataset:</p> <p>Gómez-Suárez, M., Laeseke, P., and Seebens, H. (submitted) A global dataset of native and alien distributions of alien species </p> <p>The code to generate the dataset is stored on Github (https://github.com/hseebens/SInAS) with releases available on Zenodo (https://doi.org/10.5281/zenodo.3763221).</p>
The Global Carbon Project's fossil CO2 emissions dataset
<p>The <a href="https://www.globalcarbonproject.org/">Global Carbon Project</a> (GCP) has been publishing estimates of global and national fossil CO2 emissions since 2001. In the first instance these were simple re-publications of data from another source, but over subsequent years refinements have been made in response to feedback and identification of inaccuracies. In this article (PDF document) we describe the history of this process leading up to the methodology used in the 2025 release of the GCP's fossil CO2 dataset.</p> <p>The fossil CO2 emissions dataset is included in both its standard, absolute form, and per capita, with associated metadata files in JSON format. A file indicating the source(s) of each data point is also provided.</p> <p>This is the initial release of the 2025 dataset.</p>
Global Claims Dataset
<p>Collection of claims collected from different fact-checking websites, covering various languages and topics. Described in "Global Claims: A Multilingual Dataset of Fact-Checked Claims with Veracity, Topic, and Salience Annotations"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.