Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

708

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

708 results for “Global dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

DOSE - Global dataset of reported subnational economic output

<p><strong>DOSE V2.11</strong><strong>&nbsp;is an update of DOSE V2. </strong>We made a few corrections and additions to the data where various users had identified gaps or inaccuracies. Please see 'DOSEV2.11_changes.pdf' for details.&nbsp;</p> <p><strong>DOSE &ndash; the MCC-PIK Database Of Sub-national Economic Output. </strong>DOSE v2 contains harmonised data on reported economic output for:</p> <ul> <li>1,661 sub-national regions</li> <li>across 83 countries</li> <li>from 1953 to 2020</li> <li>with sectoral detail for the agricultural, manufacturing and services sectors.</li> </ul> <p>To avoid interpolation, values were assembled from numerous statistical agencies, yearbooks and the literature and harmonised for both aggregate and sectoral output. In addition to regional economic output in local currency units (LCU) at current market prices as collected from the original data sources, DOSE contains&nbsp;per capita estimates in LCU and US dollars at both, current and 2015 market prices to enable comparison across time and space. Population data,&nbsp;market exchanges rates and deflator data used to generate them are included as well. Moreover, we provide temporally and spatially consistent data for regional boundaries, enabling matching with geo-spatial data such as climate observations. Annual temperature and precipitation data for each region are already included.&nbsp;Overall, DOSE provides the opportunity for detailed analyses of economic development at the subnational level, consistent with reported values.</p> <p>A peer-reviewed data descriptor with detailed documentation of&nbsp;the different data assembling, processing and validation steps as well as illustrative plots of the data set's coverage and examples for its application can be found here:&nbsp;</p> <p>L. Wenz, R.D. Carr, N. Koegel, M. Kotz, M. Kalkuhl.&nbsp;<a href="https://rdcu.be/dfTPH">DOSE &ndash; Global data set of reported sub-national economic output</a>. Nature Scientific Data. 2023.&nbsp;https://doi.org/10.1038/s41597-023-02323-8</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Data for: Downscaled gridded global dataset for Gross Domestic Product (GDP) per capita at purchasing power parity (PPP) over 1990-2022

<p>This dataset provides a gridded dataset for GDP per capita at purchasing power parity (PPP) downscaled to an admin 2 level (43,501 admin units). The dataset is based on reported subnational admin data (from 89 countries and 2,708 subnational units) and spans three decades from 1990 to 2022.&nbsp;</p> <p>The dataset is presented in details in the following publication.&nbsp;<strong><em>Please cite this paper when using data.&nbsp;</em></strong></p> <p>Kummu, M., Kosonen, M. &amp; Masoumzadeh Sayyar, S. 2025. Downscaled gridded global dataset for gross domestic product (GDP) per capita PPP over 1990&ndash;2022. Scientific Data 12: 178. <a href="https://doi.org/10.1038/s41597-025-04487-x" target="_blank" rel="noopener">https://doi.org/10.1038/s41597-025-04487-x</a></p> <p><strong>Code is available</strong> at: <a href="https://github.com/mattikummu/griddedGDPpc" target="_blank" rel="noopener">https://github.com/mattikummu/griddedGDPpc&nbsp;</a></p> <p>&nbsp;</p> <p><strong>The following data is given (formats in brackets)</strong></p> <ul> <li>GDP per capita (PPP) at admin 0 level (national) (GeoTIFF, gpkg, csv)</li> <li>GDP per capita (PPP) at admin 1 level (at the level of reporting, either admin 1 level or admin 0 level) (GeoTIFF, gpkg, csv)</li> <li>GDP per capita (PPP) at admin 2 level (downscaled from admin 1 level) (GeoTIFF, gpkg, csv)</li> <li>Total GDP (PPP), downscaled admin 2 level GDP per capita (PPP) multiplied by gridded population count, with three resolutions: 30 arc-sec, 5 arc-min, and 30 arc-min (GeoTIFF)&nbsp;</li> <li>Input data for the script that was used to generate the data above (code_input_data.zip). Code available at https://github.com/mattikummu/griddedGDPpc&nbsp;</li> </ul> <p><strong>Files are named as follows</strong><br><em>Format</em>: raster data (GeoTIFF) starts with rast_*, polygon data (gpkg) with polyg_*, and tabulated with tabulated_*.&nbsp;<br><em>Admin levels:</em> adm0 for admin 0 level, adm1 for admin 1 level, and adm2 for admin 2 level<br><em>Product type:</em> GDP per capita at purchasing power parity (PPP): _gdp_perCapita_; and total GDP at purchasing power parity (PPP): _gdp_tot_</p> <p>&nbsp;</p> <p><strong>Metadata&nbsp;</strong></p> <p><em>Grids for GDP per capita data:</em></p> <p>Resolution: 5 arc-min (0.083333333 degrees) &nbsp;(for admin 2 level also 30 arc-min, 0.5 degree, resolution is provided)</p> <p>Spatial extent: Lon: -180, 180; -90, 90 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: Multiband geotiff; each band for each year over 1990-2022&nbsp;</p> <p>Unit: USD in 2017 international dollars</p> <p>&nbsp;</p> <p><em>Grids for total GDP:</em></p> <p>Resolution: 30 arc-sec, 5 arc-min or 30 arc-min</p> <p>Spatial extent: Lon: -180, 180; -90, 90 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: Multiband geotiff; each band for each year over 1990-2022 (5 arc-min, 30 arc-min) or for each five years 1990, 1995, ... 2015, 2020 (30 arc-sec)</p> <p>Unit: USD in 2017 international dollars</p> <p>&nbsp;</p> <p><em>Geospatial polygon (gpkg) files:&nbsp;</em></p> <p>Spatial extent:&nbsp;-180, 180; -90, 83.67 (xmin, xmax, ymin, ymax)&nbsp;</p> <p>Temporal extent: annual over 1990-2022</p> <p>Coordinate ref system: EPSG:4326 - WGS 84&nbsp;</p> <p>Format: gkpk&nbsp;</p> <p>Unit: &nbsp;USD in 2017 international dollars</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

A Global Multi-Source Tropical Cyclone Precipitation (MSTCP) Dataset

<p>Tropical cyclone precipitation (TCP) is a key diagnostic in the context of atmospheric science, hazard, risk and flood research. This dataset provides estimates of TCP from global datasets. The various TCP metrics reported were estimated through the analysis of the global Multi-Source Weighted-Ensemble Precipitation (MSWEP) precipitation product and the International Best Track Archive for Climate Stewardship (IBTrACS) version 4. There are two main files that comprise the dataset. The main dataset file includes information on the mean and maximum TCP found within 500 km of each storm centre as well as the rainfall area and radius of maximum rain. The second file includes the estimates of azimuthally averaged precipitation using a 10 km bin spacing which is useful for analyses of the storm-scale structure of precipitation.</p>

openNov 2023View details →
zenodo40/100

Globe230k: A Benchmark Dense-Pixel Annotation Dataset for Global Land Cover Mapping

<p>We (Intelligent Mining and Analysis of Remote Sensing big data, IMARS) create a large-scale annotated dataset (Globe230k) for land use/land cover (LULC) mapping, which is annotated on Google Earth image of 1 m spatial resolution. Globe230k is annotated by numerous experts and students major in survey and mapping after necessary training, through visual interpretation on very high-resolution images, as well as in-situ field survey, under the guidance of the organized annotation pipeline. Globe230k has three superiorities:</p> <p>1) Large scale: the Globe230k includes 232,819&nbsp;annotated images with the size of 512x512 and spatial resolution of 1 m, with more than 3x1010 annotated pixels,&nbsp;and&nbsp;it includes&nbsp;10 first-level categories.&nbsp;</p> <p>2) Rich diversity: the annotated images are sampled from worldwide regions, with coverage area of over 60,000 km2, indicating a high variability and diversity.&nbsp;Besides, in order to ensure the category balance, we intentionally give more chance to the rare categories to be sampled, such as wetland, ice/snow, etc.</p> <p>3) Multi-modal: Globe230k not only contains RGB bands, but also include other important features for Earth system research, such as Normalized differential vegetation index (NDVI), digital elevation model (DEM), vertical-vertical polarization (VV) bands, vertical-horizontal polarization (VH) bands, which can facilitate the multi-modal data fusion research. Due to the large size of the multi-modal dataset (DEM 1.91G, NDVI 164G, VVVH 372G), these dataset are stored on Baidu Yunpan, the download link is :https://pan.baidu.com/s/12AKbiqOXSf4fnm7mYkCE0g?pwd=230k, the extraction code is 230k.</p> <p>The image patches and their corresponding annotated patches are respectively stored in "image_patch.zip" and "label_patch.zip" file. The RGB image is in forms of ".jpg", with size of 512x512, the pixel value is ranged from 0-255. The annotated patches is in forms of ".png", also with size of 512x512, the pixel value is ranged from 1-10, which respectively represent 1#cropland, 2#forest, 3#grass, 4#shrubland, 5#wetland, 6#water, 7#tundra, 8#impervious, 9#bareland, 10#ice/snow. The corresponding DEM, NDVI and VVVH patches are all in form of ".tif", with size of 512x512 (due to the different resolution of DEM, NDVI and VVVH patches, they are all uniformly resized to the same scale as the image patch).&nbsp;</p> <p>The total 232,819 pairs are officially divided into training set, validation set, and test set, based on ratio of 7:1:2, which can be find in "train_num.txt","val_num.txt","test_num.txt" file. Based on this division, the official baseline accuracy of several state-of-the-art semantic segmentation can be found in the related arcticle (https://spj.science.org/doi/10.34133/remotesensing.0078).</p> <p>We hope it can&nbsp;be used as a benchmark to promote further development of global land cover mapping and semantic segmentation algorithm development.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

The global 30-m human settlement dataset from 1985 to 2020

<p>GHS30 is the first global fine human settlement extents derived from the global 30 m impervious-surface (ISA) dynamic dataset and the global residential population of GHS-POP. It utilizes a refined classification system containing 4 human settlement types and covers the time span from 1985 to 2020 with an update cycle of 5-year. In specific, the global human settlement patches were delineated from the ISA data using a developed model that consisted of kernel density estimation, initial boundaries delineation, and morphological processing. The classification schema oriented to the typologies of human settlement patches were defined by combining the size of population and the population density within each patch that calculated from GHS-POP.</p> <p>The GEE code used has been deposited in GitHub at&nbsp;<a href="https://github.com/oucong1004/globalHumanSettlement">https://github.com/oucong1004/globalHumanSettlement</a>.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Early Eocene Global Vegetation Modern Plant Distribution Dataset

<p>Early Eocene Global Vegetation Modern Plant Distribution Dataset&nbsp;</p> <p>Global occurances for early Eocene fossil plant Nearest Living Relatives (NLRs) from the Global Biodiversity Information Facility (GBIF: https://www.gbif.org/), used for palaeocliate reconstruction.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Spatial dataset on global forest loss and gain across decades from 1960 to 2019

<p>This spatial dataset was used for creating Fig. 1a in the following research article.</p> <p>Estoque RC,&nbsp;<span>Dasgupta R,&nbsp;</span><span>Winkler K,</span><span> Avitabile V, </span><span>Johnson BA,&nbsp;</span><span>Myint SW,&nbsp;</span><span>Gao Y,&nbsp;</span><span>Ooba M,&nbsp;</span><span>Murayama Y, </span>and <span>Lasco RD</span> (2022). Spatiotemporal pattern of global forest change over the past 60 years and the forest transition theory. <em>Environmental Resesearch Letters, </em><strong>17</strong>:084022. https://doi.org/10.1088/1748-9326/ac7df5</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

A global land-use data cube 1992-2020 based on the Human Appropriation of Net Primary Production: Dataset 2

<p>This dataset is part of the LUIcube, a global dataset on land-use at 30 arcsecond spatial resolution. The LUIcube includes information on area, the change in NPP due to land conversions (HANPP<sub>luc</sub>), the harvested NPP (including losses, HANPP<sub>harv</sub>), and the NPP remaining in ecosystems after harvest (NPP<sub>eco</sub>) for 32 land-use classes in annual time-steps from 1992 to 2020. A detailed description of the LUIcube is available in the accompanying publication.</p> <p>The layers of land-use areas are provided in square kilometers (km&sup2;) per grid cell. All NPP flows are provided in tC/yr per grid cell. Adding HANPP<sub>harv</sub> to NPP<sub>eco</sub> results in the actual NPP available before harvest (NPP<sub>act</sub>=NPP<sub>eco</sub>+HANPP<sub>harv</sub>), and adding HANPP<sub>luc</sub> to NPP<sub>act</sub> results in the potential NPP available in the hypothetical absence of land use (NPP<sub>pot</sub>=NPP<sub>act</sub>+HANPP<sub>luc</sub>) for the given land-use class. Area-intensive values (in gC/m&sup2;/yr) can be calculated by dividing the NPP flows by the area of the respective land-use class per grid cell. HANPP in % of NPP<sub>pot</sub> can be calculated by summing up HANPP<sub>harv</sub> and HANPP<sub>luc</sub> and dividing it by NPP<sub>pot</sub>. Areas and NPP flows of land-use classes can be aggregated to calculate their overall HANPP.&nbsp;</p> <p>This Zenodo repository provides data on following land-use classes: grazing land characterized by open wooded lands (GL-owl)</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Dataset for 'Global declines in net primary production underestimated by climate models'

<p>The dataset provided here is to be used in conjuction with the JuPyTer notebook provided here: https://github.com/tjryankeogh/global_npp_trends/tree/main</p> <p>&nbsp;</p> <p>Download the file and uncompress in a root directory where there is a folder 'FIGURES'. When running the notebook make sure to change this root directory when importing packages.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Global ML-ready dataset for mining areas in satellite images

<p>This dataset is a global resource for machine learning applications in mining area detection and semantic segmentation on satellite imagery. It contains Sentinel-2 satellite images and corresponding mining area masks + bounding boxes for 1,210 sites worldwide. Ground-truth masks are derived from&nbsp;<a href="https://doi.org/10.1594/PANGAEA.942325" target="_blank" rel="noopener">Maus et al. (2022)</a> and <a href="https://doi.org/10.5281/zenodo.6806817" target="_blank" rel="noopener">Tang et al. (2023)</a>, and validated through manual verification to ensure accurate alignment with Sentinel-2 imagery from specific timestamps.&nbsp;</p> <p>The dataset includes three mask variants:</p> <ul> <li>Masks exclusively from Maus et al. (n=1,090)</li> <li>Masks exclusively from Tang et al. (n=817)</li> <li>A preferred mask selected from either Maus or Tang based on alignment quality determined during manual review (n=1,210).</li> </ul> <p>Each tile corresponds to a 2048x2048 pixel Sentinel-2 image, with metadata on mine type (surface, placer, underground, brine &amp; evaporation) and scale (artisanal, industrial). For convenience, the preferred mask dataset is already split into training (75%), validation (15%), and test (10%) sets.&nbsp;</p> <p>Furthermore, dataset quality was validated by re-validating test set tiles manually and correcting any mismatches between mining polygons and visually observed true mining area in the images, resulting in the following estimated quality metrics:&nbsp;</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Combined</td> <td>Maus</td> <td>Tang</td> </tr> <tr> <td>Accuracy</td> <td>99.78</td> <td>99.74</td> <td>99.83</td> </tr> <tr> <td>Precision</td> <td>99.22</td> <td>99.20</td> <td>99.24</td> </tr> <tr> <td>Recall</td> <td>95.71</td> <td>96.34</td> <td>95.10</td> </tr> </tbody> </table> <p>Note that the dataset does not contain the Sentinel-2 images themselves but contains a reference to specific Sentinel-2 images. Thus, for any ML applications, the images must be persisted first. For example, Sentinel-2 imagery is available from Microsoft's Planetary Computer and filterable via STAC API: <a href="https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a" target="_blank" rel="noopener">https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a</a>. Additionally, the temporal specificity of the data allows integration with other imagery sources from the indicated timestamp, such as Landsat or other high-resolution imagery.</p> <p>Source code used to generate this dataset and to use it for ML model training is available at&nbsp;<a href="https://github.com/SimonJasansky/mine-segmentation" target="_blank" rel="noopener">https://github.com/SimonJasansky/mine-segmentation</a>. It includes useful Python scripts, e.g. to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/05_persist_pixels_masks.py">download Sentinel-2 images via STAC API</a>, or to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/06_make_chips.py">divide tile images (2048x2048px) into smaller chips (e.g. 512x512px)</a>.&nbsp;</p> <p>A database schema, a schematic depiction of the dataset generation process, and a map of the global distribution of tiles are provided in the accompanying images.&nbsp;</p>

opencc-by-sa-4.0Nov 2024View details →
zenodo40/100

GlobalHighCO: Global Daily Seamless 1 km Ground-Level CO Dataset over Land (2018–Present)

<p>GlobalHighCO is part of a series of long-term, seamless, global, high-resolution, and high-quality datasets of air pollutants over land (i.e., GlobalHighAirPollutants, GHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived gapless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) global ground-level CO dataset over land <strong>from 2019 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.93 and a root-mean-square error (RMSE) of 0.21 mg m<sup>-3</sup> on a daily basis.</p> <p><strong>More GHAP datasets for different air pollutants are available at:&nbsp;<a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

GlobalHighO₃: Global Daily Seamless 10 km Ground-Level O₃ Dataset over Land (2000–Present)

<p>GlobalHighO<sub>3</sub> is part of a series of long-term, seamless, global, high-resolution, and high-quality datasets of air pollutants over land (i.e., GlobalHighAirPollutants, GHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the first big data-derived gapless (spatial coverage = 100%) daily, monthly, and yearly 10 km (i.e., D10K, M10K, and Y10K) global ground-level maximum daily 8-hour average (MDA8) O<sub>3</sub>&nbsp;dataset over land <strong>from 2000 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.86 and a root-mean-square error (RMSE) of 6.25 ppb on a daily basis.</p> <p><strong>More GHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Coding frame and dataset for the study: Land-use governance: The interplay of social, market, and policy drivers – A global systematic review

<p>This file entails the coding frame and dataset used to conduct the systematic literature review "Land-use governance: The interplay of social, market, and policy drivers &ndash; A global systematic review".</p> <p>This study was first published as Chapter 2 of the PhD Dissertation "From soil to society - Rethinking governance for multifunctional land use and management" (Elsa L. Dingkuhn, 2025), and in a modified form in the journal Earth System Governance (Dingkuhn et al. 2025).</p> <p>The file consists of three sheets:<br>- Coding frame: Includes coding instructions and definitions used to extract and categorize the data.<br>- Variables: A list of dataset variables with explanations.<br>- List of included studies: The 81 studies from which the data was sourced.<br>- Data_List of observations: The dataset itself, consisting of 718 observations extracted from the included studies.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Global surface water quality datasets under uncertain climate and socio-economic change, derived from the dynamical surface water quality model (DynQual) at 5 arcmin spatial resolution

<pre>Global ~10km (5 arcmin) surface water quality data from the dynamical surface water quality model (DynQual) from 2005-2100, with annual and monthly temporal resolution. Simulations are made under three combined climate and socio-economic scenarios (SSP1-RCP2.6; SSP3-RCP7.0 and SSP5-RCP8.5) and using five general circulation model (GFDL-ESM4; UKESM1-0-LL; MPI-ESM1-2-hr; IPSL-CM6A-LR and MRI-ESM2-0), following the ISIMIP3b protocol (<a href="https://protocol.isimip.org/#/ISIMIP3b">https://protocol.isimip.org/#/ISIMIP3b</a>). Output data are provided at annual and monthly temporal resolution over WorldClim time periods (2005-2020; 2021-2040; 2041-2060; 2061-2080; 2081-2100). Output data includes: - Discharge (m<sup>3</sup> s<sup>-1</sup>) - Water temperature (K)<br>- Total dissolved solids (TDS) load (g s<sup>-1</sup>)<br>- Biological oxygen demand (BOD) load (g s<sup>-1</sup>)<br>- Fecal coliform (FC) load (million cfu s<sup>-1</sup>) - Salinity; as indicated by TDS concentrations (mg l<sup>-1</sup>) - Organic pollution; as indicated by BOD concentrations (mg l<sup>-1</sup>) - Pathogen/bacterial pollution; as indicated by FC concentrations (cfu 100ml<sup>-1</sup>)<br><br>Note. A minimum discharge threshold of 0.1 m<sup>3</sup> s<sup>-1</sup> was used when computing TDS, BOD and FC concentrations, as uncertainties in absolute values of water availabilities have large impacts on resulting in-stream concentrations. Concentrations in these gridcells are assigned as NA.<br><br>Full time series of these variables at 30 arcmin (0.5 degree) can be found at: <a href="https://zenodo.org/records/14677534">https://zenodo.org/records/14677534</a>.</pre>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Global urban tree LAI/SAI dataset for urban climate modeling

<p>This dataset is the first global urban tree LAI/SAI product at a 500-meter resolution, specifically designed for urban climate modeling to simulate the tree's effects in urban environments. It covers the period from 2000 to 2022 and was developed by using a reprocessed MODIS LAI product with a Random Forest model, demonstrating high accuracy.</p> <p>The original product is a netCDF4 file that has been compressed into three tar.gz files: global_15s.tar.gz, global_0.05.tar.gz, and global_0.5.tar.gz, with resolutions of 500 m, 0.05&deg;, and 0.5&deg;, respectively. Each netCDF4 file in the compressed archive is named Global_UrbanTree_LAI_XX_YYYY.nc, where XX represents the resolution and YYYY represents the year. Each file contains monthly LAI/SAI data for that year, with data dimensions of mon x lat x lon.</p> <p>For version 3 of the LAI data, we replaced the meteorological data from WorldClim v2 with WorldClim&nbsp; v2.1 during model training. This version of the dataset has undergone peer review.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Supplementary dataset for "Global water gaps under future warming levels"

<p>The folder contains water gap data relative to the paper:<br>Rosa, L., Sangiorgio, M. Global water gaps under future warming levels. Nat Commun 16, 1192 (2025). https://doi.org/10.1038/s41467-025-56517-2</p> <p>All water gaps data are in km3/yr.</p> <p><br>Gridded data(NetCDF at 0.5&deg;)<br>&nbsp;- baseline<br>&nbsp; &nbsp;water_gap_baseline.nc: Baseline water gap in the period 2001-2010<br>&nbsp;- 1.5&deg;C warming (5 models + average)<br>&nbsp; &nbsp;water_gap_15C_average.nc: Water gap under 1.5&deg;C warming (multi-model average)<br>&nbsp; &nbsp;water_gap_15C_h08_ipsl-cm6a-lr.nc: Water gap under 1.5&deg;C warming (h08 + ipsl-cm6a-lr)<br>&nbsp; &nbsp;water_gap_15C_h08_mri-esm2-0.nc: Water gap under 1.5&deg;C warming (h08 + mri-esm2-0)<br>&nbsp; &nbsp;water_gap_15C_h08_ukesm1-0-ll.nc: Water gap under 1.5&deg;C warming (h08 + ukesm1-0-ll)<br>&nbsp; &nbsp;water_gap_15C_h08_mpi-esm1-2-hr.nc: Water gap under 1.5&deg;C warming (h08 + mpi-esm1-2-hr)<br>&nbsp; &nbsp;water_gap_15C_h08_gfdl-esm4.nc: Water gap under 1.5&deg;C warming (h08 + gfdl-esm4)<br>&nbsp;- 3&deg;C warming (5 models + average)<br>&nbsp; &nbsp;water_gap_3C_average.nc: Water gap under 3&deg;C warming (multi-model average)<br>&nbsp; &nbsp;water_gap_3C_h08_ipsl-cm6a-lr.nc: Water gap under 3&deg;C warming (h08 + ipsl-cm6a-lr)<br>&nbsp; &nbsp;water_gap_3C_h08_mri-esm2-0.nc: Water gap under 3&deg;C warming (h08 + mri-esm2-0)<br>&nbsp; &nbsp;water_gap_3C_h08_ukesm1-0-ll.nc: Water gap under 3&deg;C warming (h08 + ukesm1-0-ll)<br>&nbsp; &nbsp;water_gap_3C_h08_mpi-esm1-2-hr.nc: Water gap under 3&deg;C warming (h08 + mpi-esm1-2-hr)<br>&nbsp; &nbsp;water_gap_3C_h08_gfdl-esm4.nc: Water gap under 3&deg;C warming (h08 + gfdl-esm4)</p> <p><br>Aggregated data (.xlsx)<br>&nbsp;- source_data.xlsx: Water gap aggregated by country and basin for all the considered scenarios (including multi-model average and agreement analysis)</p> <p>Note: the global water gap obtained by summing all the countries/basins is not completely equivalent to the sum of all the pixels because some pixels' center is outside the polygon of the corresponding country/basin (differences in the order of 1km/yr3, &lt;0.3%).&nbsp;See sheet "Figure 3" A249:N251 (countries) and "Figure 5" A235:N237 (basins) for further details.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Improving 30-meter global impervious surface area (GISA) mapping: New method and dataset

<p>Timely and accurate monitoring of impervious surface areas (ISA) is crucial for effective urban planning and sustainable development. Recent advances in remote sensing technologies have enabled global ISA mapping at fine spatial resolution (&lt;30 m) over long time spans (&gt;30 years), offering the opportunity to track global ISA dynamics. However, existing 30 m global long-term ISA datasets suffer from omission and commission issues, affecting their accuracy in practical applications. To address these challenges, we proposed a novel global longterm ISA mapping method and generated a new 30 m global ISA dataset from 1985 to 2021, namely GISA-new. Specifically, to reduce ISA omissions, a multi-temporal Continuous Change Detection and Classification (CCDC) algorithm that accounts for newly added ISA regions (NA-CCDC) was proposed to enhance the diversity and representativeness of the training samples. Meanwhile, a multi-scale iterative (MIA) method was proposed to automatically remove global commissions of various sizes and types. Finally, we collected two independent test datasets with over 100,000 test samples globally for accuracy assessment. Results showed that GISA-new out performed other existing global ISA datasets, such as GISA, WSF-evo, GAIA, and GAUD, achieving the highest overall accuracy (93.12 %), the lowest omission errors (10.50 %), and the lowest commission errors (3.52 %). Furthermore, the spatial distribution of global ISA omissions and commissions was analyzed, revealing more mapping uncertainties in the Northern Hemisphere. In general, the proposed method in this study effectively addressed global ISA omissions and removed commissions at different scales. The generated high-quality GISAnew can serve as a fundamental parameter for a more comprehensive understanding of global urbanization.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

GTSM-ERA5-E dataset - Data underlying the paper "Global dataset of storm surges and extreme sea levels for 1950-2024 based on the ERA5 climate reanalysis"

<p>Extreme sea levels, generated by storm surges and high tides, have the potential to cause coastal flooding and erosion. Global datasets are instrumental for mapping of extreme sea levels and associated societal risks. Harnessing the backward extension of the ERA5 reanalysis, we present a dataset containing the statistics of water levels based on a global hydrodynamic model (GTSMv3.0) covering the period 1950-2024. This is an extension of a previously published dataset for 1979-2018 <a href="https://www.frontiersin.org/articles/10.3389/fmars.2020.00263/full" target="_blank" rel="noopener">(Muis et al. 2020)</a>. The timeseries (10-min, hourly mean and daily maxima) are available via the Climate Data Store of ECMWF at DOI: 10.24381/cds.a6d42d60. Using this extended ERA5 dataset, we calculate percentiles and estimate extreme water levels for various return periods globally. The percentiles dataset includes the 1, 5, 10, 25, 50, 75, 90, 95 and 99th percentiles. The extreme water levels include return values for 1, 2, 5, 10, 25, 50, 75 and 100 years, and they are estimated using POT-GPD method applied with a threshold of 99th percentile of the timeseries and using a 72-hour window for declustering peak events, and MLE method for fitting the GPD parameters. The parameters (shape, scale and location) are also supplied with this dataset.</p> <p>Validation of the underlying timeseries and the statistical values shows that there is a good agreement between observed and modelled sea levels, with the level of agreement being very similar to that of the previously published dataset. &nbsp;The extended 75-year dataset allows for a more robust estimation of extremes, often resulting in smaller uncertainties than its 40-year precursor. The present dataset can be used in global assessments of flood risk, climate variability and climate changes.</p> <p>Global modelling of water levels and extreme value analysis are associated with a number of uncertainties and limitations, that are particularly important to consider when conducting local assessments. Please refer to the Usage Notes in the corresponding manuscript (Aleksandrova et al. 2025, paper currently under review) for an overview of limitations.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Dataset for ´´A New Detailed Global Map of Lunar Light Plains´´ research article

<p>The shapefiles (.shp) provided in this repository are the datasets for the paper &acute;A new detailed global map of lunar light plains&acute; published in PSJ journal Special Issue.&nbsp;</p> <p>These shapefiles can be directly imported in ArcMap/ArcPRO. The third dataset is a .tif or image of the global map for a fast and easy overview.</p> <p>Two geomorphologic maps of lunar light plains are provided as described in the article: one with an FeO wt% cut off of about 12 wt% (Area_lightplains), and the other around 8 wt% (Area_LPFeOLow).&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

A Global Dataset of Standardized Moisture Anomaly Index Incorporating Snow Dynamics (SZIsnow) from 1948 to 2010

<p>The SZI<sub>snow</sub> dataset was calculated based on systematic physical fields from the Global Land Data Assimilation System Version 2 (GLDAS-2) with the Noah land surface model. This SZI<sub>snow</sub> dataset considers different physical water-energy processes, especially snow processes. The evaluation shows the dataset is capable of investigating different types of droughts across different timescales. The assessment also indicates that the dataset has an adequate performance to capture droughts across different spatial scales. The consideration of snow processes improved the capability of SZI<sub>snow</sub>, and the improvement is evident over snow-covered areas (e.g., Arctic region) and high-altitude areas (e.g., Tibet Plateau). Moreover, the analysis also implies that SZI<sub>snow</sub> dataset is able to well capture large-scale drought events across the world. This drought dataset has high application potential for monitoring, assessing, and supplying information on drought, and also can serve as a valuable resource for drought studies.</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record