Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
134
datasets available to search
ShareScore release 0.9.0
Dataset results
134 results for “global estimates”
Estimated population exposed to floods with a global flood model (CaMa-Flood v4.00)
<p>This is a repository for data and codes to analyze the population exposed to floods at grid level (0.25deg) and country level. The flood hazard map is modeled from the global hydrodynamic model - Catchment-based Macro-scale Floodplain (CaMa-Flood) model. </p>
Data from: Estimating maize harvest index and nitrogen concentrations in grain and residue using globally available data
<p class="MsoNormal">Estimates of crop nitrogen (N) uptake and offtake are critical in estimating N balances, N use efficiencies and potential losses to the environment. Calculation of crop N uptake and offtake requires estimates of crop product yield (e.g. grain or beans) and crop residue yield (e.g. straw or stover) and the N concentration of both components. Yields of crop products are often reasonably well known, but those of crop residues are not. While the harvest index (HI) can be used to interpolate the quantity of crop residue from available data on crop product yields, harvest indices are known to vary across locations, as do N concentrations of residues and crop products. The increasing availability of crop data and advanced statistical and machine learning methods present us with an opportunity to move towards more locally relevant estimates of crop harvest index and N concentrations using more readily available data. This dataset includes maize field experiment data. It is a culmination of summary statistic data collected from the literature as well as raw data requested from various researchers and organisations from around the world. These data will enable more locally relevant estimates of crop nutrient offtake, nutrient balances and nutrient use efficiency at national, regional or global levels, as part of strategies towards more sustainable nutrient management.</p>
Global wetland CH4 emissions estimated by LPJ-wsl model for 1980-2021
<p>This dataset contains two files that summarize the regional wetland methane emissions simulated by LPJ-wsl model for the period of 2000-2021. two estimates based on two climate forcing datasets, ground-based CRU and reanalysis-based MERRA2. The ground-based input dataset is a monthly climatic observation based on meteorological stations developed by Climatic Research Unit, University of East Anglia. The reanalysis-based climate dataset is a daily climatic dataset from 1 hourly reanalysis Modern-Era Retrospective analysis for Research and Applications Version 2. For more details about LPJ-wsl model, see Zhang et al., (2017, 2018).</p> <p>References:</p> <p>Zhang, Z., Zimmermann, N.E., Stenke, A., Li, X., Hodson, E.L., Zhu, G., Huang, C., Poulter, B., 2017. Emerging role of wetland methane emissions in driving 21st century climate change. Proceedings of the National Academy of Sciences 114, 9647–9652. <a href="https://doi.org/10.1073/pnas.1618765114">https://doi.org/10.1073/pnas.1618765114</a></p> <p>Zhang, Z., Zimmermann, N.E., Calle, L., Hurtt, G., Chatterjee, A., Poulter, B., 2018. Enhanced response of global wetland methane emissions to the 2015–2016 El Niño-Southern Oscillation event. Environmental Research Letters 13, 074009. <a href="https://doi.org/10.1088/1748-9326/aac939">https://doi.org/10.1088/1748-9326/aac939</a></p>
A Global Gridded Municipal Water Withdrawal Estimation Method Using Aggregated Data and Artificial Neural Network
<p>Global gridded municipal water withdrawal estimations for the following WST paper.</p> <p>Jiabao Yan, Shaofeng Jia; A global gridded municipal water withdrawal estimation method using aggregated data and artificial neural network. <em>Water Science Technology</em>, 2023; 87 (1): 251–274. <a href="https://doi.org/10.2166/wst.2022.399" target="_blank" rel="noopener">https://doi.org/10.2166/wst.2022.399</a></p> <p>The representative year of the data is 2015, and the unit of the data is in millimeters (mm).</p>
The role of OCO-3 XCO2 retrievals in estimating global terrestrial net ecosystem exchanges
<p>1.<strong>Exp_OCO3.zip </strong> includes regional posterior carbon fluxes from the assimilation using OCO-3 observations.</p> <p>2.<strong>Exp_OCO2.zip </strong> includes regional posterior carbon fluxes from the assimilation using OCO-2 observations.</p> <p>3.<strong>Exp_OCO3&2.zip </strong> includes regional posterior carbon fluxes from the joint assimilation using OCO-3 and OCO-2 observations together.</p> <p>4.<strong>posterior.fluxes.Exp_OCO3.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the assimilation using OCO-3 observations during August 2019 to December 2022.</p> <p>5.<strong>posterior.fluxes.Exp_OCO2.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the assimilation using OCO-2 observations during August 2019 to December 2022.</p> <p>6.<strong>posterior.fluxes.Exp_OCO3&2.nc </strong>includes information on the spatial distribution of annual as well as monthly posterior carbon fluxes from the joint assimilation using OCO-3 and OCO-2 observations together during August 2019 to December 2022.</p> <p>7.<strong>evaluation_result.txt</strong> includes the results of the evaluation of posterior carbon fluxes using independent CO2 observations from 66 surface flask sites.</p> <p> </p> <p> </p> <p> </p>
Estimating transpiration globally by integrating the Priestley-Taylor model with neural networks
Open the record for dataset details and reuse information.
Improved estimation of global soil moisture of different shared socioeconomic pathways for 2016-2099
<p>We design a novel Transformer SM Simulation Net (TSMSNet) to conduct global monthly 0.5°×0.5° SM simulation of SSP1-2.6, SSP2-4.5, and SSP5-8.5 from 2016 to 2099. Nine qualified future SM datasets, along with corresponding spatial distribution dataset of error parameters, and geographic background data are selected as model inputs. The learning target is calculated through merging the merits from Soil Moisture Active Passive (SMAP) and European Centre for Medium-Range Weather Forecast Reanalysis v5-Land (ERA5-Land) SM. The results indicate the TSMSNet simulated SM (R = 0.68, ubRMSE = 0.045 m<sup>3</sup>/m<sup>3</sup>) performs notable superiority in matching both temporal variation and absolute value of in-situ measurements across different landcover and climate regions. Besides, the TSMSNet simulated SM could favorably match the long-term trend of learning target. TSMSNet simulated SM has an overwhelming drying trend during 2016 to 2099. The decline magnitude rises accompanied by SSP changing from sustainable pathway to fossil-fueled development. The areas with significant drying trend mainly distributed in plateau and mid-latitude region. In terms of land cover types, evident drying trends are found in cropland and forest. SM shows faster descent rate in livable areas than unlivable areas. In summary, we develop a reliable future SM dataset, that is expected to act as a valuable reference for understanding future water cycle patterns.</p>
Geospatial micro-estimates of slum populations in 129 Global South countries using machine learning and public data
<p><span>Reliable estimation of populations living in slums or slum-like conditions is crucial for urban planning, humanitarian resource allocation, and human well-being improvement. We generate the micro-estimate of slum population at a neighborhood level (~</span><span>3.63 arc-minutes</span><span>, preserving the privacy of vulnerable people) for 129 Global South countries in 2018. The estimates are built based on the Sustainable Development Goals 11.1 indicator framework and machine learning algorithms to heterogeneous data from household-based surveys and satellite images, as well as grided population data. Our integrated regional models show strong predictive capabilities for cluster-level slums proxy, explaining 82% to 96% of the variation in ground-truth surveys conducted in Global South countries, with root mean squared error ranging from 4.85% to 10.47%. The models perform match or surpass benchmarks established by previous studies.</span><span> </span><span>Cross-comparison with independent data sources at multi-scales suggest that our approach can yield reliable and consistent slum population estimates.</span></p>
Data files for "Estimating a Social Cost of Carbon for Global Energy Consumption"
<p>Data files for "Estimating a Social Cost of Carbon for Global Energy Consumption".</p> <p>Findings of the paper can be replicated using these data files, along with code at https://github.com/ClimateImpactLab/energy-code-release-2020/.</p> <p> </p>
Harmonising the land-use flux estimates of global models and national inventories for 2000-2020: background data
<p>This online repository includes all the relevant data used in the paper "Harmonizing the land-use flux estimates of global models and national inventories for 2000-2020" (Grassi et al. 2023), plus some additional methodological information, organised in the following files:</p> <p>1) "<strong>Global model</strong><strong>s</strong> <strong>land CO2 </strong><strong>data 2000-2020</strong>" (MS Excel Format), including for each country data for:</p> <p>a. Land-use CO2 fluxes from each of three Bookkeeping Models (BMs) used, and for different categories (net LULUCF, deforestation, forest, other transitions, organic soils). </p> <p>b. The ensemble mean of the ‘natural terrestrial sink’ estimated by 16 Dynamic Global Vegetation Models (DGVMs), filtered with maps of intact/non-intact forest.</p> <p>The global model data included here are consistent with those included in the Global Carbon Budget 2022 (Friedlingstein et al., 2022).</p> <p>2) “<strong>National inventories LULUCF data 2000-2020</strong>” (version Dec 2022, MS Excel Format), including a comprehensive collection of LULUCF CO2 data based on countries' submissions to the United Nations Framework Convention on Climate Change (UNFCCC). The data here represent a slight update of the dataset included in Grassi et al. (2022).</p> <p>3) “<strong>Processing steps for DGVM results”,</strong> describing the protocol used to filter the results of DGVMs with maps of intact/non-intact forest and further details on the maps (PDF Format). </p> <p>4) “<strong>Intact and non-intact forest maps</strong>”, available in two files with different resolutions (0.5 and 0.05 degrees) in NetCDF format. Grassi et al. (2023) used the 0.5 degree resolution.</p> <p>5) "<strong>IntactAndNonIntactForest_0.5deg_script.js</strong>", the Google Earth Engine Java script to produce the forest maps (.js/text format)</p> <p>For further details, please refer to:</p> <p>Grassi et al. (2023) Harmonising the land-use flux estimates of global models and national inventories for 2000-2020. Earth Syst. Sci. Data.</p> <p>Other references:</p> <p>Friedlingstein et al. (2022) Global Carbon Budget 2022, Earth Syst. Sci. Data, 14, 4811–4900.</p> <p>Grassi et al (2022) Carbon fluxes from land 2000–2020: bringing clarity to countries' reporting. Earth Syst. Sci. Data, 14, 4643-4666.</p> <p> </p> <p> </p>
ODYM-RECC Copper dataset, used for sector-level estimates for global future copper demand and the potential for resource efficiency
<p>Complete Model database with 105 parameter files used in the ODYM-RECC Copper model (github: https://github.com/SteffiKlose/ODYM-RECC-Copper.git) used for Sector-level estimates for global future copper demand and the potential for resource efficiency (<a href="https://doi.org/10.1016/j.resconrec.2023.106941">https://doi.org/10.1016/j.resconrec.2023.106941</a>)</p> <p>This dataset is based on the Database of the ODYM-RECC v2.4 model, used for the GLOBAL case study on material efficiency and climate change mitigation (https://doi.org/10.5281/zenodo.4671644)</p>
Estimating global GPP from the plant functional type perspective using a machine learning approach
<p><span>The long-term monitoring of gross primary production (GPP) is crucial to the assessment of the carbon cycle of terrestrial ecosystems. In this study, a well-known machine learning model (Random Forest, RF) is established to reconstruct the global GPP dataset named ECGC_GPP. The model distinguished nine functional plant types, including C3 and C4 crops, using eddy fluxes, meteorological variables, and leaf area index as training data of the RF model. Based on ERA5_Land and the corrected GEOV2 data, the global monthly GPP dataset at a 0.05-degree resolution from 1999 to 2019 was estimated. The results showed that the RF model could explain 74.81% of the monthly variation of GPP in the testing dataset, of which the average contribution of Leaf Area Index (LAI) reached 41.73%. The average annual and standard deviation of GPP during 1999–2019 were 117.14 ± 1.51 Pg C yr<sup>-1</sup>, with an upward trend of 0.21 Pg C yr<sup>-2</sup> (<em>p</em> < 0.01). By using the plant functional type classification, the underestimation of cropland is improved. Therefore, ECGC_GPP provides reasonable global spatial patterns and long-term trends of annual GPP.</span></p>
WOMBAT v2.0 estimates of global GPP, respiration, and air-sea fluxes
These files contain estimates of global CO2 fluxes, split into GPP, respiration, and air-sea components produced by the WOMBAT v2.0 flux-inversion system (see https://arxiv.org/abs/2210.10479). These fluxes are further decomposed into trend and seasonality. The file WOMBAT_v2_CO2_gridded_climatology_samples.nc4 contains the estimated spatial fields for each beta in the paper that define the trend/seasonality of the fluxes. The bottom-up estimates are provided for each beta, as well as samples from the posterior distribution. A second file, WOMBAT_v2_CO2_gridded_flux_samples.nc4, contains bottom-up estimates and posterior samples for the flux component fields. These are split into components and the parts of the decomposition, so for example the linear component of GPP is in gpp_linear_bottom_up/gpp_linear_posterior for the bottom-up/posterior. The different fields can be summed to get meaningful quantities, such as the NEE (sum of all GPP and respiration parts), or the net flux (sum of all parts). The final file, samples-LNLGIS.rds, contains samples from the posterior distribution of the model parameters. This file format may be read using the readRDS function in the R programming language.
Supplementary files for "Data-driven estimation of nitric oxide emissions from global soils based on dominant vegetation covers"
<p>In-situ observations collected from publications, data-driven model codes, DNDC simulation files, and the supporting data for all figures of this study are uploaded. In-situ observations included 1,356 observations of soil nitric oxide (NO) emissions from 192 sites, including 1,032 for cropland soils from 70 sites, 114 for grassland soils from 36 sites, and 208 for forest soils from 86 sites. Data-driven models provided three machine learning methods, including random forest (RF), generalized boosted regression model (GBM), and radial basis function (RBF). The DNDC simulation files included simulation files of 51 selected sites.</p>
Most global gauging stations present biased estimations of total catchment discharge
<p>This dataset includes the data generated for the main figures, the codes of the Mabcd model used to simulate catchment hydrological processes in this study, and the neural network model (NN2) trained to evaluate the Budyko parameter ω.</p>
Data from: Estimating potential global sources and secondary spread of freshwater invasions under historical and future climates
<p>Aim: We employ a climate-matching method to evaluate potential source regions of freshwater invasive species to an introduced region and their potential secondary spread under historical and future climates.</p> <p>Location: Global source regions, with primary introductions to the Laurentian Great Lakes and secondary introductions throughout North America</p> <p>Methods: We conducted a climate-match analysis using the CLIMATE algorithm to estimate global source freshwater ecoregions under historical and future climates with an ensemble of general circulation models for climate change scenario SSP5-8.5. Given existing research, we use a climate match of ≥ 71.7% between ecoregions to indicate climatic conditions that will not inhibit the survival of introduced freshwater organisms. Further, we estimate the secondary spread of freshwater invaders to the ecoregions of North America under historical and future climates.</p> <p>Results: We identified 54 global freshwater ecoregions with a climate match ≥ 71.7% to the recipient Laurentian Great Lakes under historical climatic conditions and 11 additional ecoregions were predicted to exceed the threshold under climate change. Three of the 11 ecoregions were located in South America, a continent where no matches existed under historical climates and eight were located in the southern United States, southern Europe, Japan, and New Zealand. Further, we identify 34 North American ecoregions of potential secondary spread of freshwater invasions from the Great Lakes under historical climatic conditions, and five ecoregions were predicted to exceed the threshold under climate change.</p> <p>Main conclusion: We provide a climate-match method that can be employed to assess the sources and spread of freshwater invasions under historical and future climate scenarios. Our climate-match method predicted increases in climate match between the recipient region and several potential source regions, and changes in areas of potential spread under climate change. The identified ecoregions are candidates for detailed biosecurity risk assessments and related management actions. The identified ecoregions are candidates for detailed biosecurity risk assessments and related management actions.</p>
Dataset associated with "Effect of sampling bias on global estimates of ocean carbon export"
<p>Dataset and Matlab code for plotting the figures in the manuscript "Effect of sampling bias on global estimates of ocean carbon export", submitted to Geophysical Research Letters.</p>
Global sea-surface DMS concentrations estimated through data-driven approaches
<p>DMS_ML_average.nc and DMS_SAT-OPT_average.nc represent the 10-year (2011-2020) average monthly seawater DMS concentrations (1°x1°) calculated using two different methods, ML and SAT-OPT, respectively.</p>
Data from: Bayesian estimation of the global biogeographical history of the Solanaceae
Open the record for dataset details and reuse information.
Data from: Estimating potential global sources and secondary spread of freshwater invasions under historical and future climates
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.