Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,009
datasets available to search
ShareScore release 0.9.0
Dataset results
2,009 results for “Global data”
JUMP - Data collection - Part II: Zonal jets using three different approaches, laboratory - Global Climate Models - observations.
<p>The formation of large scale structures in three-dimensional (3D) turbulent flows. How small-scale dynamics organize in turbulent flows to grow large scale coherent circulation? is at the heart of fundamental studies in fluid dynamics. It appears to be equally important for our understanding of atmospheric dynamics, oceanography, meteorology and more generally geophysical fluid dynamics. Here, we deliver a data collection that <strong>(1)</strong> gathers measurements of 3D turbulent flows that emulate planetary atmospheres of the gas giants. Turbulent flows are explored using three different approaches, laboratory experiments, numerical simulations and direct planetary observations. All data set are computed in order to easily extract flow properties, i.e. high resolution maps of the different velocity components and flow vorticity (useful for further diagnostic). The data collected are fully discribed in Cabanes et al GRL (2020) "Revealing the intensity of turbulent energy transfer in planetary atmospheres" and can be used to compute <strong>(2)</strong> theoretical diagnostics with the numerical codes that allow to reveal the physical meaning of flow measurements. Numerical codes are available on https://github.com/scabanes</p> <p>We deliver (1) data collection and (2) numerical codes in the following files attached:</p> <p>(1) Data collection:</p> <ul> <li>A PDF file named <strong>JUMP-zonal-jets-data-collection-GRL.pdf</strong> that describes the following data files and nomenclature.</li> <li>A zip File of the velocity fields in the lab, interpolated on Polar and Cartesian grids <ul> <li><strong>JUMP-JetsInTheLab.zip</strong></li> </ul> </li> <li>A netcdf file of velocity fields of our Saturn reference simulation <ul> <li><strong>uvData-SRS-istep-312000-nstep-50-niz-12.nc</strong></li> </ul> </li> <li>Two netcdf files of velocity fields from Cassini observations of Jupiter<strong> </strong> <ul> <li><strong>uvData-JupObs-istep-0-nstep-4-niz-1.nc</strong></li> <li><strong>StatisticalData-JupObs.nc</strong></li> </ul> </li> <li>A zip file of potential vorticity profiles for Saturn and Jupiter observations <ul> <li><strong>IPV-QGPV-Jupiter-Saturn.zip</strong></li> </ul> </li> </ul> <p>(2) Numerical codes:</p> <ul> <li>Codes for statistical analysis in spherical geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FPOST&sa=D&sntz=1&usg=AFQjCNFuDU0eij4XGxQfReO92CHfJz6PBA">https://github.com/scabanes/POST</a></li> <li>Codes for statistical analysis in cylindrical geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FJUMP&sa=D&sntz=1&usg=AFQjCNGUQ1YIFhSxBAg4Hl_5gOLB_4LxLA">https://github.com/scabanes/JUMP</a></li> <li>Codes for statistical analysis in cartesian geometry on Github. --> <a href="https://www.google.com/url?q=https%3A%2F%2Fgithub.com%2Fscabanes%2FJUMP&sa=D&sntz=1&usg=AFQjCNGUQ1YIFhSxBAg4Hl_5gOLB_4LxLA">https://github.com/scabanes/JUMP</a></li> </ul> <p> </p> <p>The purpose of this data collection is to reveal statistical properties of planetary flows. By computing the same analysis on different data sets the researcher allows direct confrontation of planetary observations with idealized laboratory and numerical models. Idealized models are specially designed to sweep on a large array of parameters in order to understand what parameters control planetary global circulation. The data collected and generated by the researcher deliver <strong>(1)</strong> velocity measurements of 3D turbulent flows using the different approaches (observations-laboratory-numerics) and <strong>(2)</strong> guidelines to compute the appropriate statistical analysis through the PTST. Here, the ground-breaking novelty is that the researcher deliver the possibility to compute statistical diagnostics adapted to the different geometries: the spherical geometry of planetary flows, i.e. 2D latitude-longitude maps, the cylindrical geometry of laboratory experiments, i.e. 2D flows in a rotating cylindrical tank, and the Cartesian geometry of idealized numerical simulations. Indeed, the math behind each statistical diagnostics must account for the different geometrical configurations in order to properly confront the different approaches. The PTST is also designed to be easily re-used by different communities such as experimentalists, numericists and atmosphericists that deal with 3D or 2D turbulent flows.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement N° 797012.</p>
Global vegetation productivity from 1981 to 2018 estimated from remote sensing data
<p>The MUltiscale Satellite remotE Sensing (MUSES) global vegetation productivity dataset includes gross primary productivity (GPP) and net primary productivity (NPP) data from 1981 to 2018. GPP and NPP were estimated with a light use efficiency (LUE) model and MUSES leaf area index (LAI) and fraction of absorbed photosynthetically active radiation (FAPAR) products.</p> <p>The MUSES product suite includes products with different spatial and temporal resolutions for parameters such as Normalized Difference Vegetation Index (NDVI), Near-Infrared Reflectance of Vegetation (NIRv), Leaf Area Index (LAI), Fraction of Absorbed Photosynthetically Active Radiation (FAPAR), Fractional Vegetation Coverage (FVC), Gross Primary Production (GPP), Net Primary Production (NPP). For more information about the MUSES products, please refer to this website (<a href="https://muses.bnu.edu.cn/">https://muses.bnu.edu.cn/</a>).</p> <p>The detail information of the MUSES 5-km global GPP and NPP products are as below:</p> <p>Name: MUSES 5-km global GPP and NPP products</p> <p>Period: 1981-2018</p> <p>Spatial resolution: 0.05°</p> <p>Temporal resolution: 8 days</p> <p>Projection: geographic latitude/longitude</p> <p>Data format: Tiff</p> <p>Data type: integer (16bit)</p> <p>Upper left coordinates: -180°E, 90°N</p> <p>Scale factor: 100</p> <p>Unit: gCm<sup>-2</sup>d<sup>-1</sup></p> <p> </p> <p><span>Citation (Please cite these papers when these data are used)</span></p> <p><span>1. Wang, J.M., Sun, R., Zhang, H.L., Xiao, Z.Q., Zhu A.R., Wang, M.J., Yu, T., Xiang, K.L.,</span><span> </span><span>New global MuSyQ GPP/NPP remote sensing products from 1981 to 2018. </span><span>IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14, 5596-5612.</span></p> <p><span>2. Wang, M.J.; Sun, R.; </span><span>Zhu, A.R.;</span><span> Xiao, Z. Q. Evaluation and Comparison of Light Use Efficiency</span><span> </span><span>and Gross Primary Productivity Using Three</span><span> </span><span>Different Approaches. <span>Remote Sensing</span>. <span>2020</span>, 12, 1003.</span></p> <p><span>3. Yu, T.; Sun, R.; Xiao, Z.Q. ;Zhang , Q.; Liu, G.; Cui, T.X.; Wang, J.M. Estimation of Global Vegetation Productivity from Global LAnd Surface Satellite Data. <span>Remote sensing. </span><span>2018, </span>10, 327.</span></p>
Data from: The impact of human mobility networks on the global spread of COVID-19
<p>This is empirical dataset from the paper "The impact of human mobility networks on the global spread of COVID-19". Specifically, the dataset includes several files: (a) the COVID-19 network - an origin/destination matrix (i.e., "covid_network.csv"); (b) the common language network - edgelist format (i.e. "edge_list_comlang.csv"); (c) the same continent network - edgelist format (i.e., "edge_list_continent.csv"; (d) the contiguity network (i.e., "edge_list_contig.csv"); (e) the migration network - edgelist format (i.e., "edge_list_migration_in.csv"; (f) the tourism network - edgelist format (i.e., edge_list_tourism_in.csv"); (g) the list of nodes (countries) corresponding to files (b)-(e) (i.e., "nodes.csv"). Additionally, we uploaded the Rcode used in the paper (i.e. "code"), as a .pdf file format, the data source for the figures included in the paper (i.e., "covid_network_matrix.csv", "matrix_migration_out.csv", "matrix_tourism.csv" - Figure 1; "Fig_2_a_matrix_comlang.csv", Fig_2_b_matrix_contig.csv", "Fig_2_c_matrix_continent.csv" - Figure 2; "Fig_3.graphmlz - Figure 3; Fig_4.graphmlz - Figure 4) and the "global network of COVID-19 onset" (an individual-level data) (i.e., "global_covid_network.csv"). </p> <p>For details, please, see the Methods section of the paper: The impact of human mobility networks on the global spread of COVID-19 (Hancean, M.-G., Slavinec, M., Perc, M). </p> <p> </p> <p> </p> <p> </p>
Data Visualization - Final Project - Global Climate Change
<p>This Project is part of the course work for Data visualization DATS 6401. In this project, I have created webpage to show data analysis on Global Climate Change. D3 & Google Visualization API is used for all visualization graphs in the webpage.</p>
Global Carbon Budget 2023, surface ocean fugactiy of CO2 (fCO2) and air-sea CO2 flux of individual global ocean biogechemical models and surface ocean fCO2-based data-products
<p><strong>Surface ocean fugacity of CO2 (fCO2) and air-sea CO2 flux data from individual Global Ocean Biogeochemistry Models (GOBMs) and surface ocean fCO2-based data-products (fCO2-products).</strong><br>There are three types of files: (1) one file per fCO2-product with gridded fields and regionally-integrated CO2 flux time-series, (2) one file per GOBM with gridded fields, and (3) one file with the regionally-integrated time-series for the GOBMs. </p><p><strong>Note: </strong>These provided gridded outputs from fCO2-based data-products and GOBMs are regridded datasets, without adjustments. <strong>The best estimates of the annual global ocean carbon sink, based on the native grids of fCO2-products and GOBMs and with the adjustments described in the Global Carbon Budget 2023 (https://doi.org/10.5194/essd-15-5301-2023), are available in the Global Carbon Budget 2023 spreadsheet.</strong></p><p>The regionally-integrated time-series are as provided by the contributing groups, i.e. integrated from their native grids. In order to reproduce Figure 13 of the Global Carbon Budget 2023 paper (https://doi.org/10.5194/essd-15-5301-2023), the river flux adjustment needs to be added to the CO2 flux estimated from the data-products (North: 0.14 GtC yr-1, Tropics: 0.42 GtC yr-1, South: 0.09 GtC yr-1, see GCB 2023 paper, section 2.5.1). The sum of the regional fluxes may differ from the global estimates as reported in the GCB spreadsheet, because some adjustments were applied only for global fluxes.</p><p><strong>What is in the files?</strong></p><p>(1) The files for the fCO2-based data-products contain the following variables (temporal resolution: monthly):<br><br>fgco2_reg: Regionally integrated air-sea CO2 flux (positive downward), monthly, for regions: north, tropics, south<br>fgco2: Flux density of the total air-sea CO2 flux (positive downward), dimensions: time, latitude, longitude<br>sfco2: Surface ocean fCO2, dimensions: time, latitude, longitude<br>area: Area per pixel, dimensions: latitude, longitude<br>area_reg: Total surface ocean area covered by native grid, for global, north, tropics, south</p><p>(2) The files for the GOBMs contain the following fields, for simulation A ('contemporary simulation', including effects of rising CO2, climate change and variability) and simulation B ('control simulation', constant CO2, no climate change and variability). Temporal resolution: monthly</p><p>fgco2: Flux density of the total air-sea CO2 flux (positive downward), dimensions: time, latitude, longitude<br>sfco2: Surface ocean fCO2, dimensions: time, latitude, longitude<br>area: Area per pixel, dimensions: latitude, longitude<br><br>(3) One file 'GCB-2023_OceanModel_RegionalBreakdown_1959-2022.nc' with the regionally-integrated CO2 flux time-series for all individual GOBMs, and for simulations A and B. Temporal resolution: annual.</p><p><strong>Fair data use statement:</strong><br>The data and model output provided on this site are freely available and were furnished by individual scientists who encourage their use.<br><strong>Citation:</strong> Please cite the Global Carbon Budget 2023 (Friedlingstein et al., 2023, ESSD, https://doi.org/10.5194/essd-15-5301-2023) for all data. In addition, please also cite the corresponding original reference for each dataset that has been used - see Table 4 in Global Carbon Budget 2023 for references of all the individual Global Ocean Biogeochemical Models and fCO2-based data-products. Further, for an overview of the Global Ocean Biogeochemical Model output, you may find it useful to cite Hauck et al. (2020, Frontiers, doi:10.3389/fmars.2020.571720).<br><strong>Acknowledgement:</strong> Please add the following text in the acknowledgement of your paper: "We acknowledge the Global Carbon Project, which is responsible for the Global Carbon Budget and we thank the ocean modeling and fCO2-mapping groups for producing and making available their model and fCO2-product output."<br><strong>Co-authorship: </strong>An invitation of co-authorship to the contributing groups is encouraged if these data are the central data set of the publication.</p><p>Besides the surface fCO2 and air-sea CO2 flux data that is made available open access, we make<strong> additional 3D output</strong> from the Global Ocean Biogeochemical models (GCB-ocean) available upon request and with its own data policy. Please refer to the Global Carbon Budget website for these additional data: https://globalcarbonbudgetdata.org/closed-access-requests.html</p>
Identification at local and global scale: a case for using the Compact URI (CURIE) for life science data
<p>Panel A) A Local Resource Identifier (LRI) is not suited to global scale identification because of inevitable collisions: “9606” corresponds to a Pubmed article, a CGNC gene, a PubChem chemical, as well as an NCBI taxon (<em>Homo sapiens</em>), a BOLD taxon (<em>Bombycilla</em> <em>cedrorum</em>), and a GRIN taxon (<em>Catha</em> <em>edulis</em>)</p> <p>Panel B) Prefixing is often used to indicate the source of an LRI, but prefixes themselves are often undocumented and collide.</p> <p>Panel C) Prefixes may exist in alternate forms. When all of the alternates are not known, collapsing equivalent identifiers is tedious and incomplete.</p> <p>Panel D) CURIE syntax addresses these issues by having a prefix whose relationship with a resolving namespace is clearly documented.</p>
Data licences and organization type of contributors to the Global Biodiversity Information Facility as of 19 January 2016
<p>Data from the Global Biodiversity Information Facility were extracted using R (version 3.2.0) on 9 July 2015 using the rgbif package (version 0.9.0) (Chamberlain, S., Ram, K., Barve, V. & Mcglinn, D. (2015) Package ‘rgbif’: Interface to the Global 'Biodiversity' Information Facility 'API' http://cran.r-project.org/web/packages/rgbif/rgbif.pdf). The ‘rights’ statements was extracted for all occurrence datasets with one or more observations. A total of 12,458 datasets were extracted, but only about 11% of the datasets have an explicit data-useage-rights statement at the dataset level. However, some datasets use the occurrence level ‘rights’ and ‘accessRights’ fields. To extract these data the rights information was obtained from the first record of each dataset where a rights statement was missing at the dataset level.</p> <p>The datasets were categorized into 13 different types depending on the origin of the observations.</p> <ol> <li>Biodiversity Information Facility or data centre</li> <li>Botanical Garden or Herbarium</li> <li>Citizen science</li> <li>Commercial</li> <li>Data publisher</li> <li>Educational</li> <li>Government</li> <li>Museum</li> <li>Network</li> <li>Parks Authority or Nature Reserve</li> <li>Research institution</li> <li>Society</li> <li>Foundations</li> </ol>
Supplementary Data: A global fit of the MSSM with GAMBIT (arXiv:1705.07917)
<p><strong>Supplementary Data</strong></p> <p><em>A global fit of the MSSM with GAMBIT </em><br> <em>arXiv:1705.07917 </em></p> <p>The files in this record contain data for the MSSM7 model considered in the GAMBIT “Round 1” weak-scale SUSY paper.</p> <p>The files consist of</p> <ul> <li>A number of YAML files corresponding to different sets of sampling parameters and/or priors</li> <li>MSSM7.yaml, a YAML file used for postprocessing</li> <li>StandardModel_SLHA2_scan.yaml, a universal YAML fragment included from other YAML files</li> <li>StandardModel_SLHA2_postprocessing.yaml, a YAML fragment included from MSSM7.yaml</li> <li>A final hdf5 file, containing the combined results of all sampling runs</li> <li>An example pip file, for producing plots from the hdf5 file using pippi</li> <li>gambit_preamble.py, a collection of python functions used for in-line data processing in the pip file</li> <li>SLHA1 and SLHA2 files for the best-fit point in each subregion of the fit. These can found inside the tarball best_fits_SLHA.tar.gz.</li> </ul> <p>The different YAML files corresponding to different samplers and/or priors follow the naming scheme MSSM7_[scanner]_[prior]_[slice]_[special].yaml , where</p> <ul> <li>scanner = Diver, MN</li> <li>prior = log, flat</li> <li>slice = nM2, pM2, Afunnel, hZfunnel, sqcoann, slcoann (positive or negative M2, A/H funnel, h/Z funnel, squark co-annihilation, slepton co-annihilation)</li> <li>special = jDE, [blank] (used pure jDE, or used the default lambdajDE)</li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>The final hdf5 results file included here was generated in the following way:</p> <ul> <li>carry out initial runs using YAML files following the naming scheme above</li> <li>combine the resulting hdf5 output files into a single file, using<br> gambit/Printers/scripts/combine_hdf5.py</li> <li>postprocess the samples to remove all points more than 5 sigma from the current best fit, using MSSM7_strip.yaml</li> <li>postprocess the samples to include a new likelihood term for LHC Run II searches, and to recompute the FlavBit likelihoods (these were buggy in a pre-release version of GAMBIT), using MSSM7.yaml .</li> </ul> </li> <li> <p>It is not necessary to repeat the steps listed in point 1 when running new scans; the LHC Run II likelihoods can be included in the original YAML file, so that no postprocessing step is required.</p> </li> <li> <p>The YAML files that we give here are updated compared to the ones that we used when generating the hdf5 file, in order to match the set of available options in the release version of GAMBIT 1.0.0. The included physics and numerics are however identical.</p> </li> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.0.0, and the pip file is tested with pippi 2.0, commit 2ab061a8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don’t expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>
Supplementary Data: Global fits of GUT-scale SUSY models with GAMBIT (arXiv:1705.07935)
<p>Supplementary Data</p> <p><em>Global fits of GUT-scale SUSY models with GAMBIT</em><br> <em>arXiv:1705.07935</em></p> <p>The files in this record contain data for the CMSSM, NUHM1 and NUHM2 models considered in the GAMBIT "Round 1" GUT-scale SUSY paper.</p> <p>For each model, there are</p> <ul> <li>A number of YAML files, each corresponding to a different set of sampling parameters and/or priors</li> <li>A set of YAML files used for postprocessing: CMSSM_intermediate.yaml, CMSSM.yaml, NUHM1.yaml and NUHM2.yaml</li> <li>A final hdf5 file, containing the combined results of all sampling runs</li> <li>An example pip file, for producing plots from the hdf5 file using pippi</li> <li>SLHA1 and SLHA2 files for the best-fit point in each subregion of the fit. These can be found inside the tarball best_fits_SLHA.tar.gz.</li> </ul> <p>The record also contains</p> <ul> <li>StandardModel_SLHA2_scan.yaml and StandardModel_SLHA2_postprocessing.yaml, two universal YAML fragments included from other yaml files</li> <li>gambit_preamble.py, a collection of python functions used for in-line data processing in the pip files</li> </ul> <p>The different YAML files corresponding to different samplers and/or priors follow the naming scheme [model]_[scanner]_[prior]_[slice]_[special].yaml, where</p> <ul> <li>model = CMSSM, NUHM1, NUHM2</li> <li>scanner = Diver, MN</li> <li>prior = log, flat</li> <li>slice = pmu, nmu (positive or negative mu)</li> <li>special = sqcoann, slcoann, [blank] (squark co-annihilation, slepton co-annihilation, or bulk)</li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>For each model, the final hdf5 results file included here was generated in the following way:</p> <ul> <li>carry out initial runs using YAML files following the naming scheme above</li> <li>combine the resulting hdf5 output files into a single file, using gambit/Printers/scripts/combine_hdf5.py</li> <li>postprocess the samples to remove all points more than 5 sigma from the current best fit, using [model]_strip.yaml</li> <li>postprocess the samples to include a new likelihood term for LHC Run II searches, and to recompute the FlavBit likelihoods (these were buggy in a pre-release version of GAMBIT). For the CMSSM, this happened in two steps, due to persistent flavour bugs, using CMSSM_intermediate.yaml and CMSSM.yaml. For the NUHM1 and NUHM2, this was done in a single step each, using NUHM1.yaml and NUHM2.yaml.</li> </ul> </li> <li> <p>It is not necessary to repeat the steps listed in point 1 when running new scans; the LHC Run II likelihoods can be included in the original YAML file, so that no postprocessing step is required.</p> </li> <li> <p>The YAML files that we give here are updated compared to the ones that we used when generating the hdf5 file, in order to match the set of available options in the release version of GAMBIT 1.0.0. The included physics and numerics are however identical.</p> </li> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.0.0, and the pip files are tested with pippi 2.0, commit 2ab061a8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file for each model is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don't expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>
GRDC-Caravan: extending the original dataset with data from the Global Runoff Data Centre
<p>Large-sample datasets are essential in hydrological science to support modelling studies and global assessments. This dataset is an extension to <em>Caravan</em>, a global community dataset of meteorological forcing data, catchment attributes, and discharge data for catchments around the world (Kratzert et al. 2023).</p> <p>The extension includes a subset of those hydrological discharge data and station-based watersheds from the Global Runoff Data Centre (GRDC), which are covered by an open data policy (Attribution 4.0 International; CC BY 4.0). In total, the dataset covers stations from 5356 catchments and 25 countries worldwide with a time series record from 1950 – 2023.</p> <p>GRDC is an international data centre operating under the auspices of the World Meteorological Organization (WMO) at the German Federal Institute of Hydrology (BfG). Established in 1988, it holds the most substantive collection of quality assured river discharge data worldwide. Primary providers of river discharge data and associated metadata are the National Hydrological and Hydro-Meteorological Services of WMO Member States.</p> <p>Reference:</p> <p>Kratzert, F., Nearing, G., Addor, N. et al. Caravan - A global community dataset for large-sample hydrology. Sci Data 10, 61 (2023). <a href="https://doi.org/10.1038/s41597-023-01975-w">https://doi.org/10.1038/s41597-023-01975-w</a></p> <p><strong>Update:</strong></p> <p>With version 0.2 a bug has been fixed that affected the time series of four bands of all GRDC gauges in the GRDC extension. The affected bands were total_precipitation, surface_net_solar_radiation, surface_net_thermal_radiation and potential_evaporation, i.e. all features that are accumulated over the day, as per definition of ERA5-Land.<br>For details look at https://github.com/kratzert/Caravan/issues/26.</p> <p>Version 0.3: Data description file added.<br><br>Version 0.4: Added FAO Penman-Monteith PET (potential_evaporation_sum_FAO_PENMAN_MONTEITH) in the meteorological forcing data and renamed the ERA5-LAND potential_evaporation band to potential_evaporation_sum_ERA5_LAND. Also added all PET-related climated indices derived with the Penman-Monteith PET band (suffix "_FAO_PM") and renamed the old PET-related indices accordingly (suffix "_ERA5_LAND").<br><br>Version 0.5: License overview of the respective countries has been added.<br>Dataset description has been modified and improved.<br><br>Version 0.6: The attribute tables are sorted alphabetically. Minor inconsistencies in the data description file have been corrected.</p> <p> </p> <p><strong>Dataset structure:</strong></p> <p>The dataset is provided in the following two file formats:<br>1. caravan-grdc-extension-csv.zip: provides the time series data as comma-separated text files (CSV) (downloadable as 8.8 GB zip archive)<br>2. caravan-grdc-extension-nc.zip: provides the time series data in the Network Common Data Form (NetCDF) (downloadable as 7.6 GB zip archive)</p> <p><strong>The data in the versions 0.1-0.3 are identical. Version 0.4 added FAO Penman-Monteith PET (potential_evaporation_sum_FAO_PENMAN_MONTEITH) and renamed the ERA5-LAND potential_evaporation band to potential_evaporation_sum_ERA5_LAND.</strong></p> <p>Further details of the structure of the dataset are described in the data description file.</p>
TCOM-HCl : Daily global gap-free stratospheric hydrogen chloride profile data set based on TOMCAT CTM and Occultation Measurements
<p>Methodology: </p> <p><span>The </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>HCl Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated HCl profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these HCl differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>HCl bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved HCl profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean HCl profiles:</span></p> <ul> <li> <p><code><span>zmhcl_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmhcl_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105–5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023</span></p>
Map data of historical global estimates of soil respiration
<p>The map data of global soil respiration converted to NetCDF format.</p><p>All open access available estimates were collated.</p><p>Shoji Hashimoto, Akihiko Ito, Kazuya Nishina (2023) "Divergent data-driven estimates of global soil respiration". Communications Earth & Environment, 4 Article number: 460</p><p><a href="https://doi.org/10.1038/s43247-023-01136-2 ">https://doi.org/10.1038/s43247-023-01136-2</a> </p><p>Refer to Table 1 for the study ID and data source or the attributions of the NetCDF file. </p>
Global glacial lake bathymetry data
<p>This dataset collects globally published bathymetric data for glacial lakes, recording attributes such as glacial lake name, location, type, year of survey, corresponding area, volume, maximum water depth, and source.</p>
Extracted raw data from: Global dominance of lianas over trees is driven by forest disturbance, climate, and topography
<p>In a meta-analysis, we use an unprecedented dataset, representing 556 unique locations worldwide, distributed across 44 countries and six continents to show for the first time that lianas (woody vines) thrive relatively better than trees when forests are disturbed, temperature increase, precipitation decrease, and particularly in tropical lowlands. We demonstrate that liana dominance can persist for decades post-disturbance and hinder the recovery of disturbed forests, especially when climate favours lianas. With implications for the global carbon sink, our findings suggest that degraded tropical forests with environmental conditions favouring lianas should be the highest priority to consider for restoration management.</p>
Global Surface Ozone Concentration Dataset 1990-2017 Mapped at Fine Resolution through the Bayesian Maximum Entropy Data Fusion of Observations and Model Output
<p>This global surface ozone concentration dataset corresponds to the data developed in this paper:</p> <p>DeLang, M. N., J. S. Becker, K.-L. Chang, M. L. Serre, O. R. Cooper, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, S. Cleland, E. Collins, M. Brauer, and J. J. West (2021) Mapping yearly fine resolution global surface ozone through the Bayesian Maximum Entropy data fusion of observations and model output for 1990-2017, <em>Environmental Science & Technology</em>, 55, 4389-4398, doi: 10.1021/acs.est.0c07742.</p> <p>Ozone concentrations are estimated as described in the paper, with output shown for the Ozone Season Daily Maximum 8-hr metric (OSDMA8) for each year between 1990 and 2017, at 0.1 degree spatial resolution. Ozone is estimated through data fusion of output from several global models, with observations of ozone collected by TOAR. The data fusion involves application of the M3Fusion method to create a multi-model composite of several global models, followed by BME data fusion, as described in the paper. </p> <p>The *.nc file contains the latitude, longitude, ozone concentration estimate, and estimated variance for each 0.1 x 0.1 degree grid cell.</p> <p>Please contact Jason West (jasonwest@unc.edu) with questions about the dataset. We'd like to hear from you to know how you're using the data!</p> <p> </p> <p> </p>
Data of paper "Global supply chains amplify economic costs of future extreme heat risk"
<p>This is the database of articles "Global supply chains amplify economic costs of future extreme heat risk". The database contains the number of deaths caused by future heat waves in regions around the world under different SSP scenarios (e.g. SSP119, SSP245, SSP585), as well as global health losses, labor losses, and indirect losses as a percentage of regional or sectoral value added under different SSP scenarios. The regions of the database are aggregated using the GTAP 141 aggregating schema.</p>
Research data supporting "Impact of global heterogeneity of renewable energy supply on heavy industrial production and green value chains"
<p>Research data supporting the peer-reviewed article "Impact of global heterogeneity of renewable energy supply on heavy industrial production and green value chains" by the same authors.</p>
Data for publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980–2020"
<p>Data to reproduce figures for the publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980–2020" (DOI: 10.1177/01655515241245952). Each file contains the data underlying the figure corresponding to the file name.</p>
Supporting data for "Global Model of Atmospheric Chlorate on Earth" by Chan et al.
<p>Model code, simulation outputs, observation tables, and Python scripts for reproducing the analysis results/ figures presented in "Global Model of Atmospheric Chlorate on Earth" by Yuk-Chun Chan et al. Please refer to the publication and readme.txt for more information.</p>
Source Data and ambient ozone dataset generated in "Substantially underestimated global health risks of current ozone pollution"
<p>Existing assessments might have underappreciated ozone-related health impacts worldwide. Here our study assesses current global ozone pollution using the high-resolution (0.05°) estimation from a geo-ensemble learning model, with key focuses on population exposure and all-cause mortality burden. Our model demonstrates strong performance, achieving a mean bias of less than -1.5 parts per billion against in-situ measurements. We estimate that 66.2% of the global population is exposed to excess ozone for short term (> 30 days per year), and 94.2% suffers from long-term exposure. Furthermore, severe ozone exposure levels are observed in Cropland areas, particularly over Asia. Importantly, the all-cause ozone-attributable deaths significantly surpass previous recognition from specific diseases worldwide. Notably, mid-latitude Asia (30°N) and the western United States show high mortality burden, contributing substantially to global ozone-attributable deaths. Our study highlights current significant global ozone-related health risks and may benefit the ozone-exposed population in the future.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.