Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,154
datasets available to search
ShareScore release 0.7.1
Dataset results
5,154 results for “RESOLUTE”
Microbial Observatory at North Temperate Lakes LTER High-resolution temporal and spatial dynamics of microbial community structure in freshwater bog lakes 2005 - 2009 original format (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-ntl/349/4. The abstract below was extracted from the Level 0 data package and is included for context: The North Temperate Lakes - Microbial Observatory seeks to study freshwater microbes over long time scales (10+ years). Observing microbial communities over multiple years using DNA sequencing allows in-depth assessment of diversity, variability, gene content, and seasonal/annual drivers of community composition. Combining information obtained from DNA sequencing with additional experiments, such as investigating the biochemical properties of specific compounds, gene expression, or nutrient concentrations, provides insight into the functions of microbial taxa. Our 16S rRNA gene amplicon datasets were collected from bog lakes in Vilas County, WI, and from Lake Mendota in Madison, WI. Ribosomal RNA gene amplicon sequencing of freshwater environmental DNA was performed on samples from Crystal Bog, North Sparkling Bog, West Sparkling Bog, Trout Bog, South Sparkling Bog, Hell’s Kitchen, and Mary Lake. These microbial time series are valuable both for microbial ecologists seeking to understand the properties of microbial communities and for ecologists seeking to better understand how microbes contribute to ecosystem functioning in freshwater.
Ramped Pyrolysis Oxidation (RPO) coupled radiocarbon (14C-DOC) and stable carbon (13C-DOC), high-resolution molecular composition (FT-ICR MS), and biodegradable dissolved organic carbon (BDOC) of groundwater, river water, and lagoon water in northeast Alaska, 2017
Supra-permafrost groundwater (SPGW), river water, and lagoon water were sampled near Kaktovik, AK to assess the reactivity and origin of dissolved organic matter (DOM) across interconnected hydrologic systems during late summer. Water samples were collected on August 17th 2017 from SPGW along the beach of Jago Lagoon (Jago GW), surface water from the Jago River’s main channel above tidal influence (Jago R), and from the water column of Kaktovik Lagoon at 2–3 m depth (KA LW). Measurements were made from grab samples for river and lagoon water, and from a composite sample for SPGW gathered from 10 individual shoreline locations. Data include dissolved organic carbon concentration (DOC, mg C L-1), Ramped Pyrolysis Oxidation (RPO) derived fraction compositions of 13C-DOC (δ13C ‰), 14C-DOC (in fraction modern), and method/instrumental error in the 14C and 13C results, and Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS) molecular composition and summarized compound classes. Biodegradable DOC (BDOC) bottle experiments were performed using all three sample types, where DOC concentration was subsequently measured at 2, 7, 14, and 28 days. FT-ICR MS composition was subsequently measured at the 28-day timepoint to track changes in molecular formulae and compound class relative abundance following biodegradation. Data from RPO serial thermal oxidation include temperature and normalized CO2 profiles for each background sample. Thermal-oxidation profiles of CO2 were transformed into non-parametric activation energy (E) distributions using an inverse model. Model output includes C mass of oxidized CO2 (µg C), Tmax (K), Emax (kJ mol-1), Emean (kJ mol-1), Estd (kJ mol-1), and p(0,E)max of user-defined sample fractions. FT-ICR MS results include a summary table of the relative abundance of compound classes (e.g., unsaturated phenolic, polyphenolic, aliphatic, condensed aromatics, peptide-like) and elemental groupings (e.g., CHO-type, CHON-type, CHOS-type, CHON
Monthly precipitation in mm at 1 km resolution (multisource average) based on SM2RAIN-ASCAT 2007-2021, CHELSA Climate and WorldClim
<p>Monthly precipitation in mm at 1 km resolution based on SM2RAIN-ASCAT 2007-2021 (<a href="https://doi.org/10.5281/zenodo.2615278">https://doi.org/10.5281/zenodo.2615278</a>). <a href="https://github.com/Envirometrix/LandGISmaps/tree/master/input_layers/clim1km">Downscaled to 1 km resolution using gdalwarp</a> (cubic splines) and combined with WorldClim (<a href="https://worldclim.org/data/worldclim21.html">https://worldclim.org/data/worldclim21.html</a>) and CHELSA Climate (<a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a>) monthly values. Final values are estimated as a simple average between the three precipitation data sources; a more objective approach would be to use training points e.g. meteo-station monthly values, then train an ensemble model using the 3 data sources as independent variables. Another global data source of precipitation images is the <a href="https://gpm.nasa.gov/data/imerg">monthly IMERGE dataset</a>, however this requires transformation and is available only for limited span of years.</p> <p>Processing steps are available <a href="https://gitlab.com/openlandmap/global-layers/-/tree/master/input_layers/SM2RAIN"><strong>here</strong></a>. Antarctica is not included. Standard deviation (sd) indicates a difference between the 3 data sources. To access and visualize maps use: <a href="https://openlandmap.org"><strong>https://openlandmap.org</strong></a>.<strong> </strong>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>clm = theme: climate,</li> <li>precipitation = variable: precipitation,</li> <li>wc.v2.1.chelsa.v2.1.sm2rain.oct = determination method: long-term average values for October,</li> <li>m = mean value,</li> <li>1km = spatial resolution / block support: 1 km,</li> <li>s0..0cm = vertical reference: land surface,</li> <li>1980..2020 = time reference: from 1980 to 2020,</li> <li>v0.3 = version number: 0.3,</li> </ul>
Cascade project at North Temperate Lakes LTER - High-resolution spatial analysis of CASCADE lakes during experimental nutrient enrichment 2015 - 2016
This dataset contains high-resolution spatio-temporal water quality data from two experimental lakes during a whole-ecosystem experiment. Through gradual nutrient addition, we induced a cyanobacteria bloom in an experimental lake (Peter Lake) while leaving a nearby reference lake (Paul Lake) as a control. Peter and Paul Lakes (Gogebic county, MI USA), were sampled using the FLAMe platform (Crawford et al. 2015) multiple times during the summers of 2015 and 2016. In 2015 nutrient additions to Peter Lake began on 1 June, and ceased on 29 June, Paul Lake was left unmanipulated. In 2016 no nutrients were added to either lake. Measurements were taken using a YSI EXO2 probe and a Garmin echoMap 50s. Sensor- data were collected continuously at 1 Hz and linked via timestamp to create spatially explicit data for each lake. Crawford, J. T., L. C. Loken, N. J. Casson, C. Smith, A. G. Stone, and L. A. Winslow. 2015. High-speed limnology: Using advanced sensors to investigate spatial variability in biogeochemistry and hydrology. Environmental Science & Technology 49:442–450.
Microbial Observatory at North Temperate Lakes LTER High-resolution temporal and spatial dynamics of microbial community structure in freshwater bog lakes 2005 - 2009 original format
The North Temperate Lakes - Microbial Observatory seeks to study freshwater microbes over long time scales (10+ years). Observing microbial communities over multiple years using DNA sequencing allows in-depth assessment of diversity, variability, gene content, and seasonal/annual drivers of community composition. Combining information obtained from DNA sequencing with additional experiments, such as investigating the biochemical properties of specific compounds, gene expression, or nutrient concentrations, provides insight into the functions of microbial taxa. Our 16S rRNA gene amplicon datasets were collected from bog lakes in Vilas County, WI, and from Lake Mendota in Madison, WI. Ribosomal RNA gene amplicon sequencing of freshwater environmental DNA was performed on samples from Crystal Bog, North Sparkling Bog, West Sparkling Bog, Trout Bog, South Sparkling Bog, Hell’s Kitchen, and Mary Lake. These microbial time series are valuable both for microbial ecologists seeking to understand the properties of microbial communities and for ecologists seeking to better understand how microbes contribute to ecosystem functioning in freshwater.
Cascade Project at North Temperate Lakes LTER – High-resolution Spatial Data for Whole Lake Experiments 2018 - 2019
Spatial measurements of water quality from Peter and Paul lakes in 2018 and 2019. In 2019, inorganic nitrogen and phosphorus were added to Peter Lake daily to cause an algal bloom while Paul Lake was an unmanipulated reference lake. In 2018, both lakes were sampled 1 time per week, while in 2019 lakes were sampled three times per week. Measurements were taken using the FLAMe sampling platform (Crawford et al. 2015, Environmental Science and Technology 49:442-450), which was driven in a grid pattern and recorded GPS coordinates and water measurements at 1Hz to create high resolution spatial maps.
Data archive for "Stochastic Super-Resolution for Downscaling Time-Evolving Atmospheric Fields with a Generative Adversarial Network"
<p>This datasets supports the paper "Stochastic Super-Resolution for Downscaling Time-Evolving Atmospheric Fields with a Generative Adversarial Network" submitted to IEEE Transactions in Geoscience and Remote Sensing. A preprint of the paper can be found here: <a href="https://arxiv.org/abs/2005.10374">https://arxiv.org/abs/2005.10374</a>. The code that uses these data is available at <a href="https://github.com/jleinonen/downscaling-rnn-gan">https://github.com/jleinonen/downscaling-rnn-gan</a>.</p> <p>The file "goes-samples-2019-128x128.nc" contains the training dataset called "GOES-COT" in the paper, consisting of cloud optical depth measurements from the GOES-16 satellite. The files "gen_weights*.nc" contain the generator weights saved at different time steps during training for the two different datasets described in the paper.<br> </p>
High resolution pond velocity measurements, Idaho21
<p>These scientific data were obtained by Jeffrey Nielson and Stephen Henderson of Washington State University, working in collaboration with Sandra Mayne, Caren Goldberg and Jeffrey Manning. High-resolution current meters were used to obtain detailed measurements of water velocity, with supporting measurements of wind velocity and water temperature profiles. Overview of observations in referenced Henderson et al. (2024) L&O paper, more details in included files. </p>
Mediterranean Cyclone tracks between 1979-2018 (40 years) from a high-resolution perspective using ECMWF ERA5 dataset
<p>The present dataset presents the trajectories of the 13,157 cyclones identified within the Mediterranean Region (MR) between 1979 and 2018 (40 years). These cyclone tracks were obtained using the new Cyclone Detection and Tracking Method (CDTM) described in Aragão e Porcù (2021) to take advantage of the recent availability of a high-resolution reanalysis dataset of ECMWF ERA5. The CDTM uses hourly data of Geopotential Height at 1000 hPa with a spatial resolution of 0.25°x0.25°, and the analysis' domain covers the area within 15°W to 48° E and 21° N to 54°N. Additionally, trying to eliminate artificial low-pressure cores, short-living thermal-lows or too weak cyclones as much as possible, the present study only considered cyclones lasting more than 24h.<br> The dataset presents hourly information for all cyclones from the cyclogenesis time to the cyclolysis time. Each record presents: [1] Cyclone ID (integer, 8 digits), [2] Cyclone centre longitude position (°E, real, 8 digits, 3 decimal digits), [3] Cyclone centre latitude position (°N, real, 8 digits, 3 decimal digits), [4] Year (integer, 4 digits), [5] Month (integer, 2 digits), [6] Day (integer, 2 digits), [7] Hour (integer, 2 digits), [9] Cyclone centre Geopotential Height at 1000 hPa (m, real, 9 digits, 3 decimal digits).<br> The analyses presented in Aragão e Porcù (2021) revealed that the proposed CDTM is capable to capture almost the totality of the observed cyclones, as well as describing its respective area of cyclogenesis, trajectories, and durations. More than an adaptation to a high-resolution dataset, the method brings as its primary contribution a suitable set of parameters to systematically identify and track the cyclonic activities in the Mediterranean, where cyclones do not have sizeable horizontal pressure gradients and present a shorter lifetime compared to open-ocean cyclones.</p> <p>Cite this article</p> <p>Aragão, L., Porcù, F. Cyclonic activity in the Mediterranean region from a high-resolution perspective using ECMWF ERA5 dataset. <em>Clim Dyn</em> (2021). https://doi.org/10.1007/s00382-021-05963-x</p>
SM2RAIN-ASCAT (2007-2021) global daily satellite rainfall including aggregated values and trend parameters as 10km resolution GeoTIFFs
<p>This is a GeoTIFF version of the <a href="http://hydrology.irpi.cnr.it/download-area/sm2rain-data-sets/">SM2RAIN-ASCAT (2007-2021): global daily satellite rainfall from ASCAT soil moisture</a> data set v1.1 (Brocca et al. 2019). Conversion steps are available <a href="https://github.com/Envirometrix/LandGISmaps/tree/master/input_layers/SM2RAIN"><strong>here</strong></a>. Few important notes:</p> <ul> <li>Daily values are stored as integers, whereas in the NetCDF the dataset is rounded to one decimal place.</li> <li>The NetCDF has also a Quality Flag for a better and more informed use of the data (here omitted).</li> <li>P05, P50 and P95 indicate quantiles derived per pixel.</li> </ul> <p>Includes also long-term trends (trend.logit.ols) which was produced by fitting regression models to de-seasonalized time-series as explained in this <strong><a href="https://gitlab.com/openlandmap/global-layers/-/blob/master/input_layers/MOD13Q1/03-data-access.ipynb">python tutorial</a></strong>. Basically models are fitted for <strong>each pixel</strong> and the model parameters are saved as images.</p> <p>Monthly averages and s.d. of precipitation are available in the files:</p> <ul> <li>clm_precipitation_sm2rain.*_m_10km_s0..0cm_2007..2021_v1.5.tif = monthly precipitation in mm,</li> <li>clm_precipitation_sm2rain.*_sd.10_10km_s0..0cm_2007..2021_v1.5.tif = standard deviation of precipitation in mm * 10 per month (multiplied by 10 so Integers can be used),</li> </ul> <p>Downscaled monthly averages (1 km) are also available (<a href="https://doi.org/10.5281/zenodo.1435912">https://doi.org/10.5281/zenodo.1435912</a>).</p> <p>To cite this data set please refer to the <strong><a href="https://doi.org/10.5281/zenodo.2591214">original copy</a></strong> of the data set.</p> <ul> <li>Brocca, L., Filippucci, P., Hahn, S., Ciabatta, L., Massari, C., Camici, S., Schüller, L., Bojkov, B., Wagner, W. (2019). <strong><a href="https://doi.org/10.5194/essd-11-1583-2019">SM2RAIN–ASCAT (2007–2018): global daily satellite rainfall data from ASCAT soil moisture observations</a></strong>. Earth Syst. Sci. Data, 11, 1583–1601. <a href="https://doi.org/10.5194/essd-11-1583-2019">https://doi.org/10.5194/essd-11-1583-2019</a></li> </ul>
Fast and long-term super-resolution imaging of ER nano-structural dynamics in living cells using a neural network
<p>Datasets acquired and generated for the manuscript "Fast and long-term super-resolution imaging of ER nano-structural dynamics in living cells using a neural network". The datasets include test, training and time series datasets each containing the raw data and the predicted data where it applies. </p>
Quality-checked, one-hour resolution cruise track of the Antarctic Circumnavigation Expedition (ACE) undertaken during the austral summer of 2016/2017.
<p><strong>Dataset abstract</strong></p> <p>The Antarctic Circumnavigation Expedition (ACE), undertaken in the austral summer of 2016/2017 recorded the cruise track using two independent geo-location instruments: one using GLobal NAvigation Satellite Systems (GLONASS; hereafter referred to as GLONASS) and another primarily using the Global Positioning System (GPS; hereafter referred to as the Trimble GPS). Daily log files were recorded in real-time from both instruments during the expedition and added to MySQL database tables. Following the expedition, quality-checking work has been undertaken to provide a one-second resolution set of positions for the cruise track. Here we present the final quality-checked dataset aggregated to a resolution of one hour. This is of use for understanding the position of the vessel to a lower precision, such as for plotting the track throughout the voyage.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_cruise_track_1hour_YYYY-MM.csv, data file, comma-separated values</li> <li>README.txt, metadata, text file</li> <li>data_file_header.txt, metadata, text file</li> </ul> <p><strong>Dataset license</strong></p> <p>This quality-checked cruise track dataset is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
2005-2099 High resolution bioclimatic variables for the surface and bottom of the Mediterranean Sea.
<p><em><span>This dataset provides annual statistical descriptors (mean, minimum, maximum, range and standard deviation) of key biogeochemical and physical variables for the Mediterranean Sea. It covers the period 2005-2099 under the RCP8.5 scenario, with a spatial resolution of 1/24 degree (~4km²). Variables include temperature, salinity, pH, water velocity, nutrients (NO3, PO4, NH4), dissolved inorganic carbon, oxygen, and net primary production. Data are available for both surface and at bathymetry level. The original projections were generated using OGSTM-BFM and MFS16 models at daily time and 1/16 degree grid resolution. We downscaled these to 1/24 degree and applied Quantile Delta Mapping bias correction using CMEMS reanalysis products for 2005-2020. The dataset is provided in a user-friendly format, making it accessible for various ecological and environmental modelling applications.</span></em></p>
Extended data for the paper: "SentemQC - A novel and cost-efficient method for quality assurance and quality control of high-resolution frequency sensor data in fresh waters"
<p>Extended data 1 to 4 for the software article:<br>SentemQC - A novel and cost-efficient method for quality assurance and quality control of high-resolution frequency sensor data in fresh waters. </p> <p>The extended data is tables and a Figure output and input from/to SentemQC runs relevant for the SentemQC paper.</p>
Annual time series of global VIIRS nighttime lights for 2000-2024 at 500-m spatial resolution extrapolated using logistic regression
<p>The <a href="https://eogdata.mines.edu/products/vnl/"><strong>Annual Visible Night Light (VNL) V2</strong></a> (VIIRS) images at 500-m spatial resolution for the period 2012 to 2024 (Elvidge et al., 2021) have been used to extrapolate the values backwards for years 2000–2011. This was done by fitting a logistic regression (per pixel) and then predicting the values for the previous years (see nightlights_stack_500m.R). After consistent time-series have been produced, I also derived the difference between year 2024 and year 2000 (nightlights.difference_viirs.v21_m_500m_s_2000_2024_go_epsg4326_v20230318.tif): this shows average rate of change for the 25 years period. Use with caution: extrapolation of values can lead to artifacts. For most of the land surface, however, it appears that the growth of night lights follows exponential growth function and hence nights in the past can be represented accurately by fitting decay / logistic regression function.</p> <p>Original values from the Annual VNL V2 product have been converted from 0–200 to 0–2000 scale and are available as Cloud-Optimized GeoTIFFs.</p> <p>Principal components (PC1, PC2, PC3, PC4) were derived using SAGA GIS (sums-of-squares-and-cross-products matrix) method. The first PC1 usually matches the long-term mean value, PC2 matches the 1st derivation in values. File "nightlights_dmsp.v10_m_1km_s_19920101_20241231_go_epsg4326_v20251006.tif" contains 33 years 1992 to 2024, but at 1 km resolution.</p> <p>To cite the Annual VNL V2, please use:</p> <ul> <li>Elvidge, C. D., Zhizhin, M., Ghosh, T., Hsu, F. C., & Taneja, J. (2021). <a href="https://doi.org/10.3390/rs13050922">Annual time series of global VIIRS nighttime lights derived from monthly averages: 2012 to 2019</a>. Remote Sensing, 13(5), 922. https://doi.org/10.3390/rs13050922</li> </ul> <p>Historic night light images (1 km resolution) are also available from <a href="https://doi.org/10.6084/m9.figshare.9828827.v10">Figshare</a>:</p> <ul> <li>Li, X., Zhou, Y., Zhao, M., & Zhao, X. (2020). <a href="https://doi.org/10.1038/s41597-020-0510-y">A harmonized global nighttime light dataset 1992–2018</a>. Scientific data, 7(1), 168. https://doi.org/10.1038/s41597-020-0510-y</li> </ul>
Predicted occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution
<p>The dataset contains predictions of occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution + all covariate layers used for modeling. Over seven million electronic health records (EHRs), among which 11,741 EHRs reported tick attachment, were used to evaluate climate, environmental and animal host factors affecting the risk of tick attachment in cats and dogs in Great Britain (GB). The tick presence/absence EHRs for dogs and cats were further overlaid with spatiotemporal time-series of climatic, vegetation, human influence, hydrological and terrain variables (slope, wetness index) to produce a spatiotemporal regression matrix; an Ensemble Machine Learning framework was used to fine-tune hyperparameters for Random Forest (classif.ranger), Gradient boosting (classif.xgboost) and GLM-net (classif.glmnet) algorithms, which were then used to produce a final ensemble meta-learner that predicts the probability of occurrence of ticks across GB with monthly intervals.</p> <ul> <li>gb1km_covariates.zip contains ALL covariate layers as GeoTIFFs (time-series) used for modeling ticks dynamics;</li> <li>data_1km_2014_M01.rds = contains all covariates for January 2014 prepared as SpatialGridDataFrame (R data object);</li> </ul> <p>Codes of files indicate e.g.:</p> <ul> <li>"monthly.tick.prob_savsnet.mar_p_1km_s_2014_2021" = monthly occurrence probability for January based on the training data from 2014 to 2021;</li> <li>"monthly.tick.prob_savsnet.oct_md_1km_s_20211001_20211031" = monthly prediction (model) error derived as the standard deviation from multiple base learners;</li> </ul> <p>The dataset is described in detail in the following publication:</p> <ul> <li>Arsevska, E., Hengl, T., Singelton, D. et al. (2023?) <strong>Risk factors for tick attachment in companion animals in Great Britain: a spatiotemporal analysis covering 2014–2021</strong>. Submitted to Parasites & Vectors (in review).</li> </ul> <p>The model summary shows:</p> <pre><code>Call: stats::glm(formula = f, family = "binomial", data = getTaskData(.task, .subset), weights = .weights, model = FALSE) Deviance Residuals: Min 1Q Median 3Q Max -1.4749 -0.0557 -0.0471 -0.0430 3.7611 Coefficients: Estimate Std. Error z value Pr(>|z|) (Intercept) -7.64495 0.02095 -364.957 < 2e-16 *** classif.ranger 4.95061 0.63615 7.782 7.13e-15 *** classif.xgboost 189.75543 5.53109 34.307 < 2e-16 *** classif.glmnet 140.24208 5.05375 27.750 < 2e-16 *** --- Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1 (Dispersion parameter for binomial family taken to be 1) Null deviance: 170604 on 7303013 degrees of freedom Residual deviance: 162571 on 7303010 degrees of freedom AIC: 162579 Number of Fisher Scoring iterations: 9</code></pre> <p><em>Acknowledgements</em>: We are grateful to data providers in veterinary practice (VetSolutions, Teleos, CVS, and other practitioners). We are grateful to the INRAE MIGALE bioinformatics facility (MIGALE, INRAE, 2020. Migale Bioinformatics Facility, doi: <a href="https://entrepot.recherche.data.gouv.fr/dataverse/migale">10.15454/1.5572390655343293E12</a>) for providing computing resources. We are also grateful for<br> the help and support provided by <a href="https://www.liverpool.ac.uk/savsnet/">SAVSNET team members</a> Bethaney Brant, Susan Bolan and Steven Smyth.<br> This study was funded mainly by a grant from the <strong>Biotechnology and Biological Sciences Research Council</strong>,<br> BB/NO19547/1 and <strong>British Small Animal Veterinary Association</strong> (BSAVA). The research was partly funded by the National Institute for <strong>Health Research Health Protection Research Unit</strong> (NIHR HPRU) in Emerging and Zoonotic Infections at the <strong>University of Liverpool</strong> in partnership with <strong>Public Health England</strong> (PHE) and <strong>Liverpool School of Tropical Medicine</strong> (LSTM). This work has been partially funded by the <em>“Monitoring outbreak events for disease surveillance in a data science context"</em> (MOOD) project from the European Union’s Horizon 2020 research and innovation program under grant agreement No. 874850 (<a href="https://mood-h2020.eu/">https://mood-h2020.eu/</a>). The views expressed are those of the authors and not necessarily those of the NHS, the NIHR, the Department of Health or Public Health England.</p>
Global distribution of predicted soil types at 1 km resolution based on the WRB 2022 classification
<p>Global maps at 1 km spatial resolution of the predicted soil types (0–100% probabilities) at 1 km resolution based on the <a href="https://www.fao.org/soils-portal/data-hub/soil-classification/world-reference-base/en/">WRB 2022</a> (<strong>World Reference Base</strong> the international standard for soil classification) classification system. The training data comes from the following 3 main sources:</p> <ol> <li>WOSIS points available via: <a href="https://www.isric.org/explore/wosis">https://www.isric.org/explore/wosis</a>;</li> <li>HWSD v2 (random draw of cca 20,000 points): <a href="https://iiasa.ac.at/models-tools-data/hwsd">https://iiasa.ac.at/models-tools-data/hwsd</a>;</li> <li>Other national datasets / data from publications and projects.</li> </ol> <p>Predictions are based on using Rando Forest algorithm as implemented in the <a href="https://www.randomforestsrc.org/">randomForestSRC package</a> with cca 190 covariate layers representing soil forming factors (CHELSA Climate, Global Lithological DB GLiM, MODIS EVI and LST long-term derivatives, Digital Terrain model parameters and similar).</p> <p>All TIF files are provided as <a href="https://www.cogeo.org/">COGs</a>, which means that you can open them directly in QGIS or similar. Publication explaining all modeling steps is pending.</p> <p>Update of the predictions takes about 4–5 hrs and will be regularly run provided that new training points are available. Disclaimer: These are initial results with limited accuracy and possible issues with quality of training points, location errors and harmonization issues. Use at own risk.</p> <p>Note: original list of soil types have been subset to classes that appear at least 10 times and at least in 2 countries. If you notice an error or artifact <strong>please report via <a href="https://github.com/OpenGeoHub/SoilTypeMapping">the Github repository</a></strong>. Help us improve this dataset by contributing training points.</p>
Marine magnetic anomaly data from high resolution surveys off the SW Portuguese coast
<p>This dataset contains <strong>magnetic anomaly grids</strong> that result from the full processing of marine magnetic data collected off the SW Portuguese coast between 2014 and 2019. A total area of ~4400 km<sup>2</sup> was surveyed with average line spacing of 1 nautic mile. Surveys covered the continental shelf and in some regions reaching up to 2500 m bathymetric levels. Total magnetic field data were acquired with a G882 Cesium vapor marine magnetometer towed, towed at sea surface.</p> <p><strong>Full processing</strong> of magnetic data included: layback correction; noise removal; IGRF subtraction; base station correction; line leveling; minimum curvature gridding. The resulting sea level magnetic anomaly grid was further processed for upward continuation and reduction to the pole, providing additional outputs. </p> <p>The following grids are provided in <strong>georeferenced geotiff format</strong>:</p> <ul> <li>Magnetic anomaly (sealevel)</li> <li>Magnetic anomaly reduced to the pole (sealevel)</li> <li>Magnetic anomaly upward continued to 200 m height </li> <li>Magnetic anomaly upward continued to 200 m height, reduced to the pole</li> <li>Magnetic anomaly upward continued to 3000 m height </li> <li>Magnetic anomaly upward continued to 3000 m height, reduced to the pole</li> </ul> <p><strong>Published in</strong>: Neres, M., P. Terrinha, J. Noiva, P. Brito, M. Rosa, L. Batista, C. Ribeiro (2023). <em>New Late Cretaceous and CAMP magmatic sources off West Iberia, from high-resolution magnetic surveys on the continental shelf.</em> <strong>Tectonics</strong>. doi: 10.1029/2022TC007637</p> <p> </p>
Sample data for "A weakly supervised framework for high resolution crop yield forecasts"
<p>This dataset includes sample data for the United States to run the weakly supervised framework as described in the paper titled <em>A weakly supervised framework for high resolution crop yield forecasts</em>, accessible at </p> <table summary="Additional metadata"> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2205.09016">https://doi.org/10.48550/arXiv.2205.09016</a></td> </tr> </tbody> </table> <p> </p> <p>The updated paper (including results from the US) is published in Environmental Research Letters:</p> <p><a href="https://doi.org/10.1088/1748-9326/acf50e">https://doi.org/10.1088/1748-9326/acf50e</a></p> <p> </p> <p>The software implementation of the machine learning baseline is available at: https://github.com/BigDataWUR/MLforCropYieldForecasting/tree/weaksup.</p> <p> </p> <p>Data</p> <p>1. County data (county-data.zip) for county-level strongly supervised models:</p> <p>* CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p> <p>* CSSF_COUNTY_US.csv: Crop productivity indicators including total above-ground production (kg ha<sup>-1</sup>), total weight of storage organs (kg ha<sup>-1</sup>), development stage (0-2). Source: de Wit et al. (2022).</p> <p>* METEO_COUNTY_US.csv: Meteo data including maximum, minimum, average daily air temperature (℃); sum of daily precipitation (PREC) (mm); sum of daily evapotranspiration of short vegetation (ET0) (Penman-Monteith, Allen et al., (1998)) (mm); climate water balance = (PREC - ET0) (mm). Source: Boogaard et al. (2022).</p> <p>* REMOTE_SENSING_COUNTY_US.csv: Fraction of Absorbed Photosynthetically Active Radiation (Smoothed) (FAPAR). Source: Copernicus GLS (2020).</p> <p>* SOIL_COUNTY_US.csv: Soil water holding capacity. Source: WISE Soil Property Database (Batjes, 2016).</p> <p>* YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p> </p> <p>2. 10-km grid data (grid-data.zip) for grid-level strongly supervised models:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>* CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level (similar to county data above).</p> <p>* METEO_GRIDs_US.csv: Meteo data at 10km grid level (similar to county data above).</p> <p>* REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level (similar to county data above).</p> <p>* SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level (similar to county data above).</p> <p>* YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021), Lobell et al. (2020).</p> <p> </p> <p>3. County labels and 10-km grid inputs (dscale-US.zip) for weak supervision:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>* CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level.</p> <p>* METEO_GRIDs_US.csv: Meteo indicators at 10km grid level.</p> <p>* REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level.</p> <p>* SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level.</p> <p>* YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021).</p> <p>* YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p>* CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p>
Marcell Experimental Forest 30-minute resolution meteorological data, 2006 - ongoing
This data publication contains 30-minute meteorological data collected from 2006 - ongoing at the Marcell Experimental Forest (MEF) in Itasca County, Minnesota, which is operated and maintained by the USDA Forest Service, Northern Research Station. Air temperature, relative humidity, wind speed and direction, photosynthetic photon flux density, soil temperature, and soil volumetric water content were measured at three meteorological monitoring stations. One station is in an upland clearing, one is under an upland forest canopy, and the third is in a peatland.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.