Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
162
datasets available to search
ShareScore release 0.9.0
Dataset results
162 results for “continental scale”
Data for: "Continental-scale patterns in diel flight timing of high-altitude migratory insects"
<p>This dataset contains the proportional migratory insect intensity and traffic data used in Haest <em>et al.</em> (2024) to quantify patterns in diel flight periodicity of migratory insects between 50-500m above ground level during March-October 2021 using a network of seventeen vertical-looking radars across Europe. Please see the Materials and Methods section in Haest <em>et al.</em> (2024) for more details on the dataset. </p>
Continental Europe surface lithology based on EGDI / OneGeology map at 1:1M scale
<p>Continental Europe surface lithology based on <strong><a href="http://www.europe-geology.eu/onshore-geology/geological-map/onegeologyeurope/">EGDI / OneGeology map</a></strong> at 1:1M scale produced by <a href="https://egdi.geology.cz/record/basic/5729ffdf-2558-48fc-a5d2-645a0a010855">GEOZS, Slovenia</a>. European datasets harvested from national WFS for geologic units or national geological units datasets, based on OneGeology and <strong><a href="https://inspire.ec.europa.eu/codelist/LithologyValue/">INSPIRE Lithology</a></strong> and Geochronologic Era URI codelists. Layers include:</p> <ul> <li>EGDI_GE_GeologicUnit_EN_1M_Surface_LithologyPolygon_v2_250m_epsg.3035.tif = original EGDI surface lithology map;</li> <li>dtm_surface.lithology_egdi.1m_c_250m_s_20000101_20221231_eu_epsg.3035_v20240530.tif = gap filled surface lithology map;</li> </ul> <p>Missing values in the original EGDI lithology map have been imputed by training a random forest classifier model based on parameters derived from DTM and soil regions map from Die Bundesanstalt für Geowissenschaften und Rohstoffe (BGR). By generating 1 million random points, geographically balanced over the whole pan-EU land area, each class in the map was covered properly. Classes whose number of samples is less than 10 were discarded from the model training. The hyperparameter tuning of the model was carried out via a Bayesian approach with a criteria to maximize accuracy of 5k-fold cross validation. The tuned random forest model achieved an accuracy of 47% (Kappa=0.43) for the testing data, 20% of the generated sample points. The lithology of Turkey, on the other hand, was digitised from the available geology map produced by the General Directorate of Mineral Research and Exploration (MTA). The available raster map was post-processed and classified as 20 lithology classes using the k-means algorithm. These classes were harmonized with the classes in the EGDI lithology map.</p> <p>Acknowledgment: GEOZS, Continental Shelf Department at the Ministry for Transport and Infrastructure.</p>
Geochemical, physicochemical, and genomic data from a continental-scale survey of microbial diversity in Antarctic soils (2003-2023)
This data package offers comprehensive insights into Antarctic soil microbial diversity and composition. From 2003 to 2023, a total of 186 samples were collected from diverse locations spanning the Antarctic Peninsula to East Antarctica, representing a wide range of environmental gradients and climatic conditions. Soils were stored at -20°C to preserve their integrity for downstream analyses. This data package integrates cultivation-independent sequencing of prokaryotic and fungal communities alongside a robust cultivation-dependent culture collection to enable direct comparisons across microbial diversity assessment methods. Accompanying geochemical, physicochemical, and environmental parameters provide critical context for biogeographical analyses, offering a valuable resource for studying microbial adaptations and community dynamics in extreme Antarctic environments.
Using RS-DAT to study continental-scale phenology at high-spatial resolution
This repository includes the notebooks employed to calculate and analyze a gridded phenological model, as computed from a set of daily meteorological variables.
Data inputs and results for "Mammal niches are not conserved over continental scales" by Goldstein et al.
<p>This data packet provides inputs and results for "Mammal niches are not conserved over continental scales" by Goldstein et al., currently in the submission process. This repository will eventually be updated to link to the published manuscript.</p> <p> </p> <p>===============================================================================<br>===============================================================================<br>Overview<br>===============================================================================<br>===============================================================================</p> <p>Data and model products associated with the manuscript "Mammal niches are not <br>conserved over continental scales" by Goldstein et al. </p> <p>Files are organized into two subdirectories. The first, "model_inputs/", <br>contains 8 data files intended to be used as part of the reproducible code <br>repository at https://github.com/dochvam/Mammal_SVCs_ISDM_reproducible. <br>The second subdirectory, "model_outputs/", contains modeled products giving<br>estimated spatially varying niche relationships and predictions of relative<br>abundance.</p> <p>Below, we describe the contents of each file type. See the main manuscript <br>for full methodology, data sources, and discussions of spatial scales.</p> <p>NOTE: Version 1 of this dataset contained some errors that have been corrected<br>in Version 2. Version 2 was used as the input dataset for the analyses in the<br>associated manuscript. Version 1 should not be used.</p> <p>===============================================================================<br>===============================================================================<br>Subdirectory 1: "model_inputs/"<br>===============================================================================<br>===============================================================================</p> <p>Two versions of each of four files are provided, corresponding to analyses <br>that do or do not consider ancient genetic lineages as potential sources of <br>spatial nonstationarity in mammal niches. Each file type is formatted the same,<br>and the versions are differentiated by either the suffix "nolineage" or <br>"lineage" in the filename. </p> <p>===============================================================================<br>File 1: gridcell_covars_lineage.csv and gridcell_covars_nolineage.csv<br>===============================================================================<br>These files are .csvs giving spatial covariate data for each scale 2 cell<br>in North America, summarized to 5000 m. All percentage values are given in <br>10ths of a percent (scale of 0-1000). The following columns are provided:</p> <p>- grid_cell: Scale 2 cell ID<br>- Arable: pct arable land (Jung et al. 2020)<br>- EVI_mean: mean enhanced vegetation index (Didan 2021)<br>- EVI_Q95: 95th quantile of EVI (Didan 2021)<br>- Forest: Pct forest cover (Jung et al. 2020)<br>- Grassland: pct grassland (Jung et al. 2020)<br>- Pastureland: pct pastureland (Jung et al. 2020)<br>- Pop_den: Human population density, from Gridded Population of the World (CIESIN 2018)<br>- Precipitation: avg annual precip. (Vega et al. 2017)<br>- Shrubland: pct shrubland (Jung et al. 2020)<br>- Temp_max: Average maximum daily temperature (Vega et al. 2017)<br>- Terrain_roughness (Amatulli et al. 2018)<br>- Wetlands: pct wetlands (Jung et al. 2020)<br>- is_land: Whether or not the cell is on land vs. ocean, used for filtering<br>- Agriculture: Pct. agricultural land (Jung et al. 2020)<br>- Pop_den_sqrt: Square root of human population density (CIESIN 2018)<br>- EVI_variability: Distance btw the 95% inner quantiles of EVI (Didan 2021)</p> <p> </p> <p>===============================================================================<br>File 2: inat_cts_lineage.csv and inat_cts_nolineage.csv<br>===============================================================================</p> <p>These files give summaries of iNaturalist sampling effort and detections<br>for target species. The following columns are provided:</p> <p>- grid_cell: Scale 3 cell ID<br>- n: Total iNaturalist effort in the cell (number of obs. of all mammals)<br>- The remaining columns are named for species. Each column gives the count<br> of observations of the species in the cell.</p> <p>===============================================================================<br>File 3: ct_datlist_lineage.RDS and ct_datlist_nolineage.RDS<br>===============================================================================</p> <p>The ct_datlist files contain R objects that are lists of lists. These objects ultimately<br>contain all of the camera detection histories and camera-level covariate data used<br>in modeling. We use the nice data type "unmarkedFrameOccu" from the unmarked R package<br>to organize these detection data.</p> <p>Each outer list is of length equal to the number of species. The ith element of each<br>list contains the following named slots:</p> <p>- species: a string giving the name of the ith species<br>- umf: an unmarkedFrameOccu object. This object has three important slots:<br> - y: a (# deployments) x (max # replicates) matrix giving 1s, 0s, or NAs indicating<br> whether the target species was observed in that 10-day window;<br> - siteCovs: a (# deployments) x 2 data frame with the following columns:<br> - site_ID: A unique ID of the exact location, shared by deployments with the same<br> coordinates<br> - subproject_ID: A unique ID indicating which camera array is associated <br> with this deployment<br> - obsCovs: a (# deployments * max # replicates) x 6 data frame with the following columns:<br> - year: the year of survey, relative to 2020 (zero-year is 2020)<br> - yday_scaled: the (scaled) Julian date of the beginning of the window<br> - yday_scaled_sq: yday_scaled^2, for use in estimating a quadratic effect<br> - log_roaddist_scaled: Scaled distance to nearest road (Meijer et al. 2018)<br> - Canopy_height_scaled: Scaled canopy height (Potapov et al. 2021)<br> - obs_len_scaled: Scaled duration of window, to account for some windows <br> being cut off at < 10 days<br>- coords: a data frame. Originally, this file gave the exact position for each camera,<br> but these exact locations have been scrubbed for privacy. See the original sources<br> cited in the manuscript for full details. This data frame contains the following column:<br> - scale2_grid_ID: the ID of the Scale-2 5000 m grid cell containing the camera</p> <p>===============================================================================<br>File 4: grid_translator_wspecs_nolineage.csv and grid_translator_wspecs_lineage.csv<br>===============================================================================</p> <p>These files are used for bookkeeping to track the relationships between the <br>three spatial scales in this study. Each row corresponds to a single "scale 2"<br>cell, giving the ID of the corresponding S3 and S4 grid and also an ID for each<br>species indicating whether and where it is in the species' range. </p> <p>Note that the scale names in the code don't match the manuscript. In the code,<br>"scale 1" is the level of an individual camera, "scale 2" is the 5 km intensity<br>grid, "scale 3" is the 50 km iNaturalist grid, and "scale 4" is the 100 km<br>SVC grid.</p> <p>The following columns are provided:<br>- scale2_grid_ID: unique ID for each cell in the 5 km intensity grid<br>- scale3_grid_ID: unique ID for each cell in the 50 km iNaturalist aggregation<br>- scale4_grid_ID: unique ID for each cell in the 100 km SVC grid<br>- GRID_ID_[species]: for each species, a column is provided on the S4 scale<br> counting each cell in the species' modeled range. NAs<br> indicate that the S2 cell defined in the row is not<br> included in the species' modeled range.</p> <p>===============================================================================<br>===============================================================================<br>Subdirectory 2: "model_outputs/"<br>===============================================================================<br>===============================================================================</p> <p>===============================================================================<br>File 1: svc_estimates.csv<br>===============================================================================</p> <p>This file gives an estimate of the effect of each covariate on each species'<br>intensity, and the uncertainty in that estimate, for each species/covariate<br>pair. Results correspond to lineage models for species with phylogeographies<br>and non-lineage species otherwise. Each row represents the effect of one <br>covariate on one species' relative intensity process within one 100 km cell g. <br>Note that many estimates of beta_g are uncertain even for strong spatial <br>effects---the model is often confident that a spatial process is supported <br>while estimates of the realized process are uncertain.</p> <p>The following columns are provided:<br>- x: the x-coordinate of the 100 km cell<br>- y: the y-coordinate of the 100 km cell<br>- species<br>- parname: the name of the covariate<br>- mean: the mean of the posterior samples of beta_g<br>- 2.5%: the 2.5th quantile of the posterior samples of beta_g<br>- 50%: the 50th quantile of the posterior samples of beta_g<br>- 97.5%: the 97.5th quantile of the posterior samples of beta_g</p> <p>The following spatial projection is used to define X/Y coordinates:<br>"+proj=aea +lat_1=20 +lat_2=60 +lat_0=40 +lon_0=-96 +x_0=0 +y_0=0 +ellps=GRS80 +datum=NAD83"</p> <p>===============================================================================<br>File 2: predicted_intensity.tif<br>===============================================================================</p> <p>This file contains a raster "brick" giving the predicted intensity surface <br>and uncertainty in this surface for each species. All predictions are generated<br>using models that do *not* account for lineage information---this means that <br>predictions for species with lineages are not from the models reported in the<br>main manuscript. The reason for this is that we found that lineages were <br>overall unsupported, so better predictions can be arrived at by excluding this<br>source of uncertainty in the underlying intensity process.</p> <p>The raster brick has 66 layers. Each layer provides either the mean predicted<br>log intensity in each grid cell across the species range or else provides<br>the standard error of that predicted log intensity. Layer names indicate output<br>type and species associated with each layer.</p> <p><br>===============================================================================<br>===============================================================================<br>References<br>===============================================================================<br>===============================================================================</p> <p>Camera data are obtained from the following sources, which can be consulted to<br>obtain the original raw camera data</p> <p>- Cove, Michael V., et al. "SNAPSHOT USA 2019: a coordinated national camera trap survey of the United States." (2021): e03353.<br>- Kays, Roland, et al. "SNAPSHOT USA 2020: A second coordinated national camera trap survey of the United States during the COVID‐19 pandemic." (2022): e3775.<br>- Shamon, H., et al. “SNAPSHOT USA 2021: A third coordinated national camera trap survey of the United States.” Ecology, 105.6 (2024): e4318.<br>- Rooney, B., et al. “SNAPSHOT USA 2019–2023: The first five years of data from a coordinated camera trap survey of the United States.” In Press (2024).<br>- Kays, Roland, et al. "Does hunting or hiking affect wildlife communities in protected areas?." Journal of Applied Ecology 54.1 (2017): 242-252.<br>- Roberts, R. California Department of Fish and Wildlife, Bobcat Program Initiative. wildlifeinsights.org (2023).<br>- Lasky, Monica, et al. "CAROLINA CRITTERS: a collection of camera trap data from wildlife surveys across North Carolina." Ecology 102.7 (2021): e03372.<br>- Forrester, T. (2000). Urban to Wild Project. http://n2t.net/ark:/63614/w12004302. Accessed via wildlifeinsights.org on 2024-08-29.<br>- McMurry, S. et al. In review (2024).<br>- Forrester, T. (2011) Okaloosa S.C.I.E.N.C.E. Project. http://n2t.net/ark:/63614/w12004287. <br>- Myers, J. (2014) Tyson Research Center ForestGEO Project. http://n2t.net/ark:/63614/w12004295.<br>- McMurry, S., and Kays, R.(2023). Calloway Forest Preserve. http://n2t.net/ark:/63614/w12006449. Accessed via Wildlife Insights on 2024-08-29.<br>- McMurry, S., Parsons, A., Lasky, M., Luongo, K., Clark, J., McShea, W., Scher, L., Kays, R., Spurlin, J., Martin, G., Frech, G., Barajas-Salazar, K., Snider, M. (2022). Last updated October 2023. Calloway Forest Preserve. http://n2t.net/ark:/63614/w12004251. Accessed via wildlifeinsights.org on 2024-08-29.<br>- Kays, R.. (2008). Last updated March 2024. Albany Area Camera Trapping Project. http://n2t.net/ark:/63614/w12003860. Accessed via wildlifeinsights.org on 2024-08-29.<br>- Kays, R., Snider, M., McMurry, S., Alyetama, M. (2024). Last updated April 2024. Pilot Mountain Density 2024. http://n2t.net/ark:/63614/w12007160. Accessed via wildlifeinsights.org on 2024-08-29.<br>- Malleshappa, V., Smithsonian, E., Kays, R., Schuttler, S. (2015). Last updated December 2022. Museums Connect Mexico. http://n2t.net/ark:/63614/w12004298. Accessed via wildlifeinsights.org on 2024-08-29.</p> <p>Covariate data are obtained from the following sources:<br>- Vega, G. C., Pertierra, L. R. & Olalla-Tárraga, M. Á. MERRAclim, a high-resolution global dataset of remotely sensed bioclimatic variables for ecological modelling. Sci. Data 4, 170078 (2017).<br>- Jung, M. et al. A global map of terrestrial habitat types. Sci. Data 7, 256 (2020).<br>- Amatulli, G. et al. A suite of global, cross-scale topographic variables for environmental and biodiversity modeling. Sci. Data 5, 180040 (2018).<br>- Center For International Earth Science Information Network-CIESIN-Columbia University. Documentation for the Gridded Population of the World, Version 4 (GPWv4), Revision 11 Data Sets. (2018) doi:10.7927/H45Q4T5F.<br>- Didan, K. MODIS/Terra Vegetation Indices 16-Day L3 Global 1km SIN Grid V061. NASA EOSDIS Land Processes Distributed Active Archive Center https://doi.org/10.5067/MODIS/MOD13A2.061 (2021).<br>- Meijer, J. R., Huijbregts, M. A. J., Schotten, K. C. G. J. & Schipper, A. M. Global patterns of current and future road infrastructure. Environ. Res. Lett. 13, 064006 (2018).<br>- Potapov, P. et al. Mapping global forest canopy height through integration of GEDI and Landsat data. Remote Sens. Environ. 253, 112165 (2021).<br>- Jensen, A. J. et al. Geographic barriers but not life history traits shape the phylogeography of North American mammals. Glob. Ecol. Biogeogr. e13875 (2024).</p> <p>iNaturalist data are obtained from inaturalist.org via the data exporter (see manuscript for details).</p>
Baseline map of 137Cs inventories in reference soil sites at the continental scales of South America
<p>This dataset contains the baseline map of <sup>137</sup>Cs inventories in reference soil sites (Bq m<sup>-2</sup>, decay-corrected to 2020) estimated by Partial Least Square Regression (PLSR) with a spatial resolution of 2 km at the continental scale of South America, as well as the prediction uncertainties of the baseline map (coefficient of variation, %).<br> Details information regarding this dataset can be found in the original publication:<br> Mapping the spatial distribution of global <sup>137</sup>Cs fallout in soils of South America as a baseline for Earth Science studies, Earth-Science Reviews, Volume 214, 2021, 103542, ISSN 0012-8252, https://doi.org/10.1016/j.earscirev.2021.103542.</p>
DATASETS and OUTCOMES - Assessment of intrinsic aquifer vulnerability at continental scale through a critical application of the DRASTIC method: the case of South America
<p>A robust and comprehensive assessment of intrinsic aquifer vulnerability at continental scale map may represent an essential initial step towards a more sustainable land-use and water management.</p> <p>This repository contains the outcomes of an intrinsic aquifer vulnerability assessment of South America, performed by the DRASTIC method. The assets included in this repository are mainly raster maps (.tif, .geotif), created and georeferenced in QGIS (v3.16). Coordinate reference system (CRS) of the dataset is WGS84.</p> <p>Technical specifications of all graphical outcomes are stored in a dedicated file (README.txt).</p>
Ecological barriers mediate spatiotemporal shifts of bird communities at a continental scale
<p>### Ecological barriers mediate spatiotemporal shifts of bird communities ###</p> <p>Marjakangas, Bosco et al. 2022</p> <p>Methods explained in the publication (open access)</p> <p>--> readme file explains how to use the data and code</p>
Data for: Characterization of large-scale preferential flow across continental United States
<p>Understanding preferential flow (PF) at large scales is critical for improving land management and groundwater (GW) quality. However, limited knowledge of this process, due to soil surface heterogeneity and observational constraints, hampers progress. In this study, we propose estimating effective PF at remote sensing footprint scale (4 – 9 km) by examining its impact on soil moisture (SM) distribution and shallow GW (SGW) table fluctuations (depth 5 m). Effective PF encompasses macropore, funnel, and finger flow pathways influencing SGW table fluctuations. We compiled daily SGW observations (2019-2021) from 19 continental US (CONUS) sites through USGS. Using inverse modeling in HYDRUS-1D, SGW data, and CHIRPS precipitation data, we inversely estimated soil hydraulic parameters of the dual porosity model (DPM) simulating vertical flow from soil surface to subsurface. Effective PF presence was inferred using three criteria: (1) daily precipitation >= the site-specific average across multiple (calibration) years, (2) daily observed SGW table increase, and (3) daily difference between observed and DPM simulated SGW tables 50% of the site-specific RMSE. Leveraging optimized DPM parameters and associated soil texture, classified PF events, and Soil Moisture Active Passive (SMAP L3E) satellite-based SM, a Random Forest algorithm with 10-fold cross validation predicted large-scale effective PF events. Results indicate seasonal dependence, with spring having the highest occurrence of PF events. The Random Forest model achieved 98% accuracy in predicting large-scale PF events, with SMAP SM and saturated hydraulic conductivity (Ks) among the 4 most impactful variables. Our approach provides a soil hydraulic property, site characteristic, soil texture and remote sensing based generalized tool to analyze large-scale effective PF.</p>
Energy input, habitat heterogeneity, and host specificity on avian haemosporidian diversity at continental scales
<p>The correct identification of biotic and abiotic drivers affecting parasite diversity and assemblage composition at different spatial scales is crucial for understanding how pathogen distribution responds to anthropogenic disturbance and climate change. Here, we used a database of avian haemosporidian parasites to identify such drivers and their effect on the taxonomic and phylogenetic diversity of genera Plasmodium, Haemoproteus, and Leucocytozoon from three zoogeographic regions. We explored how parasite diversity is related to energy input (i.e., temperature, precipitation, and potential evapotranspiration [PET]), to habitat heterogeneity (i.e., climatic seasonality, vegetation density, ecosystem heterogeneity, human disturbance, and host richness), and to a novel assemblage-level metric related to parasite niche overlap (degree of generalism). We found that the relative importance of the predictors differed between the three studied parasite genera and across diversity metrics. Among the most consistent predictors, host richness was positively related to the taxonomic diversity of the three genera. Energy input and human footprint explained the phylogenetic diversity of Haemoproteus. Finally, the degree of generalism explained the diversity of Plasmodium and Leucocytozoon. Our results suggest that different dimensions of haemosporidian diversity are shaped by energy input, host heterogeneity, and assembly processes related to parasite resource use within local parasite assemblages.</p>
Data and scripts associated with "Rock weathering controls soil carbon storage potential at a continental scale"
<p>The zipped folder contains the following files:</p> <p>CONUS_mineral_grid.csv [gridded soil mineralogy maps, with coordinates representing cell centers]</p> <p>CONUS_grid.csv [blank grid, used for reproducing the analysis]</p> <p>NCSS_datamerge_071321.R [core script used for running the analyses presented in the associated publication]</p> <p>apply_depthwtavg.R [wrapper function for depth weighted averaging]</p> <p>attach_ENV.R [function to attach climate data]</p> <p>bootstrap_functions.R [functions for spatial bootstrap statistics]</p> <p>depthweightavg.R [function for calculating depth weighted averages]</p> <p>get_MWBM.R [function for compiling climate data into averages]</p> <p>get_SLP.R [function for reading and pre-processing the NASGLP dataset</p> <p>SLP_IDW.R [inverse distance weighting function interpolating the NASGLP data at query points]</p> <p>weather_calcs.R [secondary calculations partitioning Al and Fe; weathering rate estimation]</p>
Upscaling soil organic carbon measurements at the continental scale using multivariate clustering analysis and machine learning
<p><strong>Data Description</strong>:</p> <p>To improve SOC estimation in the United States, we upscaled site-based SOC measurements to the continental scale using multivariate geographic clustering (MGC) approach coupled with machine learning models. First, we used the MGC approach to segment the United States at 30 arc second resolution based on principal component information from environmental covariates (gNATSGO soil properties, WorldClim bioclimatic variables, MODIS biological variables, and physiographic variables) to 20 SOC regions. We then trained separate random forest model ensembles for each of the SOC regions identified using environmental covariates and soil profile measurements from the International Soil Carbon Network (ISCN) and an Alaska soil profile data. We estimated United States SOC for 0-30 cm and 0-100 cm depths were 52.6 + 3.2 and 108.3 + 8.2 Pg C, respectively.</p> <p>Files in collection (32):</p> <p>Collection contains 22 soil properties geospatial rasters, 4 soil SOC geospatial rasters, 2 ISCN site SOC observations csv files, and 4 R scripts</p> <p>gNATSGO TIF files:</p> <p>├── available_water_storage_30arc_30cm_us.tif [30 cm depth soil available water storage]<br> ├── available_water_storage_30arc_100cm_us.tif [100 cm depth soil available water storage]<br> ├── caco3_30arc_30cm_us.tif [30 cm depth soil CaCO3 content]<br> ├── caco3_30arc_100cm_us.tif [100 cm depth soil CaCO3 content]<br> ├── cec_30arc_30cm_us.tif [30 cm depth soil cation exchange capacity]<br> ├── cec_30arc_100cm_us.tif [100 cm depth soil cation exchange capacity]<br> ├── clay_30arc_30cm_us.tif [30 cm depth soil clay content]<br> ├── clay_30arc_100cm_us.tif [100 cm depth soil clay content]<br> ├── depthWT_30arc_us.tif [depth to water table]<br> ├── kfactor_30arc_30cm_us.tif [30 cm depth soil erosion factor]<br> ├── kfactor_30arc_100cm_us.tif [100 cm depth soil erosion factor]<br> ├── ph_30arc_100cm_us.tif [100 cm depth soil pH]<br> ├── ph_30arc_100cm_us.tif [30 cm depth soil pH]<br> ├── pondingFre_30arc_us.tif [ponding frequency]<br> ├── sand_30arc_30cm_us.tif [30 cm depth soil sand content]<br> ├── sand_30arc_100cm_us.tif [100 cm depth soil sand content]<br> ├── silt_30arc_30cm_us.tif [30 cm depth soil silt content]<br> ├── silt_30arc_100cm_us.tif [100 cm depth soil silt content]<br> ├── water_content_30arc_30cm_us.tif [30 cm depth soil water content]<br> └── water_content_30arc_100cm_us.tif [100 cm depth soil water content]</p> <p>SOC TIF files:</p> <p>├──30cm SOC mean.tif [30 cm depth soil SOC]<br> ├──100cm SOC mean.tif [100 cm depth soil SOC]<br> ├──30cm SOC CV.tif [30 cm depth soil SOC coefficient of variation]<br> └──100cm SOC CV.tif [100 cm depth soil SOC coefficient of variation]</p> <p>site observations csv files:</p> <p>ISCN_rmNRCS_addNCSS_30cm.csv 30cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p>ISCN_rmNRCS_addNCSS_100cm.csv 100cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p><br> <strong>Data format</strong>:</p> <p>Geospatial files are provided in Geotiff format in Lat/Lon WGS84 EPSG: 4326 projection at 30 arc second resolution.</p> <p><strong>Geospatial projection</strong>: </p> <pre><code>GEOGCS["GCS_WGS_1984", DATUM["D_WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["Degree",0.017453292519943295]] (base) [jbk@theseus ltar_regionalization]$ g.proj -w GEOGCS["wgs84", DATUM["WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["degree",0.0174532925199433]] </code></pre> <p> </p>
DNA metabarcoding captures different macroinvertebrate biodiversity than morphological identification approaches across a continental scale
<p><span>DNA-based aquatic biomonitoring methods show promise to provide rapid, standardized, and efficient biodiversity assessment to supplement and in some cases replace current morphology-based approaches that are often less efficient and can produce inconsistent results. Despite this potential, broad-scale adoption of DNA-based approaches by end-users remains limited, and studies on how these two approaches differ in detecting aquatic biodiversity across large spatial scales are lacking. Here, we present a comparison of DNA metabarcoding and morphological identification, leveraging national-scale, open-source, ecological datasets from the National Ecological Observatory Network (NEON). Across 24 wadeable streams in North America with 179 paired sample comparisons, we found that DNA metabarcoding detected twice as many unique taxa than morphological identification overall. The two approaches showed poor congruence in detecting the same taxa, averaging 59%, 35%, and 23% of shared taxa detected at the order, family, and genus levels, respectively. Importantly, the two approaches detected different proportions of indicator taxa like %EPT and %Chironomidae. DNA metabarcoding detected far fewer Chironomid and Trichopteran taxa than morphological identification, but more Ephemeropteran and Plecopteran taxa, a result likely due to primer choice. Overall, our results showed that DNA metabarcoding and morphological identification detected different benthic macroinvertebrate communities. Despite these differences, our results relating watershed-scale and local-scale abiotic variables to invertebrate community structure from both methods produced similar results. This suggests that DNA and morphological approaches are both suitable for use in basic and applied ecological research. Further refinement of DNA metabarcoding protocols, primers, and reference libraries–as well as more standardized, large-scale comparative studies–may improve our understanding of taxonomic agreement and data linkages between DNA metabarcoding and morphological approaches. </span></p>
Data for: Characterization of large-scale preferential flow across continental United States
Open the record for dataset details and reuse information.
DNA metabarcoding captures different macroinvertebrate biodiversity than morphological identification approaches across a continental scale
Open the record for dataset details and reuse information.
Energy input, habitat heterogeneity, and host specificity drive avian haemosporidian diversity at continental scales
Open the record for dataset details and reuse information.
Fine-grain predictions are key to accurately represent continental-scale biodiversity patterns
Open the record for dataset details and reuse information.
Climate more important than soils for predicting forest biomass at the continental scale
<p>Above-ground biomass in forests is critical to the global carbon cycle as it stores and sequesters carbon from the atmosphere. Climate change will disrupt the carbon cycle hence understanding how climate and other abiotic variables determine forest biomass at broad spatial scales is important for validating and constraining Earth System models and predicting the impacts of climate change on forest carbon stores. We examined the importance of climate and soil variables to explaining above-ground biomass distribution across the Australian continent using publicly available biomass data from 3130 mature forest sites, in 6 broad ecoregions, encompassing tropical, subtropical, and temperate biomes. We used the Random Forest algorithm to test the explanatory power of 14 abiotic variables (8 climate, 6 soil) and to identify the best-performing models based on climate-only, soil-only, and climate plus soil. The best performing models explained ~50% of the variation (climate-only: <i>R<sup>2</sup></i> = 0.47 ± 0.04, and climate plus soils: <i>R<sup>2</sup></i> = 0.49 ± 0.04). Mean temperature of the driest quarter was the most important climate variable, and bulk density was the most important soil variable. Climate variables were consistently more important than soil variables in combined models, and model predictive performance was not substantively improved by the inclusion of soil variables. This result was also achieved when the analysis was repeated at the ecoregion scale. Predicted forest above-ground biomass ranged from 18 to 1066 Mg ha<sup>-1</sup>, often under-predicting measured above-ground biomass, which ranged from 7 to 1500 Mg ha<sup>-1</sup>. This suggested that other non-climate, non-edaphic variables impose a substantial influence on forest above-ground biomass, particularly in the high biomass range. We conclude that climate is a strong predictor of above-ground biomass at broad spatial scales and across large environmental gradients, yet to predict forest above-ground biomass distribution under future climates, other non-climatic factors must also be identified.</p>
Data from: The role of diversification in the continental scale community assembly of the American oaks (Quercus)
Premise of the study: Evolutionary and biogeographic history, including past environmental change and diversification processes, are likely to have influenced the expansion, migration, and extinction of populations, creating evolutionary legacy effects that influence regional species pools and the composition of communities. We consider the consequences of the diversification process in shaping trait evolution and assembly of oak-dominated communities throughout the continental U.S. Methods: Within the US oaks, we tested for phylogenetic and functional trait patterns at different spatial scales, taking advantage of a dated phylogenomic analysis of American oaks and the US Forest Service Forest Inventory Analysis. Key Results: We find 1) phylogenetic overdispersion at small grain sizes throughout the US across all spatial extents and 2) a shift from overdispersion to clustering with increasing grain sizes. Leaf traits have evolved in a convergent manner, and these traits are clustered in communities at all spatial scales, except in the far west, where species with contrasting leaf types co-occur. Conclusions: Our results support the hypotheses that 1) interspecific interactions were important in parallel adaptive radiation of the genus into a range of habitats across the continent and 2) that the diversification process is a critical driver of community assembly. Functional convergence of complementary species from distinct clades adapted to the same local habitats is a likely mechanism that allows distantly related species to coexist. Our findings contribute to an explanation of the long-term maintenance of high oak diversity and the dominance of the oak genus in North America.
Data from: Mast seeding patterns are asynchronous at a continental scale
<p>Resource pulses are rare events with a short duration and high magnitude that drive the dynamics of both plant and animal populations and communities. Mast seeding is perhaps the most common type of resource pulse that occurs in terrestrial ecosystems, is characterized by the synchronous and highly variable production of seed crops by a population of perennial plants, is widespread both taxonomically and geographically, and is often associated with nutrient scarcity. The rare production of abundant seed crops (mast events) that are orders of magnitude greater than crops during low seed years leads to high reproductive success in seed consumers and has cascading impacts in ecosystems. Although it has been suggested that mast seeding is potentially synchronized at continental scales, studies are largely constrained to local areas covering tens to hundreds of kilometres. Furthermore, summer temperature, which acts as a cue for mast seeding, shows patterns at continental scales manifested as a juxtaposition of positive and negative anomalies that have been linked to irruptive movements of boreal seed-eating birds. Here, we show a breakdown in synchrony of mast seeding patterns across space, leading to asynchrony at the continental scale. In an analysis of synchrony for a transcontinental North America tree species spanning distances of greater than 5,200 km, we found that mast seeding patterns were significantly asynchronous at distances of greater than 2,000 km apart (all <i>P</i> < 0.05). Other studies have shown declines in synchrony across distance, but not asynchrony. Spatiotemporal variation in summer temperatures at the continental scale drives patterns of synchrony in mast seeding, and we anticipate that this affects the spatial dynamics of numerous seed-eating communities, from insects to small mammals to the large-scale migration patterns of boreal seed-eating birds.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.