Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10,119
datasets available to search
ShareScore release 0.9.0
Dataset results
10,119 results for “supported”
GRiMeDB: a comprehensive global database of methane concentrations and fluxes in fluvial ecosystems with supporting physical and chemical information
The Global River Methane Database (GriMeDB) is a compilation of measurements of CH4 concentrations and fluxes for flowing water environments derived from publications, reports, data repositories, and other outlets between 1973 and 2021. Assembly of GRiMeDB was motivated by the goal of having a centralized, standardized resource to facilitate further studies of CH4 pattern and process in flowing water systems, upscaling efforts, and identification of tendencies in when, where, and how CH4 has been sampled in streams and rivers across the world. Thus, CH4 data are supported by concurrent observations (as available) of aquatic CO2, N2O, temperature, conductivity, pH, dissolved oxygen, nitrogen, phosphorus, organic carbon, and discharge, along with site data (latitude, longitude, elevation, and [as available]: stream order, elevation, channel slope, catchment size, and codes for distinct or disturbed channel types). GRiMeDB includes over 24,000 records of CH4 concentration and greater than 8,000 flux measurements from over 5,000 unique sites, most of which are resolved to the daily time scale.
Supporting data for “Climate Intervention through Stratospheric Aerosol Injection may partially mitigate marine heatwaves"
Although climate intervention aims to lower the global average temperature, the potential impact of Stratospheric Aerosol Injection on marine heatwaves (MHW) has not been thoroughly examined. This spatial dataset provides global and regional MHW metrics—such as frequency, maximum intensity, and duration—from the Community Earth System Model, version 2 (CESM2), using the baseline scenario SSP2-4.5, referred to as a no climate intervention scenario, and the ARISE-SAI ensemble. The ARISE-SAI model uses the SSP2-4.5 scenario, introducing stratospheric aerosol injection at approximately 21 km in 2035, aiming to keep global mean surface air temperature near 1.5°C for ARISE-SAI-1.5 and near 1.0°C for ARISE-SAI-1.0 above pre-industrial levels. The dataset includes global MHW properties for the historical period (1990-2009), the current period under SSP2-4.5 emission scenario (2015-2034), and future scenarios under SSP2-4.5, ARISE-SAI-1.5, and ARISE-SAI-1.5 for 2050-2059 and 2060-2069.
Data and Code in support of Caterpillar abundance in a northern hardwood forest: exogenous effects, endogenous feedbacks, and multidecadal trends.
In this study, we analyzed caterpillar abundance and biomass measured over 50 years (1970 - 2021) in the Hubbard Brook Experimental Forest, New Hampshire, USA. We tested mechanisms for determination of caterpillar abundance that included weather, host plant quality, and predator abundance. This dataset includes data, R code, and spatial files supporting this study. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Datasets of "Carbide coating on nickel to enhance the stability of supported metal nanoclusters" Nanoscale, 2022, 14, 3589-3598
<p>These are the datasets related to the publication "Carbide coating on nickel to enhance the stability of supported metal nanoclusters", Nanoscale, 2022, 14, 3589-3598 (<a href="https://doi.org/10.1039/D1NR06485A">https://doi.org/10.1039/D1NR06485A</a>). They are saved as NeXus/HDF5 files according to the nxstm NeXus application definition (<a href="https://doi.org/10.5281/zenodo.5792930">https://doi.org/10.5281/zenodo.5792930</a>).</p>
Energy Cycle Characteristics for 5G/6G Networks Supported by RES, UAVs, and RISs
<h2><strong>Overview</strong></h2> <p>The following dataset presents the energy cycle characteristics for 5G/6G mobile systems supported by Renewable Energy Sources (RES) and/or Unmanned Aerial Vehicles (UAVs) and Reconfigurable Intelligent Surfaces (RISs). In addition, within the dataset, the energy gain related to the engagement of RES within the Radio Access Network (RAN) has also been distinguished.</p> <h2><strong>Scenario</strong></h2> <p>The considered network scenario includes 8 three- (<em>_results_gcas.csv</em>) or one-cell (<em>_results_scas.csv</em> & <em>_results_kras.csv</em>) base stations (BSs) placed within the Poznan city (surroundings of the old market) and supported by Renewable Energy Sources — photovoltaic panels (PVs) and/or wind turbines (WTs). The aforementioned base stations can be treated as stationary towers or mobile access points (e.g., drones/UAVs). Those latter have been additionally equipped with RIS devices, which are able to reflect and manipulate a radio signal to influence occurrences such as interferences, coverage, or human exposure. However, the use of RISs has been taken into account only to evaluate the impact of the engagement of such devices on the energy side of the mobile system, omitting the changes in radio characteristics. The network traffic has been assumed to be fixed (64 mobile users (UEs) with 100 Mbps downlink — DL, and 25 Mbps uplink — UL, per each), however, its density in specific parts of the city is modeled randomly for each simulation run. The simulation runs have been performed for 4 dates (vernal equinox, summer solstice, autumn equinox, winter solstice), each one from a different season of the year. The aim of such an approach was to highlight the impact of the time of the day and the year on the energy gain obtained thanks to enabling RES generators. The weather conditions assumed within the simulation are typical for the climate in Poland. </p> <h2><strong>Methodology</strong></h2> <p>The energy-cycle calculations (system's power consumption, renewable energy production, and excessive energy storage) have been based on the mathematical formulas from the scientific literature and performed within the digital simulation runs by using the Green Radio Access Network Design (GRAND) tool (developed by teams from the Ghent University & Poznan University of Technology). The UE-BS association process within the mobile system has been done by doing multi-objective optimization using the Gurobi software, which has taken into account parameters like path loss, predicted power consumption of BSs, and guaranteed DL & UL bit rates for UEs.</p> <h2><strong>Simulation setup</strong></h2> <p>The setup of the input parameters for used mathematical models (power consumption, energy generation, energy storage) has been done in accordance with the values attached within the delivered literature positions (cited within the publications included in the <em>Related works</em> section of the following dataset) and adjusted to the considered study. Furthermore, the data used to model the network environment (building distribution, coverage area, base stations' locations) as well as to predict weather conditions are the real data (for the year 2022) collected by the city hall of Poznan, one of the Polish mobile operators, and weather stations placed in Poznan, respectively. The number of simulation runs performed has been equal to 10 (each run has included energy-cycle calculations for 4 seasons of the year), with the time step of a single run set to 1 hour of the day.</p> <h2><strong>Results</strong></h2> <p>The results of the aforementioned investigations have been included in the attached files, which can be described as follows:</p> <h3><strong>File <em>_results_gcas.csv</em></strong></h3> <p>The first column denotes the date (season of the year), for which the values have been obtained. The columns from second to fifth present observed values of the State of Charge (SoC) of a battery system (in %) for a single network cell on average in a time step. Those columns are the obtained values for the RAN, in which no RES, only PVs, only WTs, and both types of RES generators have been enabled, respectively. </p> <h3><strong>Files <em>_results_scas.csv</em> & <em>_results_kras.csv</em></strong></h3> <p>The first column denotes the date (season of the year), for which the values have been obtained. The second and third columns denote the number of drone base station (DBS) exchanges within the wireless system on average in a particular time step, where no RES and only PVs are enabled, respectively. The fourth and fifth columns present the conventional (fossil-fuels-based) energy consumption (in kWh) for the whole system in a specific time step, in which no RES and only PVs are engaged for all the access nodes. The sixth column is the energy savings (in kWh) related to the use of RES generators within the mobile network. Furthermore, the seventh and eighth columns represent the amount of renewable energy harvested from the solar radiation in total and the peak value of this amount observed during the entire day, respectively.</p> <h2><strong>Acknowledgment</strong></h2> <p>More details about the conducted studies have been described within the attached papers (<em>Related works</em> section). The data has been collected within the COST CA10210 INTERACT. M. Deruyck is a Post-Doctoral Fellow of the FWO-V (Research Foundation – Flanders, ref: 12Z5621N). The work (including the following dataset preparation) by A. Samorzewski and A. Kliks was realized within project no. 2021/43/B/ST7/01365 funded by the National Science Center in Poland.</p>
iEEG Data for "Functional Group Bridge for Simultaneous Regression and Support Estimation"
<p>The repository contains analysis scripts and data used in Wang Z, Magnotti J, Beauchamp MS, Li M. Functional Group Bridge for Simultaneous Regression and Support Estimation, 2020. The data contains high-gamma brain responses across 8 subjects from "congruency" audio-visual experiment.</p>
Reservoir ecosystems support large pools of fish biomass
<p>Supplemental data for "Reservoir ecosystems support large pools of fish biomass".</p> <p>Parisek, C.A., F.A., De Castro, J.D. Colby, G.R. Leidy, S. Sadro, A.L. Rypel. Reservoir ecosystems support large pools of fish biomass. <em>Sci Rep</em> <strong>14</strong>, 9428 (2024). https://doi.org/10.1038/s41598-024-59730-z</p>
Temperature and Climate Attribution estimates supporting "Human Fingerprints on Daily Temperatures in 2022" (2x2 degrees, 2022)
<p>These data support the publication of "Human Fingerprints on Daily Temperatures in 2022" published in the <a href="https://www.ametsoc.org/index.cfm/ams/publications/bulletin-of-the-american-meteorological-society-bams/explaining-extreme-events-from-a-climate-perspective/">BAMS-EEE special issue</a> in 2024 (DOI: <a href="https://doi.org/10.1175/BAMS-D-23-0264.1">10.1175/BAMS-D-23-0264.1</a>). Included are:</p> <ul> <li>Temperatures: <strong>Gilfordetal2024_BAMS-EEE_T2022.nc</strong></li> <li>Attributions estimates (Climate Shift Index and Change in Information due to Perspective): <strong>Gilfordetal2024_BAMS-EEE_ChIP2022.nc</strong></li> </ul> <p>And an accompanying land-sea mask from ERA5 (<strong>Gilfordetal2024_BAMS-EEE_LandSeaMask.nc</strong>). All data values valid for the 2022 calendar year and interpolated to a 2x2 degrees spatial grid to support the study's analysis.</p> <p>For more information on this dataset or to follow up, please contact Daniel Gilford (<a href="mailto:dgilford@climatecentral.org" target="_blank" rel="noopener">dgilford@climatecentral.org</a>).<br><br><em>Funding for this work was provided by the Bezos Earth Fund, The Schmidt Family Foundation, High Meadows Foundation, and the William and Flora Hewlett Foundation.</em></p>
Supporting Data for Crawford et al. 2024, Effects of Cropland Abandonment on Biodiversity
<p><strong>This archive contains derived and supporting data products to support:</strong></p> <blockquote>Crawford CL*, Wiebe RA, Yin H, Radeloff VC, and Wilcove DS. 2024. Effects of cropland abandonment on biodiversity. <em>Nature Sustainability.</em> In press.</blockquote> <p>*Contact Christopher L. Crawford at ccrawford@alumni.princeton.edu with any questions.</p> <p>A public Zenodo archive of the Github repository containing analysis scripts developed for this project (https://github.com/chriscra/biodiversity_abandonment) can be found here: <a href="https://doi.org/10.5281/zenodo.13777205">10.5281/zenodo.13777205</a></p> <p>This analysis builds on: Crawford, C. L., Yin, H., Radeloff, V. C. & Wilcove, D. S. Rural land abandonment is too ephemeral to provide major benefits for biodiversity and climate. <em>Science Advances </em>8, 1–13 (2022). Data and scripts from Crawford et al. 2022 are archived and publicly available at Zenodo (https://doi.org/10.1126/sciadv.abm8999).</p> <p>The annual land cover maps (1987-2017, 30 meter resolution) that underlie our analysis were developed on Google Earth Engine using publicly available Landsat satellite imagery (Yin et al. 2020, Remote Sensing of Environment, https://doi.org/10.1016/j.rse.2020.111873).<br>These annual land cover maps, along with other derived data that were produced by Crawford et al. 2022, are archived and publicly available at Zenodo (https://doi.org/10.5281/zenodo.5348287).</p> <p>This archive includes important derived data products created for Crawford et al. 2024. Note that these and other project data are described in detail in **util/_util_files.R** (https://github.com/chriscra/biodiversity_abandonment). This is a convenience script that loads many of the relevant input and derived data that are used throughout the project. The primary required data files for reproducing this work are archived here, but "_util_files.R" also includes information about where additional files can be accessed (if external, e.g., https://doi.org/10.5281/zenodo.5348287) or created across the various .R and .Rmd files in this repository (e.g., "habitats.Rmd" chunk {r land-cover-of-abn-pixels}).</p> <p>Naming conventions for sites and raster files follow Crawford et al. 2022, as described here: https://doi.org/10.5281/zenodo.5348287</p> <p><strong>Site file names correspond to the following geographic locations:</strong><br>belarus = Vitebsk, Belarus / Smolensk, Russia<br>bosnia_herzegovina = Bosnia & Herzegovina<br>chongqing = Chongqing, China<br>goias = Goiás, Brazil<br>iraq = Iraq<br>mato_grosso = Mato Grosso, Brazil<br>nebraska = Nebraska / Wyoming, USA<br>orenburg = Orenburg, Russia / Uralsk, Kazakhstan<br>shaanxi = Shaanxi/Shanxi, China<br>volgograd = Volgograd, Russia<br>wisconsin = Wisconsin, USA</p> <h1><strong>This archive includes the following files:</strong></h1> <ul> <li>site_df.csv</li> <li>crop_to_abn_iucn_observed.zip</li> <li>crop_to_abn_iucn_potential.zip</li> <li>max_abn_lcc_iucn.zip</li> <li>max_abn_lcc_iucn_potential.zip</li> <li>lcc_iucn_habitat.zip</li> <li>lcc_iucn_habitat_potential.zip</li> <li>frag_df.csv</li> <li>frag_hypo_no_abn_2017_df.csv</li> <li>iucn_lc_crosswalk.csv</li> <li>habitat_age_req_coded.csv</li> <li>centroids_df.csv</li> <li>aoh_l.parquet</li> <li>aoh_feols.parquet</li> <li>aoh_start_end_l.parquet</li> <li>aoh_change_df.parquet</li> <li>aoh_est_change_tmp_all.csv</li> <li>aoh_obs_change_tmp_all.csv</li> <li>taxonomy_df.parquet</li> <li>final_species_list.csv</li> <li>trait_mod_df_modx1.rds</li> </ul> <h3>site_df.csv</h3> <p>A list of site names and related metadata describing our study sites, taken from https://zenodo.org/records/5348287</p> <h2>Derived habitat rasters:</h2> <h3>crop_to_abn_iucn_observed.zip (Calculation 1a)<br>crop_to_abn_iucn_potential.zip (Calculation 1b)<br>max_abn_lcc_iucn.zip (Calculation 2a)<br>max_abn_lcc_iucn_potential.zip (Calculation 2b)<br>lcc_iucn_habitat.zip (Calculation 3a)<br>lcc_iucn_habitat_potential.zip (Calculation 3b)</h3> <p>These maps show IUCN Level 2 habitat types (Jung et al. 2020) interpolated onto the land cover classes in the Yin et al. (2020) abandonment maps at multiple spatial and temporal extents, which serve as inputs for the three primary calculations in our manuscript. Accompanying each calculation is a corresponding map for a scenarios in which no abandoned croplands were recultivated over the course of the time series (marked as "potential"). Each .zip file contains maps for each of 11 sites.</p> <p><strong>Calculation 1. </strong>This calculation isolates the direct effect of abandonment on habitat availability, by comparing the habitat provided before and after abandonment. These "crop_to_abn_iucn" maps show IUCN Level 2 habitats in cropland pixels that experienced abandonment, including the abandonment period as well as the immediately preceding period of cultivation (to allow for a proper before and after comparison). As a result, these maps show only habitat provided by croplands when they were actively cultivated, abandoned, or, where appropriate, recultivated, which allows for a proper before and after comparison. These maps are created in the script "cluster/noncrop_precrop_mask.R".</p> <p><strong>Calculation 2.</strong> This calculation considered changes in habitat that took place exclusively in pixels that experienced abandonment at some point during the time series (following Calculation 1), but expanded to track changes across our entire time series, from 1987 through 2017, in order to account for any land cover that was cleared for agriculture prior to abandonment. These "max_abn_lcc_iucn" maps therefore show IUCN Level 2 habitat types for each pixel that was abandoned at any point during the time series, across the full time series. These maps were created in the script "habitats.Rmd" code chunks {r mask-lcc-iucn-habitat-to-abn} and {r *potential_max}. </p> <p><strong>Calculation 3. </strong>This calculation tracks habitat area provided by every pixel throughout the entire spatial and temporal extent (1987-2017), in order to place abandonment into the context of broader land-cover change dynamics like ongoing cropland expansion taking place alongside of abandonment. These "lcc_iucn" maps therefore show the IUCN Level 2 habitat types for each pixel at each site in each year of our time series. These maps were created in the script "habitats.Rmd" code chunks {r lcc-iucn-habitat-composite} and {r *potential-lcc-full} and the script "cluster/potential_full_iucn.R".</p> <p>Some analyses require these .tif files (manipulated as SpatRasters using {terra}, https://rspatial.org/terra/) to be converted to tabular format (data.tables, via {data.table} (https://rdatatable.gitlab.io/data.table/) and saved as .parquet files (via {arrow}, https://arrow.apache.org/docs/r/). This can be accomplished via scripts "cluster/save_spatraster_as_dt.R" and "cluster/save_parquet.R."</p> <h3><br>frag_df.csv<br>frag_hypo_no_abn_2017_df.csv</h3> <p>These tabular files contain derived fragmentation statistics calculated using the {landscapemetrics} R package (https://r-spatialecology.github.io/landscapemetrics/). The second file contains metrics for a scenario in which no croplands were abandoned through the year 2017, in order to assess the effect cropland abandonment on landscape configuration. Each file contains 11 columns: </p> <ol> <li>"layer" -- the spatial raster layer for which the metric is calculated, corresponding to a year.</li> <li>"level" -- the level at which the metric is calculated, in our case, the land cover "class."</li> <li>"class" -- corresponding the to land cover class for which the metric is calculated (1 = non-vegetation, 2 = woody vegetation [i.e., forest], 3 = cropland, and 4 = herbaceous vegetation [i.e., grassland]).</li> <li>"id" -- An unused field containing NA values.</li> <li>"metric" -- the specific term used for each metric by {landscapemetrics} ("area_mn", "clumpy", or "para_mn").</li> <li>"value" -- the numerical value of the statistic.</li> <li>"name" -- the name of the landscape metric being calculated ("patch area," "clumpiness index," or "perimeter-area ratio").</li> <li>"type" -- the broad type of metric being calculated ("area and edge metric," "aggregation metric," or "shape metric").</li> <li>"function_name" -- the name of the {landscapemetrics} function used to calculate the statistic.</li> <li>"site" -- the site (out of 11 study sites) for which this statistic was calculated.</li> <li>"year" -- the year corresponding to the metric statistic, between 1987-2017 (including 1986-2018 for Nebraska and 1987-2018 for Wisconsin)<br>Additional details on these metrics can be found at https://r-spatialecology.github.io/landscapemetrics/.</li> </ol> <p>The spatial IUCN data underlying our analyses (species range maps) are available upon request from BirdLife International (http://datazone.birdlife.org/species/requestdis) and IUCN (https://www.iucnredlist.org/resources/spatial-data-download). Tabular species assessment data (including habitat and elevation preferences) are freely available from IUCN (https://www.iucnredlist.org/). Here we share three IUCN-related data files that serve as important inputs throughout our analyses:</p> <h3>iucn_lc_crosswalk.csv</h3> <p>This tabular file outlines the crosswalk between the 4 land cover classes in Yin et al. 2020 and the IUCN Level 2 habitat types mapped by Jung et al. 2020. It contains five columns:</p> <ol> <li>"map_code" -- the habitat code corresponding to Jung et al. (2020).</li> <li>"Coarse_Name" -- the broad Level 1 habitat grouping.</li> <li>"lc" -- the corresponding land cover type from Yin et al. (2020) (1 = non-vegetation, 2 = woody vegetation [i.e., forest], 3 = cropland, and 4 = herbaceous vegetation [i.e., grassland]).</li> <li>"IUCNLevel" -- the full IUCN Level 2 habitat type name. </li> <li>"code" -- the IUCN Level 2 habitat code. </li> </ol> <h3>habitat_age_req_coded.csv</h3> <p>This tabular file lists whether each species was determined (by R. Alex Wiebe [AW] and Christopher L. Crawford [CLC]) to be a "mature forest obligate" (i.e., requiring forest older than 30 years, our time series length) or not. Species determined to be "mature forest obligate" species were excluded from our final analysis. The file includes 11 columns: </p> <ol> <li>"vert_class" -- Vertebrate class ("bird" or "mam" [mammal])</li> <li>"binomial" -- Species' binomial scientific name containing genus and species.</li> <li>"common_names" -- Species' common names listed by IUCN.</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"water_obl" -- Whether a species is determined to be a "water obligate" species (1) or not (0). Some species were marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty. Note: this field was not used in the analysis.</li> <li>"habitat" -- The description of the species' habitat, drawn from individual IUCN assessments (see https://www.iucnredlist.org/).</li> <li>"site_presence" -- Where each species is present across our 11 study sites.</li> <li>"suitable_habitats" -- A list of IUCN Level 2 habitat types consider suitable habitat by each species.</li> <li>"major_habitats" -- A list of IUCN Level 2 habitat types listed as having "Major Importance" for that species.</li> <li>"coder" -- The author that assigned the mature forest obligate and water obligate codes ("AW" = R. Alex Wiebe, "CLC" = Christopher L. Crawford).</li> <li>"Chris_notes" -- A text field contains notes on coding process.</li> </ol> <h3>centroids_df.csv</h3> <p>This is a simple tabular dataset containing the longitude and latitude of the centroid of each bird and mammal species' range that overlaps with one of my sites. Columns include "binomial," which lists each species binomial scientific name, "centroid_longitude," and centroid_latitude." Centroid positions were calculated in QGIS using species range files from IUCN and BirdLife International.</p> <h3><br>aoh_l.parquet</h3> <p>This tabular file contains the raw AOH results produced using the script "cluster/aoh.R." This file contains the area of each suitable IUCN Level 2 habitat for each bird and mammal species at each site in each year of our time series (1987-2017), calculated across a range of calculations and scenarios. This file includes the primary data that serve as inputs for much of the rest of the analysis. The overall area of habitat for each species in each year at each site (a tabular data file named "aoh") summed across suitable habitat types and filtered to include or exclude passage areas for migratory birds, is calculated from "aoh_l" in the "AOH.Rmd" script in code chunks "filter-aoh-suitability-by-season" and "**calculate-aoh" (similarly to other derived datasets that serve as inputs for various parts of the analysis). This "aoh" file provides input data for the linear models used to extract AOH trends and test for significance. "aoh_l.parquet" includes 20 columns: </p> <ol> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"year" -- Year for which AOH is calculated (1987-2017).</li> <li>"map_code" -- Code indicating the IUCN Level 2 habitat associated with the area statistic. See "iucn_lc_crosswalk.csv."</li> <li>"season" -- Seasonal code indicating the season in which a species considers the habitat to be suitable, drawn from IUCN. Codes are: 1 ("Resident"), 2 ("Breeding") (2), "Non-breeding Season" (3), Passage (4), and Seasonal Occurrence Uncertain (5)</li> <li>"area" -- Area of Habitat, in hectares (ha).</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"IUCN_aoh_ha" -- [Unused] A preliminary summation of all habitat area for each species in each year, prior to filtering. We did not use this field in our analysis. Our final AOH calculation involved first filtering out mismatched season and habitat suitability combinations.</li> <li>"time" -- The time required for the area of habitat calculation (in seconds).</li> <li>"className" -- Vertebrate class: "AMPHIBIA," "AVES," or "MAMMALIA."</li> <li>"category" -- Duplicate field for IUCN Red List Category, unused.</li> <li>"core_index" -- An index used to assign specific AOH calculations to run in parallel across multiple computing cores on Princeton's High-Performance Computing Cluster.</li> <li>"total_range_area" -- The species total range area, in square kilometers (km^2), calculated across all range polygons for each species provided by IUCN and BirdLife International. See "cluster/calc_range_area.R."</li> <li>"range_size_quantile" -- A numerical index representing global species range size quantiles, within each class. Values range from 0 (the smallest global range within a class) to 1 (the largest global range within a class). These quantiles are used to define "small-ranged species," as species with global range sizes smaller than the median global range size in their class. See "cluster/calc_range_area.R."</li> <li>"water_obl" -- Whether a species is determined to be a "water obligate" species (1) or not (0). Some species were marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty. Note: this field was not used in the analysis. Drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"coder" -- The author that assigned the mature forest obligate and water obligate codes ("AW" = R. Alex Wiebe, "CLC" = Christopher L. Crawford). Drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> </ol> <h3><br>aoh_feols.parquet</h3> <p>This tabular data contains the results of linear regressions predicting area of habitat as a function of time. We parameterized models for each species in each site for each of the 6 AOH calculation types described above and in Crawford et al. 2024 (Calculations 1a, 1b, 2a, 2b, 3a, and 3b). We used the R package {fixest} to parameterize these ordinary least squares (OLS) linear regressions, using the Newey-West estimator to calculate standard errors. We used the R package {broom} to extract ("tidy") the model coefficient estimates and statistics. See "AOH.Rmd" chunk {r **feols}. This file includes 20 columns:</p> <ol> <li>"term" -- The name of the regression term: "(Intercept)" or slope ("year0").</li> <li>"estimate" -- The estimated value of the regression term.</li> <li>"std.error" -- The standard error of the regression term.</li> <li>"statistic" -- The value of a T-statistic to use in a hypothesis that the regression term is non-zero.</li> <li>"p.value" -- The two-sided p-value associated with the observed statistic.</li> <li>"conf.low" -- Lower bound on the confidence interval for the estimate (in our case 5%).</li> <li>"conf.high" -- Upper bound on the confidence interval for the estimate (in our case, 95%).</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"n_obs" -- The number of observations included in the model run.</li> <li>"n_unique_obs" -- The number of unique observations included in the model run (used to exclude species with constant AOH).</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"start_year" -- The first year for which this species has area of habitat at this site (i.e., the first observation included in the model).</li> <li>"end_year" -- The last year for which this species has area of habitat at this site (i.e., the last observation included in the model).</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> </ol> <p><br><strong>Two files contain model effect sizes for AOH models:</strong></p> <h3>aoh_start_end_l.parquet</h3> <p>This tabular data file contains observed effect sizes: the observed change in AOH for each species at each site, in each calculation, derived directly from observations from the start and end of the time series. These data are calculated in "AOH.Rmd" chunk: {r observed-change-in-aoh-by-window-size}. This data serves as direct input for the file "aoh_obs_change_tmp_all" (see below), which is the primary input for the traits linear models in our analysis (see "traits.Rmd", "_util_files.R"). This file contains 24 columns:</p> <ol> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"start" -- The mean area of habitat (AOH), in hectares (ha), at the "start" of the time series, as calculated across the number of years specified in "window_size."</li> <li>"start_year" -- The year of the first AOH observation.</li> <li>"end" -- The mean area of habitat (AOH), in hectares (ha), at the "end" of the time series, as calculated across the number of years specified in "window_size."</li> <li>"end_year" -- The year of the last AOH observation.</li> <li>"window_size" -- The number of years across which "start" and "end" AOH values are averaged (e.g., if "window_size" is 5, "start" is then the mean AOH across the first 5 years of observations, and "end" is the mean AOH across the last 5 years of observations).</li> <li>"abs_change" -- The absolute change in AOH, calculated as the difference between the mean AOH at the end of the time series and the mean AOH at the start of the time series (i.e., end - start).</li> <li>"prop_change" -- The proportional change in AOH, calculated as the absolute change in AOH divided by the AOH value at the start of the time series (i.e., abs_change/start).</li> <li>"percent_change" -- The percent change in AOH, calculated as 100 times the proportional change in AOH (i.e., 100 * prop_change).</li> <li>"ratio" -- The ratio of the mean AOH at the end of the time series to the mean AOH at the start of the time series (i.e., end/start).</li> <li>"ratio_mod" -- A modified ratio of the ending AOH to the starting AOH, for which ratio values less than 1 are replaced by additive inverse of the reciprocal value (i.e., 1/ratio * -1). Ratios greater than 1 are left the same.</li> <li>"abs_change_as_prop_site_area" -- The absolute change in AOH as a proportion of site area (i.e., abs_change / total_site_area_ha_2017).</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"total_site_area_ha_2017" -- The total site area (ha) in 2017. (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"area_ever_abn_ha" -- The total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017). (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"trend" -- The overall trend in AOH ("gain," "loss," or "no trend"), determined by the sign of slope coefficients and statistical significance at p < 0.05.</li> <li>"factor_change" -- The factor change in AOH, calculated as either the proportional change in AOH (i.e., prop_change) for values greater than 0, or as the reciprocal of the proportional change in AOH (i.e., 1/prop_change) for values greater than 0.</li> </ol> <h3>aoh_change_df.parquet</h3> <p>This tabular data file contains effect sizes estimated from linear regression coefficients (i.e., slopes and intercepts), calculated in "AOH.Rmd" chunk {r estimated-changes-aoh-change-df}. This file contains 29 columns:</p> <ol> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"est_type" -- The estimate type, whether the estimated model slope ("estimate") or the lower ("conf.low") or upper ("conf.high") bounds of the 95% confidence interval around the slope estimate.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"start_year" -- The year of the first AOH observation.</li> <li>"end_year" -- The year of the last AOH observation.</li> <li>"slope" -- The model estimated slope value.</li> <li>"intercept" -- The model estimated intercept value.</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"total_site_area_ha_2017" -- The total site area (ha) in 2017. (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"area_ever_abn_ha" -- The total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017). (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"trend" -- The trend in AOH experienced by the species at this site for this aoh_type calculation ("gain," "loss," or "no trend"), determined by the sign of slope coefficients and assigning statistical significance when p < 0.05.</li> <li>"n_trends" -- The number of distinct trends in AOH experienced by the species across all of the sites overlapping with its range, including this site.</li> <li>"trend_types" -- The types of trends in AOH experienced by this species across all sites overlapping with its range (some combination of "gain", "loss", and/or "no trend").</li> <li>"overall_trend" -- The overall trend in AOH experienced by this species across all sites overlapping with its range ("gain" - experiencing "gain" trends at all occurring sites; "loss" - experiencing "loss" trends at all occurring sites; "no trend" - experiencing "no trend" at all occurring sites; "weak gain" - experiencing "gain" trends at some sites and "no trend" at others; "weak loss" - experiencing "loss" trends at some sites and "no trend" at others; or "context dependent" - experienced "gain" trends at some sites and "loss" trends at other sites [referred to as "mixed" effects in Crawford et al. 2024])</li> <li>"trend_direction" -- The general direction of the trend in AOH for the species across all occurring sites ("gain" when overall_trend is either "gain" or "weak_gain"; "loss" when overall_trend is either "loss" or "weak_loss"; "context dependent" when overall_trend is "context dependent" [i.e., "mixed" effects]; and "no trend" when overall_trend is "no trend").</li> <li>"trend_consistency" -- An indication of how consistent the trend in AOH is across all occurring sites ("consistent" if overall_trend is "gain" or "loss"; "weak" if "weak_gain" or "weak_loss"; and "opposite" if "context dependent" [i.e., "mixed" effects]).</li> <li>"time_range" -- The number of years for which the species has AOH observations at this site for this aoh_type calculations.</li> <li>"aoh_start_est" -- The estimated AOH at the start of the time series, calculated from linear regression slope and intercept coefficients.</li> <li>"aoh_end_est" -- The estimated AOH at the end of the time series, calculated from linear regression slope and intercept coefficients.</li> <li>"abs_change" -- The absolute change in estimated AOH over the course of the time series (i.e., aoh_end_est - aoh_start_est).</li> <li>"abs_change_as_prop_site_area" -- The absolute change in estimated AOH as a proportion of site area (i.e., abs_change / total_site_area_ha_2017).</li> <li>"ratio_change" -- The ratio of the estimated AOH at the end of the time series to the estimated AOH at the start of the time series (i.e., aoh_end_est / aoh_start_est).</li> <li>"prop_change" -- The proportional change in estimated AOH, calculated as the absolute change in estimated AOH divided by the estimated AOH value at the start of the time series (i.e., abs_change / aoh_start_est).</li> <li>"factor_change" -- The factor change in estimated AOH, calculated as either the proportional change in estimated AOH (i.e., prop_change) for values greater than 0, or as the reciprocal of the proportional change in estimated AOH (i.e., 1/prop_change) for values greater than 0.</li> <li>"percent_change" -- The percent change in estimated AOH, calculated as 100 times the proportional change in estimated AOH (i.e., 100 * prop_change).</li> </ol> <h3>taxonomy_df.parquet</h3> <p>This tabular data file contains basic taxonomic information used in the analysis, including 10 columns:</p> <ol> <li>"vert_class" -- Vertebrate class ("bird," birds; or "mam," mammals).</li> <li>"binomial" -- Species binomial scientific name, drawn from IUCN or BirdLife International.</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"order" -- Taxonomic order.</li> <li>"family" -- Taxonomic family.</li> <li>"n_sp_in_family_sample" -- The number of species contained in the family included in our analysis.</li> <li>"order_common" -- A common name to refer to the order.</li> <li>"family_common" -- A common name to refer to the family.</li> <li>"n_in_family" -- The total number of species contained in the family globally.</li> <li>"threatened" -- Whether a species is considered threatened with extinction (i.e., is listed as "Critically Endangered," "Endangered," or "Vulnerable" on the IUCN Red List).</li> </ol> <h3><br>aoh_obs_change_tmp_all.csv<br>aoh_est_change_tmp_all.csv</h3> <p>These two tabular data files contain data used as inputs for the linear models involved in our traits analysis exploring how species' responses to cropland abandonment are affected by habitat suitabilities and other traits. The key variables are the response variables for our models ("binary_gain_v_loss", "abs_change_percent_site", and "log(ratio)") and predictor variables c("forest_occ", "savanna_occ", "shrubland_occ", "grassland_occ", "wetlands_occ", "rocky_occ", "caves_occ", "desert_occ", "urban_occ", "arable_occ", "n_suitable_habitats_lvl2", "vert_class", "threatened", "Trophic_level", "log10(Body_mass_g)", "log10(total_range_area)", "abs(centroid_latitude)", and "max_abn_ext_percent_site"). Further details are contained in "traits.Rmd"</p> <p>These two files are developed from "aoh_start_end_l" and "aoh_change_df," but filtered to include only birds and mammals, to exclude passage areas from AOH calculations, to exclude mature forest obligate species, and to use only a window_size of 5 years (for "aoh_obs_change_tmp_all") and model estimates (rather than 95% confidence interval bounds, for "aoh_est_change_tmp_all"). </p> <p><strong>aoh_obs_change_tmp_all.csv contains 63 columns.</strong></p> <ul> <li>Columns 1-24 match "aoh_start_end_l". </li> <li>Columns 25-31 match "taxonomy_df" columns 3 through 10.</li> <li>Columns 32-34: "Body_mass_g" (species body mass, in grams), "Trophic_level" (whether a species is a "Carnivore", a "Herbivore," or an "Omnivore"), and "Habitat_breadth_IUCN" (the number of IUCN Level 2 habitats a species can occupy) were taken from from Etard et al. 2020 (https://doi.org/10.1111/geb.13184)</li> <li>Column 35: "total_range_area" -- drawn from "aoh_l," see above</li> <li>Columns 36-37: "centroid_longitude" and "centroid_latitude" are drawn from "centroids_df," see above.</li> <li>Columns 38-50 are Boolean variables that indicate whether a species can occupy a specific IUCN Level 1 habitat type (i.e., whether IUCN lists that Level 1 habitat as suitable for the species). These variables are as follows, with the IUCN Level 1 habitat code listed in brackets: "forest_occ" [1], "savanna_occ" [2], "shrubland_occ" [3], "grassland_occ" [4], "wetlands_occ" [5], "rocky_occ" [6], "caves_occ" [7], "desert_occ" [8], "marine_intertidal_occ"[12], "marine_coastal_occ" [13], "artificial_terrestrial_occ" [14], "artificial_aquatic_occ" [15], and "introduced_occ" [16].</li> <li>Columns 51-52 represent the number of IUCN Level 1 ("n_suitable_habitats") and IUCN Level 2 ("n_suitable_habitats_lvl2") habitats a species has listed as suitable habitats by IUCN, respectively.</li> <li>Columns 53-56 are Boolean variables indicating whether a species can occupy a subset of IUCN Level 2 habitats, which are listed in brackets: "arable_occ" [14.1 Arable Land]; "farmland_occ" [14.1 Arable Land, 14.2 Pastureland, or 14.4 Rural Gardens]; "ag_occ" (duplicate of "farmland_occ"); "urban_occ" [14.5 Urban Areas].</li> <li>Columns 57-58 represent the maximum spatial extent of abandonment at a give site (i.e., the area of all lands that were abandoned at least once during the time series), whether divided by site area ("max_abn_extent_div_site_area," i.e,. area_ever_abn_ha / total_site_area_ha_2017) or as a percent of site area ("max_abn_ext_percent_site").</li> <li>Column 59 is "abs_change_percent_site," calculated as 100 * abs_change_as_prop_site_area.</li> <li>Columns 60-63 are binary values (1 or 0) indicating the whether the species experienced statistically significant gains in AOH ("binary_trend_gain"), statistically significant losses in AOH ("binary_trend_loss"), no trend in AOH ("binary_trend_no_trend"). Column 63 ("binary_gain_v_loss") is a binary value assigning a value of 1 for gains, 0 for losses, and NA for other values.</li> </ul> <p><br><strong>aoh_est_change_tmp_all.csv contains 70 columns:</strong></p> <ul> <li>Columns 1-29 match "aoh_change_df".</li> <li>Columns 30-37 match "taxonomy_df" columns 3 through 10.</li> <li>Columns 38-40: "Body_mass_g" (species body mass, in grams), "Trophic_level" (whether a species is a "Carnivore", a "Herbivore," or an "Omnivore"), and "Habitat_breadth_IUCN" (the number of IUCN Level 2 habitats a species can occupy) were taken from from Etard et al. 2020 (https://doi.org/10.1111/geb.13184)</li> <li>Column 41: "total_range_area" -- drawn from "aoh_l," see above</li> <li>Columns 42-43: "centroid_longitude" and "centroid_latitude" are drawn from "centroids_df," see above.</li> <li>Columns 44-56 are Boolean variables that indicate whether a species can occupy a specific IUCN Level 1 habitat type (i.e., whether IUCN lists that Level 1 habitat as suitable for the species). These variables are as follows, with the IUCN Level 1 habitat code listed in brackets: "forest_occ" [1], "savanna_occ" [2], "shrubland_occ" [3], "grassland_occ" [4], "wetlands_occ" [5], "rocky_occ" [6], "caves_occ" [7], "desert_occ" [8], "marine_intertidal_occ"[12], "marine_coastal_occ" [13], "artificial_terrestrial_occ" [14], "artificial_aquatic_occ" [15], and "introduced_occ" [16].</li> <li>Columns 57-58 represent the number of IUCN Level 1 ("n_suitable_habitats") and IUCN Level 2 ("n_suitable_habitats_lvl2") habitats a species has listed as suitable habitats by IUCN, respectively.</li> <li>Columns 59-62 are Boolean variables indicating whether a species can occupy a subset of IUCN Level 2 habitats, which are listed in brackets: "arable_occ" [14.1 Arable Land]; "farmland_occ" [14.1 Arable Land, 14.2 Pastureland, or 14.4 Rural Gardens]; "ag_occ" (duplicate of "farmland_occ"); "urban_occ" [14.5 Urban Areas].</li> <li>Columns 63-64 represent the maximum spatial extent of abandonment at a give site (i.e., the area of all lands that were abandoned at least once during the time series), whether divided by site area ("max_abn_extent_div_site_area," i.e,. area_ever_abn_ha / total_site_area_ha_2017) or as a percent of site area ("max_abn_ext_percent_site").</li> <li>Columns 65-68 are binary values (1 or 0) indicating the whether the species experienced statistically significant gains in AOH ("binary_trend_gain"), statistically significant losses in AOH ("binary_trend_loss"), no trend in AOH ("binary_trend_no_trend"). Column 63 ("binary_gain_v_loss") is a binary value assigning a value of 1 for gains, 0 for losses, and NA for other values.</li> <li>Column 69, "slope_prop_site", is the estimated linear regression coefficient, or slope, as a proportion of site area, calculated as slope / total_site_area_ha_2017. </li> <li>Column 70 is "abs_change_percent_site," calculated as 100 * abs_change_as_prop_site_area.</li> </ul> <h3><br>final_species_list.csv</h3> <p>The final list of bird and mammal species included in our analysis, including the vertebrate class ("vert_class") and binomial species scientific name ("binomial") along with the overall response to cropland abandonment ("overall_trend"), the sites where that species had AOH affected by cropland abandonment ("sites"), the IUCN Red List Category ("redlistCategory"), the "obligate_type" (i.e., whether a species is a mature forest obligate, or not), the "range_size_quantile" (ranking species by global geographic range size), and "common_names". Note that mature forest obligates were excluded from our final results. These columns match the definitions included above.</p> <h3>trait_mod_df_modx1.rds</h3> <p>This R data file contains the results of our regression models run in "traits.Rmd" code chunk "*many-models", which is where our three traits linear regression models are run. These data are contained in the form of a nested tibble, or a set of tibbles nested within columns of a tibble (see: https://tidyr.tidyverse.org/articles/nest.html). These data include the input data ("data"), resulting models ("model"), model coefficients ("tidy"), regression tables ("gt"), and diagnostic statistics ("glance") for our many model runs across different response variables ("response") and AOH calculations ("aoh_type"). See "traits.Rmd" code chunk "*many-models" for more information.</p>
CFMDG: a Coastal Flood Modelling Dataset in Gâvres (France) to support risk prevention and metamodels development
<p>Along most of the coastal areas, detailed coastal flood observations (e.g. inland water depths) are scarce, and when they are available, this for a limited number of events. Given recent scientific advances, <strong>coastal flooding</strong> events can be properly modelled, even in complex environments and under the action of wave overtopping, and thus provide detailed information. However, such models are computationally expensive, which prevents their use for instance for forecasting and warning. At the same time, metamodelling techniques have been explored for coastal hydrodynamics and have shown promising results. Metamodels are functions that aim to reproduce the behaviour of a “true” model (e.g., a numerical hydrodynamic model) for given input variables (for instance, offshore conditions). Within the RISCOPE research project (<a href="http://perso.math.univ-toulouse.fr/riscope">https://perso.math.univ-toulouse.fr/riscope</a>/) aiming at exploring to which extent such metamodelling techniques may allow to forecast coastal floods with a good accuracy, a <strong>simulated flood database</strong> has been built for the site of Gâvres (France), characterised by a significant effect of wave overtopping processes.</p> <p>The <strong>CFMDG dataset </strong>compiles a set of post-processed coastal flood simulations on the site of Gâvres. The dataset includes 250 scenarios. Each scenarios is defined by 6h time series centered on high tide, with one time series per forcing variables. The forcing variables (called X) are: local relative mean sea-level, tide, atmospheric storm surge, the offshore wave characteristics and the offshore wind. These scenarios combine past real (flood and no flood) events in the 1900-2021 time span with extreme statistics based events, and some complementary fictive events. The post-processed outputs (called Y) includes, for each scenario, the maximal flooded area (m²) and the maximal water depth (m) in each of the 64 618 inland model grid points.</p> <p>The modelling chain that allowed building this dataset relies on the joint use of a spectral wave model (WW3) to propagate the waves to the coast, and a non-hydrostatic wave-flow model (SWASH) to simulate the nearshore hydrodynamics and the flooding. The spatial and temporal resolution of the SWASH configuration validated on the Gâvres site are respectively 3 m and more than 10Hz. All the results are obtained for a Digital Elevation Model corresponding to the 2018 configuration of the site. </p> <p>Such type of dataset is of use for local knowledge, risk prevention, metamodel testing/training, and local coastal flood forecast. </p> <p>Part of this dataset has already been used in (<a href="http://www.mdpi.com/2077-1312/9/11/1191">Idier et al., 2021</a>; <a href="http://www.sciencedirect.com/science/article/pii/S0951832021006293?via%3Dihub">López-Lopera et al., 2021</a>; <a href="https://hal.science/hal-02536624">Betancourt et al., 2022</a>), to develop metamodels and set up a coastal flood forecast and early warning prototype.</p> <p>We hope and expect that making this dataset accessible will trigger further developments/investigations for improving risk knowledge on the considered site as well as methodological developments on machine-learning/metamodel-based techniques to support flood forecast.</p> <p>The table below summarizes the variables contained in the dataset, for each scenario.</p> <table> <tbody> <tr> <td> <p><strong>Variable name</strong></p> </td> <td> <p><strong>Description and unit </strong></p> </td> <td> <p><strong>Comment</strong></p> </td> </tr> <tr> <td> <p>Scenario n°</p> </td> <td> <p>Number of the scenario.</p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p><strong>INPUTS (X)</strong></p> </td> </tr> <tr> <td> <p>NM</p> </td> <td> <p>Relative mean sea level, referenced to the French vertical datum (m, IGN69)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>T</p> </td> <td> <p>Tidal water level (m), referenced to the relative mean sea level</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>S</p> </td> <td> <p>Atmospheric storm surge (m)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>Hs</p> </td> <td> <p>Significant wave height (m)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>Tp</p> </td> <td> <p>Wave peak period (s)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>Dp</p> </td> <td> <p>Wave peak direction (° in nautical convention)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>U</p> </td> <td> <p>Wind speed (m/s)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>DU</p> </td> <td> <p>Wind direction (° in nautical convention)</p> </td> <td> <p>Time series over 6h</p> </td> </tr> <tr> <td> <p>t</p> </td> <td> <p>Relative time centered on the high tide of each event (min)</p> </td> <td> <p>Not Concerned</p> </td> </tr> <tr> <td> <p>High Tide date</p> </td> <td> <p>UTC date for scenarios corresponding to past real events</p> </td> <td> <p>Not Concerned</p> </td> </tr> <tr> <td> <p><strong>OUTPUTS (Y)</strong></p> </td> </tr> <tr> <td> <p>Smax</p> </td> <td> <p>Maximum flooded area during the event (m²)</p> </td> <td> <p>Post-processed scalar output</p> </td> </tr> <tr> <td> <p>Hmax</p> </td> <td> <p>Maximum water depth reached during the event (m), provided for each inland location</p> </td> <td> <p>Post-processed functional (map) output</p> </td> </tr> <tr> <td> <p>longitude</p> </td> <td> <p>Longitude (°, WGS84)</p> </td> <td> <p>For each inland location point</p> </td> </tr> <tr> <td> <p>latitude</p> </td> <td> <p>Latitude (°, WGS84)</p> </td> <td> <p>For each inland location point</p> </td> </tr> <tr> <td> <p>XL93</p> </td> <td> <p>Longitude (m, Lambert 93)</p> </td> <td> <p>For each inland location point</p> </td> </tr> <tr> <td> <p>YL93</p> </td> <td> <p>Latitude (m, Lambert 93)</p> </td> <td> <p>For each inland location point</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p><br> </p>
Photosynthetic quotients in aquatic ecosystems: data and code supporting Trentman et al. 2023 manuscript in L&O Letters
This study provides a summary of the mismatch between our current knowledge and the application of the photosynthetic quotient (PQ). We use data from the Upper Clark Fork River (UCFR) as a case study example of how the PQ may vary in space and time based on environmental conditions. Surface water sample measurements of dissolved oxygen (DO), temperature (T), nutrients (NO3-N, NH4-N, SRP), and several metabolism indicators are represented in this data product. Figures represent data from two sites on the mainstem of the Upper Clark Fork River (UCFR) over a roughly two-year period, from 2019 to 2021. Some measurements are derived from existing data products or manuscripts, including DOT (Valett, et al., 2023); nutrients (H. M. Valett, Dec. 2, 2022, pers. comm); air pressure (Deer Lodge Weather Station, 2023); underlying data for Trentman et al. (2023) Figure 2 and Figure 4e and 4f (via Burris, 1981); and SI-Figure2 USGS gage data (USGS, 2023). Products unique to this data product include metabolism data (Trentman, et al., 2023 (Figure 5)), chamber data supporting Trentman, et al., (2023) Figure 6, and code simulations/data. All analytes and variables are documented in the project data dictionary. For details on data collection methods, see the methods section, the manuscript, and/or referenced data products.
Data in support of 'Mechanistic insights into plant community responses to environmental variables: genome size, cellular nutrient investments, and metabolic trade-offs.'
Data was collected to examine whether and how the plant genome size (GS) influences traits (stomata size, stomata density, cellular and tissue level carbon (C), nitrogen (N), and phosphorus (P) contents) and metabolic-tradeoffs (of photosynthesis, evapotranspiration, water-use, efficiency) of plants in treatment plots in which nothing, N, P, or NP had been annually added. Data was collected from ~500 plants from seven grassland sites that are all part of the Nutrient Network (https://nutnet.org), a globally distributed experiment in which plots have different nutrient amendment treatments that are administered identically to allow cross-site comparisons of the effects of nutrients on biodiversity patterning. The sites chosen varied along a North-South latitude, longitude, mean annual precipitation (MAP) and mean annual temperature (MAT) gradient.
Interagency Ecological Program: Water quality, fish, and zooplankton monitoring and modeling to support the 2018 Suisun Marsh Salinity Control Gates Summer Action
In summer 2018 we used a unique water control structure in the San Francisco Estuary (SFE) to direct a managed flow pulse into Suisun Marsh, one of the largest contiguous tidal marshes on the west coast of the United States. The action was designed to increase habitat suitability for the endangered Delta Smelt Hypomesus transpacificus, a small osmerid fish endemic to the upper SFE. The approach was to operate the Suisun Marsh Salinity Control Gates (SMSCG) in conjunction with increased Sacramento River tributary inflow to direct an estimated 160 x 10^6 m3 pulse of low salinity water into Suisun Marsh during August, a critical time period for juvenile Delta Smelt rearing. This dataset includes physical and biological monitoring data collected for the action. Datasets include Delta Smelt catch from the USFWS Enhanced Delta Smelt Monitoring program, zooplankton and Microcystis abundance from the Environmental Monitoring Program, historic Delta Smelt catch from the Summer Townet Survey, Delta Outflow from the Dayflow model, extent of Delta Smelt habitat from the UnTRIM Bay-Delta model, and water quality (Salinity, Temperature, Chlorophyll, and Turibidity) collected at continuous sondes at three locations. These data are associated with the manuscript "Evaluation of a large-scale flow manipulation to the upper San Francisco Estuary: Response of habitat conditions for an endangered native fish," by Dr. Ted Sommer, et al. 2020 PLOS One, in review.
Hubbard Brook Experimental Forest: Data in support of Territory sizes and patterns of habitat use by forest passerines over five decades: Ideal free or ideal despotic?, Zammarelli et al.
In this study, we analyzed territory sizes of seven migratory songbirds occupying a 10-hectare plot in the Hubbard Brook Experimental Forest, New Hampshire, USA over a 52-year period (1969-2021). All species varied in abundance over the duration of the study, some dramatically. Changes in territory sizes were inversely related to changes in abundance within the study plot despite differences in habitat preference, supporting the ideal free distribution. Territory sizes varied two-fold within a year across species. Results contribute to understanding how variation in territory size relates to 1) how habitat use changes with bird abundance, 2) the evolution of territory size, and 3) the role of territoriality in population dynamics. This dataset includes data, R code, and spatial files supporting this study. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station. Associated datasets in the data catalog: Holmes, R.T., N.L. Rodenhouse, and M.T. Hallworth. 2022. Bird Abundances at the Hubbard Brook Experimental Forest (1969-present) and on three replicate plots (1986-2000) in the White Mountain National Forest ver 8. Environmental Data Initiative. https://doi.org/10.6073/pasta/6422a72893616ce9020086de5a5714cd (Accessed 2023-12-17). Zammarelli, M.B. and R.T. Holmes. 2023. Hubbard Brook Experimental Forest: 10-ha bird plot territory maps, 1969 - 2021 ver 1. Environmental Data Initiative. https://doi.org/10.6073/pasta/df93595ba8df60570d472f6e6f58839e (Accessed 2024-01-11).
Data package supporting manuscript "Widespread Heterogeneity in Density-Dependent Mortality of Nearshore Fishes"
This repository contains the complete data synthesis and analysis pipeline for a global meta-analysis on density-dependent mortality in reef fishes. We estimated mortality parameters (α and β) from >30 ecological studies and explored how ecological traits, experimental methods, and phylogenetic history explain variation in density dependence. It comprises eight data tables in csv format, three .tre files for phylogenetic trees (see method document for data sources), and the zipped code folder (including 12 R scripts) to ensure transparent, end-to-end reproducibility of data processing, analysis, and visualization. This package supports the manuscript “Widespread Heterogeneity in Density-Dependent Mortality of Nearshore Fishes” by Stier & Osenberg (Ecology Letters).
Developmental change in prefrontal cortex recruitment supports the emergence of value-guided memory
Open the record for dataset details and reuse information.
Visual and auditory brain areas share a representational structure that supports emotion perception: fMRI data
Open the record for dataset details and reuse information.
Supporting data for Becker et al. PNAS (doi: 10.1073/pnas.1912921117)
<p>Ganges-Brahmaputra-Meghna delta Relative Water Level (RWL) Reconstruction v.1.1</p> <p>The paper describing the regional Relative Water Level Reconstruction (1968-2012) is Becker et al. PNAS (doi: 10.1073/pnas.1912921117)</p> <p>Each of the 6 files contains the monthly interannual mean of the relative water-level reconstruction for one of the 6 defined regions (R1 to R6) over 1968-2012.</p> <p> </p>
Computational Supporting Information for How Chemical Environment Activates Anthralin and Molecular Oxygen for Direct Reaction
<p>The updated version of the dataset contains all original computational results, including validation of the level of theory, molecular structures, and analysis spreadsheets that are in support of our experimental observations of spontaneous reactivity of anthralin/dithranol molecule with molecular oxygen without any catalyst or co-substrate.<br> The paper was published in Journal of Organic Chemistry, 2020, 85(2), 1315–1321 (DOI: 10.1021/acs.joc.9b03133).</p> <p>In the meantime, the science was also also presented at the 8th ELSI Symposium, Tokyo Institute of Technology, Tokyo (Japan); February 3-7, 2020 in the context of molecular catalysis and their role in the chemical evolution of the building blocks of life.</p> <p>This version also has an important update that is being exclusively published here on Zenodo. The selected level of theory (MN15 functional with triple-zeta quality basis set supplemented with BOTH diffuse and polarization basis functions) is further confirmed to be one of the most reasonable one among 98 commonly used functionals.</p>
Derived Data supporting "On the Seasonal Cycles of Tropical Cyclone Potential Intensity" (Gilford et al. 2017, JoC)
<p>Derived monthly mean tropical cyclone potential intensities (and associated variables) using the Bister and Emanuel 2002 PI algorithm, ftp://texmex.mit.edu/pub/emanuel/TCMAX; from MERRA2 (averaged over 1980-2016) and ERA-I data (averaged over 1980-2013), on 2.5x2.5 degree grids and with the ERA-I land-sea mask already applied. This data supported the publication of Gilford et al. (2017, JoC). When using this data, please include the citation:</p> <p>Daniel M. Gilford, Susan Solomon, and Kerry Emanuel, 2017: On the Seasonal Cycles of Tropical Cyclone Potential Intensity. <em>J. Climate, </em><strong>30</strong>, 6085–6096. doi: <a href="http://journals.ametsoc.org/doi/10.1175/JCLI-D-16-0827.1">10.1175/JCLI-D-16-0827.1</a>.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.