Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
154
datasets available to search
ShareScore release 0.9.0
Dataset results
154 results for “Cropland”
Land transformation on multi-decadal timescales reveals expanding croplands and settlements at the expense of tree-covered areas and mangroves in Nigeria
<p>A comparative assessment of the change patterns was conducted for seven categories using multi-decadal timescales in seven agroecological zones during three time-intervals (i.e., 1986 – 2000, 2000 – 2013, and 2013 – 2022). These selected periods cover important epochs in Nigeria’s recent history. To examine how much humans have appropriated natural cover (HANLC) in Nigeria over the last four decades, we differentiated natural covers (e.g., tree-covered areas, grasslands, wetlands, and waterbodies) from human activity-related uses (e.g., cropland, artificial surfaces and otherland). To identify trajectories of changes signifying human appropriation of land cover, we evaluated the drivers and processes underlying these major transitions, 1) Natural regeneration and afforestation, 2) Cropland expansion, and 3) Settlement and infrastructure development. Cropland expansion is Nigeria’s most widespread change process with much loss of croplands related to natural regeneration and settlement expansion. The transition matrix is provided showing the extent of land-cover changes in Nigeria over almost four decades (1986 - 2022). Major land cover transitions in each agroecological zone is presented. Analysis of land cover change in each agroecological zone is over 100% when areas of persistence (i.e., areas of no change) are not considered in the analysis.</p>
NECCPB-1: The first cropland parcel boundary dataset from meter-level imagery of Northeast China
<p>The Northeast China Plain is one of the world's three largest black soil regions, characterized by high organic matter content, rich nutrients, and strong water retention capabilities. Suitable climate conditions and abundant rainfall promote the growth of crops such as corn, soybeans, and rice, making it one of the main grain production bases in China, accounting for about one-fifth of the country's grain output. The grain production in the Northeast China black soil region is crucial for food security in China and globally. This area's farmland parcels are the basic units of agricultural production and the cornerstone of precision agriculture management, providing detailed information on cultivated land location, boundaries, shape, and area. Utilizing this parcel-scale information, governments and farm managers can devise more precise planting strategies and optimize management methods, thereby enhancing the quality and productivity of crops, ensuring a continuous food supply, and promoting sustainable agricultural development.</p> <p>The first cropland parcel boundary dataset from meter-level imagery of Northeast China (NECCPB-1) was developed based on deep learning models and a custom-designed automatic parcel merging strategy. A total of 10.22 TB of very-high-resolution (VHR) imagery was downloaded and uploaded, covering the entire region of Northeast China and an area of 1,240,000 km². After further removal of non-cropland regions based on phenological differences, 32,395,946 parcels were obtained.</p> <p> Rigorous validation using manually drawn reference parcels demonstrated that this dataset had high accuracy in parcel delineation (Extraction Precision, EP: 0.85) and high consistency with the reference parcels (|Completeness Deviation|, |CompD|: 0.02; Intersection over Union, IoU: 0.90). Further comparison with official Third Survey reports confirmed the high reliability of the NECCPB-1 dataset, which exhibited an average relative difference of -3.6% and an absolute relative difference of 9.8%.</p> <p>A series of cross-validations with seven widely used cropland datasets (ESA_GLC10, ESRI_GLC10, FROM_GLC10, CLD10, GLAD250, GFSAD30, and SinoLC-1). The recall, precision, and F1 scores of the NECCPB-1 were calculated as 0.91, 0.93, and 0.92, respectively, using publicly validated sample points of land cover. Moreover, NECCPB-1 performed best regarding cropland completeness, achieving an intersection ratio (IR) of 0.94, as calculated using the reference parcels.</p> <p>Due to the extensive size of the dataset and potential policy considerations, access to the data will be granted based on specific inquiries. Please contact us at zhengjia@iga.ac.cn or guotianhao@iga.ac.cn for further details. Please indicate your purpose and other details. Thank you!</p>
Supporting Data for Crawford et al. 2024, Effects of Cropland Abandonment on Biodiversity
<p><strong>This archive contains derived and supporting data products to support:</strong></p> <blockquote>Crawford CL*, Wiebe RA, Yin H, Radeloff VC, and Wilcove DS. 2024. Effects of cropland abandonment on biodiversity. <em>Nature Sustainability.</em> In press.</blockquote> <p>*Contact Christopher L. Crawford at ccrawford@alumni.princeton.edu with any questions.</p> <p>A public Zenodo archive of the Github repository containing analysis scripts developed for this project (https://github.com/chriscra/biodiversity_abandonment) can be found here: <a href="https://doi.org/10.5281/zenodo.13777205">10.5281/zenodo.13777205</a></p> <p>This analysis builds on: Crawford, C. L., Yin, H., Radeloff, V. C. & Wilcove, D. S. Rural land abandonment is too ephemeral to provide major benefits for biodiversity and climate. <em>Science Advances </em>8, 1–13 (2022). Data and scripts from Crawford et al. 2022 are archived and publicly available at Zenodo (https://doi.org/10.1126/sciadv.abm8999).</p> <p>The annual land cover maps (1987-2017, 30 meter resolution) that underlie our analysis were developed on Google Earth Engine using publicly available Landsat satellite imagery (Yin et al. 2020, Remote Sensing of Environment, https://doi.org/10.1016/j.rse.2020.111873).<br>These annual land cover maps, along with other derived data that were produced by Crawford et al. 2022, are archived and publicly available at Zenodo (https://doi.org/10.5281/zenodo.5348287).</p> <p>This archive includes important derived data products created for Crawford et al. 2024. Note that these and other project data are described in detail in **util/_util_files.R** (https://github.com/chriscra/biodiversity_abandonment). This is a convenience script that loads many of the relevant input and derived data that are used throughout the project. The primary required data files for reproducing this work are archived here, but "_util_files.R" also includes information about where additional files can be accessed (if external, e.g., https://doi.org/10.5281/zenodo.5348287) or created across the various .R and .Rmd files in this repository (e.g., "habitats.Rmd" chunk {r land-cover-of-abn-pixels}).</p> <p>Naming conventions for sites and raster files follow Crawford et al. 2022, as described here: https://doi.org/10.5281/zenodo.5348287</p> <p><strong>Site file names correspond to the following geographic locations:</strong><br>belarus = Vitebsk, Belarus / Smolensk, Russia<br>bosnia_herzegovina = Bosnia & Herzegovina<br>chongqing = Chongqing, China<br>goias = Goiás, Brazil<br>iraq = Iraq<br>mato_grosso = Mato Grosso, Brazil<br>nebraska = Nebraska / Wyoming, USA<br>orenburg = Orenburg, Russia / Uralsk, Kazakhstan<br>shaanxi = Shaanxi/Shanxi, China<br>volgograd = Volgograd, Russia<br>wisconsin = Wisconsin, USA</p> <h1><strong>This archive includes the following files:</strong></h1> <ul> <li>site_df.csv</li> <li>crop_to_abn_iucn_observed.zip</li> <li>crop_to_abn_iucn_potential.zip</li> <li>max_abn_lcc_iucn.zip</li> <li>max_abn_lcc_iucn_potential.zip</li> <li>lcc_iucn_habitat.zip</li> <li>lcc_iucn_habitat_potential.zip</li> <li>frag_df.csv</li> <li>frag_hypo_no_abn_2017_df.csv</li> <li>iucn_lc_crosswalk.csv</li> <li>habitat_age_req_coded.csv</li> <li>centroids_df.csv</li> <li>aoh_l.parquet</li> <li>aoh_feols.parquet</li> <li>aoh_start_end_l.parquet</li> <li>aoh_change_df.parquet</li> <li>aoh_est_change_tmp_all.csv</li> <li>aoh_obs_change_tmp_all.csv</li> <li>taxonomy_df.parquet</li> <li>final_species_list.csv</li> <li>trait_mod_df_modx1.rds</li> </ul> <h3>site_df.csv</h3> <p>A list of site names and related metadata describing our study sites, taken from https://zenodo.org/records/5348287</p> <h2>Derived habitat rasters:</h2> <h3>crop_to_abn_iucn_observed.zip (Calculation 1a)<br>crop_to_abn_iucn_potential.zip (Calculation 1b)<br>max_abn_lcc_iucn.zip (Calculation 2a)<br>max_abn_lcc_iucn_potential.zip (Calculation 2b)<br>lcc_iucn_habitat.zip (Calculation 3a)<br>lcc_iucn_habitat_potential.zip (Calculation 3b)</h3> <p>These maps show IUCN Level 2 habitat types (Jung et al. 2020) interpolated onto the land cover classes in the Yin et al. (2020) abandonment maps at multiple spatial and temporal extents, which serve as inputs for the three primary calculations in our manuscript. Accompanying each calculation is a corresponding map for a scenarios in which no abandoned croplands were recultivated over the course of the time series (marked as "potential"). Each .zip file contains maps for each of 11 sites.</p> <p><strong>Calculation 1. </strong>This calculation isolates the direct effect of abandonment on habitat availability, by comparing the habitat provided before and after abandonment. These "crop_to_abn_iucn" maps show IUCN Level 2 habitats in cropland pixels that experienced abandonment, including the abandonment period as well as the immediately preceding period of cultivation (to allow for a proper before and after comparison). As a result, these maps show only habitat provided by croplands when they were actively cultivated, abandoned, or, where appropriate, recultivated, which allows for a proper before and after comparison. These maps are created in the script "cluster/noncrop_precrop_mask.R".</p> <p><strong>Calculation 2.</strong> This calculation considered changes in habitat that took place exclusively in pixels that experienced abandonment at some point during the time series (following Calculation 1), but expanded to track changes across our entire time series, from 1987 through 2017, in order to account for any land cover that was cleared for agriculture prior to abandonment. These "max_abn_lcc_iucn" maps therefore show IUCN Level 2 habitat types for each pixel that was abandoned at any point during the time series, across the full time series. These maps were created in the script "habitats.Rmd" code chunks {r mask-lcc-iucn-habitat-to-abn} and {r *potential_max}. </p> <p><strong>Calculation 3. </strong>This calculation tracks habitat area provided by every pixel throughout the entire spatial and temporal extent (1987-2017), in order to place abandonment into the context of broader land-cover change dynamics like ongoing cropland expansion taking place alongside of abandonment. These "lcc_iucn" maps therefore show the IUCN Level 2 habitat types for each pixel at each site in each year of our time series. These maps were created in the script "habitats.Rmd" code chunks {r lcc-iucn-habitat-composite} and {r *potential-lcc-full} and the script "cluster/potential_full_iucn.R".</p> <p>Some analyses require these .tif files (manipulated as SpatRasters using {terra}, https://rspatial.org/terra/) to be converted to tabular format (data.tables, via {data.table} (https://rdatatable.gitlab.io/data.table/) and saved as .parquet files (via {arrow}, https://arrow.apache.org/docs/r/). This can be accomplished via scripts "cluster/save_spatraster_as_dt.R" and "cluster/save_parquet.R."</p> <h3><br>frag_df.csv<br>frag_hypo_no_abn_2017_df.csv</h3> <p>These tabular files contain derived fragmentation statistics calculated using the {landscapemetrics} R package (https://r-spatialecology.github.io/landscapemetrics/). The second file contains metrics for a scenario in which no croplands were abandoned through the year 2017, in order to assess the effect cropland abandonment on landscape configuration. Each file contains 11 columns: </p> <ol> <li>"layer" -- the spatial raster layer for which the metric is calculated, corresponding to a year.</li> <li>"level" -- the level at which the metric is calculated, in our case, the land cover "class."</li> <li>"class" -- corresponding the to land cover class for which the metric is calculated (1 = non-vegetation, 2 = woody vegetation [i.e., forest], 3 = cropland, and 4 = herbaceous vegetation [i.e., grassland]).</li> <li>"id" -- An unused field containing NA values.</li> <li>"metric" -- the specific term used for each metric by {landscapemetrics} ("area_mn", "clumpy", or "para_mn").</li> <li>"value" -- the numerical value of the statistic.</li> <li>"name" -- the name of the landscape metric being calculated ("patch area," "clumpiness index," or "perimeter-area ratio").</li> <li>"type" -- the broad type of metric being calculated ("area and edge metric," "aggregation metric," or "shape metric").</li> <li>"function_name" -- the name of the {landscapemetrics} function used to calculate the statistic.</li> <li>"site" -- the site (out of 11 study sites) for which this statistic was calculated.</li> <li>"year" -- the year corresponding to the metric statistic, between 1987-2017 (including 1986-2018 for Nebraska and 1987-2018 for Wisconsin)<br>Additional details on these metrics can be found at https://r-spatialecology.github.io/landscapemetrics/.</li> </ol> <p>The spatial IUCN data underlying our analyses (species range maps) are available upon request from BirdLife International (http://datazone.birdlife.org/species/requestdis) and IUCN (https://www.iucnredlist.org/resources/spatial-data-download). Tabular species assessment data (including habitat and elevation preferences) are freely available from IUCN (https://www.iucnredlist.org/). Here we share three IUCN-related data files that serve as important inputs throughout our analyses:</p> <h3>iucn_lc_crosswalk.csv</h3> <p>This tabular file outlines the crosswalk between the 4 land cover classes in Yin et al. 2020 and the IUCN Level 2 habitat types mapped by Jung et al. 2020. It contains five columns:</p> <ol> <li>"map_code" -- the habitat code corresponding to Jung et al. (2020).</li> <li>"Coarse_Name" -- the broad Level 1 habitat grouping.</li> <li>"lc" -- the corresponding land cover type from Yin et al. (2020) (1 = non-vegetation, 2 = woody vegetation [i.e., forest], 3 = cropland, and 4 = herbaceous vegetation [i.e., grassland]).</li> <li>"IUCNLevel" -- the full IUCN Level 2 habitat type name. </li> <li>"code" -- the IUCN Level 2 habitat code. </li> </ol> <h3>habitat_age_req_coded.csv</h3> <p>This tabular file lists whether each species was determined (by R. Alex Wiebe [AW] and Christopher L. Crawford [CLC]) to be a "mature forest obligate" (i.e., requiring forest older than 30 years, our time series length) or not. Species determined to be "mature forest obligate" species were excluded from our final analysis. The file includes 11 columns: </p> <ol> <li>"vert_class" -- Vertebrate class ("bird" or "mam" [mammal])</li> <li>"binomial" -- Species' binomial scientific name containing genus and species.</li> <li>"common_names" -- Species' common names listed by IUCN.</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"water_obl" -- Whether a species is determined to be a "water obligate" species (1) or not (0). Some species were marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty. Note: this field was not used in the analysis.</li> <li>"habitat" -- The description of the species' habitat, drawn from individual IUCN assessments (see https://www.iucnredlist.org/).</li> <li>"site_presence" -- Where each species is present across our 11 study sites.</li> <li>"suitable_habitats" -- A list of IUCN Level 2 habitat types consider suitable habitat by each species.</li> <li>"major_habitats" -- A list of IUCN Level 2 habitat types listed as having "Major Importance" for that species.</li> <li>"coder" -- The author that assigned the mature forest obligate and water obligate codes ("AW" = R. Alex Wiebe, "CLC" = Christopher L. Crawford).</li> <li>"Chris_notes" -- A text field contains notes on coding process.</li> </ol> <h3>centroids_df.csv</h3> <p>This is a simple tabular dataset containing the longitude and latitude of the centroid of each bird and mammal species' range that overlaps with one of my sites. Columns include "binomial," which lists each species binomial scientific name, "centroid_longitude," and centroid_latitude." Centroid positions were calculated in QGIS using species range files from IUCN and BirdLife International.</p> <h3><br>aoh_l.parquet</h3> <p>This tabular file contains the raw AOH results produced using the script "cluster/aoh.R." This file contains the area of each suitable IUCN Level 2 habitat for each bird and mammal species at each site in each year of our time series (1987-2017), calculated across a range of calculations and scenarios. This file includes the primary data that serve as inputs for much of the rest of the analysis. The overall area of habitat for each species in each year at each site (a tabular data file named "aoh") summed across suitable habitat types and filtered to include or exclude passage areas for migratory birds, is calculated from "aoh_l" in the "AOH.Rmd" script in code chunks "filter-aoh-suitability-by-season" and "**calculate-aoh" (similarly to other derived datasets that serve as inputs for various parts of the analysis). This "aoh" file provides input data for the linear models used to extract AOH trends and test for significance. "aoh_l.parquet" includes 20 columns: </p> <ol> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"year" -- Year for which AOH is calculated (1987-2017).</li> <li>"map_code" -- Code indicating the IUCN Level 2 habitat associated with the area statistic. See "iucn_lc_crosswalk.csv."</li> <li>"season" -- Seasonal code indicating the season in which a species considers the habitat to be suitable, drawn from IUCN. Codes are: 1 ("Resident"), 2 ("Breeding") (2), "Non-breeding Season" (3), Passage (4), and Seasonal Occurrence Uncertain (5)</li> <li>"area" -- Area of Habitat, in hectares (ha).</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"IUCN_aoh_ha" -- [Unused] A preliminary summation of all habitat area for each species in each year, prior to filtering. We did not use this field in our analysis. Our final AOH calculation involved first filtering out mismatched season and habitat suitability combinations.</li> <li>"time" -- The time required for the area of habitat calculation (in seconds).</li> <li>"className" -- Vertebrate class: "AMPHIBIA," "AVES," or "MAMMALIA."</li> <li>"category" -- Duplicate field for IUCN Red List Category, unused.</li> <li>"core_index" -- An index used to assign specific AOH calculations to run in parallel across multiple computing cores on Princeton's High-Performance Computing Cluster.</li> <li>"total_range_area" -- The species total range area, in square kilometers (km^2), calculated across all range polygons for each species provided by IUCN and BirdLife International. See "cluster/calc_range_area.R."</li> <li>"range_size_quantile" -- A numerical index representing global species range size quantiles, within each class. Values range from 0 (the smallest global range within a class) to 1 (the largest global range within a class). These quantiles are used to define "small-ranged species," as species with global range sizes smaller than the median global range size in their class. See "cluster/calc_range_area.R."</li> <li>"water_obl" -- Whether a species is determined to be a "water obligate" species (1) or not (0). Some species were marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty. Note: this field was not used in the analysis. Drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"coder" -- The author that assigned the mature forest obligate and water obligate codes ("AW" = R. Alex Wiebe, "CLC" = Christopher L. Crawford). Drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> </ol> <h3><br>aoh_feols.parquet</h3> <p>This tabular data contains the results of linear regressions predicting area of habitat as a function of time. We parameterized models for each species in each site for each of the 6 AOH calculation types described above and in Crawford et al. 2024 (Calculations 1a, 1b, 2a, 2b, 3a, and 3b). We used the R package {fixest} to parameterize these ordinary least squares (OLS) linear regressions, using the Newey-West estimator to calculate standard errors. We used the R package {broom} to extract ("tidy") the model coefficient estimates and statistics. See "AOH.Rmd" chunk {r **feols}. This file includes 20 columns:</p> <ol> <li>"term" -- The name of the regression term: "(Intercept)" or slope ("year0").</li> <li>"estimate" -- The estimated value of the regression term.</li> <li>"std.error" -- The standard error of the regression term.</li> <li>"statistic" -- The value of a T-statistic to use in a hypothesis that the regression term is non-zero.</li> <li>"p.value" -- The two-sided p-value associated with the observed statistic.</li> <li>"conf.low" -- Lower bound on the confidence interval for the estimate (in our case 5%).</li> <li>"conf.high" -- Upper bound on the confidence interval for the estimate (in our case, 95%).</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"n_obs" -- The number of observations included in the model run.</li> <li>"n_unique_obs" -- The number of unique observations included in the model run (used to exclude species with constant AOH).</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"start_year" -- The first year for which this species has area of habitat at this site (i.e., the first observation included in the model).</li> <li>"end_year" -- The last year for which this species has area of habitat at this site (i.e., the last observation included in the model).</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> </ol> <p><br><strong>Two files contain model effect sizes for AOH models:</strong></p> <h3>aoh_start_end_l.parquet</h3> <p>This tabular data file contains observed effect sizes: the observed change in AOH for each species at each site, in each calculation, derived directly from observations from the start and end of the time series. These data are calculated in "AOH.Rmd" chunk: {r observed-change-in-aoh-by-window-size}. This data serves as direct input for the file "aoh_obs_change_tmp_all" (see below), which is the primary input for the traits linear models in our analysis (see "traits.Rmd", "_util_files.R"). This file contains 24 columns:</p> <ol> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"start" -- The mean area of habitat (AOH), in hectares (ha), at the "start" of the time series, as calculated across the number of years specified in "window_size."</li> <li>"start_year" -- The year of the first AOH observation.</li> <li>"end" -- The mean area of habitat (AOH), in hectares (ha), at the "end" of the time series, as calculated across the number of years specified in "window_size."</li> <li>"end_year" -- The year of the last AOH observation.</li> <li>"window_size" -- The number of years across which "start" and "end" AOH values are averaged (e.g., if "window_size" is 5, "start" is then the mean AOH across the first 5 years of observations, and "end" is the mean AOH across the last 5 years of observations).</li> <li>"abs_change" -- The absolute change in AOH, calculated as the difference between the mean AOH at the end of the time series and the mean AOH at the start of the time series (i.e., end - start).</li> <li>"prop_change" -- The proportional change in AOH, calculated as the absolute change in AOH divided by the AOH value at the start of the time series (i.e., abs_change/start).</li> <li>"percent_change" -- The percent change in AOH, calculated as 100 times the proportional change in AOH (i.e., 100 * prop_change).</li> <li>"ratio" -- The ratio of the mean AOH at the end of the time series to the mean AOH at the start of the time series (i.e., end/start).</li> <li>"ratio_mod" -- A modified ratio of the ending AOH to the starting AOH, for which ratio values less than 1 are replaced by additive inverse of the reciprocal value (i.e., 1/ratio * -1). Ratios greater than 1 are left the same.</li> <li>"abs_change_as_prop_site_area" -- The absolute change in AOH as a proportion of site area (i.e., abs_change / total_site_area_ha_2017).</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> <li>"common_names" -- Species' common names listed by IUCN, drawn directly from "habitat_age_req_coded.csv" (see above).</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"total_site_area_ha_2017" -- The total site area (ha) in 2017. (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"area_ever_abn_ha" -- The total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017). (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"trend" -- The overall trend in AOH ("gain," "loss," or "no trend"), determined by the sign of slope coefficients and statistical significance at p < 0.05.</li> <li>"factor_change" -- The factor change in AOH, calculated as either the proportional change in AOH (i.e., prop_change) for values greater than 0, or as the reciprocal of the proportional change in AOH (i.e., 1/prop_change) for values greater than 0.</li> </ol> <h3>aoh_change_df.parquet</h3> <p>This tabular data file contains effect sizes estimated from linear regression coefficients (i.e., slopes and intercepts), calculated in "AOH.Rmd" chunk {r estimated-changes-aoh-change-df}. This file contains 29 columns:</p> <ol> <li>"run_index" -- An index used to easily pull observations for each model run. There is one index for each unique species at each site, in each of the aoh_types, calculated including and excluding passage areas.</li> <li>"est_type" -- The estimate type, whether the estimated model slope ("estimate") or the lower ("conf.low") or upper ("conf.high") bounds of the 95% confidence interval around the slope estimate.</li> <li>"vert_class" -- Vertebrate class ("amp," amphibians; "bird," birds; or "mam," mammals). Note that only birds and mammals were included in our final analysis.</li> <li>"site" -- One of our 11 study sites (see above).</li> <li>"start_year" -- The year of the first AOH observation.</li> <li>"end_year" -- The year of the last AOH observation.</li> <li>"slope" -- The model estimated slope value.</li> <li>"intercept" -- The model estimated intercept value.</li> <li>"aoh_type" -- A label indicating the temporal and spatial scale at which AOH is calculated: "crop_abn_iucn" (Calc. 1a), "crop_abn_potential_iucn" (Calc. 1b), "max_abn_iucn" (Calc. 2a), "max_potential_abn_iucn" (Calc. 2b), "full_iucn" (Calc. 3a), and "full_potential_iucn" (Calc. 3b). "abn_iucn" and "potential_abn_iucn" correspond to calculations that only capture habitat following abandonment (i.e., not including habitat provided by croplands prior to abandonment); these calculations are not included in our final analysis.</li> <li>"passage_type" -- Whether a model run includes passage areas ("include_passage") or does not include passage areas ("exclude_passage") when calculating area of habitat (AOH) for migratory birds.</li> <li>"binomial" -- Species binomial scientific name.</li> <li>"mature_forest_obl" -- Whether a species is determined to be a "mature forest obligate" species (1) or not (0), drawn directly from "habitat_age_req_coded.csv" (see above). Some species are marked as 0.9, 0.75, 0.25, or 0.1 as an indication of some uncertainty, but these were rounded to the nearest integer for the final analysis.</li> <li>"total_site_area_ha_2017" -- The total site area (ha) in 2017. (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"area_ever_abn_ha" -- The total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017). (Drawn directly from "area_summary_df," from https://zenodo.org/records/5348287)</li> <li>"trend" -- The trend in AOH experienced by the species at this site for this aoh_type calculation ("gain," "loss," or "no trend"), determined by the sign of slope coefficients and assigning statistical significance when p < 0.05.</li> <li>"n_trends" -- The number of distinct trends in AOH experienced by the species across all of the sites overlapping with its range, including this site.</li> <li>"trend_types" -- The types of trends in AOH experienced by this species across all sites overlapping with its range (some combination of "gain", "loss", and/or "no trend").</li> <li>"overall_trend" -- The overall trend in AOH experienced by this species across all sites overlapping with its range ("gain" - experiencing "gain" trends at all occurring sites; "loss" - experiencing "loss" trends at all occurring sites; "no trend" - experiencing "no trend" at all occurring sites; "weak gain" - experiencing "gain" trends at some sites and "no trend" at others; "weak loss" - experiencing "loss" trends at some sites and "no trend" at others; or "context dependent" - experienced "gain" trends at some sites and "loss" trends at other sites [referred to as "mixed" effects in Crawford et al. 2024])</li> <li>"trend_direction" -- The general direction of the trend in AOH for the species across all occurring sites ("gain" when overall_trend is either "gain" or "weak_gain"; "loss" when overall_trend is either "loss" or "weak_loss"; "context dependent" when overall_trend is "context dependent" [i.e., "mixed" effects]; and "no trend" when overall_trend is "no trend").</li> <li>"trend_consistency" -- An indication of how consistent the trend in AOH is across all occurring sites ("consistent" if overall_trend is "gain" or "loss"; "weak" if "weak_gain" or "weak_loss"; and "opposite" if "context dependent" [i.e., "mixed" effects]).</li> <li>"time_range" -- The number of years for which the species has AOH observations at this site for this aoh_type calculations.</li> <li>"aoh_start_est" -- The estimated AOH at the start of the time series, calculated from linear regression slope and intercept coefficients.</li> <li>"aoh_end_est" -- The estimated AOH at the end of the time series, calculated from linear regression slope and intercept coefficients.</li> <li>"abs_change" -- The absolute change in estimated AOH over the course of the time series (i.e., aoh_end_est - aoh_start_est).</li> <li>"abs_change_as_prop_site_area" -- The absolute change in estimated AOH as a proportion of site area (i.e., abs_change / total_site_area_ha_2017).</li> <li>"ratio_change" -- The ratio of the estimated AOH at the end of the time series to the estimated AOH at the start of the time series (i.e., aoh_end_est / aoh_start_est).</li> <li>"prop_change" -- The proportional change in estimated AOH, calculated as the absolute change in estimated AOH divided by the estimated AOH value at the start of the time series (i.e., abs_change / aoh_start_est).</li> <li>"factor_change" -- The factor change in estimated AOH, calculated as either the proportional change in estimated AOH (i.e., prop_change) for values greater than 0, or as the reciprocal of the proportional change in estimated AOH (i.e., 1/prop_change) for values greater than 0.</li> <li>"percent_change" -- The percent change in estimated AOH, calculated as 100 times the proportional change in estimated AOH (i.e., 100 * prop_change).</li> </ol> <h3>taxonomy_df.parquet</h3> <p>This tabular data file contains basic taxonomic information used in the analysis, including 10 columns:</p> <ol> <li>"vert_class" -- Vertebrate class ("bird," birds; or "mam," mammals).</li> <li>"binomial" -- Species binomial scientific name, drawn from IUCN or BirdLife International.</li> <li>"redlistCategory" -- IUCN Red List Category: "Extinct," "Extinct in the Wild," "Critically Endangered," "Endangered," "Vulnerable," "Near Threatened," "Least Concern," "Data Deficient," or "Not Evaluated."</li> <li>"order" -- Taxonomic order.</li> <li>"family" -- Taxonomic family.</li> <li>"n_sp_in_family_sample" -- The number of species contained in the family included in our analysis.</li> <li>"order_common" -- A common name to refer to the order.</li> <li>"family_common" -- A common name to refer to the family.</li> <li>"n_in_family" -- The total number of species contained in the family globally.</li> <li>"threatened" -- Whether a species is considered threatened with extinction (i.e., is listed as "Critically Endangered," "Endangered," or "Vulnerable" on the IUCN Red List).</li> </ol> <h3><br>aoh_obs_change_tmp_all.csv<br>aoh_est_change_tmp_all.csv</h3> <p>These two tabular data files contain data used as inputs for the linear models involved in our traits analysis exploring how species' responses to cropland abandonment are affected by habitat suitabilities and other traits. The key variables are the response variables for our models ("binary_gain_v_loss", "abs_change_percent_site", and "log(ratio)") and predictor variables c("forest_occ", "savanna_occ", "shrubland_occ", "grassland_occ", "wetlands_occ", "rocky_occ", "caves_occ", "desert_occ", "urban_occ", "arable_occ", "n_suitable_habitats_lvl2", "vert_class", "threatened", "Trophic_level", "log10(Body_mass_g)", "log10(total_range_area)", "abs(centroid_latitude)", and "max_abn_ext_percent_site"). Further details are contained in "traits.Rmd"</p> <p>These two files are developed from "aoh_start_end_l" and "aoh_change_df," but filtered to include only birds and mammals, to exclude passage areas from AOH calculations, to exclude mature forest obligate species, and to use only a window_size of 5 years (for "aoh_obs_change_tmp_all") and model estimates (rather than 95% confidence interval bounds, for "aoh_est_change_tmp_all"). </p> <p><strong>aoh_obs_change_tmp_all.csv contains 63 columns.</strong></p> <ul> <li>Columns 1-24 match "aoh_start_end_l". </li> <li>Columns 25-31 match "taxonomy_df" columns 3 through 10.</li> <li>Columns 32-34: "Body_mass_g" (species body mass, in grams), "Trophic_level" (whether a species is a "Carnivore", a "Herbivore," or an "Omnivore"), and "Habitat_breadth_IUCN" (the number of IUCN Level 2 habitats a species can occupy) were taken from from Etard et al. 2020 (https://doi.org/10.1111/geb.13184)</li> <li>Column 35: "total_range_area" -- drawn from "aoh_l," see above</li> <li>Columns 36-37: "centroid_longitude" and "centroid_latitude" are drawn from "centroids_df," see above.</li> <li>Columns 38-50 are Boolean variables that indicate whether a species can occupy a specific IUCN Level 1 habitat type (i.e., whether IUCN lists that Level 1 habitat as suitable for the species). These variables are as follows, with the IUCN Level 1 habitat code listed in brackets: "forest_occ" [1], "savanna_occ" [2], "shrubland_occ" [3], "grassland_occ" [4], "wetlands_occ" [5], "rocky_occ" [6], "caves_occ" [7], "desert_occ" [8], "marine_intertidal_occ"[12], "marine_coastal_occ" [13], "artificial_terrestrial_occ" [14], "artificial_aquatic_occ" [15], and "introduced_occ" [16].</li> <li>Columns 51-52 represent the number of IUCN Level 1 ("n_suitable_habitats") and IUCN Level 2 ("n_suitable_habitats_lvl2") habitats a species has listed as suitable habitats by IUCN, respectively.</li> <li>Columns 53-56 are Boolean variables indicating whether a species can occupy a subset of IUCN Level 2 habitats, which are listed in brackets: "arable_occ" [14.1 Arable Land]; "farmland_occ" [14.1 Arable Land, 14.2 Pastureland, or 14.4 Rural Gardens]; "ag_occ" (duplicate of "farmland_occ"); "urban_occ" [14.5 Urban Areas].</li> <li>Columns 57-58 represent the maximum spatial extent of abandonment at a give site (i.e., the area of all lands that were abandoned at least once during the time series), whether divided by site area ("max_abn_extent_div_site_area," i.e,. area_ever_abn_ha / total_site_area_ha_2017) or as a percent of site area ("max_abn_ext_percent_site").</li> <li>Column 59 is "abs_change_percent_site," calculated as 100 * abs_change_as_prop_site_area.</li> <li>Columns 60-63 are binary values (1 or 0) indicating the whether the species experienced statistically significant gains in AOH ("binary_trend_gain"), statistically significant losses in AOH ("binary_trend_loss"), no trend in AOH ("binary_trend_no_trend"). Column 63 ("binary_gain_v_loss") is a binary value assigning a value of 1 for gains, 0 for losses, and NA for other values.</li> </ul> <p><br><strong>aoh_est_change_tmp_all.csv contains 70 columns:</strong></p> <ul> <li>Columns 1-29 match "aoh_change_df".</li> <li>Columns 30-37 match "taxonomy_df" columns 3 through 10.</li> <li>Columns 38-40: "Body_mass_g" (species body mass, in grams), "Trophic_level" (whether a species is a "Carnivore", a "Herbivore," or an "Omnivore"), and "Habitat_breadth_IUCN" (the number of IUCN Level 2 habitats a species can occupy) were taken from from Etard et al. 2020 (https://doi.org/10.1111/geb.13184)</li> <li>Column 41: "total_range_area" -- drawn from "aoh_l," see above</li> <li>Columns 42-43: "centroid_longitude" and "centroid_latitude" are drawn from "centroids_df," see above.</li> <li>Columns 44-56 are Boolean variables that indicate whether a species can occupy a specific IUCN Level 1 habitat type (i.e., whether IUCN lists that Level 1 habitat as suitable for the species). These variables are as follows, with the IUCN Level 1 habitat code listed in brackets: "forest_occ" [1], "savanna_occ" [2], "shrubland_occ" [3], "grassland_occ" [4], "wetlands_occ" [5], "rocky_occ" [6], "caves_occ" [7], "desert_occ" [8], "marine_intertidal_occ"[12], "marine_coastal_occ" [13], "artificial_terrestrial_occ" [14], "artificial_aquatic_occ" [15], and "introduced_occ" [16].</li> <li>Columns 57-58 represent the number of IUCN Level 1 ("n_suitable_habitats") and IUCN Level 2 ("n_suitable_habitats_lvl2") habitats a species has listed as suitable habitats by IUCN, respectively.</li> <li>Columns 59-62 are Boolean variables indicating whether a species can occupy a subset of IUCN Level 2 habitats, which are listed in brackets: "arable_occ" [14.1 Arable Land]; "farmland_occ" [14.1 Arable Land, 14.2 Pastureland, or 14.4 Rural Gardens]; "ag_occ" (duplicate of "farmland_occ"); "urban_occ" [14.5 Urban Areas].</li> <li>Columns 63-64 represent the maximum spatial extent of abandonment at a give site (i.e., the area of all lands that were abandoned at least once during the time series), whether divided by site area ("max_abn_extent_div_site_area," i.e,. area_ever_abn_ha / total_site_area_ha_2017) or as a percent of site area ("max_abn_ext_percent_site").</li> <li>Columns 65-68 are binary values (1 or 0) indicating the whether the species experienced statistically significant gains in AOH ("binary_trend_gain"), statistically significant losses in AOH ("binary_trend_loss"), no trend in AOH ("binary_trend_no_trend"). Column 63 ("binary_gain_v_loss") is a binary value assigning a value of 1 for gains, 0 for losses, and NA for other values.</li> <li>Column 69, "slope_prop_site", is the estimated linear regression coefficient, or slope, as a proportion of site area, calculated as slope / total_site_area_ha_2017. </li> <li>Column 70 is "abs_change_percent_site," calculated as 100 * abs_change_as_prop_site_area.</li> </ul> <h3><br>final_species_list.csv</h3> <p>The final list of bird and mammal species included in our analysis, including the vertebrate class ("vert_class") and binomial species scientific name ("binomial") along with the overall response to cropland abandonment ("overall_trend"), the sites where that species had AOH affected by cropland abandonment ("sites"), the IUCN Red List Category ("redlistCategory"), the "obligate_type" (i.e., whether a species is a mature forest obligate, or not), the "range_size_quantile" (ranking species by global geographic range size), and "common_names". Note that mature forest obligates were excluded from our final results. These columns match the definitions included above.</p> <h3>trait_mod_df_modx1.rds</h3> <p>This R data file contains the results of our regression models run in "traits.Rmd" code chunk "*many-models", which is where our three traits linear regression models are run. These data are contained in the form of a nested tibble, or a set of tibbles nested within columns of a tibble (see: https://tidyr.tidyverse.org/articles/nest.html). These data include the input data ("data"), resulting models ("model"), model coefficients ("tidy"), regression tables ("gt"), and diagnostic statistics ("glance") for our many model runs across different response variables ("response") and AOH calculations ("aoh_type"). See "traits.Rmd" code chunk "*many-models" for more information.</p>
Indicative distribution map for Ecosystem Functional Group T7.1 Annual croplands
<p>This archive contains indicative distribution maps and profiles for <strong>T7.1 Annual croplands</strong>, a ecosystem functional group (EFG, level 3) of the <a href="https://global-ecosystems.org/">IUCN Global Ecosystem Typology</a> (v2.0). Please refer to Keith <em>et al.</em> (2020) for details.</p> <p>The descriptive profiles provide brief summaries of key ecological traits and processes, maps are indicative of global distribution patterns, and are not intended to represent fine-scale patterns. The maps show areas of the world containing major (value of 1, coloured red) or minor occurrences (value of 2, coloured yellow) of each ecosystem functional group. Minor occurrences are areas where an ecosystem functional group is scattered in patches within matrices of other ecosystem functional groups or where they occur in substantial areas, but only within a segment of a larger region. Given bounds of resolution and accuracy of source data, the maps should be used to query which EFG are likely to occur within areas, rather than which occur at particular point locations. Detailed methods and references for the maps are included in the profile (xml format).</p>
Global maps of area suitability for solar photovoltaics on rooftops and cropland
<p>This dataset consists of the following global maps:</p> <p>1. The suitable area for installing agri-photovoltaics or agrivoltaics on agricultural land cover in about 10 km x 10 km resolution for three scenarios of policy and social acceptance (optimistic, neutral and conservative)</p> <p>2. The percentage of PV-suitable building area of total area in 10km x 10km resolution</p> <p>3. The area suitability of rooftop-photovoltaics in a 10km x 10km resolution </p> <p>The maps are based on the land cover data from the Copernicus Land Monitoring Service from 2019 and the Open Street Map accessed in 2021.</p> <p>Updates from Version 1.0 :</p> <p>- New maps of agrivoltaic suitability due to updated methodology : inclusion of regional crop distribution </p>
Landsat-based dataset for mapping annual center-pivot irrigated cropland in Brazil
<p>Center-pivot irrigated cropland (CPIC) is a critical component of irrigation and plays an essential role in improving water use efficiency and increasing food production. To automatically extract the spatial distribution of CPIC in Brazil based on the remote sensing technology, we constructed a training dataset that supports the semantic segmentation models.</p><p>The dataset were built with the <a href="https://www.sciencedirect.com/topics/earth-and-planetary-sciences/landsat-5">Landsat 5</a> , 7 and 8<a href="https://www.sciencedirect.com/topics/earth-and-planetary-sciences/landsat-7"> </a> images as well as the CPIC maps from <a href="https://metadados.snirh.gov.br/geonetwork/srv/por/catalog.search#/metadata/e2d38e3f-5e62-41ad-87ab-990490841073">ANA reference</a> data. We used the Landsat images in 2005, 2010 and 2015 to build the dataset.</p><p>The samples in train_images and train_masks were used to train and valid the Convolutional Neural Network models; </p><p>The samples in valid_data were used to test the model's prediction accuracy.</p><p>Pixels with values 255 and 0 in the mask samples represent the CPIC and background categories.</p><p><strong>For technical details that used to create the dataset, please refer to </strong><i><strong>https://doi.org/</strong></i><strong>10.1016/j.isprsjprs.2023.10.007.</strong></p>
Extended Data on China's 30-m Annual Cropland Dataset for 1990–2023 (CACD-v1)
<h2>CACD 2022 and 2023 are now available!</h2> <p>The 30-m Annual Cropland Dataset of China (CACD) provides long-term, high-resolution maps of cropland extent across the country and has been widely applied in diverse studies. To meet the growing needs of the research community, we have extended the dataset to include the years 2022 and 2023, reprocessing all spatial tiles on the Google Earth Engine platform. In this updated version, minor methodological adjustments were introduced to further improve classification accuracy. For details on the mapping procedures and performance, please refer to the attached document.</p> <p> </p> <p>Data description</p> <p>*Data format: GeoTIFF (.tif)</p> <p>*Pixel size: 30 m (∼ 0.00027°)</p> <p>*Projection: EPSG: 4326 (WGS84)</p> <p>*Values: 1 denotes cropland and 0 denotes non-cropland</p> <p> </p> <p>Reference: Ying Tu, Shengbiao Wu, Bin Chen, Qihao Weng, Yuqi Bai, Jun Yang, Le Yu, and Bing Xu*. A 30 m annual cropland dataset of China from 1986 to 2021. <em>Earth System Science Data</em> 16 (2024): 2297–2316. https://doi.org/10.5194/essd-16-2297-2024</p>
Database to: Cover crops affect pool specific soil organic carbon in cropland – A meta‐analysis
<p>Database to a meta-analysis studying the effects of cover crops on the mineral-associated organic carbon pool (MAOC), the particulate organic carbon pool (POC) and the microbial biomass carbon pool (MBC). Consists of:<br>1. information on the database<br>2. legend<br>3. list of included studies, all extracted data necessary for response ratio calculation and moderator analysis, and additional information</p>
The impact of Syrian Civil War on cultivated cropland dynamics
<p>The data needed to replicate paper --'The impact of Syrian Civil War on cultivated cropland dynamics' </p>
A 1 km global cropland dataset from 10000 BCE to 2100 CE
<p>This dataset is the 1 km global cropland dataset from 10000 BCE to 2100 CE. It contains a total of 131 global cropland maps at 1 km resolution from past to future. The time-step intervals are 1000 years for 10000 BCE-1 CE, 100 years for 1 CE-1700 CE, and 10 years for 1700 CE-2100 CE. After 2010 CE, eight future SSP-RCP scenarios are provided. The map values indicate the proportion of cropland within 1×1 km grid cell.</p> <p>This dataset can also be viewed online at <a href="https://cbw.users.earthengine.app/view/globalcroplanddataset">https://cbw.users.earthengine.app/view/globalcroplanddataset</a></p> <p><strong>Citations:</strong></p> <p>When using this dataset, please cite both the dataset and the following data description article:</p> <p><em>Cao, B., Yu, L., Li, X., Chen, M., Li, X., Hao, P., and Gong, P.: A 1 km global cropland dataset from 10 000 BCE to 2100 CE, Earth Syst. Sci. Data, 13, 5403–5421, https://doi.org/10.5194/essd-13-5403-2021, 2021. </em></p>
Annual maps of cropland abandonment, land cover, and other derived data for time-series analysis of cropland abandonment
<p>This archive contains raw annual land cover maps, cropland abandonment maps, and accompanying derived data products to support:</p> <blockquote> <p>Crawford C.L., Yin, H., Radeloff, V.C., and Wilcove, D.S. 2022. Rural land abandonment is too ephemeral to provide major benefits for biodiversity and climate. <em>Science Advances</em> <a href="https://doi.org/10.1126/sciadv.abm8999">doi.org/10.1126/sciadv.abm8999</a><em>.</em></p> </blockquote> <p>An archive of the analysis scripts developed for this project can be found at: <a href="https://github.com/chriscra/abandonment_trajectories">https://github.com/chriscra/abandonment_trajectories</a> (<a href="https://doi.org/10.5281/zenodo.6383127">https://doi.org/10.5281/zenodo.6383127</a>).</p> <p>Note that the label "_2022_02_07" in many file names refers to the date of the primary analysis. "dts” or “dt” refer to “data.tables," large .csv files that were manipulated using the data.table package in R (Dowle and Srinivasan 2021, <a href="http://r-datatable.com/">http://r-datatable.com/</a>). “Rasters” refer to “.tif” files that were processed using the raster and terra packages in R (Hijmans, 2022; <a href="https://rspatial.org/terra/">https://rspatial.org/terra/</a>; <a href="https://rspatial.org/raster">https://rspatial.org/raster</a>).</p> <p>Data files fall into one of four categories of data derived during our analysis of abandonment: <strong>observed</strong>, <strong>potential</strong>, <strong>maximum</strong>, or <strong>recultivation</strong>. Derived datasets also follow the same naming convention, though are aggregated across sites. These four categories are as follows (using “age_dts” for our site in Shaanxi Province, China as an example):</p> <ol> <li><strong>observed</strong> abandonment identified through our primary analysis, with a threshold of five years. These files do not have a specific label beyond the description of the file and the date of analysis (e.g., shaanxi_age_2022_02_07.csv);</li> <li><strong>potential</strong> abandonment for a scenario without any recultivation, in which abandoned croplands are left abandoned from the year of initial abandonment through the end of the time series, with the label “_potential” (e.g., shaanxi_potential_age_2022_02_07.csv);</li> <li><strong>maximum</strong> age of abandonment over the course of the time series, with the label “_max” (e.g., shaanxi_max_age_2022_02_07.csv);</li> <li><strong>recultivation </strong>periods, corresponding to the lengths of recultivation periods following abandonment, given the label “_recult” (e.g., shaanxi_recult_age_2022_02_07.csv).</li> </ol> <p> </p> <p><strong>This archive includes multiple .zip files, the contents of which are described below:</strong></p> <ul> <li><strong>age_dts.zip</strong> - Maps of abandonment age (i.e., how long each pixel has been abandoned for, as of that year, also referred to as length, duration, etc.), for each year between 1987-2017 for all 11 sites. These maps are stored as .csv files, where each row is a pixel, the first two columns refer to the x and y coordinates (in terms of longitude and latitude), and subsequent columns contain the abandonment age values for an individual year (where years are labeled with "y" followed by the year, e.g., "y1987"). Maps are given with a latitude and longitude coordinate reference system. Folder contains observed age, potential age (“_potential”), maximum age (“_max”), and recultivation lengths (“_recult”) for all sites. Maximum age .csv files include only three columns: x, y, and the maximum length (i.e., “max age”, in years) for each pixel throughout the entire time series (1987-2017). Files were produced using the custom functions "cc_filter_abn_dt()," “cc_calc_max_age()," “cc_calc_potential_age(),” and “cc_calc_recult_age();” see "_util/_util_functions.R."</li> <li><strong>age_rasters.zip</strong> - Maps of abandonment age (i.e., how long each pixel has been abandoned for), for each year between 1987-2017 for all 11 sites. Maps are stored as .tif files, where each band corresponds to one of the 31 years in our analysis (1987-2017), in ascending order (i.e., the first layer is 1987 and the 31st layer is 2017). Folder contains observed age, potential age (“_potential”), and maximum age (“_max”) rasters for all sites. Maximum age rasters include just one band (“layer”). These rasters match the corresponding .csv files contained in "age_dts.zip.”</li> <li><strong>derived_data.zip</strong> - summary datasets created throughout this analysis, listed below.</li> <li><strong>diff.zip</strong> - .csv files for each of our eleven sites containing the year-to-year lagged differences in abandonment age (i.e., length of time abandoned) for each pixel. The rows correspond to a single pixel of land, and the columns refer to the year the difference is in reference to. These rows do not have longitude or latitude values associated with them; however, rows correspond to the same rows in the .csv files in "input_data.tables.zip" and "age_dts.zip." These files were produced using the custom function "cc_diff_dt()" (much like the base R function "diff()"), contained within the custom function "cc_filter_abn_dt()" (see "_util/_util_functions.R"). Folder contains diff files for observed abandonment, potential abandonment (“_potential”), and recultivation lengths (“_recult”) for all sites.</li> <li><strong>input_dts.zip</strong> - annual land cover maps for eleven sites with four land cover classes (see below), adapted from Yin et al. 2020 <em>Remote Sensing of Environment </em>(<a href="https://doi.org/10.1016/j.rse.2020.111873">https://doi.org/10.1016/j.rse.2020.111873</a>)<em>. </em>Like “age_dts,” these maps are stored as .csv files, where each row is a pixel and the first two columns refer to x and y coordinates (in terms of longitude and latitude). Subsequent columns contain the land cover class for an individual year (e.g., "y1987"). Note that these maps were recoded from Yin et al. 2020 so that land cover classification was consistent across sites (see below). This contains two files for each site: the raw land cover maps from Yin et al. 2020 (after recoding), and a “clean” version produced by applying 5- and 8-year temporal filters to the raw input (see custom function “cc_temporal_filter_lc(),” in “_util/_util_functions.R” and “1_prep_r_to_dt.R”). These files correspond to those in "input_rasters.zip," and serve as the primary inputs for the analysis.</li> <li><strong>input_rasters.zip</strong> - annual land cover maps for eleven sites with four land cover classes (see below), adapted from Yin et al. 2020 <em>Remote Sensing of Environment. </em>Maps are stored as ".tif" files, where each band corresponds one of the 31 years in our analysis (1987-2017), in ascending order (i.e., the first layer is 1987 and the 31st layer is 2017). Maps are given with a latitude and longitude coordinate reference system. Note that these maps were recoded so that land cover classes matched across sites (see below). Contains two files for each site: the raw land cover maps (after recoding), and a “clean” version that has been processed with 5- and 8-year temporal filters (see above). These files match those in "input_dts.zip."</li> <li><strong>length.zip</strong> - .csv files containing the length (i.e., age or duration, in years) of each distinct individual period of abandonment at each site. This folder contains length files for observed and potential abandonment, as well as recultivation lengths. Produced using the custom function "cc_filter_abn_dt()" and “cc_extract_length();” see "_util/_util_functions.R."</li> </ul> <p><strong>derived_data.zip</strong> contains the following files:</p> <ul> <li>"<strong>site_df.csv</strong>" - a simple .csv containing descriptive information for each of our eleven sites, along with the original land cover codes used by Yin et al. 2020 (updated so that all eleven sites in how land cover classes were coded; see below).</li> <li><strong>Primary derived datasets </strong>for both observed abandonment (“area_dat”) and potential abandonment (“potential_area_dat”). <ul> <li><strong>area_dat</strong> - Shows the area (in ha) in each land cover class at each site in each year (1987-2017), along with the area of cropland abandoned in each year following a five-year abandonment threshold (abandoned for >=5 years) or no threshold (abandoned for >=1 years). Produced using custom functions "cc_calc_area_per_lc_abn()" via "cc_summarize_abn_dts()". See scripts "cluster/2_analyze_abn.R" and "_util/_util_functions.R."</li> <li><strong>persistence_dat</strong> - A .csv containing the area of cropland abandoned (ha) for a given "cohort" of abandoned cropland (i.e., a group of cropland abandoned in the same year, also called "year_abn") in a specific year. This area is also given as a proportion of the initial area abandoned in each cohort, or the area of each cohort when it was first classified as abandoned at year 5 ("initial_area_abn"). The "age" is given as the number of years since a given cohort of abandoned cropland was last actively cultivated, and "time" is marked relative to the 5th year, when our five-year definition first classifies that land as abandoned (and where the proportion of abandoned land remaining abandoned is 1). Produced using custom functions "cc_calc_persistence()" via "cc_summarize_abn_dts()". See scripts "cluster/2_analyze_abn.R" and "_util/_util_functions.R." This serves as the main input for our linear models of recultivation (“decay”) trajectories.</li> <li><strong>turnover_dat</strong> - A .csv showing the annual gross gain, annual gross loss, and annual net change in the area (in ha) of abandoned cropland at each site in each year of the time series. Produced using custom functions "cc_calc_abn_diff()" via "cc_summarize_abn_dts()" (see "_util/_util_functions.R"), implemented in "cluster/2_analyze_abn.R." This file is only produced for observed abandonment.</li> </ul> </li> <li><strong>Area summary files </strong>(for observed abandonment only) <ul> <li><strong>area_summary_df</strong> - Contains a range of summary values relating to the area of cropland abandonment for each of our eleven sites. All area values are given in hectares (ha) unless stated otherwise. It contains 16 variables as columns, including 1) "site," 2) "total_site_area_ha_2017" - the total site area (ha) in 2017, 3) "cropland_area_1987" - the area in cropland in 1987 (ha), 4) "area_abn_ha_2017" - the area of cropland abandoned as of 2017 (ha), 5) "area_ever_abn_ha" - the total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017), 6) "total_crop_extent_ha" - the total area of those pixels that were classified as cropland at least once during the time series, 7) "total_area_abn_remaining_2017" - duplicate of "area_abn_ha_2017," the area abandoned as of 2017 (ha), taken from "area_recult_threshold," 8) "total_initial_area_abn" - the sum of the initial area of each cohort of abandonment when it is first classified as "abandoned," i.e., at the 5 year mark (note that this is cumulative, and because it counts those pixels that were abandoned more than once, it is therefore larger than "area_ever_abn_ha"), taken from "area_recult_threshold" 9) "total_area_abn_recultivated_2017" - the area of abandoned land that was recultivated as of 2017 (cumulatively, i.e., "total_initial_area_abn" - "area_abn_ha_2017"), taken from "area_recult_threshold," 10) "proportion_recultivated" - the proportion of all abandoned cropland (including multiple periods per pixel) that was recultivated by 2017, taken from "area_recult_threshold," 11) "area_2017_as_prop_site" - area abandoned as of 2017 as a proportion of the total site area, 12) "area_2017_as_prop_total_crop" - area abandoned as of 2017 as a proportion of the total crop extent, 13) "area_2017_as_prop_crop87" - area abandoned as of 2017 as a proportion of cropland area in 1987, 14) "area_ever_abn_as_prop_site" - area ever abandoned as a proportion of the total site area, 15) "area_ever_abn_as_prop_total_crop" - area ever abandoned as a proportion of the total crop extent, 16) "area_ever_abn_as_prop_crop87" - area ever abandoned as a proportion of cropland area in 1987. See script "1_summary_stats.Rmd."</li> <li><strong>area_recult_threshold</strong> - Contains data on the proportion of observed abandoned cropland area that is recultivated by the end of our time series. This includes the area of abandoned cropland as of 2017 ("total_area_abn_remaining_2017") and the sum of the initial area of each cohort of abandonment when it is first classified as abandoned (at year 5; "total_initial_area_abn"). This "total_initial_area_abn" is cumulative, and allows for pixels that were abandoned multiple times during the time series to be counted multiple times. The difference between these two columns yields the "total_area_abn_recultivated_2017," which in turn is used to calculate the "proportion_recultivated," and the (ascending) "order" of sites based on this proportion. This file includes recultivation stats for each site for three abandonment definitions: 5, 7, and 10 years. See script "1_summary_stats.Rmd."</li> <li><strong>abn_lc_area_2017</strong> - Contains the number of pixels and corresponding area (in ha) of abandoned cropland in the year 2017 at each site, according to the land cover class (either woody vegetation [2], or herbaceous vegetation [4]) and the age in 2017 (5 to 30 years). See script "cluster/6_lc_of_abn.R."</li> <li><strong>abn_prop_lc_2017 </strong>- Contains the number of pixels and corresponding area (ha) of cropland abandoned in the year 2017 in each land cover type (woody vegetation [2], or herbaceous vegetation [4]). It also shows this area as a proportion of the total area abandoned at each site (i.e., in either land cover class: 2 or 4). See script "cluster/6_lc_of_abn.R."</li> </ul> </li> <li><strong>Carbon</strong> <ul> <li><strong>carbon_df </strong>– contains the observed and potential carbon accumulation in abandoned croplands in each site in each year (in Mg C), for two abandonment thresholds: 5 years (our default abandonment definition) and 1 year (i.e., no threshold). Each data point corresponds to one of two scenarios (“type” column), either “observed” or “potential.” Carbon accumulation figures are for both the sum of forest and soil carbon at each site in a given year. Carbon accumulation is listed in three columns: 1) “C_up_to_20” contains the total carbon accumulated in those abandoned croplands with abandonment durations between 5 and 20 years. 2) “C_21_30” contains the total carbon accumulation in croplands with durations between 21 and 30 years, which are differentiated in order to account for non-linear carbon accumulation rates in soils over time, and 3) “total_C_Mg” contains the sum of the previous two columns, representing the total carbon accumulated across all abandoned croplands in each year.</li> <li><strong>soc_mean</strong> – contains mean soil organic carbon accumulation rates for years 1-20 and years 21-80, derived from Sanderman et al. 2020 (in Mg C; <a href="https://doi.org/10.7910/DVN/HA17D3">https://doi.org/10.7910/DVN/HA17D3</a>). These values correspond to accumulation rates in croplands upon abandonment and regeneration to natural vegetation (Sanderman et al. 2020’s “rewilding” scenario). These mean values are calculated across those pixels identified as cropland by Sanderman et al. 2020 at each site. Mean values in year 20 and 80 are contained in columns “mean_soc_20” and “mean_soc_80” respectively, and the annualized rate over the first 20 years and the subsequent years 21 through 80 are contained in columns “mean_annual_soc_1_20” and “mean_annual_soc_21_80” respectively.</li> </ul> </li> <li><strong>Decay model data</strong> – two R data files containing data products for our linear models of abandonment recultivation trajectories. <ul> <li><strong>decay_endpoints_files</strong> – an R data file (.rds) containing seven data products produced as part of our common endpoint analysis, which calculated mean trajectories for each site across a range of common endpoints, ensuring that means were based on coefficient estimates derived from a consistent number of observations for each cohort. These files are: <ul> <li><strong>common_endpoint_dat – </strong>a .csv containing subsets of “persistence_dat” for each “endpoint” (7 through 29).</li> <li><strong>endpoint_n – </strong>a .csv describing, for each endpoint, the corresponding number of observations per cohort (“n_obs”), the number of cohorts (“n_cohorts”), the total number of observations across cohorts included (“total_obs”), and the cohorts that meet the endpoint threshold (“cohorts”).</li> <li><strong>coef_l3_endpoints – </strong>corresponding model coefficients for our primary model (“l3”) parameterized by the range of subsets across endpoints.</li> <li><strong>augment_endpoints – </strong>fitted values (i.e., model predictions) for linear models produced across the full range of endpoint subsets.</li> <li><strong>fitted_endpoints – </strong>a simplified .csv containing the mean linear and log coefficients for each site at each endpoint, and the corresponding predicted proportion remaining abandoned through time (based on the “age,” or duration, of abandonment).</li> <li><strong>time_to_endpoints – </strong>a .csv containing, for mean trajectories for each endpoint at each site, the estimated time required for a given amount of abandoned cropland in a cohort to be recultivated (deciles, 10% through 100%).</li> <li><strong>endpoint_half_lives – </strong>a .csv containing the half-lives calculated for the mean trajectories for each endpoint at each site.</li> </ul> </li> <li><strong>decay_mod_archive</strong> - an R data file (.rds) containing eleven data products derived from linear models of abandonment recultivation ("decay"): <ul> <li><strong>lm_mega_lin_log_lin_l</strong> – the primary linear model produced in our analysis. This model is referred to as “lin_log_lin” (or “l3”) because the model predicts linear persistence (“lin”) as a function of a log term of time (“log”) and a linear term of time (“lin”). “mega” refers to the fact that this model is run for the full dataset, pooled across all 11 sites.</li> <li><strong>coef_l3_mega</strong> – a .csv containing model coefficients for our primary linear model of recultivation (“lin_log_lin”, or “l3”), with a single row each for the linear term of time and the log term of time, for 26 cohorts at 11 sites.</li> <li><strong>mean_coef_l3_mega</strong> – a data frame containing the mean coefficient values for the log and linear terms of time across cohorts at each site. This also contains the mean of the low and high coefficient estimates, based on the 95% confidence interval.</li> <li><strong>half_lives_all_cohorts_l3</strong> – half-lives calculated for each cohort at each site, for our primary model.</li> <li><strong>half_life_mean_coefs_l3</strong> – half-lives calculated based on the mean trajectory for each site (based on the mean log coefficients and mean linear coefficients across all cohorts), for our primary model.</li> <li><strong>mod_AIC_mega</strong> – Akaike Information Criterion (AIC) values for all tested model specifications.</li> <li><strong>fitted_combo</strong> – fitted values (i.e., model predictions) for our primary model (“l3”) and a series of alternative model specifications (“l3_trim” – excluding cohorts with fewer than 5 observations; “lin_log” – a model including only one log time term; “log2_lin” – in which the log of persistence is predicted by log and linear time terms; and “l3_no_cohort” – our primary model, predicting linear persistence as a function of log time and linear time, but without cohort-level fixed effects).</li> <li><strong>time_to_combo</strong> – contains the estimated time required for a certain amount of abandoned cropland in a cohort to be recultivated (deciles, 10% through 100%). See script "2_decay_models.Rmd." These values are calculated for a range of alternative model specifications ("l3_trim", “lin_log”, "log2_lin", and "l3_no_cohort"; see above).</li> </ul> </li> </ul> </li> <li><strong>Length data</strong> – includes “_distill_df” files and “mean_length_df” files for observed, potential, and recultivation. <ul> <li><strong>length_distill_df</strong> - .csvs containing the number ("freq") of abandonment periods of a specific "length" of time (i.e., age) at each site over the course of the entire time series. Derived from the "length" files in "length.zip." See script "cluster/5_distill_lengths.R."</li> <li><strong>mean_length_df</strong> - .csvs with the mean, median, and standard deviation, for each site, for both "all" lengths or just the "max" length per pixel, and for a range of abandonment definitions (1, 3, 5, 7, and 10 years). Derived from "length_distill_df." See script "1_summary_stats.Rmd."</li> </ul> </li> <li><strong>Duration summary files</strong> – includes “summary_stats_all_sites” and “summary_stats_all_sites_pooled,” for observed and potential abandonment, and recultivation periods following abandonment. <ul> <li><strong>“summary_stats_all_sites”</strong> - A simple .csv derived from "mean_length_df" files containing summary stats across the 11 sites. This includes the mean of the mean abandonment duration ("length", in years) for each of our 11 sites ("mean_of_means"), the standard deviation of these site mean abandonment lengths ("sd_of_means"), the mean of the standard deviation at each site ("mean_of_sds"), the mean median ("mean_of_medians"), and the mean number of abandonment periods ("mean_n_abn_periods"). Note that length "all" indicates that these stats account for all periods (including multiple per pixel), rather than just the max duration per pixel. See script "1_summary_stats.Rmd."</li> <li><strong>“summary_stats_all_sites_pooled”</strong> - A summary .csv similar to "summary_stats_all_sites," but calculated by pooling all distinct periods of abandonment across all eleven sites, and then calculating the mean, median, and standard deviation of abandonment duration. See script "1_summary_stats.Rmd."</li> </ul> </li> <li><strong>Comparing annual approach to identifying abandonment to a two-timepoint (“2yr”) approach:</strong> <ul> <li><strong>abn_2yr_ages_df</strong> - Contains the age of former croplands identified as "abandoned" using a two-timepoint method (i.e., 2017 - 1987), where age values (as of 2017) are derived from our map of abandonment identified using the full annual time series. This includes the area in hectares (ha), in each age class (along with the number of pixels), at each of our 11 sites. This dataset is used to calculate the percent of cropland "abandonment" identified using the two-year method that is actually too "young," i.e., less than 5 years old, and therefore not truly abandonment according to our five-year abandonment definition</li> <li><strong>abn_2yr_overestimation</strong> - Compares the area (in hectares) of cropland abandonment at each site identified with our full annual time series (and a five-year abandonment definition) and the "abandonment" identified using a two-timepoint method (2017-1987). This also includes the percent difference in area between the two methods, the Jaccard similarity of the areas identified as abandonment, and the percent of "young" (i.e., <5-year-old) "abandonment" identified by the two-timepoint method.</li> </ul> </li> </ul> <p><strong>Input land cover maps:</strong></p> <p>As noted, the file "input_rasters.zip" contain the raw annual land cover maps for eleven sites generated by:</p> <blockquote> <p>Yin, H., A. Brandão, J. Buchner, D. Helmers, B. G. Iuliano, N. E. Kimambo, K. E. Lewińska, E. Razenkova, A. Rizayeva, N. Rogova, S. A. Spawn, Y. Xie, and V. C. Radeloff. 2020. Monitoring cropland abandonment with Landsat time series. <em>Remote Sensing of Environment</em> 246:111873. https://doi.org/10.1016/j.rse.2020.111873</p> </blockquote> <p>These land cover maps served as raw inputs for this project and form the basis of the analysis.</p> <p>All land cover maps have a resolution of 30-m and exist for each year from 1987 through 2017. The exceptions are Nebraska / Wyoming (1986-2018) and Wisconsin (1987-2018); these additional years were excluded from our analysis of abandonment duration.</p> <p><strong>Land cover categories in these maps are coded as follows:</strong></p> <ol> <li>Non-vegetated area (e.g., water, urban, barren land)</li> <li>Woody vegetation (e.g., forests)</li> <li>Cropland</li> <li>Herbaceous vegetation (e.g., grassland)</li> </ol> <p><strong>Site file names correspond to the following geographic locations:</strong></p> <ul> <li>belarus = Vitebsk, Belarus / Smolensk, Russia</li> <li>bosnia_herzegovina = Bosnia & Herzegovina</li> <li>chongqing = Chongqing, China</li> <li>goias = Goiás, Brazil</li> <li>iraq = Iraq</li> <li>mato_grosso = Mato Grosso, Brazil</li> <li>nebraska = Nebraska / Wyoming, USA</li> <li>orenburg = Orenburg, Russia / Uralsk, Kazakhstan</li> <li>shaanxi = Shaanxi/Shanxi, China</li> <li>volgograd = Volgograd, Russia</li> <li>wisconsin = Wisconsin, USA</li> </ul> <p>This dataset is minimally altered from Yin et al. 2020. However, land cover codes were updated for five sites (Iraq, Nebraska/Wyoming, Orenburg/Uralsk, Volgograd, and Wisconsin) in order to maintain consistency in how land cover was coded across all sites. The original land cover codes (matching Yin et al. 2020) are described in the file "site_df.csv" and are as follows:</p> <ol> <li>Iraq: 1 Non-vegetated; 2 Cropland; 3 Woody; 4 Herbaceous</li> <li>Nebraska / Wyoming (USA): 1 Cropland; 2 Woody; 3 Non-vegetated; 4 Herbaceous</li> <li>Orenburg, Russia / Uralsk, Kazakhstan: 1 Non-vegetated; 2 Cropland; 3 Herbaceous; 4 Woody</li> <li>Volgograd (Russia): 1 Non-vegetated; 2 Cropland; 3 Herbaceous; 4 Woody</li> <li>Wisconsin (USA): 1 Cropland; 2 Herbaceous; 3 Woody; 4 Non-vegetated</li> </ol>
"Agricultural trade and its impacts on cropland use and the global loss of species habitat." - Supplementary data
<p>This dataset and code is part of the following publication:<br> Schwarzmueller, F. & Kastner, T (2022), Agricultural trade and its impact on cropland use<br> and the global loss of species' habitats. Sustainability Science, doi: 10.1007/s11625-022-01138-7<br> </p> <p>There are three zip-folders accompanying this publication:</p> <p>Code.zip contains all the R-Scripts and input files neccessary for the calculation that were written by the authors.</p> <p>Data.zip contains the FAO-input data (as dowloaded in 2021). This exact data is not available anymore from the FAOSTAT website, which is why we included it in this repository.</p> <p>TradeMatrixFeed_import_dry_matter_1986-2013.zip contains the results from the calculation as shown in the paper.</p>
Location, biophysical and agronomic parameters for croplands in Northern Ghana
<p>We present a dataset describing (i) crop locations, (ii) biophysical parameters and (iii) crop yield and biomass was collected in 2020 and 2021 in Ghana, mostly focusing on maize in northern Ghana. The dataset contains repeated multiple measurements of leaf area index (LAI), leaf chlorophyll concentration over a large number of maize fields, as well as associated grain yield, biomass and polygons that delineate the fields.</p>
A map of active cropland and short-term fallows across Northern Mozambique derived from PlanetScope data
<p><strong>Overview</strong></p> <p>A map of smallholder-dominated landscapes covering the provinces Niassa, Zambezia, Cabo Delgado, and Nampula in Northern Mozambique. The map includes active cropland and short-term fallows as separate classes, as well as five land cover classes (herbaceous vegetation, open woodlands, closed woodlands, non-vegetated land, water). The map is based on PlanetScope mosaics and consequently comes at 4.77m spatial resolution.</p> <p>The download contains the following files:</p> <ul> <li>ps_lc_nmoz.tif / .qml: land cover map and associated QGIS style file</li> <li>ps_lc_nmoz_probmargins.tif / .qml: probability margins and associated QGIS style file</li> <li>training.gpkg: training samples with class labels</li> <li>LICENSE.pdf: NICFI data program user license</li> </ul> <p><strong>Map accuracy</strong></p> <p>We conducted an area-adjusted accuracy assessment based on a stratified random sample, which yielded important insights regarding accuracies and error types. The area-adjusted overall accuracy of the map is 88.9%, but users should be aware of the most important error types:</p> <ul> <li>Active cropland were overestimated, whereas local topographical depressions with moist soils, and regions with exposed soils/rocks and sparse vegetation cover were found to be falsely classified.</li> <li>Short-term fallows were underestimated, particularly in regions with high growth rates and extensive land management, such as parts of the northern and north-eastern study region.</li> </ul> <p><strong>Further resources</strong></p> <p>The production of this map was made possible through the <a href="https://www.planet.com/nicfi/">NICFI data program</a>, providing the PlanetScope mosaics and the Google Earth Engine cloud computing platform for preprocessing of the satellite data and classification. As such, the use of the map falls under the <a href="https://assets.planet.com/docs/Planet_ParticipantLicenseAgreement_NICFI.pdf">NICFI data program license agreement</a> included in the download. The code for preprocessing the PlanetScope mosaics is based on the Google Earth Engine Python API and made available at <a href="https://github.com/philipperufin/eepypr/">https://github.com/philipperufin/eepypr/</a>.</p> <p>We advise map users to read the <a href="https://eartharxiv.org/repository/view/3174/">preprint</a> or the <a href="https://doi.org/10.1016/j.jag.2022.102937">open access paper</a> for detailed insights. In case of questions please consult these resources or contact the lead author of the work.</p>
Soil Carbon Dynamics in Soybean Cropland and Forests in Mato Grosso, Brazil
<p>These files contain the carbon content, radiocarbon, and stable isotope data for soils collected to 2 m deep in forest and soybean cropland in Mato Grosso, Brazil. </p>
Global cropland extent (fractions) annual 2000-2022 at 250 m and 1 km
<p>Global cropland extent annual for 2000-2022 based on the <a href="https://glad.umd.edu/dataset/croplands">Potapov et al. (2021)</a>. Cropland defined as: land used for annual and perennial herbaceous crops for human consumption, forage (including hay), and biofuel. Perennial woody crops, permanent pastures, and shifting cultivation are excluded from the definition. The original 30-m resolution data (0/1 values) was interpolated from time-series 2003, 2007, 2011, 2015, 2019 to annual values 2000 to 2022 using linear interpolation. All values shown are in principle fractions 0-100%. The 30-m and 100-m resoluton images (COGs) are too large for Zenodo but you can access them from URLs in the filenames_openlandmap_cropland.txt file. See for example (drop the URL in QGIS):</p> <ul> <li>https://s3.eu-central-1.wasabisys.com/openlandmap/layers30m/cropland_glad.potapov.et.al_p_30m_s_20030101_20031231_go_epsg.4326_v20240624.tif (3.2GB)</li> <li>https://s3.eu-central-1.wasabisys.com/openlandmap/layers100m/cropland_glad.potapov.et.al_p_100m_s_20030101_20031231_go_epsg.4326_v20240624.tif (2.3GB)</li> </ul> <p><strong>Disclaimer</strong>: linear interpolation has limited accuracy and is basically only used to gap-fill the missing years. The remaining missing values in the maps can be ALL consider to be 0 value for cropland. A more detailed up-to-date cropland map of the world is provided by <a href="https://doi.org/10.5194/essd-15-5491-2023">van Tricht et al., (2023)</a>, however only single year (2021) has been mapped at 10-m resolution within the <a href="../doi/10.5281/zenodo.7875104">WorldCereal project</a>.</p> <p>The temporal interpolation was implemented using terra package ii.e. using the following fuction:</p> <pre><code>library(terra) y.l = c(2003, 2007, 2011, 2015, 2019) out.years = 2000:2022 i = parallel::mclapply(y.l, function(x){system(paste0('gdal_translate Global_cropland_', x, '.vrt Global_cropland_', x, '.tif -co TILED=YES -co BIGTIFF=YES -co COMPRESS=DEFLATE -co ZLEVEL=9 -co BLOCKXSIZE=1024 -co BLOCKYSIZE=1024 -co NUM_THREADS=8 -co SPARSE_OK=TRUE -a_nodata 255 -scale 0 1 0 100 -ot Byte'))}, mc.cores = length(y.l)) ## land mask at 1 deg (100x100km) ---- x = parallel::mclapply(y.l, function(x){system(paste0("gdal_translate Global_cropland_", x, ".tif Global_cropland_", x, "_1d.tif -tr 1 1 -r average -co BIGTIFF=YES -ot Byte -co NUM_THREADS=10"))}, mc.cores = length(y.l)) ## 2 hrs g1 = terra::rast(paste0("Global_cropland_", y.l, "_1d.tif")) gs = sum(g1, na.rm=TRUE) plot(gs) gs.p <- as.polygons(gs, values = TRUE, extent=FALSE, dissolve=FALSE, na.rm=TRUE) ## Input layers: r = terra::rast(paste0("Global_cropland_", y.l, ".tif")) int.mc = function(r, tile, y.l, out.years=2000:2022){ bb = paste(as.vector(ext(tile)), collapse = ".") if(any(!file.exists(paste0("./tmp/", out.years, "/Global_cropland_", out.years, "_", bb, ".tif")))){ r.t = terra::crop(r, ext(tile)) ## each tile is 16M pixels r.x = as.data.frame(r.t, xy=TRUE, na.rm=FALSE) rs = rowSums(r.x[,-c(1:2)], na.rm=TRUE) ## if sum is == 0 means no cropland throughout the time-series sel = which(rs>0) ## extract complete values: r.x0 = r.x[sel,-c(1:2)] r.x0[is.na(r.x0)] = 0 ## interpolate between values: t1s = as.data.frame(t(apply(r.x0, 1, function(y){ try( approx(y.l, as.vector(y), xout=out.years, rule=2)$y ) }))) ## write to GeoTIFFs t1s$x <- r.x$x[sel]; t1s$y <- r.x$y[sel] ## convert to RasterLayer: r.x = rast(t1s[,c("x","y",paste0("V", 1:length(out.years)))], type="xyz", crs="+proj=longlat +datum=WGS84 +no_defs") for(j in 1:length(out.years)){ writeRaster(r.x[[j]], filename=paste0("./tmp/", out.years[j], "/Global_cropland_", out.years[j], "_", bb, ".tif"), gdal=c("COMPRESS=DEFLATE"), datatype='INT1U', NAflag=0, overwrite=FALSE) } } } ## test it: #int.mc(r, tile=gs.p[1000], y.l) ## run in parallel ---- ## takes 12 hrs... 1TB RAM i = parallel::mclapply(sample(1:length(gs.p)), function(x){try( int.mc(r, tile=gs.p[x], y.l) )}, mc.cores = 70) </code></pre>
High resolution cropland agreement map (30 m) circa 2020
<p>Accurate and precise measurements of global cropland extent are needed for monitoring the sustainability of agriculture at all scales. Recent advancement in remote sensing and land cover mapping methods have greatly increased the ability to estimate cropland area distribution and trends. Here the FAO presents a map of cropland agreement produced by consolidating information at pixel level from six high-resolutions maps for <em>circa </em>2020. The following six high resolution layers were used: ESRI 10 meter LU/LC, FROM-GLC, GLAD, GLC-FCS30, Globeland30 and Worldcover.</p> <p>Two bands are included in the dataset:</p> <ol> <li>Simple agreement (values between 1 and 6)</li> <li>Detailed agreement (values between 1 and 63)</li> </ol> <p>The map, developed in the Google Earth Engine platform, combines the 6 land cover/cropland layers to show their cropland agreement on pixel level at a spatial resolution of 30 meters. The simple agreement has pixel values that range from 1 (only 1 dataset classifies as cropland) to 6 (all datasets agree on presence of cropland). Pixels with a value of 0 indicate pixels where all datasets agree on absence of cropland. The second band includes a detailed agreement, showing which combination of the 6 datasets classify a pixel as cropland. The overview table (<em>DetailedAgreement_LookupTable.xlsx</em>) shows what the pixel values of this detailed agreement (from 1 to 63) correspond to.</p> <p>The dataset has been uploaded in 16 tiles, in the preview below and in the file "A<em>CroplandAgreement_30m_Tiles.png</em>" the extent of each tile can be found.</p> <p>For more information on FAO statistics on land cover and land use:</p> <p>FAO. 2022. <em>Land use statistics and indicators. Global, regional and country trends, 2000–2020</em>. FAOSTAT Analytical Brief, no. 48. Rome. <a href="https://doi.org/10.4060/cc0963en">https://doi.org/10.4060/cc0963en</a></p> <p>FAO. 2021. <em>Land cover statistics. Global, regional and country trends, 2000–2019</em>. FAOSTAT Analytical Brief Series No. 37. Rome. </p>
Refined Cropland Data Layer (R-CDL)
<p>A decision tree algorithm was employed to refine Cropland Data Layer (CDL) using spatial and temporal information. The Refined Cropland Data Layer (R-CDL) could be used as an alternative to researchers as it provides more accurate cropland information. Annual RCDL maps were produced for the contiguous United States from the year 2017 to 2021.</p> <p>To explore the data online and access the web-based services, please visit the project homepage: https://cloud.csiss.gmu.edu/icrop/</p> <p>To read our full paper on RCDL, please visit: https://www.nature.com/articles/s41597-022-01169-w</p> <p>Cite this article:</p> <p>Lin, L., Di, L., Zhang, C. <em>et al.</em> Validation and refinement of cropland data layer using a spatial-temporal decision tree algorithm. <em>Sci Data</em> <strong>9, </strong>63 (2022). https://doi.org/10.1038/s41597-022-01169-w</p> <p>Note:</p> <p>2020 RCDL was reproduced with the Re-released CDL on February 1, 2022. More information could be found at: https://www.nass.usda.gov/Research_and_Science/Cropland/SARS1a.php</p> <p>2021 RCDL was released with the official release of 2021 CDL on February 14, 2022</p> <p>2022 RCDL was released on March 24, 2022</p>
A cropland cover dataset for the middle and lower reaches of the Yellow River over the past millennium
<p>We developed a new gridding allocation model for croplands with unequal weight factors and reconstructed 58 time-point cropland cover maps at a 10-km resolution for the past millennium in the middle and lower reaches of the Yellow River in China. The cropland dataset can be used to model past climate change, estimate carbon emissions, assess human-activity-induced ecological effects, and improve global historical land use scenarios.</p>
Soil respiration dataset from abandoned croplands across China
Soil respiration, a critical component of the global carbon cycle, is highly sensitive to warming. Agricultural soils, including abandoned croplands, are large sources of carbon dioxide (CO2) to the atmosphere. Here, we report the responses of soil respiration and its components to warming from abandoned croplands spanning a large range in latitude (22.3 to 46.6°N) and elevation (2 to 3734 m) across China from three-year (2019 to 2021) in situ warming experiments. This dataset is a collation of soil respiration, heterotrophic respiration, and autotrophic respiration and their temperature sensitivity along with information on microclimates (e.g., soil temperature), plant biomasses, and soil carbon components (e.g., SOC) and quality under climate warming.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.