Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,609
datasets available to search
ShareScore release 0.7.1
Dataset results
1,609 results for “Crops”
Managing Crop Yield Risk at the Kellogg Biological Station, Hickory Corners, MI (2022 to 2023)
Dataset Abstract As farmers adapt to changing climate, they modify practices and technologies to manage evolving risk. Adaptive changes may be as small as adjusting a crop insurance coverage level or as large as investing in an irrigation system. Farmer attitudes toward risk and their subjective perceptions of the evolving probability distributions of crop yields drive adaptation decisions. To understand climate change adaptation behavior by farmers, we undertook the study “Elicitation and Estimation of Risk Preference and Subjective Probabilities to Understand Farmer Decisions on Climate Change Adaptation.” We interviewed 44 Michigan corn and soybean farmers to elicit mathematical expressions of their risk attitudes. During the interviews, each completed two sets of lottery choices, the first using 25 general risky gambles and the second using 18 risky gambles in a crop farming context that enable econometric estimation of risk attitudes (using variants of Expected Utility Theory). Next, they answered questions about corn yield probability distributions over the past ten years and the next ten years (triangular distributions of minimum, most likely, and maximum values) with no water management, irrigation, tile drainage, and drought-resistant seed. After that, they reported on water management investments that they have made in past and intend to make in future. Finally, they provided background information about themselves and their farms. This study (MSU Study ID: STUDY00007871) was submitted to the Michigan State University Institutional Review Board (IRB) by principal investigator Scott Swinton. On July 5, 2022, it was determined to be exempt under 45 CFR 46.104(d) 3(i)(B). Data collection took place during September 2022 through March 2023. Farmer respondents completed the survey instrument on Qualtrics with assistance from graduate students in Agricultural, Food, and Resource Economics at Michigan State University at various MSU Extension offices and restaura
Reduced erosion augments soil carbon storage under cover crops
This dataset comprises field measurements of soil organic carbon erosion and soil organic carbon stock from 152 paired control and cover crop treatments, collected from 57 published studies worldwide. It also provides related information on the collected study sites, including climate (mean annual temperature and mean annual precipitation), geography (slope and altitude), soil properties (silt+clay and SOC concentration), and agricultural management (cover crop species, tillage intensity and experimental duration). Furthermore, it includes the estimated effect sizes of soil organic carbon erosion reduction induced by cover crops in agricultural lands at the global scale.
GHG Dataset for the frontiers publication "Soil Nitrous Oxide Emission and Methane Exchange from Diversified Cropping Systems in Pannonian Region"
<p>GHG Dataset used in the Frontiers Publication "Soil Nitrous Oxide Emission and Methane Exchange from Diversified Cropping Systems in Pannonian Region". Additionally including CO2 besides N2O and CH4. Includes 3 cropping seasons.</p> <p>The data is also available online on the GHG flux visualisation and calculation tool "gasflxvis": https://sae-interactive-data.ethz.ch/gasflxvis/</p> <p>Further details on the calulation are provided both on gasflxvis and the Frontiers publication. Calculation procedure according the following PLOS ONE publication: http://dx.doi.org/10.1371/journal.pone.0200876</p>
Sample data for "A weakly supervised framework for high resolution crop yield forecasts"
<p>This dataset includes sample data for the United States to run the weakly supervised framework as described in the paper titled <em>A weakly supervised framework for high resolution crop yield forecasts</em>, accessible at </p> <table summary="Additional metadata"> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2205.09016">https://doi.org/10.48550/arXiv.2205.09016</a></td> </tr> </tbody> </table> <p> </p> <p>The updated paper (including results from the US) is published in Environmental Research Letters:</p> <p><a href="https://doi.org/10.1088/1748-9326/acf50e">https://doi.org/10.1088/1748-9326/acf50e</a></p> <p> </p> <p>The software implementation of the machine learning baseline is available at: https://github.com/BigDataWUR/MLforCropYieldForecasting/tree/weaksup.</p> <p> </p> <p>Data</p> <p>1. County data (county-data.zip) for county-level strongly supervised models:</p> <p>* CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p> <p>* CSSF_COUNTY_US.csv: Crop productivity indicators including total above-ground production (kg ha<sup>-1</sup>), total weight of storage organs (kg ha<sup>-1</sup>), development stage (0-2). Source: de Wit et al. (2022).</p> <p>* METEO_COUNTY_US.csv: Meteo data including maximum, minimum, average daily air temperature (℃); sum of daily precipitation (PREC) (mm); sum of daily evapotranspiration of short vegetation (ET0) (Penman-Monteith, Allen et al., (1998)) (mm); climate water balance = (PREC - ET0) (mm). Source: Boogaard et al. (2022).</p> <p>* REMOTE_SENSING_COUNTY_US.csv: Fraction of Absorbed Photosynthetically Active Radiation (Smoothed) (FAPAR). Source: Copernicus GLS (2020).</p> <p>* SOIL_COUNTY_US.csv: Soil water holding capacity. Source: WISE Soil Property Database (Batjes, 2016).</p> <p>* YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p> </p> <p>2. 10-km grid data (grid-data.zip) for grid-level strongly supervised models:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>* CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level (similar to county data above).</p> <p>* METEO_GRIDs_US.csv: Meteo data at 10km grid level (similar to county data above).</p> <p>* REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level (similar to county data above).</p> <p>* SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level (similar to county data above).</p> <p>* YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021), Lobell et al. (2020).</p> <p> </p> <p>3. County labels and 10-km grid inputs (dscale-US.zip) for weak supervision:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>* CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level.</p> <p>* METEO_GRIDs_US.csv: Meteo indicators at 10km grid level.</p> <p>* REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level.</p> <p>* SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level.</p> <p>* YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021).</p> <p>* YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p>* CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p>
Influence of soil amendment and crop species on nutrient cycling in a St. Paul urban garden, 2017-2023
An experiment was conducted from 2017-2023 at the University of St. Thomas research garden (Saint Paul, MN) to determine rates of nutrient recycling and loss from compost applied to urban gardens. Thirty-two 4 m2 study plots received one of six different soil amendment treatments, with four different crops growing on each plot. Meteorological data includes hourly measurements of rainfall, solar radiation, temperature and relative humidity, and wind speed and direction, from June 2017-October 2023. Hourly soil moisture measurements were recorded at depths of 10 cm, 20 cm, and 30 cm, from June-December 2021, June-October 2022, and June-October 2023. Annual crop harvest totals from each subplot are reported for 2017-2023. Leachate was collected from lysimeters installed in each of the 132 subplots weekly from June-October of each year (2017-2023), recording total volume. Leachate subsamples were analyzed for NO3-N, NH4-N, and PO4-P. Soil samples were collected at the beginning and end of the growing season in 2017, and every two weeks during the growing season from 2018-2023, and analyzed for pH, organic matter, Bray-1 extractable P, available K, nitrate, and ammonium, at the University of Minnesota Analytical Research Laboratory.
Agronomic Yields in Row Crop Agriculture at the Kellogg Biological Station, Hickory Corners, MI (1989 to 2021)
Dataset AbstractThis data set contains information about agronomic yields for the Main Cropping System Experiment which include treatments 1-4 (corn – wheat – soybean rotations) and after 1994 treatment 6 (alfalfa). Agronomic yields are measured during normal crop harvest; yields are determined by machine harvesters appropriate to each crop as described in the Agronomic protocol.original data source http://lter.kbs.msu.edu/datasets/23
Leaf Area Index on the GLBRC Biofuel Cropping System Experiment at the Kellogg Biological Station, Hickory Corners, MI (2009 to 2017)
Dataset AbstractThe leaf area index was measured to estimate the phenology and growth patterns of the different biofuel crops.original data source http://lter.kbs.msu.edu/datasets/225
SBC LTER: REEF: Net primary production, growth and standing crop of Macrocystis pyrifera in Southern California
The giant kelp Macrocystis pyrifera forms subtidal forests on shallow reefs in temperate regions of the world. It is one of the fastest-growing multicellular autotrophs on Earth and its high productivity supports diverse marine food webs. In 2008, we published a method for estimating biomass and net primary production (NPP) of giant kelp along with five years of data, to provide a more integrated measure of NPP than those yielded by previous methods. Our method combines monthly field measurements of standing crop and loss rates with a model of kelp biomass dynamics to estimate instantaneous mass-specific growth rates and NPP for each season of each year. We have since improved our approach to account for several previously unresolved sources of biomass loss. These improvements have led to a near doubling of our prior estimates of growth and NPP. At our site with the most persistent stand of giant kelp, NPP averages ~5.2 kg dry mass m-2 y-1 and results from the rapid growth (~3.5% per day) of a relatively small standing biomass (~ 0.4 kg dry mass m-2 on average) that turns over ~ 12 times annually. Here we provide revised estimates of seasonal biomass, growth and NPP for the five years covered by our previous publication (2002-2006), along with an additional data collect since then (2007-present). We also present updated relationships for predicting giant kelp biomass and NPP from much more easily obtained measurements of frond density. These data can be used to understand the mechanisms that drive variation in giant kelp NPP at a wide range of temporal scales.
A global dataset gathering 37 field experiments involving cereal-legume intercrops and their corresponding sole crops.
<p>The overall description of the dataset is reported in the <strong>data_report.pdf</strong> file. The methodology for data curation and tidying is published in Peer Community Journal (<a href="https://doi.org/10.24072/pcjournal.389">Mahmoud2024</a>).</p> <p>This dataset gathers the results of 37 field experiments, which involved cereal-legume intercrops and their corresponding sole crops. The field experiments were carried in 5 European countries (France, Denmark, Italy, Germany and England) from 2001 to 2017. The dataset includes:</p> <ul> <li>5 legume species , <em>i.e.</em> chickpea (<em>Cicer arietinum</em> L.), faba bean (<em>Vicia faba</em> L.), lentil (<em>Lens culinaris</em> Med.), lupin (<em>Lupinus albus</em> L.) and pea (<em>Pisum sativum</em> L.),</li> <li>3 cereal species, <em>i.e.</em> barley (<em>Hordeum vulgare</em> L.), durum wheat (<em>Triticum turgidum</em> L.) and soft wheat (<em>Triticum aestivum</em> L.), </li> <li>8 resulting intercrops, <em>i.e.</em> i) barley associated with faba bean, lupin or pea, ii) durum wheat associated with chickpea, faba bean or pea, and iii) soft wheat associated with lentil or pea. </li> </ul> <p>In total, the dataset contains 299 sole crop and 308 intercrop experimental units, one given experimental unit being defined as the unique combination of {site, year, crop management}, with the crop management including species and cultivar choice as well as agricultural interventions (sowing conditions, inputs).</p> <p>The global dataset includes four tables, all sharing a common identifier (experiment_id):</p> <ul> <li>data_trials.csv: the global features describing the experimental sites,</li> <li>data_management.csv: the agricultural management actions carried out on each of the experimental sites,</li> <li>data_traits.csv: measured plant and crop characteristics,</li> <li>data_climate.csv: climate for the experimental sites, retrieved from NASA POWER API.</li> </ul> <p>Additionally, a metadata file is provided (<strong>metadata.xlsx</strong>), describing the table to which the variables belong (variable_type, i.e. trials, management, traits or climate), their name (variable_name), their significance (description) and their unit (unit). Finally, a table including the original references related to experimental files gathered (<strong>references.xlsx</strong>) is also provided.</p> <p>Data providers and field experiments: Laurent Bedoussac, Eric Justes, Etienne-Pascal Jour- net, Christophe Naudin, Henrik Hauggaard-Nielsen, Erik Steen Jensen, Elise Pelzer, Guénaëlle Corre-Hellou, Bochra Kammoun, Loic Viguier, Romain Barillot, Antoine Couëdel, Philippe Hinsinger</p> <p>Database and management: Noémie Gaudio, Rémi Mahmoud, Pierre Casadebaig</p>
Crop and soil measurements of quinoa in Morocco and Belgium (SALAD project)
<p>Crop and soil measurements of quinoa used to calibrate the SWAP-WOFOST model for the SALAD project (https://www.saline-agriculture.com/en).</p> <p>The data were collected from two locations:</p> <p>1) Laayoune, Southern Morocco: ICBA-Q5 quinoa variety, grown in 2021 under irrigation with saline water at levels of 4, 12, and 20 dS/m (https://doi.org/10.3389/fpls.2023.1143170)</p> <p><br>2) Merelbeke, Belgium: Bastille quinoa variety, grown in 2018, 2019, 2022, and 2023 under rainfed and non-saline conditions (https://www.quinoalokaal.be/nl/, https://doi.org/10.3390/plants10122689, https://doi.org/10.3390/plants11030265)</p> <p> </p>
A global dataset of specialty crop biomass and N2O emissions
<div> <p>We reviewed global field studies of vineyard, orchard, and vegetable cropping systems, which were also included in a meta-analysis (<a href="https://doi.org/10.1111/gcb.17233">https://doi.org/10.1111/gcb.17233</a>). We narrowed down the studies to those with field measurements of adequate variables (biomass C, N, and N<sub>2</sub>O) covering at least one growing season. As a result, cumulative N₂O emission measurements (per growing rotation, season, or year), along with biomass data of different plant organs from the same regions, were compiled for grape (<em>Vitis vinifera</em>), almond [<em>Prunus dulcis</em> (Mill.) D.A. Webb], peach (<em>Prunus persica</em> L.), walnut (<em>Juglans regia</em>), lettuce (<em>Lactuca sativa</em>), broccoli (<em>Brassica oleracea</em> var. <em>italica </em>P.), cauliflower (<em>Brassica oleracea</em> var. <em>botrytis </em>L.), and tomato (<em>Lycopersicon esculentum</em> L.) planting system. These observations were collected from fields spanning seven Koppen-Geiger climate types and five countries (the United States, Germany, Spain, France, and Australia). When only dry mass was measured, biomass C content for aboveground vegetable crops and berry fruit was assumed at 43%; nut fruit and woody organs of orchard tree at 48%. Area-weighted averages of N<sub>2</sub>O emissions were used (tree/vine row and interrow).</p> <p> </p> <p>Corresponding author: Mu Hong (mu.hong@colostate.edu)</p> </div> <p> </p>
3C dataverse: Community capitals, cover crops, & conservation agriculture in the U.S. corn-soybean belt, version 2.2
<p><strong>What? </strong></p> <p>A dataset containing 315 total variables from 33 secondary sources. There are 262 unique variables, and 53 variables that have the same measurement but are reported for a different year; e.g. average farm size in 2017 (CapitalID: N27a) and 2022 (N27b). Variables were grouped by the community capital framework's seven capitals—Natural (96 total variables), Cultural (38), Human (39), Social (40), Political (18), Financial (67), & Built (15)—and temporally and thematically ordered. The geographic boundary is NOAA NCEI's corn and soybean belt (figure below), which stretches across 18 states and includes N=860 counties/observations. Cover crop data for the 80 Crop Reporting Districts in the boundary are also included for 2015-2021.</p> <p><strong>Why? </strong></p> <p>Comprehensively assessing how community capital clustered variables, for both farmers and nonfarmers, impact conservation practices (and perennial groundcover) over time helps to examine county-level farm conservation agriculture practices in the context of community development. We contribute to the robust U.S. cover crop literature a better understanding of how overarching cultural, social, and human factors influence conservation agriculture practices to encourage better farm management practices. Analyses of this Dataverse will be presented as recomendations for farmers, nonfarmers, ag-adjacent stakeholders, and community leaders.</p> <p><strong>How? </strong></p> <p>Variables used in this dataset range 20 years, from 2004-2023, though primary analyses focus on data collected between 2017-2024, primarily 2017 and 2022 (NASS Ag Census years). First, JAM-K requested, accessed, and downloaded data, most of which was already publically available. Next, JAM-K cleaned the data and aggregated into one dataset, and made it publically available on Google Drive and Zenodo. </p> <p><strong>What is 'new' or corrected in version 2.2? </strong></p> <p><em>Edited/amended</em>: Carroll, KY is now spelled correctly (two 'l's, not one); variable names, full and abbreviated, were updated to include the data year; Pike County's (IL) FIPS has been corrected from its wrong 17153 (same as Pulaski County) to 17149 (correct fips), and all Pike County (IL) data has been correctly amended; Farming dependent (ERS) updated for all variables; Data for built capital variables irrCorn17, irrSoy17, irrHcrp17, tractor17, and combine17 were incorrect for v.1, but were corrected for v.2; Several variable labels aggregated by Wisconsin University's Population Health Institute's County Health Rankings and Roadmaps were corrected to have the data's original source and years included, rather than citing CHR&R as the source (except for CHR&R's originally-produced values such as quartiles or rank scores); variables were reorganized by hypothesized community capital clusters (Natural -> Built), and temporally within each cluster. </p> <p><em>Added</em>: 55 variables, mostly from the 2022 Ag Census, and v 2.2 added a .pdf file with descriptives of data sources and years, and a .sav file. </p> <p><em>Omitted</em>: Four variables deemed irrelevant to the study; V1 codebook's "years internally available" column. Variable herbac22 for 55079, Milwaukee, WI, incorrectly had the value 2,049.612. That value was correctly changed to missing, with no data in the cell.</p> <p><strong>CRediT</strong>: conceptualization, CBF, JAM-K; methodology, JAM-K; data aggregation and curation, JAM-K; formal analysis, JAM-K; visualization, JAM-K; supervision, CBF; funding acquisition, CBF; project administration, CBF; resources, CBF, JAM-K</p> <p><strong>Acknowledgements</strong>: This research was funded by the Agriculture and Food Research Initiative Competitive Grant No. 2021-68012-35923 from the United States Department of Agriculture National Institute for Food and Agriculture. Any opinions, findings, conclusions, or recommendations expressed in this presentation are those of the authors and do not necessarily reflect the view of the U.S. Department of Agriculture. Much thanks to Corteva for granting data access of OpTIS 2.0 (2005-2019), and Austin Landini for STATA code and visualization assistance. </p>
Density independent prey choice, taxonomy, life history and web characteristics determine the diet and biocontrol potential of spiders (Linyphiidae and Lycosidae) in cereal crops - Dataset
<p>Materials and Methods</p> <p>Fieldwork</p> <p>Money spiders (Araneae: Linyphiidae) and wolf spiders (Araneae: Lycosidae) were the two most common families present in these field surveys, so were prioritised for collection. Spiders were visually located along transects in two adjacent barley fields at Burdons Farm, Wenvoe in South Wales (51°26'24.8"N, 3°16'17.9"W) and collected from occupied webs and the ground, between April and September 2018. Surveys and sampling were conducted five days per week across this period. Each transect was adjacent to a randomly selected tramline and they were distributed across the entire field. The areas searched were 4 m<sup>2</sup> quadrats at least 10 m apart and all observed linyphiids and lycosids were collected in approximately 15-minute searches. The spiders included in this study were taken from 64 locations across 24 days (Supplementary Table 3) along the aforementioned transects. Spiders were individually placed into 1.5 ml microcentrifuge tubes containing 100 % ethanol using an aspirator, regularly changing meshing, at least every five spiders, to limit potential cross-contamination between spiders (spiders were also subsequently washed during transferral to fresh ethanol at the identification and, separately, dissection stages). Linyphiids occupying webs were prioritised for collection, but ground-active linyphiid spiders were also collected. For each spider taken from a web, the height of the web from the ground and its approximate dimensions were recorded, the latter calculated as approximate web area. Spiders were taken to Cardiff University, transferred to fresh ethanol, adults identified to species-level and juveniles to genus, and stored at -80 °C in 100 % ethanol until subsequent DNA extraction. To obtain data on local prey density, 4 m<sup>2</sup> of ground and crop stems were suction sampled using a ‘G-vac’ for 30 seconds at each quadrat from which spiders were collected, with the collected material emptied into a bag, any organisms immediately killed with ethyl-acetate and material frozen for storage before sorting into 70 % ethanol in the lab.</p> <p>All invertebrates were identified to family level due to the restriction of many of the metabarcoding-derived dietary data to this level, and the difficulty associated with finer taxonomic resolution of many taxa. Exceptions included springtails of the superfamily Sminthuroidea (Sminthuridae and Bourletiellidae, which were often indistinguishable following suction sampling and preservation due to the fine features necessary to distinguish them) which were left at super-family, mites (many of which were immature or in poor condition) which were identified to order level and wasps of the superfamily Ichneumonoidea (which were identified no further due to obscurity of wing venation due to damage).</p> <p> </p> <p>Extraction and high-throughput sequencing of spider gut DNA</p> <p>Given their prevalence in field collections, dietary analysis was carried out for the linyphiid genera <em>Erigone</em>, <em>Tenuiphantes</em>, <em>Bathyphantes</em> and <em>Microlinyphia </em>(Araneae: Linyphiidae), and the Lycosidae genus <em>Pardosa</em>. Spiders were transferred to and washed in fresh 100 % ethanol to reduce external contaminants prior to identification via morphological key <sup>1</sup>. Abdomens were removed from spiders and again washed in and transferred to fresh 100 % ethanol. DNA was extracted from the abdomens via Qiagen TissueLyser II and DNeasy Blood & Tissue Kit (Qiagen) as per the manufacturer protocol, but with an extended lysis time of 12 hours to account for the complex and branched gut system in spider abdomens <sup>2</sup>. At least one extraction negative (blank tubes treated identically to samples) was included per 12 spiders (each extraction typically contained 24 spiders, thus two extraction negatives), which was included in subsequent PCR and high-throughput sequencing to detect instances of lab/reagent contamination.</p> <p>For amplification of DNA, two primer pairs were used. BerenF-LuthienR <sup>3</sup> amplified a broad range of invertebrates including spiders, and TelperionF-LaureR, amplified a range of invertebrates but fewer spiders (modified from TelperionF-LaurelinR <sup>3</sup> via one base-pair change from Laurelin; 5’-ggrtawacwgttcawccagt-3’). Primers were labelled with unique 10 bp molecular identifier tags (MID-tags) so that each individual had a unique pairing of forward and reverse tags for identification of each spider post-sequencing. PCR reactions of 25 µl contained 12.5 µl Qiagen PCR Multiplex kit, 0.2 µmol (2.5 µl of 2 µM) of each primer and 5 µl template DNA. Reactions were carried out in the same thermocycler, optimised via temperature gradient, with an initial 15 minutes at 95 °C, 35 cycles of 95 °C for 30 seconds, the primer-specific annealing temperature for 90 seconds and 72 °C for 90 seconds, respectively, followed by a final extension at 72 °C for 10 minutes. BerenF-LuthienR and TelperionF-LaureR used annealing temperatures of 52 °C and 42 °C, respectively.</p> <p>Within each PCR 96-well plate, 12 negative controls (extraction and PCR), 2 blank controls and 2 positive controls were included (i.e. 80 samples per plate), based on Taberlet <em>et al. </em>(2018). Positive controls were mixtures of invertebrate DNA comprised of non-native Asiatic species in four different proportions (Supplementary Table 1) and blanks were empty wells within each plate to identify tag-jumping into unused MID-tag combinations. PCR negative controls were DNase-free water treated identically to DNA samples. A negative control was present for each MID-tag to identify any contamination of primers. All PCR products were visualised in a 2 % agarose gel with SYBRSafe (Thermo Fisher Scientific, Paisley, UK) and placed in categories based on their relative brightness. The concentration of these brightness categories was quantified via Qubit dsDNA High-sensitivity Assay Kits (Thermo Fisher Scientific, Waltham, MA, USA) with at least three representatives of each category per plate. The PCR products were then proportionally pooled according to these concentrations. Each pool was cleaned via SPRIselect beads (Beckman Coulter, Brea, USA), with a left-side size selection using a 1:1 ratio (retaining ~300-1000 bp fragments). The concentration of the pooled DNA was then determined via Qubit dsDNA High-sensitivity Assay Kits and pooled together into one library per primer pair. Library preparation for Illumina sequencing was carried out on the cleaned libraries via NEXTflex Rapid DNA-Seq Kit (Bioo Scientific, Austin, USA) and samples were sequenced on an Illumina MiSeq via a V3 chip with 300-bp paired-end reads (expected capacity ≤25,000,000 reads). Bioinformatic analysis followed (Drake et al., 2021; Supplementary Information 1).</p> <p> </p> <p>Statistical analysis</p> <p>All analyses were conducted in R v4.0.0 <sup>6</sup>. Initial multivariate analyses used binary data (i.e., presence/absence) given the various problems inherent to quantifying metabarcoding data <sup>7,8</sup>. Prey species that occurred only once across all of the dietary samples were removed before further analyses to prevent outliers skewing the results, which is particularly problematic for non-metric multidimensional scaling. Spider diets were compared between variables using multivariate generalized linear models (MGLMs) via ‘manyglm’ in the ‘mvabund’ package <sup>9</sup> with a binomial error family and Monte Carlo resampling. Model independent variables included spider genus, spider life stage (juvenile or adult, the latter defined by fully developed genitalia), spider sex and all two-way interactions between these variables. Pairwise two-way interactions were also included between the aforementioned variables and Julian day to account for how seasonality may affect these relationships.</p> <p>Coarse dietary differences were visualised by non-metric multidimensional scaling (NMDS) via metaMDS in the ‘vegan’ package <sup>10</sup> with Jaccard distance in two dimensions and 999 tries. For NMDS, outliers (usually samples containing rare taxa) were identified by plotting and subsequently removed to facilitate separation of samples and achieve minimum stress. For visualisation of the effect of categorical variables against the dietary NMDS, spider plots were created using ‘ordispider’ with ‘ggplot’ and the ‘RColorBrewer’ ‘Accent’ colour palette <sup>11</sup>. Spider diet was compared against web characteristics for spiders for which both data were available using the MGLM process outlined above, but with starting models containing web height, web area, an interaction between the two, and pairwise interactions between genus, life stage and sex with the two web variables. This model used the same binomial error family as above, but with a ‘cloglog’ link function. For visualisation of the effect of continuous variables against the NMDS, surf plots were created with scaled coloured contours using the function “ordisurf” of the “ggplot” package in R.</p> <p>All prey taxa were classified as agricultural pests, natural enemies or excluded from subsequent analyses of intraguild predation and biocontrol (Supplementary Table 2). Intraguild predation and biocontrol variables were created by counting the number of natural enemy taxa, and, separately, of agriculturally relevant “pest” taxa (taxa containing species that commonly adversely affect agricultural productivity; Supplementary Table 2) in each spider’s diet. These resultant count data (effectively the diversity of pests and natural enemies predated by each individual spider) were separately analysed against spider genus, life stage and sex via GLM. “Site” (denoting the 4 m<sup>2</sup> area from which spiders were collected within fields) was initially included as a random effect in generalized linear mixed-models, but no significant effect was observed when comparing this model against a standard GLM via a likelihood ratio test of nested models using the ‘lrtest’ command in the ‘lmtest’ package <sup>12</sup>. Standard GLMs were thus used to avoid issues relating to singularity in the mixed models. The assumptions for the resultant Poisson error family GLMs were tested using the “testResiduals” function of the ‘DHARMa’ package <sup>13</sup>. Intraguild predation and biocontrol differences between significant terms were visualised using violin plots with the quartiles, median and 95 % upper limit annotated using the ‘geom_violin’ function in ‘ggplot2’.</p> <p><em>In situ</em> spider prey choice was analysed using network-based null models in the ‘econullnetr’ package <sup>14</sup> with the ‘generate_null_net’ command, visually represented with the ‘plot_preferences’ command. Binary dietary data were used alongside suction sample count data to represent prey availability. These suction sample data, as described above, were collected at the same sites as the spiders three days after spider collection. Prior to the taxonomic prey choice analysis, an hemipteran identified no further than order level through dietary analysis was removed due to the inability to pair it to any present prey taxa with certainty. Standardised effect sizes (SES) were extracted for all comparisons for each individual spider and compared between genera, life stages and sexes using permutational multivariate analysis of variance (PerMANOVA) using the ‘adonis’ function of the ’vegan’ package with 9999 permutations and a Euclidean distance matrix to determine overall differences in prey choice.</p> <p> </p> <p>References</p> <p>1. Roberts, M. J. <em>The Spiders of Great Britain and Ireland (Compact Edition)</em>. (Harley Books, 1993).</p> <p>2. Krehenwinkel, H., Kennedy, S., Pekár, S. & Gillespie, R. G. A cost-efficient and simple protocol to enrich prey DNA from extractions of predatory arthropods for large-scale gut content analysis by Illumina sequencing. <em>Methods Ecol. Evol.</em> <strong>8</strong>, 126–134 (2017).</p> <p>3. Cuff, J. P. <em>et al.</em> Money spider dietary choice in pre- and post-harvest cereal crops using metabarcoding. <em>Ecol. Entomol.</em> <strong>46</strong>, 249–261 (2021).</p> <p>4. Taberlet, P., Bonin, A., Zinger, L. & Coissac, E. <em>Environmental DNA</em>. (Oxford University Press, 2018).</p> <p>5. Drake, L. E. <em>et al.</em> An assessment of minimum sequence copy thresholds for identifying and reducing the prevalence of artefacts in dietary metabarcoding data. <em>Methods Ecol. Evol.</em> <strong>in press</strong>, (2021).</p> <p>6. R Core Team. R: A language and environment for statistical computing. (2020).</p> <p>7. Deagle, B. E., Thomas, A. C., Shaffer, A. K. & Trites, A. W. Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count? <em>Mol. Ecol. Resour.</em> <strong>13</strong>, 620–633 (2013).</p> <p>8. Deagle, B. E. <em>et al.</em> Counting with DNA in metabarcoding studies: How should we convert sequence reads to dietary data? <em>Mol. Ecol.</em> <strong>28</strong>, 391–406 (2019).</p> <p>9. Wang, Y., Naumann, U., Wright, S. T. & Warton, D. I. mvabund – an R package for model-based analysis of multivariate abundance data. <em>Methods Ecol. Evol.</em> <strong>3</strong>, 471–474 (2012).</p> <p>10. Oksanen, J. <em>et al.</em> vegan: Community Ecology Package. (2016).</p> <p>11. Neuwirth, E. RColorBrewer: ColorBrewer palettes. (2014).</p> <p>12. Zeileis, A. & Hothorn, T. Diagnostic checking in regression relationships. <em>R News</em> <strong>2</strong>, 7–10 (2002).</p> <p>13. Hartig, F. DHARMa: residual diagnostics for hierarchical (multi-level/mixed) regression models. (2020).</p> <p>14. Vaughan, I. P. <em>et al.</em> econullnetr: an r package using null models to analyse the structure of ecological networks and identify resource selection. <em>Methods Ecol. Evol.</em> <strong>9</strong>, 728–733 (2018).</p>
Data for "Temperate Regenerative Agriculture practices increase soil carbon but not crop yield – a meta-analysis"
<p>Supplementary Files for systematic review and meta-analysis: Temperate Regenerative Agriculture practices increase soil carbon but not crop yield – a meta-analysis</p> <p> </p>
A Novel Crop Shortlisting Method for Sustainable Agricultural Diversification Across EU (Italy)
<p>In order to shortlist possible options from a pool of 2700 crops, a crop-climate-soil matching ex-ercise was performed across Italian territory and crops with more than 70% suitability where chosen for further analysis. In the second phase, a multicriteria ranking index was employed to assign ranks to chosen crops of 4 main types; (i) cereals and pseudocereals, (ii) legumes, (iii) starchy roots/ tubers and (iv) vegetables. In order to provide a comprehensive analysis, major crops that are grown in the region where also included in the analysis. The results of evaluation of 4 major criteria (a) calorie and nutrition demand b) functions and uses c) availability and acces-sibility to their genomic material d) possession of adaptive traits, and e) physiological traits) re-vealed the potential for teff, faba bean, cowpea, green arrow arum, Jerusalem artichoke, Fig-leaved Gourd and Watercress. </p>
Datasets for paper 'Cabello, V., Renner, A., Giampietro, M. 2019. Relational analysis of the resource nexus in arid land crop production. Advances in Water Resources 130:258-629'
<p>Datasets produced for the paper Cabello, V., Renner, A., Giampietro, M. 2019.<em> </em>Relational analysis of the resource nexus in arid land crop production. <em>Advances in Water Resources </em>130:258-269</p>
Data from: Carbon and Water Balances in a Watermelon Crop Mulched with Biodegradable Films in Mediterranean Conditions at Extended Growth Season Scale
<p><span>Abstract</span></p> <p><span>The uploaded data are relative to the investigation around (i) the carbon source/sink nature and, further, (ii) the water and carbon balances, of a drip-irrigated and mulched watermelon. The crop was cultivated under the semi-arid climate of the Apulia region, in south Italy.</span></p> <p><span>The used mulching films were biodegradable as indicate by the producer; plants and some non-standard fruits were left on the soil as green manure after harvesting, thus, the experiment spanned from planting to the subsequent crop (6 months of continuous measurement from June to November 2023). </span></p> <p><span>The results detailed in the original publication indicate that mulching films contribute to carbon sequestration in the soil (+19.3 gC m<sup>−2</sup>). However, this mulched watermelon represents a net carbon source, with a net biome exchange, as loss from ecosystems, equal to +230 gC m<sup>−2</sup>. This is primarily due to the substantial amount of carbon exported through marketable fruits. Fixed water scheduling led to water waste through deep percolation (approximately 1/6 of the water supplied), which also contributed to the loss of organic carbon via leaching (−4.3 gC m<sup>−2</sup>). </span></p> <p><span> </span></p> <p><span>Methods</span></p> <p><span>Site and crop</span></p> <p><span>The field site was at the CREA-AA Research Unit experimental farm located in southern Italy (Rutigliano–Bari, 41 01’ N, 17°01’ E, altitude 147 m a.s.l.)., characterized by a Mediterranean semi-arid climate (average annual rainfall of 535 mm). The soil is classified as Lithic Rhodoxeralf, with a clay texture, stable structure, shallow profile (0.6–1.1 m) and rapid drainage due to an underlying cracked limestone subsoil. The SOC content averages around 12.0 g kg<sup>−1</sup>. The field capacity and the permanent wilting point volumetric water contents are 0.36 and 0.21 m<sup>3</sup> m<sup>−3</sup>, respectively; with a bulk density of 1.15 Mg m<sup>−3</sup>, the available soil water ranges from 80 to 140 mm.</span></p> <p><span>The studied watermelon crop (seedless var. Lion king), followed a broccoli cabbage crop harvested in April and partially incorporated (0.81 kg m<sup>−2</sup> of fresh biomass in a soil layer depth of 0.30 m, corresponding to 0.69 kgH2O m<sup>−2</sup>) as green manure on 25 May 2023. Main tillage at medium depth ploughing (0.30 m) and seedbed preparation were performed between 25 and 30 May 2023; the biodegradable film mulch (model PC 100 d8, BASF, Italy, 1 m width) was applied on 1 June 2023. On the same day, driplines (2.1 Lh<sup>−1</sup> emitters, 0.60 m apart) and the main organic fertilization (Orga-Kem 6.11.8 + 11CaO, 300 kg ha<sup>−1</sup>) were also applied. The watermelon plants were transplanted on 9 June at a spacing of 2.70 m between rows and 1 m between plants, covering an area of about 4.0 ha, with a density of approximately 3200 plants ha<sup>−1</sup>. Every 6 rows, the inter-row distance was 5 m to facilitate machinery passage. The first irrigation was performed the day before planting. Crop management adhered to the usual treatments in the area including mechanical weed removal every 4 weeks, irrigation around three times per week to maintain optimal soil water conditions and monthly fertigation (ammonium sulphate 50 kg ha<sup>−1</sup>, magnesium nitrate 30 kg ha<sup>−1</sup>, calcium nitrate 60 kg ha<sup>−1</sup>, mycorrhizae 20 kg ha<sup>−1</sup>). The scalar harvest of marketable fruits occurred between 28 and 31 August 2023. After harvesting, on 25 September 2023, the fresh plant residues (0.6 kg m<sup>−2</sup> of fresh biomass, corresponding to 0.49 kgH2O m<sup>−2</sup>), unharvested fruits (4.0 kg m<sup>−2</sup> of fresh material, corresponding to 3.7 kgH2O m<sup>−2</sup>) and the mulching film were chopped by a tractor shredder and ploughed in two steps, on 2 and 13 October 2023, to a soil depth of 0.30 m. Measurements concluded at the end of November 2023, when tillage for the new winter crop commenced.</span></p> <p><span> </span></p> <p><span>Measurements of H<sub>2</sub>O and CO<sub>2</sub> fluxes; partitioning in evaporation, transpiration, photosynthesis and respiration</span></p> <p><span>The eddy covariance technique was employed to monitor water vapor (H<sub>2</sub>O) and carbon dioxide (CO<sub>2</sub>) fluxes. The equipment comprised a three-dimensional sonic anemometer (uSonic 3 Scientific, Metek GmbH, 25337 Elmshorn, Germany) and a fast response open-path infrared gas analyzer (LI-7500, Li-COR Inc., Lincoln, NE, USA). The three wind components, sonic temperature and atmospheric concentrations of CO<sub>2</sub> and H<sub>2</sub>O were continuously measured at 1.5 m above the crop canopy, with the sensor height adjusted to follow crop growth, reaching a maximum of 1.75 m. </span></p> <p><span>Data were recorded at a frequency of 10 Hz on a dedicated computer using the MeteoFlux software (Servizi Territorio, S.n.c., Cinisello Balsamo, Italy) and were stored on an hourly scale. Post-processing and computation of hourly fluxes of H<sub>2</sub>O (mmol m<sup>−2</sup> s<sup>−1</sup>) and CO<sub>2</sub> (</span>μ<span>mol m<sup>−2</sup> s<sup>−1</sup>) were conducted using EddyPro software, v7.0.9 (</span><a href="http://www.licor.com/eddypro"><span>http://www.licor.com/eddypro</span></a><span>), applying 60 min block averaging, double coordinate rotation, the statistical test, the maximum cross-covariance method, and the WPL density correction.</span></p> <p><span>H<sub>2</sub>O and CO<sub>2</sub> fluxes were partitioned into transpiration, evaporation, photosynthesis and respiration, respectively, using the flux variance similarity method. This method utilizes the Monin–Obukhov similarity theory to separate stomatal (photosynthesis, Fp, and transpiration, Ft) from non-stomatal (respiration, Fr, and evaporation, Fe) processes (Palatella et al., 2014). the H<sub>2</sub>O and CO<sub>2</sub> EC fluxes were partitioned using an adaptation of the code in Phyton provided by (Skaggs et al., 2018) and downloaded from <span> </span></span><a href="https://github.com/usda-arsussl/fluxpart"><span>https://github.com/usda-arsussl/fluxpart</span></a><span> (V0.2.10).</span></p>
Crop-specific global fertilizer application rates from "Closing yield gaps through nutrient and water management"
<p>Crop-specific global maps of N, P2O5, and K2O fertilizer application rates circa the year 2000 from the following paper:</p> <p>Mueller, ND, JS Gerber, M Johnston, DK Ray, N Ramankutty, and JA Foley. 2012. Closing yield gaps through nutrient and water management. <em>Nature</em> <strong>490</strong>: 254–257</p> <p>Data are provided at five arc-minute resolution and are saved as netcdf files. Fertilizer application rates are estimated from reconciling various national and subnational data sources. See the Supplementary Information from the 2012 paper for a full description of data sources and methods. Data quality for each grid cell is described in a map layer. Files containing the text "totalcons" sum nutrient consumption across crops per grid cell, using crop harvested areas from Monfreda et al. 2008 Global Biogeochemical Cycles. For maize, wheat, and soybean N application rates, additional maps and csv files (containing the text "politboundaries") identify the political units around the world containing unique information. Crops and crop group categories are consistent with those utilized in Monfreda et al. 2008 Global Biogeochemical Cycles.</p>
Harmonised LUCAS database classified by crop sequence type
<p>Assessing the benefits of crop diversification – a pillar of the agroecological transition – on a large scale requires a description of current crop sequences as a baseline, which is lacking at the scale of the European Union (EU). This work is based on the Harmonised LUCAS in-situ land cover and use database for field surveys from 2006 to 2018 in the European Union (doi: <a href="http://doi.org/10.2905/f85907ae-d123-471f-a44a-8cca993485a2">10.2905/f85907ae-d123-471f-a44a-8cca993485a2)</a> to fill this gap, We completed this dataset with a crop sequence type information for each point under non-perennial agricultural land cover in 2012, 2015 and 2018.</p> <p>The dataset lucas_classified.csv includes 31 159 points. Variables "point_id", "nuts0", "nuts2", "th_lat", "th_long", "LC1_2012", "LC1_2015", "LC1_2018" are inherited from the Harmonised LUCAS databse. Variables "cereals", "corn", "rapeseed", "sunflower", "pulses", "rootCrops", "forageLeg", "grassland" correspond to the temporal frequencies of respectively cereals, corn, rapeseed, sunflower, pulses, root crops, forage legumes and grassland within the 2012, 2015 and 2018 crop sequence for each point. Variable "crop_sequence_type" is the crop sequence type assigned to each point, among eight options: cereals, corn and cereals, forage legumes and cereals, pulses and cereals, rapeseed and cereals, root crops and cereals, sunflower and cereals, temporary grasslands.</p> <p>This dataset could be used to map current dominant crop sequences in the European Union, as illustrated in the map attached, and to assess the benefits of future crop diversification.</p> <p> </p>
Data from Sand aggradation alters biofilm standing crop and metabolism in a low-gradient Lake Superior tributary
We conducted a comparative study of biofilm standing crop and metabolism in the Salmon Trout River, a tributary of Lake Superior where watershed disturbances have led to 3-fold increases in streambed fine sediments, predominately sand, in the past decade. We compared biofilm standing crop and metabolism rates using light–dark chambers in reaches where substrate consisted of predominately exposed rock or sand substrates. This data archive includes rates of primary production and respiration, biomass measurements from chambers, and benthic standing crop and water chemistry data collected from the same river sites over the course of a summer. All data were published in Journal of Great Lakes research in 2015, https://doi.org/10.1016/j.jglr.2015.09.004
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.