Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,062
datasets available to search
ShareScore release 0.7.1
Dataset results
13,062 results for “prediction”
Spartina alterniflora above- and belowground biomass predictions and inundation intensity as estimated by the Belowground Ecosystem Resiliency Model for U.S. Georgia marshes from 2014 to 2023.
We applied the Belowground Ecosystem Resiliency Model (BERM) to estimate monthly aboveground biomass (AGB) and belowground biomass (BGB) in U.S. Georgia Spartina alterniflora marshes from 2014 to 2023 at 30 m scale. This application involved BERM version 2.0 (https://doi.org/10.5281/zenodo.13306821), which was built using data in the PLT-GCET-2308 dataset (https://dx.doi.org/10.6073/pasta/4a0b715104849d98320fcc34e7cd63a4). Data sources for BERM application included Landsat-8/9, NOAA CO-OPS Station ID: 8670870, Daymet, and USGS 3DEP 2018 DEM. Download and processing steps are described in the BERM code and in metadata methods section. Specific descriptions of data processing are available in model code: https://doi.org/10.5281/zenodo.13306821. Data provided here include model output of AGB estimates, BGB estimates, and calculated inundation intensity. See "Data reporting" method in the metadata for description of data files. For logisitical purposes here we present only select data from the model input and output. All model input data sources as listed in the abstract are publicly available. Model calibration data and code are published as well. Additional predictions not published here include foliar chlorophyll, foliar nitrogen, and leaf area index.
Fluxes project at North Temperate Lakes LTER: Predicting Peat Depth in a North Temperate Lake District 2008
Peat deposits contain on the order of 1/6 of the Earth's terrestrial fixed carbon (C), but uncertainty in peat depth precludes precise estimates of peat C storage. To assess peat C in the Northern Highlands Lake District (NHLD), a approximately 7000 square km region in northern Wisconsin, United States, with 20 percent peatland by area, we sampled 21 peatlands. In each peatland, peat depth (including basal organic lake sediment, where present) was measured on a grid and interpolated to calculate mean depth. Our study addressed three questions: (1) How spatially variable is peat depth? (2) To what degree can mean peat depth be predicted from other field measurements (water chemistry, water table depth, vegetation cover, slope) and/or remotely sensed spatial data? (3) How much C is stored in NHLD peatlands? Site mean peat depth ranged from 0.1 to 5.1 m. Most of the peatlands had been formed by the in-filling of small lake basins (terrestrialization), and depths up to 15 m were observed. Mean peat depth for small peat basins could be best predicted from basin edge slope at the peatland/upland interface, either measured in the field or calculated from digital elevation (DEM) data (Adj. R2 = 0.70). Upscaling using the DEM-based regression gave a regional mean peat depth of 2.1 plus or minus 0.2 m (including approximately 0.1 to 0.4 m of organic lake sediment) and 144 plus or minus 21 Tg-C in total. As DEM data are widely available, this technique has the potential to improve C storage estimates in regions with peatlands formed primarily by terrestrialization. Number of sites: 21 Sampling Frequency: once for each site
S71 | CECSCREEN | HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites
<p>This is the collection associated with list S71 CECSCREEN HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites<strong> </strong>on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>CECScreen is part of the HBM4EU project (coord. UBA) > WP16 "emerging chemicals" (lead INRA, JP Antignac/L Debrauwer) > Task 16.1 (lead IRAS, J Vlanderen / R Vermeulen) > Main contributor (J Meijer) > Involved Partners (M Lamoree, T Hamers, S Hutinet, A, Covaci, C Huber, M Krauss, DI Walker, EL Schymanski). Further details in Meijer et al (2021) DOI: <a href="https://doi.org/10.1016/j.envint.2021.106511">10.1016/j.envint.2021.106511</a>. Dataset DOI: <a href="https://doi.org/10.5281/zenodo.3956586">10.5281/zenodo.3956586</a>.</p> <p>Update 23/7/2020 (v0.1.1): updated MetFrag files to remove elements causing errors (Os, Pd, Ag, Be). Update 8 Nov 2022 (v0.1.2) removed new lines in several synonyms as detected at BioHackEU22.</p>
Cross-phyla protein annotation by structural prediction and alignment
<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is typically limited to a few model organisms. In non-model species, the sequence-based prediction of gene orthology can be used to infer protein identity, however this approach loses predictive power at longer evolutionary distances. Here we propose a workflow for protein annotation using structural similarity, exploiting the fact that similar protein structures often reflect homology and are more conserved than protein sequences.</p> <p><strong>Results:</strong> We propose a workflow of openly available tools for the functional annotation of proteins via structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with known homology in >90% cases, and annotates an additional 50% of the proteome beyond standard sequence-based methods. We uncover new functions for sponge cell types, including extensive FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes, proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>
LAGOS - Predicted and observed maximum depth values for lakes in a 17-state region of the U.S.
This dataset includes predicted and observed values of maximum depth for lakes in the upper Midwest and northeast United States. All observed values came from LAGOS ver 1.040.0 (LAke multi-scaled GeOSpatial and temporal database), an integrated database of lake ecosystems (Soranno et al. 2015). LAGOS contains a complete census of lakes great than or equal to 4 ha with corresponding geospatial information for a 17-state region of the U.S., and a subset of the lakes has observational data on morphometry and chemistry. Approximately 40 different sources of data were compiled for this dataset and were mostly generated by government agencies (state, federal, tribal) and universities. Here, observed maximum depth values (n = 8164) were used to train and validate a predictive mixed effects model for lake depth using terrestrial and lake morphology as predictors (Oliver et al., submitted). Predicted values (n = 50 607) generated by the model had a root mean squared error of 7.1 m. This research was supported by the NSF Macrosystem Biology awards 1065786, 1065818, and 1065649.
Simulated NGS read datasets for bacterial pathogenic potential prediction
<p>## Predicting pathogenic potentials from NGS reads: novel bacterial species</p> <p>This repository contains simulated Illumina read datasets for bacterial pathogenic potential prediction and associated metadata extracted from the IMG Database (https://img.jgi.doe.gov/). The reads are 250bp long and were simulated with Mason (https://www.seqan.de/apps/mason/) from genomes downloaded from NCBI. The training-validation-test split was done on the species level to ensure "novelty" of validation and test species. The training sets contain 10 million reads per class, validation sets - 1.25 million reads per class, and test sets - 1.25 million paired reads per class. Additional, imbalanced training sets contain 2.5 million "nonpathogenic" and 17.5 million "pathogenic" reads, keeping the mean covarage constant for all species. The temporal benchmark test set contains reads from 3 additional pathogenic species in the Pantoea genus.</p> <p>## Predicting pathogenic potentials from NGS reads: novel strains of known species</p> <p>The BacPaCS datasets contain reads simulated from the dataset compiled by Barash et al. (https://doi.org/10.1093/bioinformatics/bty928). It this case, the training-validation-test split was done on the strain level (so different strains of the same species may be present in all three sets).</p>
S38 | SOLNSLMCTPS | SOLUTIONS Predicted Transformation Products by LMC
<p>This is the collection associated with list S38 SOLNSLMCTPS on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>S38 | SOLNSLMCTPS | <strong>SOLUTIONS Predicted Transformation Products by LMC</strong></p> <p>Predicted Transformation Products calculated by LMC during the SOLUTIONS project, interactive table available <a href="https://www.normandata.eu/solutions/modelsTransformationProducts.php">here</a>.</p> <p>14/11/19 update: added CSV version. 9/7/2025: fixed several corrupt SMILES and added InChIKeys to XLSX/CSV. Note that the author had to be changed to the University to satisfy Zenodo upload requirements, the original authors were listed as <a href="https://oasis-lmc.org/about/contacts.aspx">LMC</a>. </p>
Pawpaws prevent predictability: A locally-dominant tree alters understory beta-diversity and community assembly
<p>Data used in "Pawpaws Prevent Predictability: A locally-dominant tree alters understory beta-diversity and community assembly" (Wassel and Myers) accepted for publication in Ecosphere.<br><br><strong>Metadata for Zenodo.pdf </strong>contains more information on the following data files including descriptions of the columns. </p> <p>The file <strong>understory_abundance_data2021.csv</strong> contains all species abundances in 1x1m plots. This data was used for analyses in publication. Each row is a plot, each column is a speceis or plot descriptor, values for columns 5 and higher are species abundances. Data was collected July-August 2021 by Anna Wassel in Missouri, USA. </p> <p>The file <strong>understory_species_list2021.csv </strong>contains a list of the species codes used in the first file with their scientific names and their status as herbs or woody. This was used to filter out herbaceous species from the data set for herbaceous-only analyses. </p> <p> </p>
FixMe: An Incremental Lightweight Method for Vulnerability Data Collection for Security Patch Prediction
<div> <div>This repository has the FixMe dataset and the source code for extracting the new dataset. is a lightweight approach for collecting code patches based on analyzing the commits of various version control systems. The practical framework is designed to generate patches across a wide array of programming languages. This open-source tool streamlines the process of gathering vulnerability records from the Common Vulnerabilities and Exposures (CVE) database through an incremental approach. By embracing an incremental methodology, we expedite the acquisition of data, ensuring the inclusion of newly identified vulnerabilities and their corresponding patch pairs. Our methodology involves extracting security issues, obtaining vulnerability-fixing commits, and retrieving relevant source code from various projects. The extracted dataset by the FixMe tool supports for the automated patch prediction, automated program repair, commit classification, vulnerability prediction and more.</div> </div>
Predicted occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution
<p>The dataset contains predictions of occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution + all covariate layers used for modeling. Over seven million electronic health records (EHRs), among which 11,741 EHRs reported tick attachment, were used to evaluate climate, environmental and animal host factors affecting the risk of tick attachment in cats and dogs in Great Britain (GB). The tick presence/absence EHRs for dogs and cats were further overlaid with spatiotemporal time-series of climatic, vegetation, human influence, hydrological and terrain variables (slope, wetness index) to produce a spatiotemporal regression matrix; an Ensemble Machine Learning framework was used to fine-tune hyperparameters for Random Forest (classif.ranger), Gradient boosting (classif.xgboost) and GLM-net (classif.glmnet) algorithms, which were then used to produce a final ensemble meta-learner that predicts the probability of occurrence of ticks across GB with monthly intervals.</p> <ul> <li>gb1km_covariates.zip contains ALL covariate layers as GeoTIFFs (time-series) used for modeling ticks dynamics;</li> <li>data_1km_2014_M01.rds = contains all covariates for January 2014 prepared as SpatialGridDataFrame (R data object);</li> </ul> <p>Codes of files indicate e.g.:</p> <ul> <li>"monthly.tick.prob_savsnet.mar_p_1km_s_2014_2021" = monthly occurrence probability for January based on the training data from 2014 to 2021;</li> <li>"monthly.tick.prob_savsnet.oct_md_1km_s_20211001_20211031" = monthly prediction (model) error derived as the standard deviation from multiple base learners;</li> </ul> <p>The dataset is described in detail in the following publication:</p> <ul> <li>Arsevska, E., Hengl, T., Singelton, D. et al. (2023?) <strong>Risk factors for tick attachment in companion animals in Great Britain: a spatiotemporal analysis covering 2014–2021</strong>. Submitted to Parasites & Vectors (in review).</li> </ul> <p>The model summary shows:</p> <pre><code>Call: stats::glm(formula = f, family = "binomial", data = getTaskData(.task, .subset), weights = .weights, model = FALSE) Deviance Residuals: Min 1Q Median 3Q Max -1.4749 -0.0557 -0.0471 -0.0430 3.7611 Coefficients: Estimate Std. Error z value Pr(>|z|) (Intercept) -7.64495 0.02095 -364.957 < 2e-16 *** classif.ranger 4.95061 0.63615 7.782 7.13e-15 *** classif.xgboost 189.75543 5.53109 34.307 < 2e-16 *** classif.glmnet 140.24208 5.05375 27.750 < 2e-16 *** --- Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1 (Dispersion parameter for binomial family taken to be 1) Null deviance: 170604 on 7303013 degrees of freedom Residual deviance: 162571 on 7303010 degrees of freedom AIC: 162579 Number of Fisher Scoring iterations: 9</code></pre> <p><em>Acknowledgements</em>: We are grateful to data providers in veterinary practice (VetSolutions, Teleos, CVS, and other practitioners). We are grateful to the INRAE MIGALE bioinformatics facility (MIGALE, INRAE, 2020. Migale Bioinformatics Facility, doi: <a href="https://entrepot.recherche.data.gouv.fr/dataverse/migale">10.15454/1.5572390655343293E12</a>) for providing computing resources. We are also grateful for<br> the help and support provided by <a href="https://www.liverpool.ac.uk/savsnet/">SAVSNET team members</a> Bethaney Brant, Susan Bolan and Steven Smyth.<br> This study was funded mainly by a grant from the <strong>Biotechnology and Biological Sciences Research Council</strong>,<br> BB/NO19547/1 and <strong>British Small Animal Veterinary Association</strong> (BSAVA). The research was partly funded by the National Institute for <strong>Health Research Health Protection Research Unit</strong> (NIHR HPRU) in Emerging and Zoonotic Infections at the <strong>University of Liverpool</strong> in partnership with <strong>Public Health England</strong> (PHE) and <strong>Liverpool School of Tropical Medicine</strong> (LSTM). This work has been partially funded by the <em>“Monitoring outbreak events for disease surveillance in a data science context"</em> (MOOD) project from the European Union’s Horizon 2020 research and innovation program under grant agreement No. 874850 (<a href="https://mood-h2020.eu/">https://mood-h2020.eu/</a>). The views expressed are those of the authors and not necessarily those of the NHS, the NIHR, the Department of Health or Public Health England.</p>
Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.
<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">"Predicting compound activity from phenotypic profiles and chemical structures"</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper's GitHub repository</a> for reproduction.</p> <p>Folders and files and are described below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p> </p>
Global distribution of predicted soil types at 1 km resolution based on the WRB 2022 classification
<p>Global maps at 1 km spatial resolution of the predicted soil types (0–100% probabilities) at 1 km resolution based on the <a href="https://www.fao.org/soils-portal/data-hub/soil-classification/world-reference-base/en/">WRB 2022</a> (<strong>World Reference Base</strong> the international standard for soil classification) classification system. The training data comes from the following 3 main sources:</p> <ol> <li>WOSIS points available via: <a href="https://www.isric.org/explore/wosis">https://www.isric.org/explore/wosis</a>;</li> <li>HWSD v2 (random draw of cca 20,000 points): <a href="https://iiasa.ac.at/models-tools-data/hwsd">https://iiasa.ac.at/models-tools-data/hwsd</a>;</li> <li>Other national datasets / data from publications and projects.</li> </ol> <p>Predictions are based on using Rando Forest algorithm as implemented in the <a href="https://www.randomforestsrc.org/">randomForestSRC package</a> with cca 190 covariate layers representing soil forming factors (CHELSA Climate, Global Lithological DB GLiM, MODIS EVI and LST long-term derivatives, Digital Terrain model parameters and similar).</p> <p>All TIF files are provided as <a href="https://www.cogeo.org/">COGs</a>, which means that you can open them directly in QGIS or similar. Publication explaining all modeling steps is pending.</p> <p>Update of the predictions takes about 4–5 hrs and will be regularly run provided that new training points are available. Disclaimer: These are initial results with limited accuracy and possible issues with quality of training points, location errors and harmonization issues. Use at own risk.</p> <p>Note: original list of soil types have been subset to classes that appear at least 10 times and at least in 2 countries. If you notice an error or artifact <strong>please report via <a href="https://github.com/OpenGeoHub/SoilTypeMapping">the Github repository</a></strong>. Help us improve this dataset by contributing training points.</p>
Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction [dataset]
<p>This dataset contains the extension of a publicly available dataset that was published initially by Ferenc et al. in their paper:</p> <p><em>“Ferenc, R.; Hegedus, P.; Gyimesi, P.; Antal, G.; Bán, D.; Gyimóthy, T. Challenging machine learning algorithms in predicting vulnerable javascript functions. 2019 IEEE/ACM 7th InternationalWorkshop on Realizing Artificial Intelligence Synergies in Software Engineering (RAISE). IEEE, 2019, pp. 8–14.”</em></p> <p>The dataset contained software metrics for source code functions written in JavaScript (JS) programming language. Each function was labeled as vulnerable or clean. The authors gathered vulnerabilities from publicly available vulnerability databases.</p> <p>In our paper entitled: “<strong>Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction</strong>” and cited as:</p> <p><em>“Kalouptsoglou I, Siavvas M, Kehagias D, Chatzigeorgiou A, Ampatzoglou A. Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction. Entropy. 2022; 24(5):651. <a href="https://doi.org/10.3390/e24050651">https://doi.org/10.3390/e24050651</a>”</em></p> <p>, we presented an extended version of the dataset by extracting textual features for the labeled JS functions. In particular, we got the dataset provided by Ferenc et al. in CSV format and then we gathered all the GitHub URLs of the dataset's functions (i.e., methods). Using these URLs, we collected the source code of the corresponding JS files from GitHub. Subsequently, by utilizing the start and end line information for every function, we cut off the code of the functions. Each function was then tokenized to construct a list of tokens per function.</p> <p>To extract text features, we used a text mining technique called sequences of tokens. As a result, we created a repository with all methods' source code, the token sequences of each method, and their labels. To boost the generalizability of type-specific tokens, all comments were eliminated, as well as all integers and strings, which were replaced with two unique IDs.</p> <p>The dataset contains 12,106 JavaScript functions, from which 1,493 are considered vulnerable.</p> <p>This dataset was created and utilized during the Vulnerability Prediction Task of the Horizon2020 IoTAC Project as training and evaluation data for the construction of vulnerability prediction models. The dataset is provided in the csv format. Each row of the csv file has the following parts:</p> <ul> <li>Label: Flag with values ‘1’ for vulnerable and ‘0’ for non-vulnerable methods</li> <li>Name: The name of the JavaScript method</li> <li>Longname: The longname of the JavaScript method</li> <li>Path: The path of the file of the method in the repository</li> <li>Full_repo_path: The GitHub URL of the file of the method</li> <li>TokenX: Each next row corresponds to each token included in the method</li> </ul>
Russo-Ukrainian War: Prediction and explanation of Twitter suspension
<p>The dataset utilized in the research paper: "Russo-Ukrainian War: Prediction and explanation of Twitter suspension" accepted to ASONAM 2023 conference. The provided dataset contains multiple extracted feature categories based on the Twitter dataset collected during the Russo Ukrainian War. The dataset dose not contain any private user information since user and tweet IDs are removed.</p>
LAGOS-NE Shallow Lakes: a dataset of lake variables and multi-scaled ecological context variables used to predict and compare trophic status and TP:CHLa relationships between shallow and non-shallow lakes in the Upper Midwest and Northeastern United States.
We conducted a macroscale study of 2,210 shallow lakes (mean depth ≤ 3m or a maximum depth ≤ 5m) in the Upper Midwestern and Northeastern U.S. We asked: What are the patterns and drivers of shallow lake total phosphorus (TP), chlorophyll a (CHLa), and TP–CHLa relationships at the macroscale, how do these differ from those for 4,360 non-shallow lakes, and do results differ by hydrologic connectivity class? To answer this question, we assembled the LAGOS-NE Shallow Lakes dataset described herein, a dataset derived from existing LAGOS-NE, LAGOS-DEPTH, and LAGOS-CLIMATE datasets. Response data variables were the median of available summer (e.g., 15 June to 15 September) values of total phosphorus (TP) and chlorophyll a (CHLa). Predictor variables were assembled at two spatial scales for incorporation into hierarchical models. At the local or lake-specific scale (including the individual lake, its inter-lake watershed [iws] or corresponding HU12 watershed), variables included those representing land use/cover, hydrology, climate, morphometry, and acid deposition. At the regional scale (e.g., HU4 watershed), variables included a smaller set of predictor variables for hydrology and land use/cover. The dataset also includes the unique identifier assigned by LAGOS-NE(lagoslakeid); the latitude and longitude of the study lakes; their maximum and mean depths along with a depth classification of Shallow or non-Shallow; connectivity class (i.e., whether a lake was classified as connected (with inlets and outlets) or unconnected (lacking inlets); and the zone id for the HU4 to which each lake belongs. Along with the database, we provide the R scripts for the hierarchical models predicting TP or CHLa (TPorCHL_predictive_model.R), and the TP—CHLa relationship (TP_CHL_CSI_Model.R) for depth and connectivity subsets of the study lakes.
Time series of in situ Uv-Vis absorbance spectra and high-frequency predictions of total and soluble Fe and Mn concentrations measured at multiple depths in Falling Creek Reservoir (Vinton, VA, USA) in 2020 and 2021
High-frequency measurements of light absorbance were collected at multiple depths in Falling Creek Reservoir (FCR; Vinton, VA, USA) using a s::can Spectrolyser UV-Visible spectrophotometer coupled with a multiplexor pumping system. The system pumps water samples from individual depths into a flow-through cuvette where the UV-vis absorbance spectra of the sample are measured by the spectrophotometer. The system used in our study collected measurements of light absorbance every 2.5 nm wavelengths from 200 nm to 732.5 nm (optical path length of 10 mm) approximately at an hourly time step for seven monitoring depths in the reservoir. Data was collected during two periods; the first deployment (16 October to 9 November 2020) was to observe changes in Fe and Mn concentrations before, during, and after reservoir fall turnover and the second deployment (26 May to 21 June 2021) was to observe the effects of engineered hypolimnetic oxygenation on Fe and Mn concentrations. Partial least squares regression models were developed to generate predictions of total and soluble Fe and Mn concentrations based on the correlation between absorbance spectra and sampling data.
Code for Random Forest models that predict pharmaceutical and water chemistry measurements in Baltimore Ecosystem Study streams
This file contains code to model the relationship between the water chemistry measurements and discharge measured as part of BES routine sampling and the pharmaceuticals measured in WY 2018. We use Random Forest models to predict 1) total (i.e., summed) concentration of the pharmaceuticals for which we screened, 2) total nutrient concentrations (TN & TP), 3) whether or not the antibiotic trimethoprim was detected in a given sample, and 4) whether or not nitrate and TP were above or below environmentally-relevant threshold concentrations. We also use RF models to predict N and P concentrations over a longer period, in order to compare models for nutrients to pharma. Code and analyses here rely on data processed in the file "BESPharma_WY2018.Rmd", published on EDI (doi:10.6073/pasta/610cb67fcbc8982c2af8ed946dce8ea5) and BES water chemistry data published on EDI (doi:10.6073/pasta/ce7f30e6013e003bfe28c5fd7d4aed23 )
Species cover, community biomass, and richness in global grasslands from NutNet (2007–2023): Dominant species predict plant richness and biomass in global grasslands
The Nutrient Network (NutNet) is a globally coordinated research initiative designed to investigate the impacts of human-driven alterations in nutrient availability and consumer presence on grassland ecosystems. Data were collected from over 130 herbaceous-dominated sites worldwide, spanning diverse environmental conditions from desert grasslands to arctic tundra. Standardized methodologies were employed across all sites to enable direct comparisons of productivity, diversity, and ecosystem responses. Experimental treatments included nutrient additions to assess co-limitation of plant growth by multiple nutrients, as well as grazer manipulations to examine their role in regulating biomass, species diversity, and community composition. By compiling these cross-site data, NutNet aims to enhance our understanding of productivity-diversity relationships and provide new insights into the ecological consequences of anthropogenic changes to nutrient cycles and food webs at a global scale.
Hubbard Brook Experimental Forest: Soil type prediction raster files
This dataset consists of raster files predicting spatial patterns in soils for the entire Hubbard Brook Experimental Forest. Eight soil units are used, following a hydropedologic approach, based on relationships between soil genetic horizon presence and thickness, and the frequency and depth of groundwater fluctuations. Nine raster files on a five-meter grid are presented, including one raster each showing the probability of presence of each of the eight soil units; the ninth raster represents the soil unit most likely to be present at each grid cell. The methods section of the metadata includes descriptions of the eight soil units and guidance for users of the model outputs. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
SBC LTER: Daily averages of modeled significant wave height (Hs) and peak wave period (Tp) in the Santa Barbara Coastal area from the Coastal Data Information Program - Monitoring and Prediction System (CDIP MOP)
From http://cdip.ucsb.edu: The Coastal Data Information Program (CDIP) is a research group at Scripps Institution of Oceanography that monitors coastal waves and nearshore sand levels on regional scales. CDIP maintains a network of optimally-placed, directional wave buoys from San Diego to Eureka. The buoy measurements are used to initialize a high spatial resolution (100m x 100m) linear spectral wave propagation model. The resulting hourly hindcasts and nowcasts of CA coastal wave conditions have a level of accuracy that is not possible with more traditional wind-wave generation models that are initialized with modeled wind fields.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.