Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,064
datasets available to search
ShareScore release 0.9.0
Dataset results
13,064 results for “Prediction”
Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements
The Forest and Biodiversity (FAB2) experiment uses native tree species in varying levels of species richness, phylogenetic diversity, and functional diversity planted in 100 m2 and 400 m2 plots at 1 m spacing, appropriate for testing long-term ecosystem consequences. FAB2 was designed and established in conjunction with a prior experiment (FAB1) in which the same set of twelve species was planted in 16 m2 plots at 0.5 m spacing. This data package examines the connections between aboveground and belowground processes in FAB2. This data package includes information on tree diversity and community composition, forest structure, forest understories, soil microbes, net nitrogen mineralization, and canopy nitrogen. A wide variety of data types are included, such as data from hyperspectral and LiDAR remote sensing, percent cover analysis, soil microbial analyses, and soil assays including C:N, pH, and net nitrogen mineralization. This data package is included in the submission of the manuscript entitled “Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements.”
Soil factors predict initial plant colonization on Puerto Rican landslides
Tropical storms are the principal cause of landslides in montane rainforests, such as the Luquillo Experimental Forest (LEF) of Puerto Rico . A storm in 2003 caused 30 new landslides in the LEF that we used to examine prior hypotheses that slope stability and organically enriched soils are prerequisites for plant colonization. We measured slope stability and litterfall in 1 m2 plots 8-13 months following landslide formation. At 13 months we also measured microtopography, soil characteristics (organic matter, particle size, total nitrogen, and water holding capacity), elevation, distance to forest edge, and canopy cover, as well as plant aboveground biomass, plant cover, and root biomass. Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.
LIDAR Derived Dune-Crest Elevation Values and Shrub Prediction Morphometrics for the Virginia Coast Reserve Barrier Islands: 2010 - 2017
This dataset includes LIDAR derived dune-crest elevation values for the VCR for 2010-2017. Dune-crest elevation values were sampled every 100 m from Smith to Cedar islands. The ArcGis Pro file includes the location of dune-crest transects. Additionally, this dataset includes island characteristics related to predicting shrub presence or absence for 2010, 2016, and 2017.
Human hippocampal replay during rest prioritizes weakly learned information and predicts memory performance
Open the record for dataset details and reuse information.
Associative Prediction of Visual Shape in the Hippocampus
Open the record for dataset details and reuse information.
Adaptive memory distortions are predicted by feature representations in parietal cortex
Open the record for dataset details and reuse information.
Enzymes from the BRENDA and CAZy databases annotated with organism growth temperatures and predicted Topt
<p>This repo is an updated version of repo <strong>Gang Li, & Martin KM Engqvist. (2019). Enzymes from the BRENDA database annotated with organism growth temperatures and predicted <em>T</em><sub>opt</sub> (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2539114. </strong></p> <p>Experimental as well as predicted organism growth temperatures were used to annotate enzymes from the BRENDA database (doi: 10.1093/nar/gky1048, https://www.brenda-enzymes.org) version 2018.2 (July 2018) and CAZy database (http://www.cazy.org/). </p> <p>An updated machine learning model was applied to predict the optimal functional temperature of enzymes from BRENDA and CAZy. </p> <p>There are four files in this repo:</p> <p>1. 'annotated_brenda.tsv' is a tab-seperated file that contains the annotated enzymes from BRENDA. There are 9 columns in the file: index column; "ec", EC number; "uniprot_id", protein id in Uniprot database; "domain", the domain of life (superkingdom), either Archaea, Bacteria, or Eukarya; "organism", species name; "ogt", optimal growth temperature of the organism; "ogt_note", whether the experimental or predicted ogt is used; "topt", the optimal functional temperature of the enzyme; "topt_note", whether the experimental or predicted topt is used.</p> <p>2. 'annotated_cazy.tsv' is a tab-seperated file that contains the annotated enzymes from CAZy. There are 12 columns in the file: index column; "family", CAZy family id; "genbank", genbank id; "Protein Name", the protein name from CAZy database; "ec", EC number; "organism", strain name; "uniprot_id", protein id in Uniprot database; "PDB/3D", structure id in PDB database; "ogt", optimal growth temperature of the organism; "ogt_note", whether the experimental or predicted ogt is used; "topt", the optimal functional temperature of the enzyme; "topt_note", whether the experimental or predicted topt is used.</p> <p>3. 'brenda.sql', which is a SQLite3 database version of 'annotated_brenda.tsv', with an additional column of enzyme sequences.</p> <p>4. 'cazy.sql', which is a SQLite3 database version of 'annotated_cazy.tsv'', with an additional column of enzyme sequences.</p> <p>The SQLite3 databases are for the Tome tool (<a href="https://github.com/EngqvistLab/Tome">https://github.com/EngqvistLab/Tome</a>), version 2.0.</p>
Predicting evaporation from mountain streams -- data set
<p>These files constitute the data sets used for the analysis, and generation of figures and tables reported in the manuscript titled "Predicting evaporation from mountain streams" by Andras J. Szeitz and R. Dan Moore. The manuscript was submitted for publication in the journal 'Hydrological Processes'.</p>
Predicting gene expression using morphological cell responses to nanotopography
<p>This dataset contains the raw files, results files and R workspace files (.RData) associated with the paper:</p> <p>Predicting gene expression using morphological cell responses to nanotopography</p> <p>Please note that this dataset is separated according to the Figure presented in the published and peer-reviewed version of the manuscript. Particular folders contain its own README file to facilitate reproduction/replication of results and figures. </p>
Supplementary data for "Machine learning-based prediction of activity and substrate specificity for OleA enzymes in the thiolase superfamily"
<p>Supplementary data for "Machine learning-based prediction of activity and substrate specificity for OleA enzymes in the thiolase superfamily"</p>
Prediction of repurposed drugs for treating lung injury in COVID-19
<p>These are output files of shared R scripts used in prediction of repurposed drugs for treating lung injury in COVID-19.</p> <p> </p> <p>R scripts are available here: https://doi.org/10.5281/zenodo.3822923</p> <p> </p> <p>Description of files:</p> <p>HCC515_6_data_for_drug.csv #Differential expression of genes in HCC515 cell at 6 h after treatment of ACE2 inhibitor</p> <p>HCC515_24_data_for_drug.csv #Differential expression of genes in HCC515 cell at 24 h after treatment of ACE2 inhibitor</p> <p>COVID19-Lung_data_for_drug.csv #Differential expression of genes in lung tissues with COVID-19</p> <p>HCC515_6_drug.csv #Drugs for HCC515 cell at 6 h after transfection of ACE2 inhibitor</p> <p>HCC515_24_drug.csv #Drugs for HCC515 cell at 24 h after transfection of ACE2 inhibitor</p> <p>COVID19-Lung_drug.csv #Drugs for lung tissuse from COVID-19 patients</p> <p>COL-3_single_treatment_response_data.csv #Differential expression of genes in HCC515 cell at 24h after treatment of COL-3</p> <p>CGP-60474_single_treatment_response_data.csv #Differential expression of genes in HCC515 cell at 24h after treatment of CGP-60474</p>
predicted microRNA target sites (miRanda)
<p>Target predictions based on the miRanda algorithm. The target sites are scored for likelihood of mRNA downregulation using mirSVR, a regression model that is trained on sequence and contextual features of the predicted miRNA::mRNA duplex. Expression profiles are derived from a comprehensive sequencing project of a large set of mammalian tissues and cell lines of normal and disease origin.</p> <p>This collection contains the following datasets from the August 2010 release of <a href="http://www.microrna.org">microRNA.org</a>:</p> <ul> <li>16228619 predicted microRNA target sites in 34911 distinct 3'UTR from isoforms of 19898 human genes</li> <li>7459149 predicted microRNA target sites in 28287 distinct 3'UTR from isoforms of 19231 mouse genes</li> <li>586068 predicted microRNA target sites in 6865 distinct 3'UTR from isoforms of 6256 rat genes</li> <li>345671 predicted microRNA target sites in 12285 distinct 3'UTR from isoforms of 10532 fruitfly genes</li> </ul>
Habitat suitability predictions for a boreal forest indicator species, the northern goshawk (Accipiter gentilis), in Central Finland
<p>This repository contains files that show optimal sites in Central Finland for the northern goshawk (<em>Accipiter gentilis</em>, hereafter goshawk), an indicator species of boreal forests with conservation values. The optimal sites were derived from the habitat suitability model outputs included in the following publication:</p> <p> </p> <p><strong>Björklund Heidi<sup>a</sup>, Parkkinen Anssi<sup>b</sup>, Hakkari Tomi<sup>c</sup>, Heikkinen Risto K.<sup>d</sup>, Virkkala Raimo<sup>d</sup>, Lensu Anssi<sup>b</sup> (2020): Predicting valuable forest habitats using an indicator species for biodiversity. Biological Conservation, </strong><a href="https://doi.org/10.1016/j.biocon.2020.108682">https://doi.org/10.1016/j.biocon.2020.108682</a> . </p> <p> </p> <p><sup>a</sup> Finnish Museum of Natural History Luomus, P.O. Box 17, FI-00014 University of Helsinki, Finland</p> <p><sup>b</sup> University of Jyvaskyla, Department of Biological and Environmental Science, P.O. Box 35, FI-40014 University of Jyvaskyla, Finland</p> <p><sup>c</sup> Centre for Economic Development, Transport and the Environment Central Finland, P.O. Box 250, FI-40101 Jyväskylä, Finland</p> <p><sup>d</sup> Finnish Environment Institute, Biodiversity Centre, Latokartanonkaari 11, FI-00790 Helsinki, Finland</p> <p> </p> <p>The files are ArcGIS compatible shape files which indicate the spatial location of the 160 m × 160 m grid cells which include forest stands projected to be either highly suitable or suitable as a nesting site for the goshawk in Central Finland. The habitat suitability models and values were developed across the study area using Maxent software. The files show those 160-m grid cells from the study area which were included in one of the following two categories: (i) cells deemed as the most optimal (with high probability of suitable conditions) for goshawk nesting with suitability index values in Maxent outputs varying between 0.92–1.00 (‘best’ goshawk squares), and (ii) cells deemed as ‘good’ goshawk squares (with Maxent suitability index values of ≥ 0.69 and < 0.92). The coordinate system for the data files is: ETRS-TM35FIN (EPSG: 3067) (or YKJ Finland/Finnish Uniform Coordinate System (EPSG: 2393)). </p> <p>Summarization of the key settings and elements of the study are provided below. A detailed treatment can now be found in the article published in Biological Conservation (Björklund et al.) for which the link is the following: <a href="https://doi.org/10.1016/j.biocon.2020.108682">https://doi.org/10.1016/j.biocon.2020.108682</a> .</p> <p> </p> <p><strong>Summary of the study</strong></p> <p>Intensive commercial use of boreal forests is an accelerating threat to forest biodiversity, highlighting the development of cost-effective tools to detect the locations valuable for conservation. We applied species distribution models (SDMs) in our study area, Central Finland, to locate the optimal nesting sites for the goshawk, an indicator bird species for biodiversity hotspots in mature boreal forests. The optimal sites (here, 160 x 160 m grid squares) for the goshawk were determined using the Maxent software. Optimal squares for the goshawk had forests with considerably high volumes of Norway spruce (<em>Picea abies</em>, hereafter spruce) covering only 3.4% of the boreal landscape, and they were located mostly outside protected areas. Many of the squares with optimal nesting forests appeared to be under threat due to recently intensified logging operations. Half of the squares were logged to some extent and 10% were already lost or notably deteriorated due to logging after 2015 for which our models were calibrated. Threats to biodiversity of mature boreal spruce forests are likely to accelerate with increasing logging pressures. Thus, there is an urgent need to secure the continuous supply of mature spruce forests in the landscape by developing a denser network of protected areas and applying measures that aid in sparing large entities of mature forest on privately-owned land. Our modelled optimal squares can be used for selection of potential areas with biodiversity values in conservation prioritization.</p> <p><strong>The study species</strong></p> <p>The goshawk is a raptor species which prefers mature forests for nesting in Europe. Old forests dominated by spruce are considered as important for the breeding success of the species particularly in northern latitudes. Thus, intensive forest management can impair the breeding possibilities of the goshawk, and changes in forest landscapes are likely to contribute to the decline of the species. For example, in Finland, the goshawk is classified as nearly threatened species. In our study, we used the goshawk as an indicator species to model the spatial locations of boreal forest with much potential for including biodiversity values. The indicator species status of the goshawk is based on earlier studies showing the close association of the goshawk with various taxa of mature spruce forest, as well as the reported declines of both the goshawk and associated species due to loggings.</p> <p><strong>Developing Maxent models for the goshawk</strong></p> <p>The location data on occupied nests of the goshawk gathered in spring and summer 2015 and 2016 in Central Finland – as a part of the Finnish Common Birds of Prey Monitoring – were related to a set of environmental predictor variables using a maximum entropy method, Maxent software, which is considered particularly useful for modelling presence-only data (such as our goshawk nest site data). In our case, the data on forest stand and tree characteristics were related using Maxent to the known nesting sites to predict suitable conditions for the species across the Central Finland. The forest data used in the modelling were extracted from the multi-source national forest inventory (MS-NFI) data sources governed by the Natural Resources Institute Finland. The MS-NFI data used in our modelling are based on field data of the 11th and 12th NFIs from 2009 to 2016 and satellite images from 2015 and 2016.</p> <p>Prior modelling, Pearson correlations were calculated between the continuous environmental variables at the nest sites. Of the highly (|r| ≥ 0.7) correlated variables, we chose those variables which are known to be important for the goshawk, which are useful for generalization in other areas, or whose impact was of specific interest. Our final selected set of predictor variables included one class variable, site fertility class, and nine continuous variables: growing stock volume of the spruce, pine, birches and other hardwood, canopy cover, canopy cover of broad-leaved trees, saw timber of other broad-leaved trees than birches, pulpwood volume of the birches, and the biomass of the stem residual of the spruce. The original MS-NFI data recorded at the resolution of 16 × 16 m were resampled to the resolution of 160 × 160 m for the Maxent models, to represent one potential nesting forest stand.</p> <p>The accuracy of Maxent models were assessed with cross-validation and associated averaged AUC-values. The relative importance of the variables was measured by variable contribution and model deterioration measures provided by Maxent. The cloglog-transformed output index values ranging from 0 to 1 described the relative suitability of the 160-m squares to goshawk nesting. Based on the index values, the squares were classified as ‘optimal’ (with index values of 0.69–1.00), ‘typical’ (0.46– <0.69) and ‘poor’ (<0.46). In addition, we divided optimal squares into ‘best’ goshawk squares (index values of 0.92–1.00 corresponding to a high probability of suitable conditions), and ‘good’ goshawk squares (index values ≥ 0.69 and < 0.92).</p> <p><strong>Maxent model outputs</strong></p> <p>Spruce volume was the most important variable in defining habitat suitability for goshawk nesting, but hardwood cover, other hardwood logs and site fertility class contributed also to some extent to habitat suitability. In Maxent outputs, the set of 160-m squares deemed as optimal for goshawk nesting included 6 895 (cover 0.9% of the study area) best goshawk squares and 19 421 (cover 2.5%) good goshawk squares. The projected best and good goshawk squares were mostly located in unprotected areas: 95.0% of the best and 96.0% of the good goshawk squares occurred completely outside protected areas. For further details concerning the data and the model outputs, see the referred article Björklund et al. (2020).</p> <p><strong>State of the optimal goshawk squares</strong></p> <p>In total, 11% of best and over 9% of good goshawk squares were severely altered due to recent harvesting, typically clear-cutting, of the forests during the time period between 2015 and 2019. Altogether, some level of logging occurred in 3 062 (44%) of best goshawk and 9 846 (51%) of good goshawk squares during the recent years. However, many of the squares still included enough unlogged area for the goshawk in 2019.</p> <p>In our article, we conclude that while most of the optimal squares for the goshawk were still preserved in 2019, they are under risk as they are mainly situated outside protected area network. This stresses the importance of conserving biodiversity with complementary measures in privately-owned managed forests. In conclusion, a denser network with more PAs for forest-dwelling species should be secured in areas with intensive forestry, e.g. in southern Finland where PAs currently cover a smaller proportion of land compared to northern Finland.</p>
The datasets used in the manuscript named "Dynamical Seasonal Prediction of Tropical Cyclone Activity Using a Global Ensemble Prediction System FGOALS-f2 V1.0"
<p>The hindcast and real-time prediction output of FGOALS-f2 V1.0 used in the study named "Dynamical Seasonal Prediction of Tropical Cyclone Activity Using a Global Ensemble Prediction System FGOALS-f2 V1.0"</p>
RNA sequencing dataset for prediction of liver hepatocellular carcinoma using SIMON analysis
<p>The LIHC dataset was used for data mining and for the generation of machine learning model for the detection of liver hepatocellular carcinoma cells (LIHC) using the SIMON platform as described in the "SIMON: open-source knowledge discovery platform" publication (<a href="https://doi.org/10.1101/2020.08.16.252767">https://doi.org/10.1101/2020.08.16.252767</a>). The LIHC dataset was obtained from the <em>GSEABenchmarkeR</em> package ( <a href="https://doi.org/10.1093/bib/bbz158">https://doi.org/10.1093/bib/bbz158</a>) and it contains RNA expression data from 374 liver hepatocellular carcinoma (LIHC) cells and 50 adjacent normal cells.</p>
Data and Code Accompanying the Study on "Benefits of reflex prediction: A case study of Western Kho-Bwa"
<p>Cite the source of the dataset as:</p> <blockquote> <p>Timotheus A. Bodt and Johann-Mattis List (to appear): Benefits of reflex prediction: A case study of Western Kho-Bwa. Diachronica.</p> </blockquote>
Associated Data: RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features
<p>Additional digital data to "RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features" (ChemRxiv preprint:<a href="https://doi.org/10.26434/chemrxiv.12636704.v1">https://doi.org/10.26434/chemrxiv.12636704</a>).</p> <p>Associated code can be found at: <a href="https://github.com/HITS-MCM/RASPDplus">https://github.com/HITS-MCM/RASPDplus</a></p> <p>Files:</p> <ul> <li>weights.tar.gz: contains the model weights of one random dataset split and its associated crossvalidation folds. Used for standard RASPD+ evaluation.</li> <li>additional_model_replicates.tar.gz: contains the remaining models trained on the full set of descriptors.</li> <li>external_test_sets.tar.gz: contains the descriptor tables for all external test sets used</li> <li>dude.tar.gz: contains the descriptor tables for and several identifier lists for evaluation on the Directory of Useful Decoys - Enhanced (DUD-E)</li> <li>run_outputs.tar.gz: Performance metric data and predicted values created during the model training and evaluation runs. Basis for the figures and metrics in the manuscript.</li> </ul> <p> </p>
Soil organic carbon stocks and trends (1984-2019) predicted at 30m spatial resolution for topsoil in natural areas of South Africa
<p>Link to scientific publication: <a href="https://doi.org/10.1016/j.scitotenv.2021.145384">https://doi.org/10.1016/j.scitotenv.2021.145384</a></p> <p>Soil organic carbon (SOC) stocks (kg C m-2) are predicted over natural areas (excluding water, urban, and cultivated) of South Africa using a machine learning workflow driven by optical satellite data and other ancillary climatic, morphometric and biological covariates. The temporal scope covers 1984-2019. The spatial scope covers 0-30cm topsoil in South Africa natural land area (84% of the country). See methodology in linked publication for details. Data are provided here at 30m spatial resolution in GeoTIFF files. There is a dataset for the long-term average SOC and trend in SOC. Each dataset is split into four files (suffix *_1, *_2 etc.) covering separate regions of South Africa for ease of download. The raster files are:</p> <ul> <li>"SOC_mean_30m..." - average of annual SOC predictions between 1984 and 2019. Values are expressed in kg C m-2</li> <li>"SOC_trend_30m..." - long-term trend in SOC derived from the Sens slope (M) across annual SOC values between 1984 and 2019. Pixel values (Y) are expressed as a percentage change over the 35 years relative to the long-term mean (X). Y = M / X * 100 * 35 years</li> </ul> <p>NB: All files are scaled by *100 and converted to floating data point to save space. To back-convert to original values, simply divide the raster values by 100.</p>
Supporting data for 'Hourly prediction of phytoplankton biomass and its environmental controls in lowland rivers'
<p>This dataset is used in the manuscript 'Hourly prediction of phytoplankton biomass and its environmental controls in lowland rivers' published in Water Resources Research. The dataset contains hourly observation of water quality in the lower Thames catchment, UK and were made available by the Environment Agency, UK. </p>
RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)
<p>RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.