Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: A new theory of plant-microbe nutrient competition resolves inconsistencies between observations and model predictions

Terrestrial plants assimilate anthropogenic CO2 through photosynthesis and synthesizing new tissues. However, sustaining these processes requires plants to compete with microbes for soil nutrients, which therefore calls for an appropriate understanding and modeling of nutrient competition mechanisms in Earth System Models (ESMs). Here, we survey existing plant-microbe competition theories and their implementations in Earth System Models (ESMs). We found no consensus regarding the representation of nutrient competition and that observational and theoretical support for current implementations are weak. To reconcile this situation, we applied the Equilibrium Chemistry Approximation (ECA) theory to plant-microbe nitrogen competition in a detailed grassland 15N tracer study and found that competition theories in current ESMs fail to capture observed patterns and the ECA prediction simplifies the complex nature of nutrient competition and quantitatively matches the 15N observations. Since plant carbon dynamics are strongly modulated by soil nutrient acquisition, we conclude that (1) predicted nutrient limitation effects on terrestrial carbon accumulation by existing ESMs may be biased and (2) our ECA-based approach may improve predictions by mechanistically representing plant-microbe nutrient competition.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Are pollination "syndromes" predictive? Asian dalechampia fit neotropical models

Using pollination-syndrome parameters and pollinator correlations with floral phenotype from the neotropics, we predicted that Dalechampia bidentata Blume (Euphorbiaceae) in southern China would be pollinated by female resin-collecting bees between 12 and 20 mm in length. Observations in southwestern Yunnan Province, China, revealed pollination by resin-collecting, female Megachile (Callomegachile) nr. faceta Bingham (Hymenoptera: Megachilidae). These bees, at 14 mm in length, were in the predicted size range, confirming the utility of syndromes and models developed in distant regions. Phenotypic-selection analyses and estimation of adaptive surfaces and adaptive accuracies together suggest that the blossoms of D. bidentata are well adapted to pollination by their most common floral visitors.

opencc-zeroDec 2010View details →
dryad32/100

Data from: A framework for developing ecosystem-specific nutrient criteria: integrating biological thresholds with predictive modeling

We present a novel ecosystem-specific framework for developing nutrient criteria from biological thresholds and predictive modeling (BTPM) and an application of this framework to lakes in Michigan, U.S. The four main components for the BTPM framework are: (1) to predict each ecosystem s expected nutrient concentration in the absence of human effects using a predictive model, (2) to identify important biological thresholds along a nutrient gradient (i.e., biological [BIO] benchmarks), (3) to determine each ecosystem s current nutrient concentration, and (4) to use the above information to derive a nutrient criterion for each ecosystem using the BTPM algorithm. The BTPM framework is extremely flexible in that it can be applied to any aquatic ecosystem type or nutrient and the four components can be implemented in a variety of ways. Our BTPM framework has two additional features: it recognizes that prior to human disturbance, ecosystems varied in their natural nutrient concentrations, and it incorporates risk into the decision-making process. In the simplest scheme, a nutrient criterion is set at a BIO benchmark greater than the expected nutrient concentration. However, to protect ecosystems more conservatively, a criterion is set at current lake nutrient concentrations if current is less than the BIO benchmark. In our application of the BTPM framework, we developed total phosphorus (TP) criteria for a diverse set of 374 lakes in MI. The expected lake TP concentrations in the absence of human effects ranged from 3 µg L-1 to 24 µg L-1, suggesting that a single criterion approach would not be appropriate.We found two predominant benchmarks in the biological data along the TP gradient, one for zooplankton metrics at 8 µg L-1, and one for phytoplankton metrics at 18 µg L-1. We present the sequence of analyses and decisions that could be used to apply this approach in a management context using Michigan lakes as an example.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Testing species assemblage predictions from stacked and joint species distribution models

Aim: Predicting the spatial distribution of species assemblages remains an important challenge in biogeography. Recently, it has been proposed to extend correlative species distribution models (SDMs) by taking into account (a) covariance between species occurrences in so-called joint species distribution models (JSDMs) and (b) ecological assembly rules within the SESAM (spatially explicit species assemblage modelling) framework. Yet, little guidance exists on how these approaches could be combined. We, thus, aim to compare the accuracy of assemblage predictions derived from stacked and from joint SDMs. Location: Switzerland Taxon Birds, tree species Methods: Based on two monitoring schemes (national forest inventory and Swiss breeding bird atlas), we built SDMs and JSDMs for tree species (at 100m resolution) and forest birds (at 1km resolution). We tested accuracy of species assemblage and richness predictions on holdout data using different stacking procedures and ecological assembly rules. Results Despite minor differences, results were consistent between birds and tree species. Cross-validated species-level model performance was generally higher in SDMs than JSDMs. Differences in species richness and assemblage predictions were larger between stacking procedures and ecological assembly rules than between stacked SDMs and JSDMs. On average, predictions were slightly better for stacked SDMs compared to JSDMs, probabilistic stacks outperformed binary stacks, and ecological assembly rules yielded best predictions. Main conclusions: When predicting the composition of species assemblages, the choice of stacking procedure and ecological assembly rule seems more decisive than differences in underlying model type (SDM vs. JSDM). JSDMs do not seem to improve community predictions compared to SDMs or improve predictions for rare species. Still, JSDMs may provide additional insights into community assembly and may help deriving hypotheses about prevailing biotic interactions in the system. We provide simple rules of thumb for choosing appropriate modelling pathways. Future studies should test these preliminary guidelines for other taxa and biogeographic realms as well as for other JSDM algorithms.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Predicting species occurrences with habitat network models

1. Biodiversity conservation requires modelling tools capable of predicting the presence or absence (i.e. occurrence-state) of species in habitat patches. Local habitat characteristics of a patch (lh), the cost of traversing the landscape matrix between patches (weighted connectivity; (wc), and the position of the patch in the habitat network topology (nt) all influence occurrence-state. Existing models are data demanding or consider only local habitat characteristics. We address these shortcomings and present a network-based modelling approach, which aims to predict species occurrence-state in habitat patches using readily available presence-only records. 2. For the tree frog Hyla arborea on the Swiss Plateau, we delineated habitat network nodes from an ensemble habitat suitability model, and used different cost surfaces to generate the edges of three networks: one limited only by dispersal distance (Uniform), another incorporating traffic, and a third based on inverse habitat suitability. For each network, we calculated explanatory variables representing the three categories (lh, wc and nt). The response variable, occurrence-state, was parametrized by a sampling-intensity procedure assessing observations of comparable species over a threshold of patch visits. The explanatory variables from the three networks and an additional non-topological model were related to the response variable with boosted regression trees. 3. The habitat network models had a similar fit; they all outperformed the non-topological model. Habitat suitability index ((lh) was the most important predictor in all networks, followed by third-order neighborhood (nt). Patch size (lh) was unimportant in all three networks. 4. We found that topological variables of habitat networks are relevant for the prediction of species occurrence-state, a step-forward from models considering only local habitat characteristics. For any habitat patch, occurrence-state is most prominently influenced by its habitat suitability, and then by the number of patches in a wide neighborhood. Our approach is generic and can be applied to multiple species in different habitats.

opencc-zeroSep 2019View details →
dryad32/100

Data from: Ecological niche modeling as a tool for prediction of the potential geographic distribution of Bacillus anthracis spores in Tanzania

Introduction: Anthrax is caused by the spore-forming, Gram-positive bacterium Bacillus anthracis. The aim of this study was to predict the potential distribution of B. anthracis in Tanzania and produce epidemiological evidence for the management of anthrax outbreaks in the country. Methods: The Maxent algorithm was used to predict areas at risk of anthrax outbreaks based on the occurrence and environmental data in Arusha and Kilimanjaro regions; the model was later transferred to predict the entire country. Seventy percent of the occurrence data were used to train the model, while 30% were used for model evaluation. Results: Four regions of northern Tanzania are predicted to have a high risk for anthrax outbreaks, while the southern and western regions had low-risk areas. Soil type (56.5%), soil pH (23.7%), and isothermally (10.4%) were the most important variables for the model prediction, and the most significant soil types were solonetz, fluvisols, and lithosols. Conclusions: A strong risk level across districts of the Tanzania mainland was identified in this study. A total of 18 districts in Tanzania Mainland are predicted to be at very high risk of an anthrax outbreak occurrence. These findings are important for policymakers to effectively mount targeted control measures for anthrax outbreaks in Tanzania.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Historical species distribution models predict species limits in western Plethodon salamanders

Allopatry is commonly used to predict boundaries in species delimitation investigations under the assumption that currently allopatric distributions are indicative of reproductive isolation; however, species ranges are known to change over time. Incorporating a temporal perspective of geographic distributions should improve species delimitation; to explore this, we investigate three species of western Plethodon salamanders that have shifted their ranges since the end of the Pleistocene. We generate species distribution models (SDM) of the current range, hindcast these models onto a climatic model 21 Ka, and use three molecular approaches to delimit species in an integrated fashion. In contrast to expectations based on the current distribution, we detect no independent lineages in species with allopatric and patchy distributions (Plethodon vandykei and Plethodon larselli). The SDMs indicate that probable habitat is more expansive than their current range, especially during the last glacial maximum (LGM) (21 Ka). However, with a contiguous distribution, two independent lineages were detected in Plethodon idahoensis, possibly due to isolation in multiple glacial refugia. Results indicate that historical SDMs are a better predictor of species boundaries than current distributions, and strongly imply that researchers should incorporate SDM and hindcasting into their investigations and the development of species hypotheses.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Seasonal difference in temporal transferability of an ecological model: near-term predictions of lemming outbreak abundances

Ecological models have been criticized for a lack of validation of their temporal transferability. Here we answer this call by investigating the temporal transferability of a dynamic state-space model developed to estimate season-dependent biotic and climatic predictors of spatial variability in outbreak abundance of the Norwegian lemming. Modelled summer and winter dynamics parametrized by spatial trapping data from one cyclic outbreak were validated with data from a subsequent outbreak. There was a distinct difference in model transferability between seasons. Summer dynamics had good temporal transferability, displaying ecological models' potential to be temporally transferable. However, the winter dynamics transferred poorly. This discrepancy is likely due to a temporal inconsistency in the ability of the climate predictor (i.e. elevation) to reflect the winter conditions affecting lemmings both directly and indirectly. We conclude that there is an urgent need for data and models that yield better predictions of winter processes, in particular in face of the expected rapid climate change in the Arctic.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Comparing the prediction of joint species distribution models with respect to characteristics of sampling data

Biotic interactions have been rarely included in traditional species distribution models, wherein Joint Species Distribution Models (JSDMs) emerge as a feasible approach to incorporate environmental factors and interspecific interactions simultaneously, making it a powerful tool for analyzing the structure and assembly processes of biotic communities. However, the predictability and statistical robustness of JSDMs are largely unknown because of the lack of research efforts for those newly developed models. This study systematically evaluated the performances of five JSDMs in predicting the occurrence and biomass of multiple species, with a particular focus on diverse characteristics of sampling data, including type of response variables, number of sampling sites, and the number of species included in models. In general, most models yielded satisfactory performances on fitting to observed data and on the estimation of environmental effects; however, they showed less well performances in evaluating species associations, and their predictability had large variations. The JSDMs showed inconsistent performances between the goodness-of-fit and predictability in cross-validation, and the Boral model was relatively robust than others. The predictability of JSDMs was less influenced by sample sizes and substantially improved by incorporating rare species. This study contributes to an appropriate model selection and application of JSDMs.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Developing allometric models to predict the individual aboveground biomass of shrubs worldwide

Aim Existing global models to predict standing woody biomass are based on trees characterized by a single principal stem, well-developed in height. However, their use in open woodlands and shrublands, characterized by multistemmed species with substantial crown development, generates a high level of uncertainty in biomass estimates. This limitation led us to i) develop global predictive models of shrub individual aboveground biomass based on simple allometric variables; ii) to compare the fit of these models with existing global biomass models; and iii) to assess whether models fit change when bioclimatic variables are considered. Location Global. Time period Present. Major taxa studied 118 species. Methods We compile a database of 3243 individuals across 49 sites distributed worldwide. Including basal diameter, height and crown diameter as predictor variables, we built potential models and compared their fit using generalized least squares. We used mixed effects models to determine if bioclimatic variables improved the accuracy of biomass models. Results Although the most important variable in terms of predictive capacity was stem basal diameter, crown diameter significantly improved the models fit, followed by height. Four models were finally chosen, with the best model combining all these variables in the same equation (R2 = 0.930, RMSE = 0.476). Selected models performed as well as established global biomass models. Including the individual bioform significantly improved the models fit. Main conclusions Basal diameter, crown diameter and height measures could be combined to provide robust AGB estimates of individual shrub species. Our study supplements well-established models developed for trees, allowing more accurate biomass estimation of multistemmed woody individuals. We further provide tools for a methodological standardization of individual biomass quantification in these species. We expect these results contribute to improve the quality of biomass estimates across ecosystems, but also to generate methodological consensus on field biomass assessments in shrubs.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Detection error influences both temporal seroprevalence predictions and risk factors associations in wildlife disease models

Understanding the prevalence of pathogens in invasive species is essential to guide efforts to prevent transmission to agricultural animals, wildlife, and humans. Pathogen prevalence can be difficult to estimate for wild species due to imperfect sampling and testing (pathogens may not be detected in infected individuals and erroneously detected in individuals that are not infected). The invasive wild pig (Sus scrofa, also referred to as wild boar and feral swine) is one of the most widespread hosts of domestic animal and human pathogens in North America. We developed hierarchical Bayesian models that account for imperfect detection to estimate the seroprevalence of five pathogens (porcine reproductive and respiratory syndrome virus, pseudorabies virus, Influenza A virus in swine, Hepatitis E virus, and Brucella spp.) in wild pigs in the United States using a dataset of over 50,000 samples across nine years. To assess the effect of incorporating detection error in models, we also evaluated models that ignored detection error. Both sets of models included effects of demographic parameters on seroprevalence. We compared our predictions of seroprevalence to 40 published studies, only one of which accounted for imperfect detection. We found a range of seroprevalence among the pathogens with a high seroprevalence of pseudorabies virus, indicating significant risk to livestock and wildlife. Demographics had mostly weak effects, indicating that other variables may have greater effects in predicting seroprevalence. Models that ignored detection error led to different predictions of seroprevalence as well as different inferences on the effects of demographic parameters. Our results highlight the importance of incorporating detection error in models of seroprevalence and demonstrate that ignoring such error may lead to erroneous conclusions about the risk associated with pathogen transmission. When using opportunistic sampling data to model seroprevalence and evaluate risk factors, detection error should be included.

opencc-zeroAug 2019View details →
dryad32/100

Data from: Ocean circulation model predicts high genetic structure in a long-lived pelagic developer

Understanding the movement of genes and individuals across marine seascapes is a long-standing challenge in marine ecology, and can inform our understanding of local adaptation, the persistence and movement of populations, and the spatial scale of effective management. Patterns of gene flow in the ocean are often inferred based on population genetic analyses coupled with knowledge of species' dispersive life histories. However, genetic structure is the result of time-integrated processes, and may not capture present-day connectivity between populations. Here we use a high-resolution oceanographic circulation model to predict larval dispersal along the complex coastline of western Canada that includes the transition between two well-studied zoogeographic provinces. We simulate dispersal in a benthic sea star with a 6-10 week pelagic larval phase, and test predictions of this model against previously observed genetic structure including a strong phylogeographic break within the zoogeographical transition zone. We also test predictions with new genetic sampling in a site within the phylogeographic break. We find that the coupled genetic and circulation model predicts the high degree of genetic structure observed in this species, despite its long pelagic duration. High genetic structure on this complex coastline can thus be explained through ocean circulation patterns which tend to retain passive larvae within 20 - 50 km of their parents, suggesting a necessity for close-knit design of Marine Protected Area networks.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models

Genomic selection have been proposed as the standard method to predict breeding values in animal and plant breeding. Although some crops have benefited from this methodology, studies in Coffea are still emerging. To date, there have been no studies of how well genomic prediction models work across populations and environments for different complex traits in coffee. Considering that predictive models are based on biological and statistical assumptions, it is expected that their performance vary depending on how well these assumptions align with the true genetic architecture of the phenotype. To investigate this, we used data from two recurrent selection populations of Coffea canephora, evaluated in two locations, and single nucleotide polymorphisms identified by Genotyping-by-Sequencing. In particular, we evaluated the performance of 13 statistical approaches to predict three important traits in the coffee — production of coffee beans, leaf rust incidence and yield of green beans. Analyses were performed for predictions within-environment, across locations and across populations to assess the reliability of genomic selection. Overall, differences in the prediction accuracy of the competing models were small, although the Bayesian methods showed a modest improvement over other methods, at the cost of more computation time. As expected, predictive accuracy for within-environment analysis, on average, were higher than predictions across locations and across populations. Our results support the potential of genomic selection to reshape traditional plant breeding schemes. In practice, we expect to increase the genetic gain per unit of time by reducing the length cycle of recurrent selection in coffee.

opencc-zeroDec 2017View details →
zenodo32/100

FIGURE 9. Maximum entropy model developed for D in A new species of Desmopachria Babington (Coleoptera: Dytiscidae) from Cuba with a prediction of its geographic distribution and notes on other Cuban species of the genus

FIGURE 9. Maximum entropy model developed for D. andreae sp. n. in Cuba. Values range from high (red areas) to low environmental suitability (blue areas).

opennotspecifiedDec 2014View details →
zenodo32/100

Model for Predicting the Hydraulic Conductivity of Frozen Soils Using the Soil Freezing Characteristic Curve

<p>This is the data used in this manuscript.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Supplemental Materials - Performance Figures for "A Model for Predicting the (re)-occurrence of a ≥40% eGFR Decline in a large Population-based cohort of Persons with or At-Risk of Chronic Kidney Disease " paper

<p>The zip file contains performance metrics figures for each dynamic Bayesian Network (DBN) model, stratified by comorbidities, race, CKD stages, and ethnicity.</p> <p>Contains:</p> <ul> <li>Stratified: Bootstrapping of 1000 iterations and 1000 samples with stratified proportions (as in the original population of the test set) of rapid eGFR decliners and non-decliners.</li> </ul> <p>&nbsp;</p> <p>Second zip contains DBN structures as matrices for 2 periods study entry to entry period and entry period to year 1 for all sites in 2 excel files.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Challenging Bug Prediction and Repair Models with Synthetic Bugs

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Predicting p53-Dependent Cell Transitions From Thermodynamic Models

<p>Matlab codes containing Grand Canonical Monte Carlo (GCMC) simulation codes and theoretical codes to study the first-order phase transition behavior of a malignant cell to a normal healthy cell. In this work, we develop thermodynamic models to determine the fate of a malignant cell as governed by the tumor suppressor p53 signaling network.</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

A predictive model of UV-A-riboflavin crosslinking treatment on porcine corneas

<p>The crosslinking technique (CXL) is an effective low-risk therapeutic treatment of keratoconus and other ectatic disorders of the human cornea. The effect of corneal CXL is to increase the stiffness of the stroma to prevent the progression of the cornea distortion. Several clinical and experimental studies have shown that the stiffening effects predominantly localise on the anterior portion of the stroma and that the in-depth stiffening distribution is highly dependent on the duration of treatment. Yet, how the stiffening effects distribute through the cornea thickness as a function of the treatment duration is an open question. Here we propose an analytical model of the stiffening profile due to CXL-treatment as a function of the irradiation time. We consider linear and nonlinear variations of the crosslinking effects across the thickness and implement them into a finite element model of the porcine cornea. We present a time-dependent in-depth stiffening profile that allows us to predict the post-operative cornea response to physiological intraocular pressure for different irradiation times. We anticipate that this predictive model will support the development of patient specific 3D models that will allow clinicians to design customised CXL treatment, thus enhancing treatment outcomes.&nbsp;</p>

openmit-licenseOct 2023View details →
zenodo32/100

Dataset for "McSnow 2.0: Explicit habit-prediction in a Lagrangian super-particle ice microphysics model"

<p>Data and plot scripts to create figures included in the paper draft &quot;McSnow 2.0: Explicit habit-prediction in a Lagrangian super-particle ice microphysics model&quot; (submitted to JAMES).</p> <p>This work has been funded by the German Science Foundation (DFG) under grant SE 1784/3-1, project ID 408011764 as part of the DFG priority program SPP 2115 on radar polarimetry.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record