Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
TECPR2 protein models predicted by three algorithms
<p>This ZIP-file contains the files used for TECPR2 protein modeling and resulting PDB (Protein Data Bank) files from three algorithms/pipelines (GalaxyWEB, trRosetta, SWISS-MODEL) used for clustering analysis and visualization in Neuser et al. ("Clinical, neuroimaging and molecular spectrum of <em>TECPR2-</em>associated hereditary sensory and autonomic neuropathy with intellectual disability").<br> We always used standard parameters and the respective top model ("model_1 | model01 | model1") for each algorithm/pipeline and/or downstream steps.</p>
Predictive multi-scale occupancy models at range-wide extents: effects of habitat and human disturbance on distributions of wetland birds
<p><span><i>Aim:</i> Predicting distributions is fundamental to ecology, yet hindered by spatially-restricted sampling, scale-dependent relationships, and detection error associated with field surveys. Predictive species distribution models (SDMs) are nonetheless vital for conservation of many species. We developed a framework for building predictive SDMs with multi-scale data, and used it to develop range-wide breeding-season SDMs for 14 marsh bird species of concern.</span></p> <p><span><i>Location: </i>USA.</span></p> <p><span><i>Methods: </i>We built SDMs using data from range-wide surveys conducted over 14 years, and habitat and disturbance covariates measured at multiple spatial scales. We built hierarchical occupancy models that included heterogeneity in detectability during sampling, and used Bayesian model selection to regulate model complexity (covariates and scales) based explicitly on spatial predictive abilities. We thus integrated model selection for optimizing out-of-sample prediction, range-wide sampling over broad conditions, multi-scale analyses and scale-optimization, and species-specific detectability for a suite of wide-ranging species. </span></p> <p><span><i>Results: </i>Distributions of marsh birds were affected by local wetland conditions, but also by agricultural, urban, and hydrologic disturbances operating from local scales (100 – 500 m) to the watershed level. Variables measuring human disturbances improved prediction for most species, and every species was affected by attributes at > 1 scale. Five species showed evidence for continental-scale range contraction during the study.</span></p> <p><span><i>Main conclusions: </i>We demonstrate how hierarchical occupancy models can be optimized for prediction across a species' range at the extent of a continent while also accounting for imperfect detection, and thus describe a generalizable approach that can be used for any species. We provide the first data-driven, empirical SDMs built at the range-wide extent for most of our 14 study species and demonstrate that previous studies focused on local distributions and the effects of fine-scale wetland vegetation missed important broad-scale drivers of occupancy for marsh birds. </span></p>
Using predictive models to evaluate the quality of a test suite at class and method level.
<p>Vídeo de apresentação para o Workshop de Teses e Dissertações do CBSoft (WTDSoft).</p>
Data from: A predictive model for improving placement of wind turbines to minimise collision risk potential for a large soaring raptor
<p><span><span><span><span><span><span><span><span><span><span><span>1. With the rapid growth of wind energy developments worldwide, it is critical that the negative impacts on wildlife are considered and mitigated. This includes minimising the numbers of large soaring raptors which are killed when they collide with wind turbines.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>2. To reduce the likelihood of raptor collisions, turbines should be placed at locations which are least used by sensitive species. For resident or breeding species, this is often delineated crudely through the use of circular buffers centred on nest sites, which assume uniform habitat use around a nest site.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>3. Using GPS tracking data together with a digital elevation model we build and cross-validate a simple generalizable model, to classify the spatial likelihood of wind turbine collisions for resident adult Verreaux's eagles in any landscape where there are known nests. We apply our methods to operational developments in South Africa to validate the model and demonstrate its ability in predicting actual collision mortalities.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>4. Our Collision Risk Potential (CRP) model included the variables distance to nest, distance to conspecific nest, slope, distance to slope and elevation. Using our model, rather than a circular buffer, resulted in ca. 4–5% improvement in eagle protection while excluding development from the same amount (but not shape) of area. For an equal level of eagle protection, our model can make ca. 20–21% more area available for wind energy development compared to a circular buffer.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>5. Exploring collisions at operational wind farms in South Africa we show that our CRP model correctly predicted 87% of known collisions, while circular buffers (5.2km radius) only captured 50% of collisions.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>6. <i>Synthesis and applications</i>: We show that by using predictive models to account for habitat use, a greater area of land can be made available for wind energy development without increased mortality risk to raptors. Our predictive model can be used to provide robust guidance on wind turbine placement in South Africa in a way which minimizes the conflict between a vulnerable raptor species and the development of renewable energy. </span></span></span></span></span></span></span></span></span></span></span></p>
Exploring the predictive value of lesion topology on motor function outcomes in a porcine ischemic stroke model
<p>Dataset to accompany manuscript currently in review at Scientific Reports (12/17/2020)</p>
Supplementary material 2 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049
Table S2. Basic information obtained for 49 exotic plants in Chile
Supplementary material 3 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049
Map of the species
Data: Experimental evidence of warming-induced disease emergence and its prediction by a trait-based mechanistic model
<p>Predicting the effects of seasonality and climate change on the emergence and spread of infectious disease remains difficult, in part because of poorly understood connections between warming and the mechanisms driving disease. Trait-based mechanistic models combined with thermal performance curves arising from the Metabolic Theory of Ecology (MTE) have been highlighted as a promising approach going forward; however, this framework has not been tested under controlled experimental conditions that isolate the role of gradual temporal warming on disease dynamics and emergence. Here, we provide experimental evidence that a slowly warming host – parasite system can be pushed through a critical transition into an epidemic state. We then show that a trait-based mechanistic model with MTE functional forms can predict the critical temperature for disease emergence, subsequent disease dynamics through time, and final infection prevalence in an experimentally warmed system of <i>Daphnia </i>and a microsporidian parasite. Our results serve as a proof of principle that trait-based mechanistic models using MTE sub-functions can predict warming-induced disease emergence in data-rich systems – a critical step towards generalizing the approach to other systems.</p>
Data from: Evaluating predictive performance of statistical models explaining wild bee abundance in a mass-flowering crop
<p>Wild bee populations are threatened by current agricultural practices in many parts of the world, which may put pollination services and crop yields at risk. Loss of pollination services can potentially be predicted by models that link bee abundances with landscape-scale land-use, but there is little knowledge on the degree to which these statistical models are transferable across time and space. This study assesses the transferability of models for wild bee abundance in a mass-flowering crop across space (from one region to another) and across time (from one year to another). The models used existing data on bumblebee and solitary bee abundance in winter oilseed rape fields, together with high-resolution land-use crop-cover and semi-natural habitats data, from studies conducted in five different regions located in four countries (Sweden, Germany, Netherlands, and the UK), in three different years (2011, 2012, 2013). We developed a hierarchical model combining all studies and evaluated the transferability using cross-validation. We found that both the landscape-scale cover of mass-flowering crops and permanent semi-natural habitats, including grasslands and forests, are important drivers of wild bee abundance in all regions. However, while the negative effect of increasing mass-flowering crops on the density of the pollinators is consistent between studies, the direction of the effect of semi-natural habitat is variable between studies. The transferability of these statistical models is limited, especially across regions, but also across time. Our study demonstrates the limits of using statistical models in conjunction with widely available land-use crop-cover classes for extrapolating pollinator density across years and regions, likely in part because input variables such as cover of semi-natural habitats poorly capture variability in pollinator resources between regions and years.</p>
Data from: Joint prediction of multiple quantitative traits using a Bayesian multivariate antedependence model
Predicting organismal phenotypes from genotype data is important for preventive and personalized medicine as well as plant and animal breeding. Although genome-wide association studies (GWAS) for complex traits have discovered a large number of trait- and disease-associated variants, phenotype prediction based on associated variants is usually in low accuracy even for a high-heritability trait because these variants can typically account for a limited fraction of total genetic variance. In comparison with GWAS, the whole-genome prediction (WGP) methods can increase prediction accuracy by making use of a huge number of variants simultaneously. Among various statistical methods for WGP, multiple-trait model and antedependence model show their respective advantages. To take advantage of both strategies within a unified framework, we proposed a novel multivariate antedependence-based method for joint prediction of multiple quantitative traits using a Bayesian algorithm via modeling a linear relationship of effect vector between each pair of adjacent markers. Through both simulation and real-data analyses, our studies demonstrated that the proposed antedependence-based multiple-trait WGP method is more accurate and robust than corresponding traditional counterparts (Bayes A and multi-trait Bayes A) under various scenarios. Our method can be readily extended to deal with missing phenotypes and resequence data with rare variants, offering a feasible way to jointly predict phenotypes for multiple complex traits in human genetic epidemiology as well as plant and livestock breeding.
Data from: Modelled three-dimensional suction accuracy predicts prey capture success in three species of centrarchid fishes
Prey capture is critical for survival, and differences in correctly positioning and timing a strike (accuracy) are likely related to variation in capture success. However, an ability to quantify accuracy under natural conditions, particularly for fishes, is lacking. We developed a predictive model of suction hydrodynamics and applied it to natural behaviours using three-dimensional kinematics of three centrarchid fishes capturing evasive and non-evasive prey. A spheroid ingested volume of water (IVW) with dimensions predicted by peak gape and ram speed was verified with known hydrodynamics for two species. Differences in capture success occurred primarily with evasive prey (64–96% success). Micropterus salmoides had the greatest ram and gape when capturing evasive prey, resulting in the largest and most elongate IVW. Accuracy predicted capture success, although other factors may also be important. The lower accuracy previously observed in M. salmoides was not replicated, but this is likely due to more natural conditions in our study. Additionally, we discuss the role of modulation and integrated behaviours in shaping the IVW and determining accuracy. With our model, accuracy is a more accessible performance measure for suction-feeding fishes, which can be used to explore macroevolutionary patterns of prey capture evolution.
Data from: A national-scale model of linear features improves predictions of farmland biodiversity
1. Modelling species distribution and abundance is important for many conservation applications, but it is typically performed using relatively coarse-scale environmental variables such as the area of broad land-cover types. Fine-scale environmental data capturing the most biologically-relevant variables have the potential to improve these models. For example, field studies have demonstrated the importance of linear features, such as hedgerows, for multiple taxa, but the absence of large-scale datasets of their extent prevents their inclusion in large-scale modelling studies. 2. We assessed whether a novel spatial dataset mapping linear and woody linear features across the UK improves the performance of abundance models of 18 bird and 24 butterfly species across 3723 and 1547 UK monitoring sites respectively. 3. Although improvements in explanatory power were small, the inclusion of linear features data significantly improved model predictive performance for many species. For some species, the importance of linear features depended on landscape context, with greater importance in agricultural areas. 4. Synthesis and applications. This study demonstrates that a national-scale model of the extent and distribution of linear features improves predictions of farmland biodiversity. The ability to model spatial variability in the role of linear features will be important in targeting agri-environment schemes to maximally deliver biodiversity benefits. Although this study focuses on farmland, data on the extent of different linear features are likely to improve species distribution and abundance models in a wide range of systems, and also can potentially be used to assess habitat connectivity. 10-Mar-2017
Data from: High invasion potential of Hydrilla verticillata in the Americas predicted using ecological niche modeling combined with genetic data
Ecological niche modeling is an effective tool to characterize the spatial distribution of suitable areas for species, and it is especially useful for predicting the potential distribution of invasive species. The widespread submerged plant Hydrilla verticillata (hydrilla) has an obvious phylogeographical pattern: Four genetic lineages occupy distinct regions in native range, and only one lineage invades the Americas. Here, we aimed to evaluate climatic niche conservatism of hydrilla in North America at the intraspecific level and explore its invasion potential in the Americas by comparing climatic niches in a phylogenetic context. Niche shift was found in the invasion process of hydrilla in North America, which is probably mainly attributed to high levels of somatic mutation. Dramatic changes in range expansion in the Americas were predicted in the situation of all four genetic lineages invading the Americas or future climatic changes, especially in South America; this suggests that there is a high invasion potential of hydrilla in the Americas. Our findings provide useful information for the management of hydrilla in the Americas and give an example of exploring intraspecific climatic niche to better understand species invasion.
Data from: Multi-perspective predictive modeling for acute kidney injury in general hospital populations using electronic medical records
Objective: Acute kidney injury (AKI) in hospitalized patients puts them at much higher risk for developing future health problems such as chronic kidney disease, stroke, and heart disease. Accurate AKI prediction would allow timely prevention and intervention. However, current AKI prediction researches pay less attention to model building strategies that meet complex clinical application scenario. This study aims to build and evaluate AKI prediction models from multiple perspectives that reflect different clinical applications. Material and Methods: A retrospective cohort of 76,957 encounters and relevant clinical variables were extracted from a tertiary care, academic hospital electronic medical record (EMR) system between November 2007 and December 2016. Five machine learning methods were used to build prediction models. Prediction tasks from four clinical perspectives with different modeling and evaluation strategies were designed to build and evaluate the models. Results: Experimental analysis of the AKI prediction models built from four different clinical perspectives suggest a realistic prediction performance in cross-validated AUC ranging from 0.720 to 0.764. Discussion: Results show that models built at admission is effective for predicting AKI events in the next day; models built using data with a fixed lead time to AKI onset is still effective in the dynamic clinical application scenario in which each patient's lead time to AKI onset is different. Conclusion: To our best knowledge, this is the first systematic study to explore multiple clinical perspectives in building predictive models for AKI in the general inpatient population to reflect real performance in clinical application.
Data from: Species' ecological functionality alters the outcome of fish stocking success predicted by a food-web model
Fish stocking is used worldwide in conservation and management but its effects on food-web dynamics and ecosystem stability are poorly known. To better understand these effects and predict the outcomes of stocking, we used an empirically validated network model of a well-studied lake ecosystem. We simulate two stocking scenarios with two native fish species valuable for fishing. In the first scenario, we stock planktivorous fish (whitefish) larvae in the ecosystem. This leads to 1% increase in adult whitefish biomasses and decreases the biomasses of the top predator (perch). In the second scenario, we also stock perch larvae in the ecosystem. This decreases the planktivorous whitefish and the oldest top predator age class biomasses, and destabilizes the ecosystem. Our results demonstrate that the effects of stocking depend on the species' position in the food web and thus cannot be assessed without considering interacting species. We further show that stocking can lead to undesired outcomes from both management and conservation perspectives. The gains of stocking can remain minor and have adverse effects on the entire ecosystem.
Data from: Posterior predictive Bayesian phylogenetic model selection
We present two distinctly different posterior predictive approaches to Bayesian phylogenetic model selection, and illustrate these methods using examples from green algal protein-coding cpDNA sequences and flowering plant rDNA sequences. The Gelfand-Ghosh (GG) approach allows dissection of an overall measure of model fit into components due to posterior predictive variance (Pm) and goodness-of-fit (Gm), which distinguishes this method from the posterior predictive P-value approach. The conditional predictive ordinate (CPO) method provides a site-specific measure of model fit useful for exploratory analyses and can be combined over sites yielding the log pseudomarginal likelihood (LPML), which is useful as an overall measure of model fit. CPO provides a useful cross-validation approach that is computationally efficient, requiring only a sample from the posterior distribution (no additional simulation is required). Both GG and CPO add new perspectives to Bayesian phylogenetic model selection based on the predictive abilities of models, and complement the perspective provided by the marginal likelihood (including Bayes Factor comparisons) based solely on the fit of competing models to observed data.
Data from: Prediction limits of mobile phone activity modelling
Thanks to their widespread usage, mobile devices have become one of the main sensors of human behaviour and digital traces left behind can be used as a proxy to study urban environments. Exploring the nature of the spatio-temporal patterns of mobile phone activity could thus be a crucial step towards understanding the full spectrum of human activities. Using 10 months of mobile phone records from Greater London resolved in both space and time, we investigate the regularity of human telecommunication activity on urban scales. We evaluate several options for decomposing activity timelines into typical and residual patterns, accounting for the strong periodic and seasonal components. We carry out our analysis on various spatial scales, showing that regularity increases as we look at aggregated activity in larger spatial units with more activity in them. We examine the statistical properties of the residuals and show that it can be explained by noise and specific outliers. Also, we look at sources of deviations from the general trends, which we find to be explainable based on knowledge of the city structure and places of attractions. We show examples how some of the outliers can be related to external factors such as specific social events.
Data from: Modeling strategic sperm allocation: tailoring the predictions to the species
Two major challenges exist when empirically testing the predictions of sperm allocation theory. First, the study species must adhere to the assumptions of the model being tested. Unfortunately, the common assumption of sperm allocation models that females mate a maximum of once or twice does not hold for many, if not most, multiply and sequentially mating animals. Second, a model's parameters, which dictate its predictions, must be measured in the study species. Common examples of such parameters, female mating frequency and sperm precedence patterns, are unknown for many species used in empirical tests. Here, we present a broadly applicable model, appropriate for multiply, sequentially mating animals, and test it in three species for which data on all the relevant parameter values are available. The model predicts that relative allocation to virgin females, compared to non-virgins, depends on the interaction between female mating rate and the sperm precedence pattern: relative allocation to virgins increases with female mating rate under first-male precedence, while the opposite is true under later-male precedence. Our model is moderately successful in predicting actual allocation patterns in the three species, including a cricket in which we measured the parameter values and performed an empirical test of allocation.
Data from: Predictions of response to temperature are contingent on model choice and data quality
The equations used to account for the temperature dependence of biological processes, including growth and metabolic rates are the foundations of our predictions of how global biogeochemistry and biogeography change in response to global climate change. We review and test the use of 12 equations used to model the temperature dependence of biological processes across the full range of their temperature response, including supra- and sub-optimal temperatures. We focus on fitting these equations to thermal response curves for phytoplankton growth, but also tested the equations on a variety of traits across a wide diversity of organisms. We found that many of the surveyed equations have comparable abilities to fit data and equally high requirements for data quality (number of test temperatures and range of response captured), but lead to different estimates of cardinal temperatures and of the biological rates at these temperatures. When these rate estimates are used for biogeographic predictions, differences between the estimates of even the best fitting models can exceed the global biological change predicted for a decade of global warming. As a result, studies of the biological response to global changes in temperature must make careful consideration of model selection and of the quality of the data used for parametrizing these models.
Data from: Mainland size variation informs predictive models of exceptional insular body size change in rodents
The tendency for island populations of mammalian taxa to diverge in body size from their mainland counterparts consistently in particular directions is both impressive for its regularity and, especially among rodents, troublesome for its exceptions. However, previous studies have largely ignored mainland body size variation, treating size differences of any magnitude as equally noteworthy. Here, we use distributions of mainland population body sizes to identify island populations as 'extremely' big or small, and we compare traits of extreme populations and their islands with those of island populations more typical in body size. We find that although insular rodents vary in the directions of body size change, 'extreme' populations tend towards gigantism. With classification tree methods, we develop a predictive model, which points to resource limitations as major drivers in the few cases of insular dwarfism. Highly successful in classifying our dataset, our model also successfully predicts change in untested cases.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.