Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
74
datasets available to search
ShareScore release 0.7.1
Dataset results
74 results for “Environment Prediction”
Data and code from: Predicting population genetic change in an autocorrelated random environment: insights from a large automated experiment
Open the record for dataset details and reuse information.
Data from: Population climatic history predicts phenotypic responses in novel environments for Arabidopsis thaliana in North America
Open the record for dataset details and reuse information.
Predicting Phenotype from Multi-Scale Genomic and Environment Data using Neural Networks and Knowledge Graphs
<p><strong>Background: To mitigate the effects of climate change on public health and conservation, we need to better understand the dynamic interplay between biological processes and environmental effects. Machine learning (ML) methods in general, and Deep Learning (DL) methods in particular, are a potential way forward because they are able to cope with the nonlinearity of natural systems. However, there are several barriers that exist, including the absence of ML-ready data. We propose to develop a machine learning framework capable of predicting phenotypes based on multi-scale data about genes and environments. A critical part of this framework are data transformation methods that map the heterogeneous input data into formats that are consumable by the ML techniques. The central hypothesis of this research is that deep learning algorithms and biological knowledge graphs will predict phenotypes more accurately across more taxa and more ecosystems than do current numerical and traditional statistical modeling methods. Our long term goal is to develop predictive analytics for organismal response to environmental perturbations using innovative data science approaches. This pilot project on predicting emergent properties of complex systems and multidimensional interactions is funded by the NSF (Award # 1939945, 1940059, 1940062, 1940330). </strong></p> <p> </p> <p><strong>Results: We have established shared project governance, communication channels, project timeline, and data and computing environment across four universities. We have successfully reached out to three other projects for broader collaboration.</strong></p>
Genomic prediction applied to multiple traits and environments in second season maize hybrids
<p>Genomic selection has become a reality in plant breeding programs with the reduction in genotyping costs. Especially in maize breeding programs, it emerges as a promising tool for predicting hybrid performance. The dynamics of a commercial breeding program involve the evaluation of several traits simultaneously in a large set of target environments. Therefore, multi-trait multi-environment (MTME) genomic prediction models can leverage these data sets by exploring the correlation between traits and Genotype-by-Environment (G×E) interaction. Herein, we assess predictive abilities of univariate and multivariate genomic prediction models in a maize breeding program. To this end, we used data from 415 maize hybrids evaluated in four years of second season field trials for the traits grain yield, number of ears and grain moisture. Genotypes of these hybrids were inferred <i>in silico</i> based on their parental inbred lines using Single Nucleotide Polymorphisms (SNPs) markers obtained via genotyping-by-sequencing (GBS). Because genotypic information was available for only 257 hybrids, we used the genomic and pedigree relationship matrices to obtain the <b>H</b> matrix for all 415 hybrids. Our results demonstrated that in the single-environment context the use of multi-trait models was always superior in comparison to their univariate counterparts. Besides that, although MTME models were not particularly successful in predicting hybrid performance in untested years, they improved the ability to predict the performance of hybrids that had not been evaluated in any environment. However, the computational requirements of this kind of model could represent a limitation to its practical implementation and further investigation is necessary.</p>
Data from: Predicting bird phenology from space: satellite-derived vegetation green-up signal uncovers spatial variation in phenological synchrony between birds and their environment
Population-level studies of how tit species (Parus spp.) track the changing phenology of their caterpillar food source have provided a model system allowing inference into how populations can adjust to changing climates, but are often limited because they implicitly assume all individuals experience similar environments. Ecologists are increasingly using satellite-derived data to quantify aspects of animals' environments, but so far studies examining phenology have generally done so at large spatial scales. Considering the scale at which individuals experience their environment is likely to be key if we are to understand the ecological and evolutionary processes acting on reproductive phenology within populations. Here, we use time series of satellite images, with a resolution of 240 m, to quantify spatial variation in vegetation green-up for a 385-ha mixed-deciduous woodland. Using data spanning 13 years, we demonstrate that annual population-level measures of the timing of peak abundance of winter moth larvae (Operophtera brumata) and the timing of egg laying in great tits (Parus major) and blue tits (Cyanistes caeruleus) is related to satellite-derived spring vegetation phenology. We go on to show that timing of local vegetation green-up significantly explained individual differences in tit reproductive phenology within the population, and that the degree of synchrony between bird and vegetation phenology showed marked spatial variation across the woodland. Areas of high oak tree (Quercus robur) and hazel (Corylus avellana) density showed the strongest match between remote-sensed vegetation phenology and reproductive phenology in both species. Marked within-population variation in the extent to which phenology of different trophic levels match suggests that more attention should be given to small-scale processes when exploring the causes and consequences of phenological matching. We discuss how use of remotely sensed data to study within-population variation could broaden the scale and scope of studies exploring phenological synchrony between organisms and their environment.
Data from: Global gradients in vertebrate diversity predicted by historical area-productivity dynamics and contemporary environment
Broad-scale geographic gradients in species richness have now been extensively documented, but their historical underpinning is still not well understood. While the importance of productivity, temperature, and a scale-dependence of the determinants of diversity is broadly acknowledged, we here argue that limitation to a single analysis scale and data pseudoreplication have impeded an integrated evolutionary and ecological understanding of diversity gradients. We develop and apply a hierarchical analysis framework for global diversity gradients that incorporates an explicit accounting of past environmental variation and provides an appropriate measurement of richness. Due to environmental niche conservatism organisms generally reside in climatically defined bioregions, or "evolutionary arenas", characterized by in situ speciation and extinction. These bioregions differ in age and their total productivity and have varied over time in area and energy available for diversification. We show that, consistently across the four major terrestrial vertebrate groups, current-day species richness of the world's main 32 bioregions is best explained by a model that integrates area and productivity over geological time together with temperature. Adding finer-scale variation in energy availability as an ecological predictor of within-bioregional patterns of richness explains much of the remaining global variation in richness at the 110km grain. These results highlight the separate evolutionary and ecological effects of energy availability and provide a first conceptual and empirical integration of the key drivers of broad-scale richness gradients. Avoiding the pseudo-replication that impedes the evolutionary interpretation of non-hierarchical macroecological analyses, our findings integrate evolutionary and ecological mechanisms at their most relevant scales and offer a new synthesis regarding global diversity gradients.
Data from: Predicting ecological and phenotypic differentiation in the wild: a case of piscivorous fish in a fishless environment
Environmental variation drives ecological and phenotypic change. How predictable is differentiation in response to environmental change? Answering this question requires the development and testing of multifarious a priori predictions in natural systems. We employ this approach using Gobiomorus dormitor populations that have colonized inland blue holes differing in the availability of fish prey. We evaluated predictions of differences in demographics, habitat use, diet, locomotor and trophic morphology, and feeding kinematics and performance between G. dormitor populations inhabiting blue holes with and without fish prey. Populations of G. dormitor independently diverged between prey regimes, with broad agreement between observed differences and a priori predictions. For example, in populations lacking fish prey, we observed male-biased sex ratios, a greater use of shallow-water habitat, and larger population diet breadths as a result of greater individual diet specialization. Furthermore, we found predictable differences in body shape, mouth morphology, suction generation capacity, strike kinematics, and feeding performance on different prey types, consistent with the adaptation of G. dormitor to piscivory when coexisting with fish prey and to feeding on small invertebrates in their absence. The results of the present study suggest great potential in our ability to predict population responses to changing environments, which is an increasingly important capability in a human-dominated, ever-changing world.
Data from: Nest size is predicted by female identity and the local environment in the blue tit, but is not related to genetic or foster mother's nest size
The potential for animals to respond to changing climates has sparked interest in intraspecific variation in avian nest structure since this may influence nest microclimate and protect eggs and offspring from inclement weather. However, there have been relatively few large-scale attempts to examine variation in nests or the determinates of individual variation in nest structure within populations. Using a set of mostly pre-registered analyses, we studied potential predictors of variation in the size of a large sample (803) of blue tit (Cyanistes caeruleus) nests across three breeding seasons at Wytham Woods, UK. Whilst our pre-registered analyses found that individual females built very similar nests across years, there was no evidence in follow-up (post hoc) analyses that their nest size correlated to that of their genetic mother or, in a cross-fostering experiment, to the nest where they were reared. In further pre-registered analyses, spatial environmental variability explained nest size variability at relatively broad spatial scales, and especially strongly at the scale of individual nestboxes. Our study indicates that nest structure is a characteristic of individuals, but is not strongly heritable, indicating that it will not respond rapidly to selection. Explaining the within-individual and within-location repeatability we observed requires further study.
Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models
Genomic selection have been proposed as the standard method to predict breeding values in animal and plant breeding. Although some crops have benefited from this methodology, studies in Coffea are still emerging. To date, there have been no studies of how well genomic prediction models work across populations and environments for different complex traits in coffee. Considering that predictive models are based on biological and statistical assumptions, it is expected that their performance vary depending on how well these assumptions align with the true genetic architecture of the phenotype. To investigate this, we used data from two recurrent selection populations of Coffea canephora, evaluated in two locations, and single nucleotide polymorphisms identified by Genotyping-by-Sequencing. In particular, we evaluated the performance of 13 statistical approaches to predict three important traits in the coffee — production of coffee beans, leaf rust incidence and yield of green beans. Analyses were performed for predictions within-environment, across locations and across populations to assess the reliability of genomic selection. Overall, differences in the prediction accuracy of the competing models were small, although the Bayesian methods showed a modest improvement over other methods, at the cost of more computation time. As expected, predictive accuracy for within-environment analysis, on average, were higher than predictions across locations and across populations. Our results support the potential of genomic selection to reshape traditional plant breeding schemes. In practice, we expect to increase the genetic gain per unit of time by reducing the length cycle of recurrent selection in coffee.
Data from: A quantitative test of the predicted relationship between countershading and lighting environment.
Countershading, a vertical luminance gradient from a dark back to a light belly, is perhaps the most common coloration phenotype in the animal kingdom. Why? We investigated whether countershading functions as self-shadow concealment (SSC) in ruminants. We calculated "optimal" countershading for SSC by measuring illumination falling onto a model ruminant as a function of time of day and lighting environment. Calibrated images of 114 species of ruminant were compared to the countershading model, and phylogenetic analyses were used to find the best predictors of coats' countershading characteristics. In many species, countershading was close to the model's prediction of "optimal" countershading for SSC. Stronger countershading was associated with increased use of open lighting environments, living closer to the equator, and small body size. Abrupt transitions from dark to light tones were more common in open lighting environments but unassociated with group size or antipredator behavior. Though the SSC hypothesis prediction for stronger countershading in diurnal species was not supported and noncountershaded or reverse-countershaded species were unexpectedly common, this basic pattern of associations is explained only by the SSC hypothesis. Despite extreme variation in lighting conditions, many terrestrial animals still find protection from predation by compensating for their own shadows.
Data from: Partner's age, not social environment, predicts extrapair paternity in wild great tits (Parus major)
An individual's fitness is not only influenced by its own phenotype, but by the phenotypes of interacting conspecifics. This is likely to be particularly true when considering fitness gains and losses caused by extrapair matings, as they depend directly on the social environment. While previous work has explored effects of dyadic interactions, limited understanding exists regarding how group-level characteristics of the social environment affect extrapair paternity (EPP) and cuckoldry. We use a wild population of great tits (Parus major) to examine how, in addition to the phenotypes of focal parents, two neighborhood-level traits – age and personality composition – predict EPP and cuckoldry. We used the well-studied trait "exploration behavior" as a measure of the reactive-proactive personality axis. Because breeding pairs inhabit a continuous "social landscape", we first established an ecologically relevant definition of a breeding "neighborhood" through genotyping parents and nestlings in a 51-ha patch of woodland and assessing the spatial predictors of EPP events. Using the observed decline in likelihood of EPP with increasing spatial separation between nests, we determined the relevant neighborhood boundaries, and thus the group phenotypic composition of an individual's neighborhood, by calculating the point at which the likelihood of EPP became negligible. We found no evidence that "social environment" effects (i.e. neighborhood age or personality composition) influenced EPP or cuckoldry. We did, however, find that a female's own age influenced the EPP of her social mate, with males paired to older females gaining more EPP, even when controlling for the social environment. These findings suggest that partner characteristics, rather than group phenotypic composition, influence mating activity patterns at the individual level.
History and environment shape spatial genetic variation and predict climate maladaptation in a narrowly distributed serotinous pine, Pinus muricata
<p><span></span></p> <p>Understanding the distribution of genetic diversity and differentiation in species with disjunct and isolated populations is critical for assessing how environment shapes genetic variation and the potential response to climate change. In contrast to the large distributions and population sizes of most pine species, <em>Pinus muricata</em> (Bishop pine) occurs in a small number of isolated and disjunct populations occupying a narrow band of environmental conditions along the coast of western North America. We used genotyping by sequencing to generate population genomic data for trees sampled from nearly all existing populations of <em>P. muricata</em> (12 populations, 213 individuals, 7,828 loci) to describe the spatial arrangement of genetic differentiation and diversity. We used genetic-environment association (GEA) analyses to quantify the contribution of environmental variables to local adaptation and spatial genetic structure. Based on these results, we quantified relative levels of potential maladaptation given future climate projections at 2041 – 2060 and 2081 – 2100. Our analyses reveal pronounced spatial genetic structure across the distribution, with most populations forming genetically identifiable groups across a latitudinal gradient, and remarkable evidence for differentiation among three proximally distributed stands on Santa Cruz Island. Despite occurring in small, isolated populations, <em>P. muricata</em> do not exhibit strongly reduced diversity. GEA analyses suggested that specific soil and climate variables have contributed to local adaptation. Genomic offset analyses suggest geographic variation in potential maladaptation, with northern populations experiencing higher levels under projected climate change. Overall, our results suggest that isolation and local adaptation have shaped genetic variation among disjunct populations, and illustrate the consequences of this variation for <em>P. muricata</em> under projected climate change.</p>
Island size predicts mammal diversity in insular environments, except for land-bridge islands
<p class="MsoBodyText"><span>Insular environments are among the most endangered ecosystems as they face a myriad of anthropogenic stressors. Forest mammals perform a wide range of ecological services, with their persistence being vital for ecosystem functionality in both natural and artificial islands. Studies revealed that shrinkage in island size usually leads to the decay of mammal species richness and abundance in patchy landscapes. However, mammal species-area (SARs) and abundance-area (AARs) relationships can differ among insular environments: oceanic, fluvial, artificial, and land-bridge islands (i.e., natural islands connected to the mainland). Large dams create vast insularized landscapes after river impoundment, leading to pervasive habitat loss and potentially causing even worse biodiversity losses than other insular systems. We conducted an extensive literature search and used meta-analysis techniques to quantify the magnitude of SAR and AAR for forest mammals across different archipelago landscapes worldwide. After a screening process, we ended up with 26 studies comprising 55 different effect sizes representing the magnitude of SARs and AARs. Our global analysis unveiled a positive relationship between effect sizes and island area, with mammal species richness and abundance increasing in fluvial, oceanic, and artificial islands accordingly with island area, but not in land-bridge islands. These results demonstrate that, except for land-bridge islands, SAR and AAR are still fair models to predict mammal diversity.<span> </span>These results could improve the prediction of SAR and AAR in insular environments under habitat loss scenarios and propose sound conservation strategies since the rate at which insular communities have been lost is presently unknown.</span></p>
Input data for predicted sedimentary environments on the Norwegian continental margin
<p>Input and output data relating to R workflow for predicting sedimentary environments on the Norwegian continental margin (https://github.com/diesing-ngu/SedEnv). The following files are included:</p><p><strong>SedEnv_4km_MaxCombArea_point_20230622.shp</strong> - Point shapefile of the response data (substrate type). Note that these data points were derived from mapped products and are not sample points as such. </p><p><strong>predictors_ngb.tif </strong>- Multi-band georeferenced TIFF-file of predictor variables</p><p><strong>predictors_description</strong>.txt - Information on variables stored in predictor_ngb.tif including units, statistics, time period and sources.</p><p><strong>GrainSizeReg_folk8_classes_2023-06-28.tif</strong> - Georeferenced TIFF-file of predicted substrate classes. Used to update the area of interest (exclude areas mapped as Rock and boulders).</p><p><strong>mud_2023-06-30.tif </strong>- Georeferenced TIFF-file of predicted mud content. Used as an additional predctor.</p>
Predicted sedimentary environments on the Norwegian continental margin
<p>Output data relating to R workflow for predicting substrate types on the Norwegian continental margin (https://github.com/diesing-ngu/SedEnv). The following files are included:</p><p><strong>SedEnv3_classes_2023-07-01.tif </strong>- Georeferenced Tiff-file of the predicted sedimentary environments</p><p><strong>SedEnv3_probabilities_2023-07-01.tif </strong>- Georeferenced Tiff-file of the prediction probabilities of the predicted sedimentary environments</p><p><strong>SedEnv3_max_probabilities_2023-07-01.tif </strong>- Georeferenced Tiff-file of the maximum probabilities, i.e., the probability of the class that was mapped. Can be used as an indicator of map confidence.</p><p><strong>SedEnv3_AOA_2023-07-01.tif </strong>- Georeferenced Tiff-file of the area of applicability of the model <a href="https://doi.org/10.1111/2041-210X.13650">(Meyer & Pebesma, 2021)</a></p><p><strong>SedEnv3_AOA_2023-07-01.shp</strong> - Same as above but as polygon shapefile</p>
Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models
Open the record for dataset details and reuse information.
Island size predicts mammal diversity in insular environments, except for land-bridge islands
Open the record for dataset details and reuse information.
Data from: A quantitative test of the predicted relationship between countershading and lighting environment.
Open the record for dataset details and reuse information.
Data from: Global gradients in vertebrate diversity predicted by historical area-productivity dynamics and contemporary environment
Open the record for dataset details and reuse information.
Data from: Predictable genome-wide sorting of standing genetic variation during parallel adaptation to basic versus acidic environments in stickleback fish
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.