Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
285
datasets available to search
ShareScore release 0.9.0
Dataset results
285 results for “evolution models”
Tempo and mode in karyotype evolution revealed by a probabilistic model incorporating both chromosome number and morphology
Open the record for dataset details and reuse information.
Data from: Joint reconstruction of divergence times and life-history evolution in placental mammals using a phylogenetic covariance model
Open the record for dataset details and reuse information.
Data from: Testing models of sex ratio evolution in a gynodioecious plant: female frequency covaries with the cost of male fertility restoration
Open the record for dataset details and reuse information.
Data from: A structured population model suggests that long life and post-reproductive lifespan promote the evolution of cooperation
Social organization correlates with longevity across animal taxa. This correlation has been explained by selection for longevity by social evolution. The reverse causality is also conceivable but has not been sufficiently considered. We constructed a simple, spatially structured population model of asexually reproducing individuals to study the effect of temporal life history structuring on the evolution of cooperation. Individuals employed fixed strategies of cooperation or defection towards all neighbours in a basic Prisoner׳s Dilemma paradigm. Individuals aged and transitioned through different life history stages asynchronously without migration. An individual׳s death triggered a reproductive event by one immediate neighbour. The specific neighbour was chosen probabilistically according to the cumulative payoff from all local interactions. Varying the duration of pre-reproductive, reproductive, and post-reproductive life history stages, long-term simulations allowed a systematic evaluation of the influence of the duration of these specific life history stages. Our results revealed complex interactions among the effects of the three basic life history stages and the benefit to defect. Overall, a long post-reproductive stage promoted the evolution of cooperation, while a prolonged pre-reproductive stage has a negative effect. In general, the total length of life also increased the probability of the evolution of cooperation. Thus, our specific model suggests that the timing of life history transitions and total duration of life history stages may affect the evolution of cooperative behaviour. We conclude that the causation of the empirically observed association of life expectancy and sociality may be more complex than previously realized.
Data from: A mathematical model of marine bacteriophage evolution
To explore how particularities of a cell-virus system affects viral evolution, we formulate a mathematical model of marine bacteriophage evolution. The intrinsic simplicity of real-life phage-bacteria systems allows to have a reasonably simple model. The model constructed in this paper is based upon Beretta-Kuang model of bacteria-phage interaction. Compared to the Beretta-Kuang model, the model assumes the existence of a multitude of viral variants which correspond to continuously distributed phenotypes. It is noteworthy that this model does not include any explicit law or mechanism of evolution; instead it is assumed, in agreement to the principles of Darwinian evolution, that evolution in this system can occur as a result of random mutations and natural selection. Simulations with a leaner fitness landscape (which is chosen for the convenience of demonstration only) show that a pulse-type traveling wave moving towards increasing Darwinian fitness appears in the phenotype space. This implies that the overall fitness of a viral quasispecies steadily increasing in time. That is, the simulations demonstrate that for an uneven fitness landscape random mutations combined with a mechanism of natural selection lead to the Darwinian evolution. It is noteworthy that in this system the speed of propagation of this wave (and hence the rate of evolution) is not constant but varies, depending on the current viral fitness and the abundance of susceptible bacteria.
Data from: Evolution of female multiple mating: a quantitative model of the "sexually-selected sperm" hypothesis
Explaining the evolution and maintenance of polyandry remains a key challenge in evolutionary ecology. One appealing explanation is the sexually-selected sperm (SSS) hypothesis, which proposes that polyandry evolves due to indirect selection stemming from positive genetic covariance with male fertilization efficiency, and hence with a male's success in post-copulatory competition for paternity. However, the SSS hypothesis relies on verbal analogy with 'sexy-son' models explaining co-evolution of female preferences for male displays, and explicit models that validate the basic SSS principle are surprisingly lacking. We developed analogous genetically-explicit individual-based models describing the SSS and 'sexy-son' processes. We show that the analogy between the two is only partly valid, such that the genetic correlation arising between polyandry and fertilization efficiency is generally smaller than that arising between preference and display, resulting in less reliable co-evolution. Importantly, indirect selection was too weak to cause polyandry to evolve in the presence of negative direct selection. Negatively-biased mutations on fertilization efficiency did not generally rescue runaway evolution of polyandry unless realized fertilization was highly skewed towards a single male, and co-evolution was even weaker given random mating-order effects on fertilization. Our models suggest that the SSS process is, on its own, unlikely to generally explain the evolution of polyandry.
Data from: A multilocus timescale for oomycete evolution estimated under three distinct molecular clock models
Background: Molecular clock methodologies allow for the estimation of divergence times across a variety of organisms; this can be particularly useful for groups lacking robust fossil histories, such as microbial eukaryotes with few distinguishing morphological traits. Here we have used a Bayesian molecular clock method under three distinct clock models to estimate divergence times within oomycetes, a group of fungal-like eukaryotes that are ubiquitous in the environment and include a number of devastating pathogenic species. The earliest fossil evidence for oomycetes comes from the Lower Devonian (~400 Ma), however the taxonomic affinities of these fossils are unclear. Results: Complete genome sequences were used to identify orthologous proteins among oomycetes, diatoms, and a brown alga, with a focus on conserved regulators of gene expression such as DNA and histone modifiers and transcription factors. Our molecular clock estimates place the origin of oomycetes by at least the mid-Paleozoic (~430-400 Ma), with the divergence between two major lineages, the peronosporaleans and saprolegnialeans, in the early Mesozoic (~225-190 Ma). Divergence times estimated under the three clock models were similar, although only the strict and random local clock models produced reliable estimates for most parameters. Conclusions: Our molecular timescale suggests that modern pathogenic oomycetes diverged well after the origin of their respective hosts, indicating that environmental conditions or perhaps horizontal gene transfer events, rather than host availability, may have driven lineage diversification. Our findings also suggest that the last common ancestor of oomycetes possessed a full complement of eukaryotic regulatory proteins, including those involved in histone modification, RNA interference, and tRNA and rRNA methylation; interestingly no match to canonical DNA methyltransferases could be identified in the oomycete genomes studied here.
Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences
Most models of nucleotide or amino acid substitution used in phylogenetic studies assume that the evolutionary process has been homogeneous across lineages and that composition of nucleotides or amino acids has remained the same throughout the tree. These oversimplified assumptions are refuted by the observation that compositional variability characterizes extant biological sequences. Branch-heterogeneous models of protein evolution that account for compositional variability have been developed, but are not yet in common use because of the large number of parameters required, leading to high computational costs and potential overparameterization. Here, we present a new branch-nonhomogeneous and nonstationary model of protein evolution that captures more accurately the high complexity of sequence evolution. This model, henceforth called Correspondence and likelihood analysis (COaLA), makes use of a correspondence analysis to reduce the number of parameters to be optimized through maximum likelihood, focusing on most of the compositional variation observed in the data. The model was thoroughly tested on both simulated and biological data sets to show its high performance in terms of data fitting and CPU time. COaLA efficiently estimates ancestral amino acid frequencies and sequences, making it relevant for studies aiming at reconstructing and resurrecting ancestral amino acid sequences. Finally, we applied COaLA on a concatenate of universal amino acid sequences to confirm previous results obtained with a nonhomogeneous Bayesian model regarding the early pattern of adaptation to optimal growth temperature, supporting the mesophilic nature of the Last Universal Common Ancestor.
Data from: The impact of rate heterogeneity on inference of phylogenetic models of trait evolution
Rates of trait evolution are known to vary across phylogenies; however, standard evolutionary models assume a homogeneous process of trait change. These simple methods are widely applied in small-scale phylogenetic studies, whereas models of rate heterogeneity are not, so the prevalence and patterns of potential rate variation in groups up to hundreds of species remain unclear. The extent to which trait evolution is modelled accurately on a given phylogeny is also largely unknown because studies typically lack absolute model fit tests. We investigated these issues by applying both rate-static and variable-rates methods on (i) body mass data for 88 avian clades of 10–318 species, and (ii) data simulated under a range of rate-heterogeneity scenarios. Our results show that rate heterogeneity is present across small-scaled avian clades, and consequently applying only standard single-process models prompts inaccurate inferences about the generating evolutionary process. Specifically, these approaches underestimate rate variation, and systematically mislabel temporal trends in trait evolution. Conversely, variable-rates approaches have superior relative fit (they are the best model) and absolute fit (they describe the data well). We show that rate changes such as single internal branch variations, rate decreases and early bursts are hard to detect, even by variable-rates models. We also use recently developed absolute adequacy tests to highlight misleading conclusions based on relative fit alone (e.g. a consistent preference for constrained evolution when isolated terminal branch rate increases are present). This work highlights the potential for robust inferences about trait evolution when fitting flexible models in conjunction with tests for absolute model fit.
Data from: Complex models of sequence evolution require accurate estimators as exemplified with the invariable site plus Gamma model
The invariable site plus Γ model is widely used to model rate heterogeneity among alignment sites in maximum likelihood and Bayesian phylogenetic analyses. The proof that the invariable site plus continuous Γ model is identifiable (model parameters can be inferred correctly given enough data) has increased the creditability of its application to phylogeny reconstruction. However, most phylogenetic software implement the invariable site plus discrete Γ model, whose identifiability is likely but unproven. How well the parameters of the invariable site plus discrete Γ model are estimated is still disputed. Especially the correlation of the fraction of invariable sites with the fractions of sites with a slow evolutionary rate is discussed as being problematic. We show that optimization heuristics as implemented in frequently used phylogenetic software cannot always reliably estimate the shape parameter, the proportion of invariable sites and the tree length. Here, we propose an improved optimization heuristic that accurately estimates the three parameters. While research efforts mainly focus on tree search methods, our results signify the equal importance of verifying and developing effective estimation methods for complex models of sequence evolution.
Data from: Using historical biogeography models to study color pattern evolution
Color is among the most striking features of organisms, varying not only in spectral properties like hue and brightness, but also in where and how it is produced on the body. Different combinations of colors on a bird's body are important in both environmental and social contexts. Previous comparative studies have treated plumage patches individually or derived plumage complexity scores from color measurements across a bird's body. However, these approaches do not consider the multivariate nature of plumages (allowing for plumage to evolve as a whole) or account for interpatch distances. Here, we leverage a rich toolkit used in historical biogeography to assess color pattern evolution in a cosmopolitan radiation of birds, kingfishers (Aves: Alcedinidae). We demonstrate the utility of this approach and test hypotheses about the tempo and mode of color evolution in kingfishers. Our results highlight the importance of considering interpatch distances in understanding macroevolutionary trends in color diversity and demonstrate how historical biogeography models are a useful way to model plumage color pattern evolution. Furthermore, they show that distinct color mechanisms (pigments or structural colors) spread across the body in different ways and at different rates. Specifically, net rates are higher for structural colors than pigment-based colors. Together, our study suggests a role for both development and selection in driving extraordinary color pattern diversity in kingfishers. We anticipate this approach will be useful for modeling other complex phenotypes besides color, such as parasite evolution across the body.
Data from: Detecting adaptive evolution in phylogenetic comparative analysis using the Ornstein-Uhlenbeck model
Phylogenetic comparative analysis is an approach to inferring evolutionary process from a combination of phylogenetic and phenotypic data. The last few years have seen increasingly sophisticated models employed in the evaluation of more and more detailed evolutionary hypotheses, including adaptive hypotheses with multiple selective optima and hypotheses with rate variation within and across lineages. The statistical performance of these sophisticated models has received relatively little systematic attention, however. We conducted an extensive simulation study to quantify the statistical properties of a class of models toward the simpler end of the spectrum that model phenotypic evolution using Ornstein–Uhlenbeck processes. We focused on identifying where, how, and why these methods break down so that users can apply them with greater understanding of their strengths and weaknesses. Our analysis identifies three key determinants of performance: a discriminability ratio, a signal-to-noise ratio, and the number of taxa sampled. Interestingly, we find that model-selection power can be high even in regions that were previously thought to be difficult, such as when tree size is small. On the other hand, we find that model parameters are in many circumstances difficult to estimate accurately, indicating a relative paucity of information in the data relative to these parameters. Nevertheless, we note that accurate model selection is often possible when parameters are only weakly identified. Our results have implications for more sophisticated methods inasmuch as the latter are generalizations of the case we study.
Data from: Modelling the co-evolution of indirect genetic effects and inherited variability
When individuals interact, their phenotypes may be affected by genes in their social partners, a phenomenon known as Indirect Genetic Effects (IGEs). In aquaculture species and some plants, competition not only affects trait levels of individuals, but also inflates variation of trait values among individuals. Variability of trait values has been studied as a quantitative trait in itself, and is often referred to as inherited variability. Although the observed phenotypic relationship between competition and variability suggests an underlying genetic relationship, models of IGE and inherited variability do not allow for such relationship. Models of trait levels show IGEs may considerably change heritable variation in trait values. Currently, we lack the tools to investigate whether this result extends to inherited variability. Here we present a model that integrates IGEs and inherited variability. In this model, the target phenotype, say growth rate, is a function of genetic and environmental effects of the focal individual and of the difference in trait values between the social partner and the focal individual, multiplied by a regression coefficient. The regression coefficient is a genetic trait which is measure of cooperation; a negative value indicates competition, a positive value cooperation, and an increasing value due to selection indicates the evolution of cooperation. Our simulations show that the model results in increased variability of body weight with increase of competition. When competition decreases, variability becomes significantly smaller. Our findings suggest we may have been overlooking an entire level of genetic variation in variability, the one due to IGEs.
Functional and ecomorphological evolution of orbit shape in Mesozoic archosaurs is driven by body size and diet: Geometric morphometric data, 3D models (stl files), FEA models (Hypermesh, Abaqus files)
<p class="MsoNormal">The orbit is one of several skull openings in the archosauromorph skull. Intuitively, it could be assumed that orbit shape would closely approximate the shape and size of the eyeball resulting in a predominantly circular morphology. However, a quantification of orbit shape across Archosauromorpha using a geometric morphometric approach demonstrates a large morphological diversity despite the fact that the majority of species retained a circular orbit. This morphological diversity is nearly exclusively driven by large (skull length > 1000 mm) and carnivorous species in all studied archosauromorph groups, but particularly prominently in theropod dinosaurs. While circular orbit shapes are retained in most herbivores and smaller species, as well as in juveniles and early ontogenetic stages, large carnivores adopted elliptical and keyhole-shaped orbits. Biomechanical modeling using finite element analysis reveals that these morphologies are beneficial in mitigating and dissipating feeding-induced stresses without additional reinforcement of the bony structure of the skull.</p>
Dataset related to article "Evolution of brain injury and neurological dysfunction after cardiac arrest in the rat – a multimodal and comprehensive model."
<p>Excel file related to the article</p>
Seismic velocity models related to the manuscript entitled "Paleo-ocean and Early Evolution of Mars Revealed by Seismic Crustal Stratigraphy"
<p>Seismic velocity models of the Martian crust.</p>
The "lastfm" data set used in the article "A comparative study of social network models: Network evolution models and nodal attribute models"
<p>This is the "lastfm" network used in the article:</p> <p>Toivonen, R., Kovanen, L., Kivelä, M., Onnela, J. P., Saramäki, J., & Kaski, K. (2009). A comparative study of social network models: Network evolution models and nodal attribute models. Social networks, 31(4), 240-254.</p> <p>doi:10.1016/j.socnet.2009.06.004</p> <p>The data set is described in the article. Please cite the original article when using this data set.</p> <p>Format of the data set is an edge list, where row in the file is an edge connecting the two nodes indicated by the two numbers separated by a whitespace. Each node number corresponds to a single account in the website.</p> <p>The original data in which this network is based on was licensed under the "Creative Commons Attribution-NonCommercial-ShareAlike 2.0 UK: England & Wales" licese, and accordinly this data set uses the same license. License available at https://creativecommons.org/licenses/by-nc/2.0/uk/</p>
Unraveling the Geodynamic Evolution of the Pre– and Early–Andean Margin: Insights from Numerical Modeling
Open the record for dataset details and reuse information.
Data from: An integrated model of phenotypic trait changes and site-specific sequence evolution
Recent years have seen a constant rise in the availability of trait data, including morphological features, ecological preferences, and life history characteristics. These phenotypic data provide means to associate genomic regions with phenotypic attributes, thus allowing the identification of phenotypic traits associated with the rate of genome and sequence evolution. However, inference methodologies that analyze sequence and phenotypic data in a unified statistical framework are still scarce. Here, we present TraitRateProp, a probabilistic method that allows testing whether the rate of sequence evolution is associated with a binary phenotypic character trait. The method further allows the detection of specific sequence sites whose evolutionary rate is most noticeably affected following the character transition, suggesting a shift in functional/structural constraints. TraitRateProp is first evaluated in simulations and then applied to study the evolutionary process of plastid plant genomes upon a transition to a heterotrophic lifestyle. To this end, we analyze 25 plastid genes across 85 orchid species, spanning different lifestyles and representing different genera in this large family of flowering plants. Our results indicate higher evolutionary rates following repeated transitions to a heterotrophic lifestyle in all but four of the loci analyzed.
Assessing the Accuracy of 2-D Planetary Evolution Models against the 3-D Sphere
<p><strong>Datasets concerning isoviscous simulations:</strong><br> Tables containing the time averaged (on the last 10% of the run) values for all the outputs and geometry studied. There is one table per Ra number with a given heating mode. In total there are 15 tables for each scenarios (i.e., three different heating modes and five different Ra numbers)</p> <p><strong>Datasets concerning temperature dependent simulations:</strong><br> Tables containing the time averaged (on the last 10% of the run) values for all the outputs and geometry studied for temperature dependent viscosity simulations. Only one Ra is investigated. In total three tables, for three heating modes.</p> <p><strong>Datasets concerning thermal evolution simulations with and without crust:</strong><br> Tables containing dimensional present day values of all the investigated outputs for different geometries and planet scenarios for cases with and without crust. In total six tables, for three planets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.