Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: Modelling tooth–prey interactions in sharks: the importance of dynamic testing
The shape of shark teeth varies among species, but traditional testing protocols have revealed no predictive relationship between shark tooth morphology and performance. We developed a dynamic testing device to quantify cutting performance of teeth. We mimicked head-shaking behaviour in feeding large sharks by attaching teeth to the blade of a reciprocating power saw fixed in a custom-built frame. We tested three tooth types at biologically relevant speeds and found differences in tooth cutting ability and wear. Teeth from the bluntnose sixgill (Hexanchus griseus) showed poor cutting ability compared with tiger (Galeocerdo cuvier), sandbar (Carcharhinus plumbeus) and silky (C. falciformis) sharks, but they also showed no wear with repeated use. Some shark teeth are very sharp at the expense of quickly dulling, while others are less sharp but dull more slowly. This demonstrates that dynamic testing is vital to understanding the performance of shark teeth.
Data from: Using multi-response models to investigate pathogen coinfections across scales: insights from emerging diseases of amphibians
1.Associations among parasites affect many aspects of host-parasite dynamics, but a lack of analytical tools has limited investigations of parasite correlations in observational data that are often nested across spatial and biological scales. 2.Here we illustrate how hierarchical, multiresponse modeling can characterize parasite associations by allowing for hierarchical structuring, offering estimates of uncertainty, and incorporating correlational model structures. After introducing the general approach, we apply this framework to investigate coinfections among four amphibian parasites (the trematodes Ribeiroia ondatrae and Echinostoma spp., the chytrid fungus Batrachochytrium dendrobatidis, and ranaviruses) and among >2000 individual hosts, 90 study sites, and five amphibian host species. 3.Ninety-two percent of sites and 80% of hosts supported two or more pathogen species. Our results revealed strong correlations between parasite pairs that varied by scale (from among hosts to among sites) and classification (microparasite versus macroparasite), but were broadly consistent across taxonomically diverse host species. At the host-scale, infection by the trematode R. ondatrae correlated positively with the microparasites, B. dendrobatidis and ranavirus, which were themselves positively associated. However, infection by a second trematode (Echinostoma spp.) correlated negatively with B. dendrobatidis and ranavirus, both at the host- and site-level scales, highlighting the importance of differential relationships between micro- and macroparasites. 4.Given the extensive number of coinfecting symbiont combinations inherent to natural systems, particularly across multiple host species, multiresponse modeling of cross-sectional field data offers a valuable tool to identify a tractable number of hypothesized interactions for experimental testing while accounting for uncertainty and potential sources of co-exposure. For amphibians specifically, the high frequency of co-occurrence and coinfection among these pathogens – each of which is known to impair host fitness or survival – highlights the urgency of understanding parasite associations for conservation and disease management.
Data from: Neurospora and the dead-end hypothesis: genomic consequences of selfing in the model genus.
It is becoming increasingly evident that adoption of different reproductive strategies, such as sexual selfing and asexuality, greatly impacts genome evolution. In this study, we test theoretical predictions on genomic maladaptation of selfing lineages using empirical data from the model fungus Neurospora. We sequenced the genomes of four species representing distinct transitions to selfing within the history of the genus, as well as the transcriptome of one of these, and compared with available data from three outcrossing species. Our results provide evidence for a relaxation of purifying selection in protein-coding genes and for a reduced efficiency of transposable element silencing by Repeat Induced Point mutation. A reduction in adaptive evolution was also identified in the form of reduced codon usage bias in highly expressed genes of selfing Neurospora, but this result may be confounded by mutational bias. Potentially counteracting these negative effects, the nucleotide substitution rate and the spread of transposons is reduced in selfing species. We suggest that differences in substitution rate relate to the absence of the asexual pathway producing conidia in selfing Neurospora. Our results support the dead-end theory and show that Neurospora genomes bear signatures of both sexual and asexual reproductive mode.
Data from: Robustness of the approximate likelihood of the protracted speciation model
The protracted speciation model presents a realistic and parsimonious explanation for the observed slowdown in lineage accumulation through time, by accounting for the fact that speciation takes time. A method to compute the likelihood for this model given a phylogeny is available and allows estimation of its parameters (rate of initiation of speciation, rate of completion of speciation, and extinction rate) and statistical comparison of this model to other proposed models of diversification. However this likelihood computation method makes an approximation of the protracted speciation model to be mathematically tractable: it sometimes counts fewer species than one would do from a biological perspective. This approximation may have large consequences for likelihood-based inferences: it may render any conclusions based on this method completely irrelevant. Here we study to what extent this approximation affects parameter estimations. We simulated phylogenies from which we reconstructed the tree of extant species according to the original, biologically meaningful protracted speciation model and according to the approximation. We then compared the resulting parameter estimates. We found that the differences were larger for high values of extinction rates and small values of speciation-completion rates. Indeed, a long speciation-completion time and a high extinction rate promote the appearance of cases to which the approximation applies. However, surprisingly, the deviation introduced is largely negligible over the parameter space explored, suggesting that this approximate likelihood can be applied reliably in practice to estimate biologically relevant parameters under the original protracted speciation model.
Data from: Model selection with overdispersed distance sampling data
1. Distance sampling (DS) is a widely-used framework for estimating animal abundance. DS models assume that observations of distances to animals are independent. Non-independent observations introduce overdispersion, causing model selection criteria such as AIC or AICc to favour overly complex models, with adverse effects on accuracy and precision. 2. We describe, and evaluate via simulation and with real data, estimators of an overdispersion factor (c ̂), and associated adjusted model selection criteria (QAIC) for use with overdispersed DS data. In other contexts, a single value of c ̂ is calculated from the "global" model, i.e., the most highly-parameterized model in the candidate set, and used to calculate QAIC for all models in the set; the resulting QAIC values, and associated ΔQAIC values and QAIC weights, are comparable across the entire set. Candidate models of the DS detection function include models with different general forms (e.g., half-normal, hazard rate, uniform), so it may not be possible to identify a single global model. We therefore propose a two-step model selection procedure by which QAIC is used to select among models with the same general form, and then a goodness-of-fit statistic is used to select among models with different forms. A drawback of this approach is that QAIC values are not comparable across all models in the candidate set. 3. Relative to AIC, QAIC and the two-step model selection procedure avoided overfitting and improved the accuracy and precision of densities estimated from simulated data. When applied to six real data sets, adjusted criteria and procedures selected either the same model as AIC or a model that yielded a more accurate density estimate in 5 cases, and a model that yielded a less accurate estimate in 1 case. 4. Many DS surveys yield overdispersed data, including cue counting surveys of songbirds and cetaceans, surveys of social species including primates, and camera-trapping surveys. Methods that adjust for overdispersion during the model selection stage of DS analyses therefore address a conspicuous gap in the DS analytical framework as applied to species of conservation concern.
Data from: Prediction limits of mobile phone activity modelling
Thanks to their widespread usage, mobile devices have become one of the main sensors of human behaviour and digital traces left behind can be used as a proxy to study urban environments. Exploring the nature of the spatio-temporal patterns of mobile phone activity could thus be a crucial step towards understanding the full spectrum of human activities. Using 10 months of mobile phone records from Greater London resolved in both space and time, we investigate the regularity of human telecommunication activity on urban scales. We evaluate several options for decomposing activity timelines into typical and residual patterns, accounting for the strong periodic and seasonal components. We carry out our analysis on various spatial scales, showing that regularity increases as we look at aggregated activity in larger spatial units with more activity in them. We examine the statistical properties of the residuals and show that it can be explained by noise and specific outliers. Also, we look at sources of deviations from the general trends, which we find to be explainable based on knowledge of the city structure and places of attractions. We show examples how some of the outliers can be related to external factors such as specific social events.
Data from: A systematic review and meta-analysis of gene therapy in animal models of cerebral glioma: why did promise not translate to human therapy?
Background: The development of therapeutics is often characterized by promising animal research that fails to translate into clinical efficacy; this holds for the development of gene therapy in glioma. We tested the hypothesis that this is because of limitations in the internal and external validity of studies reporting the use of gene therapy in experimental glioma. Method: We systematically identified studies testing gene therapy in rodent glioma models by searching three online databases. The number of animals treated and median survival were extracted and studies graded using a quality checklist. We calculated median survival ratios and used random effects meta-analysis to estimate efficacy. We explored effects of study design and quality and searched for evidence of publication bias. Results: We identified 193 publications using gene therapy in experimental glioma, including 6,366 animals. Overall, gene therapy improved median survival by a factor of 1.60 (95% CI 1.53–1.67). Study quality was low and the type of gene therapy did not account for differences in outcome. Study design characteristics accounted for a significant proportion of between-study heterogeneity. We observed similar findings in a data subset limited to the most common gene therapy. Conclusion: As the dysregulation of key molecular pathways is characteristic of gliomas, gene therapy remains a promising treatment for glioma. Nevertheless, we have identified areas for improvement in conduct and reporting of studies, and we provide a basis for sample size calculations. Further work should focus on genes of interest in paradigms recapitulating human disease. This might improve the translation of such therapies into the clinic.
Data from: Finding candidate genes under positive selection in non-model species: examples of genes involved in host specialization in pathogens
Numerous genes in diverse organisms have been shown to be under positive selection, especially genes involved in reproduction, adaptation to contrasting environments, hybrid inviability, and host-pathogen interactions. Looking for genes under positive selection in pathogens has been a priority in efforts to investigate coevolution dynamics and to develop vaccines or drugs. To elucidate the functions involved in host specialization, here we aimed at identifying candidate sequences that could have evolved under positive selection among closely related pathogens specialized on different hosts. For this goal, we sequenced ca. 17,000-32,000 ESTs from each of four Microbotryum species, which are fungal pathogens responsible for anther smut disease on host plants in the Caryophyllaceae. Forty-two of the 372 predicted orthologous genes showed significant signal of positive selection, which represents a good number of candidate genes for further investigation. Sequencing 16 of these genes in 9 additional Microbotryum species confirmed that they have indeed been rapidly evolving in the pathogen species specialized on different hosts. The genes showing significant signals of positive selection were putatively involved in nutrient uptake from the host, secondary metabolite synthesis and secretion, respiration under stressful conditions and stress response, hyphal growth and differentiation, and regulation of expression by other genes. Many of these genes had transmembrane domains and may therefore also be involved in pathogen recognition by the host. Our approach thus revealed fruitful and should be feasible for many non-model organisms for which candidate genes for diversifying selection are needed.
Data from: Auxotrophy and intra-population complementary in the 'interactome' of a cultivated freshwater model community
Microorganisms are usually studied either in highly complex natural communities or in isolation as monoclonal model populations that we manage to grow in the laboratory. Here, we uncover the biology of some of the most common and yet-uncultured bacteria in freshwater environments using a mixed culture from Lake Grosse Fuchskuhle. From a single shotgun metagenome of a freshwater mixed culture of low complexity, we recovered four high-quality metagenome-assembled genomes (MAGs) for metabolic reconstruction. This analysis revealed the metabolic interconnectedness and niche partitioning of these naturally dominant bacteria. In particular, vitamin- and amino acid biosynthetic pathways were distributed unequally with a member of Crenarchaeota most likely being the sole producer of vitamin B12 in the mixed culture. Using coverage-based partitioning of the genes recovered from a single MAG intrapopulation metabolic complementarity was revealed pointing to 'social' interactions for the common good of populations dominating freshwater plankton. As such, our MAGs highlight the power of mixed cultures to extract naturally occurring 'interactomes' and to overcome our inability to isolate and grow the microbes dominating in nature.
Data from: Effect of detection heterogeneity in occupancy-detection models: an experimental test of time-to-first-detection methods
Imperfect detection can bias estimates of site occupancy in ecological surveys but can be corrected by estimating detection probability. Time-to-first-detection (TTD) occupancy models have been proposed as a cost-effective survey method that allows detection probability to be estimated from single site visits. Nevertheless, few studies have validated the performance of occupancy-detection models by creating a situation where occupancy is known, and model outputs can be compared with the truth. We tested the performance of TTD occupancy models in the face of detection heterogeneity using an experiment based on standard survey methods to monitor koala (Phascolarctos cinereus) populations in Australia. Known numbers of koala faecal pellets were placed under trees, and observers, uninformed as to which trees had pellets under them, carried out a TTD survey. We fitted five TTD occupancy models to the survey data, each making different assumptions about detectability, to evaluate how well each estimated the true occupancy status. Relative to the truth, all five models produced strongly biased estimates, overestimating detection probability and underestimating the number of occupied trees. Despite this, goodness-of-fit tests indicated that some models fitted the data well, with no evidence of model misfit. Hence, TTD occupancy models that appear to perform well with respect to the available data may be performing poorly. The reason for poor model performance was unaccounted for heterogeneity in detection probability, which is known to bias occupancy-detection models. This poses a problem because unaccounted for heterogeneity could not be detected using goodness-of-fit tests and was only revealed because we knew the experimentally determined outcome. A challenge for occupancy-detection models is to find ways to identify and mitigate the impacts of unobserved heterogeneity, which could unknowingly bias many models.
Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model
The fossil record is a rich source of information about biological diversity in the past. However, the fossil record is not only incomplete but has also inherent biases due to geological, physical, chemical and biological factors. Our knowledge of past life is also biased because of differences in academic and amateur interests and sampling efforts. As a result, not all individuals or species that lived in the past are equally likely to be discovered at any point in time or space. To reconstruct temporal dynamics of diversity using the fossil record, biased sampling must be explicitly taken into account. Here, we introduce an approach that uses the variation in the number of times each species is observed in the fossil record to estimate both sampling bias and true richness. We term our technique TRiPS (True Richness estimated using a Poisson Sampling model) and explore its robustness to violation of its assumptions via simulations. We then venture to estimate sampling bias and absolute species richness of dinosaurs in the geological stages of the Mesozoic. Using TRiPS, we estimate that 1936 (1543–2468) species of dinosaurs roamed the Earth during the Mesozoic. We also present improved estimates of species richness trajectories of the three major dinosaur clades: the sauropodomorphs, ornithischians and theropods, casting doubt on the Jurassic–Cretaceous extinction event and demonstrating that all dinosaur groups are subject to considerable sampling bias throughout the Mesozoic.
Data from: A multilocus timescale for oomycete evolution estimated under three distinct molecular clock models
Background: Molecular clock methodologies allow for the estimation of divergence times across a variety of organisms; this can be particularly useful for groups lacking robust fossil histories, such as microbial eukaryotes with few distinguishing morphological traits. Here we have used a Bayesian molecular clock method under three distinct clock models to estimate divergence times within oomycetes, a group of fungal-like eukaryotes that are ubiquitous in the environment and include a number of devastating pathogenic species. The earliest fossil evidence for oomycetes comes from the Lower Devonian (~400 Ma), however the taxonomic affinities of these fossils are unclear. Results: Complete genome sequences were used to identify orthologous proteins among oomycetes, diatoms, and a brown alga, with a focus on conserved regulators of gene expression such as DNA and histone modifiers and transcription factors. Our molecular clock estimates place the origin of oomycetes by at least the mid-Paleozoic (~430-400 Ma), with the divergence between two major lineages, the peronosporaleans and saprolegnialeans, in the early Mesozoic (~225-190 Ma). Divergence times estimated under the three clock models were similar, although only the strict and random local clock models produced reliable estimates for most parameters. Conclusions: Our molecular timescale suggests that modern pathogenic oomycetes diverged well after the origin of their respective hosts, indicating that environmental conditions or perhaps horizontal gene transfer events, rather than host availability, may have driven lineage diversification. Our findings also suggest that the last common ancestor of oomycetes possessed a full complement of eukaryotic regulatory proteins, including those involved in histone modification, RNA interference, and tRNA and rRNA methylation; interestingly no match to canonical DNA methyltransferases could be identified in the oomycete genomes studied here.
Data from: Modeling angle-resolved photoemission of graphene and black phosphorus nano structures
Angle-resolved photoemission spectroscopy (ARPES) data on electronic structure are difficult to interpret, because various factors such as atomic structure and experimental setup influence the quantum mechanical effects during the measurement. Therefore, we simulated ARPES of nano-sized molecules to corroborate the interpretation of experimental results. Applying the independent atomic-center approximation, we used density functional theory calculations and custom-made simulation code to compute photoelectron intensity in given experimental setups for every atomic orbital in poly-aromatic hydrocarbons of various size, and in a molecule of black phosphorus. The simulation results were validated by comparing them to experimental ARPES for highly-oriented pyrolytic graphite. This database provides the calculation method and every file used during the work flow.
Data from: Recasting the dynamic equilibrium model through a functional lens: the interplay of trait-based community assembly and climate
1. According to the dynamic equilibrium hypothesis (DEH), plant species richness is locally controlled by productivity and disturbance. Given that regional conditions widely affect local environmental variables such as soil nutrient availability, the DEH predictions could be improved by considering how climate influences local controls of species richness. Further, a trait-based approach to community assembly has the potential to reveal a deeper, mechanistic understanding of species richness variation across environments. Here we bring together DEH and trait-based community assembly expectations to examine if and how local relationships between diversity, disturbance and productivity are affected by habitat filtering and regional climate. 2. We specifically tested how gradients of local nutrient availability and disturbance intensity interact with climatic conditions to drive the species richness of grassland communities. Further, we recast the DEH through a functional lens by exploring how disturbance-diversity and nutrient availability-diversity relationships are shaped by the functional space occupied by species in a community and species packing within this functional space. 3. The functional space occupied by co-occurring species and the way they are functionally packed are quantified using multi-trait indices calculated with five core plant functional traits. Working with grassland communities spread across differing regional climatic conditions, we used mixed models to test if the variation in taxonomic and functional metrics corresponded to the dynamic equilibrium model's predictions as well as to determine the relationship between those metrics. 4. Contrary to the expectations based on the relation between species richness and the functional components considered, taxonomic and functional metrics did not vary in accordance along environmental gradients. Climate strongly interacted with the local environment to modulate local diversity patterns, sometimes even inversing a given trend and falsifying the DEH predictions. 5. Synthesis. Our findings quantitatively highlight the interplay between regional and local environmental gradients in driving community assembly. We demonstrate that, depending on climatic conditions, observed patterns of both taxonomic and functional community composition can be opposite to expected productivity-diversity and disturbance-diversity relationships. This emphasizes the relevance of multi-faceted studies of biodiversity and the need for a more systematic quantification of regional controls in community assembly studies.
Data from: Mathematical modelling of the vitamin C clock reaction
Chemical clock reactions are characterised by a relatively long induction period followed by a rapid `switchover' during which the concentration of a \emph{clock chemical} rises rapidly. In addition to their interest in chemistry education, these reactions are relevant to industrial and biochemical applications. A substrate-depletive, non-autocatalytic clock reaction involving household chemicals (vitamin C, iodine, hydrogen peroxide and starch) is modelled mathematically via a system of nonlinear ordinary differential equations. Following dimensional analysis the model is analysed in the phase plane and via matched asymptotic expansions. Asymptotic approximations are found to agree closely with numerical solutions in the appropriate time regions. Asymptotic analysis also yields an approximate formula for the dependence of switchover time on initial concentrations and the rate of the slow reaction. This formula is tested via `kitchen sink chemistry' experiments, and is found to enable a good fit to experimental series varying in initial concentrations of both iodine and vitamin C. The vitamin C clock reaction provides an accessible model system for mathematical chemistry.
Data from: Fear of the human "super predator" far exceeds the fear of large carnivores in a model mesocarnivore
The fear (perceived predation risk) large carnivores inspire in mesocarnivores can affect ecosystem structure and function, and loss of the "landscape of fear" large carnivores create adds to concerns regarding the worldwide loss of large carnivores. Fear of humans has been proposed to act as a substitute, but new research identifies humans as a "super predator" globally far more lethal to mesocarnivores, and thus presumably far more frightening. Although much of the world now consists of human-dominated landscapes, there remains relatively little research regarding how behavioral responses to humans affect trophic networks, to the extent that no study has yet experimentally tested the relative fearfulness mesocarnivores demonstrate in reaction to humans versus nonhuman predators. Badgers (Meles meles) in Britain are a model mesocarnivore insofar as they no longer need fear native large carnivores (bears, Ursus arctos; wolves, Canis lupus) and now perhaps fear humans more. We tested the fearfulness badgers demonstrated to audio playbacks of extant (dog) and extinct (bear and wolf) large carnivores, and humans, by assaying the suppression of foraging behavior. Hearing humans affected latency to feed, vigilance, foraging time, number of feeding visits, and number of badgers feeding. Hearing dogs and bears had far lesser effects on latency to feed, and hearing wolves had no effects. Our results indicate fear of humans evidently cannot substitute for the fear large carnivores inspire in mesocarnivores because humans are perceived as far more frightening, which we discuss in light of the recovery of large carnivores in human-dominated landscapes.
Data from: Modeling the internet of things, self-organizing and other complex adaptive communication networks: a cognitive agent-based computing approach
Background: Computer Networks have a tendency to grow at an unprecedented scale. Modern networks involve not only computers but also a wide variety of other interconnected devices ranging from mobile phones to other household items fitted with sensors. This vision of the "Internet of Things" (IoT) implies an inherent difficulty in modeling problems. Purpose: It is practically impossible to implement and test all scenarios for large-scale and complex adaptive communication networks as part of Complex Adaptive Communication Networks and Environments (CACOONS). The goal of this study is to explore the use of Agent-based Modeling as part of the Cognitive Agent-based Computing (CABC) framework to model a Complex communication network problem. Method: We use Exploratory Agent-based Modeling (EABM), as part of the CABC framework, to develop an autonomous multi-agent architecture for managing carbon footprint in a corporate network. To evaluate the application of complexity in practical scenarios, we have also introduced a company-defined computer usage policy. Results: The conducted experiments demonstrated two important results: Primarily CABC-based modeling approach such as using Agent-based Modeling can be an effective approach to modeling complex problems in the domain of IoT. Secondly, the specific problem of managing the Carbon footprint can be solved using a multiagent system approach.
Data from: Improving estimates of environmental change using multilevel regression models of Ellenberg indicator values
Ellenberg indicator values (EIVs) are a widely used metric in plant ecology comprising a semi-quantitative description of species' ecological requirements. Typically, point estimates of mean EIV scores are compared to infer differences in the environmental conditions structuring plant communities – particularly in resurvey studies with no historical environmental data available. However, the use of point estimates as a basis for inference does not take into account variance among species EIVs within sampled plots, and gives equal weighting to means calculated from sites with differing numbers of species. We present a set of multilevel models – fitted with and without group-level predictors – to improve precision and accuracy of site mean EIV scores, and to provide more reliable inference on changing environmental conditions over spatial and temporal gradients in re-visitation studies. We compare multilevel model performance to GLMM's fitted to point estimates of site mean EIVs. We also test the reliability of this method to improve inferences with incomplete species lists in some or all sample sites. Hierarchical modelling led to more accurate and precise estimates of site-level differences in mean EIV scores between time-periods, particularly for datasets with incomplete records of species occurrence. They also revealed directional environmental change within ecological habitat types, which estimates from GLMM's were inadequate to detect. Multilevel models also highlighted a prominent role of hydrological differences as a driver of community change in our case study, which traditional use of EIVs failed to reveal. We have demonstrated that multilevel modelling of EIVs allows for a nuanced estimation of environmental change underlying ecological communities from plant assemblage data, leading to a better understanding of temporal dynamics of ecosystems. Further, the ability of these methods to perform well with missing data should increase the total set of historical data which can be used to this end.
Data from: Visual modelling supports the potential for prey detection by means of diurnal active photolocation in a small cryptobenthic fish
Active sensing has been well documented in animals that use echolocation and electrolocation. Active photolocation, or active sensing using light, has received much less attention, and only in bioluminescent nocturnal species. However, evidence has suggested the diurnal triplefin Tripterygion delaisi uses controlled iris radiance, termed ocular sparks, for prey detection. While this form of diurnal active photolocation was behaviourally described, a study exploring the physical process would provide compelling support for this mechanism. In this paper, we investigate the conditions under which diurnal active photolocation could assist T. delaisi in detecting potential prey. In the field, we sampled gammarids (genus Cheirocratus) and characterized the spectral properties of their eyes, which possess strong directional reflectors. In the laboratory, we quantified ocular sparks size and their angle-dependent radiance. Combined with environmental light measurements and known properties of the visual system of T. delaisi, we modeled diurnal active photolocation under various scenarios. Our results corroborate that diurnal active photolocation should help T. delaisi detect gammarids at distances relevant to foraging, 4.5 cm under favourable conditions and up to 2.5 cm under average conditions. To determine the prevalence of diurnal active photolocation for micro-prey, we encourage further theoretical and empirical work.
Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences
Most models of nucleotide or amino acid substitution used in phylogenetic studies assume that the evolutionary process has been homogeneous across lineages and that composition of nucleotides or amino acids has remained the same throughout the tree. These oversimplified assumptions are refuted by the observation that compositional variability characterizes extant biological sequences. Branch-heterogeneous models of protein evolution that account for compositional variability have been developed, but are not yet in common use because of the large number of parameters required, leading to high computational costs and potential overparameterization. Here, we present a new branch-nonhomogeneous and nonstationary model of protein evolution that captures more accurately the high complexity of sequence evolution. This model, henceforth called Correspondence and likelihood analysis (COaLA), makes use of a correspondence analysis to reduce the number of parameters to be optimized through maximum likelihood, focusing on most of the compositional variation observed in the data. The model was thoroughly tested on both simulated and biological data sets to show its high performance in terms of data fitting and CPU time. COaLA efficiently estimates ancestral amino acid frequencies and sequences, making it relevant for studies aiming at reconstructing and resurrecting ancestral amino acid sequences. Finally, we applied COaLA on a concatenate of universal amino acid sequences to confirm previous results obtained with a nonhomogeneous Bayesian model regarding the early pattern of adaptation to optimal growth temperature, supporting the mesophilic nature of the Last Universal Common Ancestor.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.