Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.9.0
Dataset results
35 results for “comparative phylogenetic methods”
FIGURE 1 in Comparative morphology of the eggs from the eight species in the genus Agathemera Stål (Insecta: Phasmatodea), through phylogenetic comparative method approach
FIGURE 1. schematic drawing of an Agathemera egg showing the different variables measured. a. dorsal view. b. lateral view. c. upper view. Abbreviations: w = width; mpl = micropylar plate width; mpl = micropylar plate length; h = capsule height; l = capsule length; opl = operculum length; opa = opercular angle; opw = operculum width; operculum height.
Data from: Rethinking phylogenetic comparative methods
As a result of the process of descent with modification, closely related species tend to be similar to one another in a myriad different ways. In statistical terms, this means that traits measured on one species will not be independent of traits measured on others. Since their introduction in the 1980s, phylogenetic comparative methods (PCMs) have been framed as a solution to this problem. In this paper, we argue that this way of thinking about PCMs is deeply misleading. Not only has this sowed widespread confusion in the literature about what PCMs are doing but has led us to develop methods that are susceptible to the very thing we sought to build defenses against --- unreplicated evolutionary events. Through three Case Studies, we demonstrate that the susceptibility to singular events is indeed a recurring problem in comparative biology that links several seemingly unrelated controversies. In each Case Study we propose a potential solution to the problem. While the details of our proposed solutions differ, they share a common theme: unifying hypothesis testing with data-driven approaches (which we term ``phylogenetic natural history'') to disentangle the impact of singular evolutionary events from that of the factors we are investigating. More broadly, we argue that our field has, at times, been sloppy when weighing evidence in support of causal hypotheses. We suggest that one way to refine our inferences is to re-imagine phylogenies as probabilistic graphical models; adopting this way of thinking will help clarify precisely what we are testing and what evidence supports our claims.
Data from: A phylogenetic comparative method for evaluating trait coevolution across two phylogenies for sets of interacting species
Evaluating trait correlations across species within a lineage via phylogenetic regression is fundamental to comparative evolutionary biology, but when traits of interest are derived from two sets of lineages that co-evolve with one another, methods for evaluating such patterns in a dual-phylogenetic context remain underdeveloped. Here we extend multivariate permutation-based phylogenetic regression to evaluate trait correlations in two sets of interacting species while accounting for their respective phylogenies. This extension is appropriate for both univariate and multivariate response data, and may utilize one or more independent variables, including environmental covariates. Imperfect correspondence between species in the interacting lineages can also be accommodated, such as when species in one lineage associate with multiple species in the other, or when there are unmatched taxa in one or both lineages. For both univariate and multivariate data, the method displays appropriate type I error, and statistical power increases with the strength of the trait covariation and the number of species in the phylogeny. These properties are retained even when there is not a 1:1 correspondence between lineages. Finally, we demonstrate the approach by evaluating the evolutionary correlation between traits in fig species and traits in their agaonid wasp pollinators. R computer code is provided.
Data from: Multivariate phylogenetic comparative methods: evaluations, comparisons, and recommendations
Recent years have seen increased interest in phylogenetic comparative analyses of multivariate datasets, but to date the varied proposed approaches have not been extensively examined. Here we review the mathematical properties required of any multivariate method, and specifically evaluate existing multivariate phylogenetic comparative methods in this context. Phylogenetic comparative methods based on the full multivariate likelihood are robust to levels of covariation among trait dimensions and are insensitive to the orientation of the dataset, but display increasing model misspecification as the number of trait dimensions increases. This is because the expected evolutionary covariance matrix (V) used in the likelihood calculations becomes more ill-conditioned as trait dimensionality increases, and as evolutionary models become more complex. Thus, these approaches are only appropriate for datasets with few traits and many species. Methods that summarize patterns across trait dimensions treated separately (e.g., SURFACE) incorrectly assume independence among trait dimensions, resulting in nearly a 100% model misspecification rate. Methods using pairwise composite likelihood are highly sensitive to levels of trait covariation, the orientation of the dataset, and the number of trait dimensions. The consequences of these debilitating deficiencies is that a user can arrive at differing statistical conclusions, and therefore biological inferences, simply from a dataspace rotation, like principal component analysis. By contrast, algebraic generalizations of the standard phylogenetic comparative toolkit that use the trace of covariance matrices are insensitive to levels of trait covariation, the number of trait dimensions, and the orientation of the dataset. Further, when appropriate permutation tests are used, these approaches display acceptable Type I error and statistical power. We conclude that methods summarizing information across trait dimensions, as well as pairwise composite likelihood methods should be avoided, while algebraic generalizations of the phylogenetic comparative toolkit provide a useful means of assessing macroevolutionary patterns in multivariate data. Finally, we discuss areas in which multivariate phylogenetic comparative methods are still in need of future development; namely highly multivariate Ornstein-Uhlenbeck models and approaches for multivariate evolutionary model comparisons.
Data from: Phylogenetic comparative methods for evaluating the evolutionary history of function-valued traits
Phylogenetic comparative methods offer a suite of tools for studying trait evolution. However, most models inherently assume fixed trait values within species. Although some methods can incorporate error around species means, few are capable of accounting for variation driven by environmental or temporal gradients, such as trait responses to abiotic stress or ontogenetic trajectories. Such traits, often referred to as function-valued or infinite-dimensional, are typically expressed as reaction norms, dose–response curves, or time plots and are described by mathematical functions linking independent predictor variables to the trait of interest. Here, I introduce a method for extending ancestral state reconstruction to incorporate function-valued traits in a phylogenetic generalized least squares (PGLS) framework, as well as extensions of this method for testing phylogenetic signal, performing phylogenetic analysis of variance (ANOVA), and testing for correlated trait evolution using recently proposed multivariate PGLS methods. Statistical power of function-valued comparative methods is compared to univariate approaches using data simulations, and the assumptions and challenges of each are discussed in detail.
Data from: Multivariate phylogenetic comparative methods: evaluations, comparisons, and recommendations
Open the record for dataset details and reuse information.
Data from: Phylogenetic comparative methods for evaluating the evolutionary history of function-valued traits
Open the record for dataset details and reuse information.
Data from: Rethinking phylogenetic comparative methods
Open the record for dataset details and reuse information.
Data from: A phylogenetic comparative method for evaluating trait coevolution across two phylogenies for sets of interacting species
Open the record for dataset details and reuse information.
Data from: A penalized likelihood framework for high- dimensional phylogenetic comparative methods and an application to new-world monkeys brain evolution
Working with high-dimensional phylogenetic comparative datasets is challenging because likelihood-based multivariate methods suffer from low statistical performances as the number of traits p approaches the number of species n and because some computational complications occur when p exceeds n. Alternative phylogenetic comparative methods have recently been proposed to deal with the large p small n scenario but their use and performances are limited. Here we develop a penalized likelihood framework to deal with high-dimensional comparative datasets. We propose various penalizations and methods for selecting the intensity of the penalties. We apply this general framework to the estimation of parameters (the evolutionary trait covariance matrix and parameters of the evolutionary model) and model comparison for the high-dimensional multivariate Brownian (BM), Early-burst (EB), Ornstein-Uhlenbeck (OU) and Pagel's lambda models. We show using simulations that our penalized likelihood approach dramatically improves the estimation of evolutionary trait covariance matrices and model parameters when p approaches n, and allows for their accurate estimation when p equals or exceeds n. In addition, we show that penalized likelihood models can be efficiently compared using Generalized Information Criterion (GIC). We implement these methods, as well as the related estimation of ancestral states and the computation of phylogenetic PCA in the R package RPANDA and mvMORPH. Finally, we illustrate the utility of the new proposed framework by evaluating evolutionary models fit, analyzing integration patterns, and reconstructing evolutionary trajectories for a high-dimensional 3-D dataset of brain shape in the New World monkeys. We find a clear support for an Early-burst model suggesting an early diversification of brain morphology during the ecological radiation of the clade. Penalized likelihood offers an efficient way to deal with high-dimensional multivariate comparative data.
Data from: A novel Bayesian method for inferring and interpreting the dynamics of adaptive landscapes from phylogenetic comparative data
Our understanding of macroevolutionary patterns of adaptive evolution has greatly increased with the advent of large-scale phylogenetic comparative methods. Widely used Ornstein-Uhlenbeck (OU) models can describe an adaptive process of divergence and selection. However, inference of the dynamics of adaptive landscapes from comparative data is complicated by interpretational difficulties, lack of identifiability among parameter values and the common requirement that adaptive hypotheses must be assigned a priori. Here we develop a reversible-jump Bayesian method of fitting multi-optima OU models to phylogenetic comparative data that estimates the placement and magnitude of adaptive shifts directly from the data. We show how biologically informed hypotheses can be tested against this inferred posterior of shift locations using Bayes Factors to establish whether our a priori models adequately describe the dynamics of adaptive peak shifts. Furthermore, we show how the inclusion of informative priors can be used to restrict models to biologically realistic parameter space and test particular biological interpretations of evolutionary models. We argue that Bayesian model-fitting of OU models to comparative data provides a framework for integrating of multiple sources of biological data–such as microevolutionary estimates of selection parameters and paleontological timeseries–allowing inference of adaptive landscape dynamics with explicit, process-based biological interpretations.
Data from: Likelihood-based parameter estimation for high-dimensional phylogenetic comparative models: overcoming the limitations of 'distance-based' methods
Open the record for dataset details and reuse information.
Data from: A penalized likelihood framework for high- dimensional phylogenetic comparative methods and an application to new-world monkeys brain evolution
Open the record for dataset details and reuse information.
Data from: A novel Bayesian method for inferring and interpreting the dynamics of adaptive landscapes from phylogenetic comparative data
Open the record for dataset details and reuse information.
FIGURE 4 in Comparative morphology of the eggs from the eight species in the genus Agathemera Stål (Insecta: Phasmatodea), through phylogenetic comparative method approach
FIGURE 4. Distribution of height/length ratio. the species are ordered following the phylogenetic relationships. Boxes represent the values between the 25 and 75 percentiles respectively, the horizontal line is the median and the point within each box is the mean and the whiskers indicate the sample range. Light grey boxplots correspond to species from clade 1 and dark grey boxplots correspond to species from clade 2. Different letters indicate significant differences (α = 0.05) in multiple comparisons. * indicates that differences are marginally significant (p = 0.043, see table 3), however the letters indicate no differences between those group, for simplicity in the coding.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.