Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
68
datasets available to search
ShareScore release 0.9.0
Dataset results
68 results for “phylogenetic comparative data”
Data from: Gill evolution in Neotropical electric fishes: Comparative phylogenetic evidence for hypoxia-driven adaptation
Open the record for dataset details and reuse information.
Data and code for: Feeding, mating, and animal wellbeing: New insights from Phylogenetic Comparative Methods
Open the record for dataset details and reuse information.
Data from: The evolution of reproductive diversity in Afrobatrachia: a phylogenetic comparative analysis of an extensive radiation of African frogs
The reproductive modes of anurans (frogs and toads) are the most diverse of terrestrial vertebrates, and a major challenge is identifying selective factors that promote the evolution or retention of reproductive modes across clades. Terrestrialized anuran breeding strategies have evolved repeatedly from the plesiomorphic fully aquatic reproductive mode, a process thought to occur through intermediate reproductive stages. Several selective forces have been proposed for the evolution of terrestrialized reproductive traits, but factors such as water systems and co-evolution with ecomorphologies have not been investigated. We examined these topics in a comparative phylogenetic framework using Afrobatrachian frogs, an ecologically and reproductively diverse clade representing more than half of the total frog diversity found in Africa (∼400 species). We infer direct development has evolved twice independently from terrestrialized reproductive modes involving subterranean or terrestrial oviposition, supporting evolution through intermediate stages. We also detect associations between specific ecomorphologies and oviposition sites, and demonstrate arboreal species exhibit an overall shift towards using lentic water systems for breeding. These results indicate that changes in microhabitat use associated with ecomorphology, which allow access to novel sites for reproductive behavior, oviposition, or larval development, may also promote reproductive mode diversity in anurans.
Data from: Rethinking phylogenetic comparative methods
As a result of the process of descent with modification, closely related species tend to be similar to one another in a myriad different ways. In statistical terms, this means that traits measured on one species will not be independent of traits measured on others. Since their introduction in the 1980s, phylogenetic comparative methods (PCMs) have been framed as a solution to this problem. In this paper, we argue that this way of thinking about PCMs is deeply misleading. Not only has this sowed widespread confusion in the literature about what PCMs are doing but has led us to develop methods that are susceptible to the very thing we sought to build defenses against --- unreplicated evolutionary events. Through three Case Studies, we demonstrate that the susceptibility to singular events is indeed a recurring problem in comparative biology that links several seemingly unrelated controversies. In each Case Study we propose a potential solution to the problem. While the details of our proposed solutions differ, they share a common theme: unifying hypothesis testing with data-driven approaches (which we term ``phylogenetic natural history'') to disentangle the impact of singular evolutionary events from that of the factors we are investigating. More broadly, we argue that our field has, at times, been sloppy when weighing evidence in support of causal hypotheses. We suggest that one way to refine our inferences is to re-imagine phylogenies as probabilistic graphical models; adopting this way of thinking will help clarify precisely what we are testing and what evidence supports our claims.
Data from: A phylogenetic comparative method for evaluating trait coevolution across two phylogenies for sets of interacting species
Evaluating trait correlations across species within a lineage via phylogenetic regression is fundamental to comparative evolutionary biology, but when traits of interest are derived from two sets of lineages that co-evolve with one another, methods for evaluating such patterns in a dual-phylogenetic context remain underdeveloped. Here we extend multivariate permutation-based phylogenetic regression to evaluate trait correlations in two sets of interacting species while accounting for their respective phylogenies. This extension is appropriate for both univariate and multivariate response data, and may utilize one or more independent variables, including environmental covariates. Imperfect correspondence between species in the interacting lineages can also be accommodated, such as when species in one lineage associate with multiple species in the other, or when there are unmatched taxa in one or both lineages. For both univariate and multivariate data, the method displays appropriate type I error, and statistical power increases with the strength of the trait covariation and the number of species in the phylogeny. These properties are retained even when there is not a 1:1 correspondence between lineages. Finally, we demonstrate the approach by evaluating the evolutionary correlation between traits in fig species and traits in their agaonid wasp pollinators. R computer code is provided.
Data from: Multivariate phylogenetic comparative methods: evaluations, comparisons, and recommendations
Recent years have seen increased interest in phylogenetic comparative analyses of multivariate datasets, but to date the varied proposed approaches have not been extensively examined. Here we review the mathematical properties required of any multivariate method, and specifically evaluate existing multivariate phylogenetic comparative methods in this context. Phylogenetic comparative methods based on the full multivariate likelihood are robust to levels of covariation among trait dimensions and are insensitive to the orientation of the dataset, but display increasing model misspecification as the number of trait dimensions increases. This is because the expected evolutionary covariance matrix (V) used in the likelihood calculations becomes more ill-conditioned as trait dimensionality increases, and as evolutionary models become more complex. Thus, these approaches are only appropriate for datasets with few traits and many species. Methods that summarize patterns across trait dimensions treated separately (e.g., SURFACE) incorrectly assume independence among trait dimensions, resulting in nearly a 100% model misspecification rate. Methods using pairwise composite likelihood are highly sensitive to levels of trait covariation, the orientation of the dataset, and the number of trait dimensions. The consequences of these debilitating deficiencies is that a user can arrive at differing statistical conclusions, and therefore biological inferences, simply from a dataspace rotation, like principal component analysis. By contrast, algebraic generalizations of the standard phylogenetic comparative toolkit that use the trace of covariance matrices are insensitive to levels of trait covariation, the number of trait dimensions, and the orientation of the dataset. Further, when appropriate permutation tests are used, these approaches display acceptable Type I error and statistical power. We conclude that methods summarizing information across trait dimensions, as well as pairwise composite likelihood methods should be avoided, while algebraic generalizations of the phylogenetic comparative toolkit provide a useful means of assessing macroevolutionary patterns in multivariate data. Finally, we discuss areas in which multivariate phylogenetic comparative methods are still in need of future development; namely highly multivariate Ornstein-Uhlenbeck models and approaches for multivariate evolutionary model comparisons.
Data from: Testing hypotheses of marsupial brain size variation using phylogenetic multiple imputations and a Bayesian comparative framework
<p>Considerable controversy exists about which hypotheses and variables best explain mammalian brain size variation. We use a new, high-coverage dataset of marsupial brain and body sizes, and the first phylogenetically imputed full datasets of 16 predictor variables, to model the prevalent hypotheses explaining brain size evolution using phylogenetically corrected Bayesian generalised linear mixed-effects modelling. Despite this comprehensive analysis, litter size emerges as the only significant predictor. Marsupials differ from the more frequently studied placentals in displaying much lower diversity of reproductive traits, which are known to interact extensively with many behavioural and ecological predictors of brain size. Our results therefore suggest that studies of relative brain size evolution in placental mammals may require targeted co-analysis or adjustment of reproductive parameters like litter size, weaning age, or gestation length. This supports suggestions that significant associations between behavioural or ecological variables with relative brain size may be due to a confounding influence of the extensive reproductive diversity of placental mammals.</p>
Data from: Phylogenetic comparative methods for evaluating the evolutionary history of function-valued traits
Phylogenetic comparative methods offer a suite of tools for studying trait evolution. However, most models inherently assume fixed trait values within species. Although some methods can incorporate error around species means, few are capable of accounting for variation driven by environmental or temporal gradients, such as trait responses to abiotic stress or ontogenetic trajectories. Such traits, often referred to as function-valued or infinite-dimensional, are typically expressed as reaction norms, dose–response curves, or time plots and are described by mathematical functions linking independent predictor variables to the trait of interest. Here, I introduce a method for extending ancestral state reconstruction to incorporate function-valued traits in a phylogenetic generalized least squares (PGLS) framework, as well as extensions of this method for testing phylogenetic signal, performing phylogenetic analysis of variance (ANOVA), and testing for correlated trait evolution using recently proposed multivariate PGLS methods. Statistical power of function-valued comparative methods is compared to univariate approaches using data simulations, and the assumptions and challenges of each are discussed in detail.
Data from: Multivariate phylogenetic comparative methods: evaluations, comparisons, and recommendations
Open the record for dataset details and reuse information.
Data from: Phylogenetic comparative methods for evaluating the evolutionary history of function-valued traits
Open the record for dataset details and reuse information.
Data from: Testing hypotheses of marsupial brain size variation using phylogenetic multiple imputations and a Bayesian comparative framework
Open the record for dataset details and reuse information.
Data from: Rethinking phylogenetic comparative methods
Open the record for dataset details and reuse information.
Data from: A phylogenetic comparative method for evaluating trait coevolution across two phylogenies for sets of interacting species
Open the record for dataset details and reuse information.
Data from: The evolution of reproductive diversity in Afrobatrachia: a phylogenetic comparative analysis of an extensive radiation of African frogs
Open the record for dataset details and reuse information.
Data from: A penalized likelihood framework for high- dimensional phylogenetic comparative methods and an application to new-world monkeys brain evolution
Working with high-dimensional phylogenetic comparative datasets is challenging because likelihood-based multivariate methods suffer from low statistical performances as the number of traits p approaches the number of species n and because some computational complications occur when p exceeds n. Alternative phylogenetic comparative methods have recently been proposed to deal with the large p small n scenario but their use and performances are limited. Here we develop a penalized likelihood framework to deal with high-dimensional comparative datasets. We propose various penalizations and methods for selecting the intensity of the penalties. We apply this general framework to the estimation of parameters (the evolutionary trait covariance matrix and parameters of the evolutionary model) and model comparison for the high-dimensional multivariate Brownian (BM), Early-burst (EB), Ornstein-Uhlenbeck (OU) and Pagel's lambda models. We show using simulations that our penalized likelihood approach dramatically improves the estimation of evolutionary trait covariance matrices and model parameters when p approaches n, and allows for their accurate estimation when p equals or exceeds n. In addition, we show that penalized likelihood models can be efficiently compared using Generalized Information Criterion (GIC). We implement these methods, as well as the related estimation of ancestral states and the computation of phylogenetic PCA in the R package RPANDA and mvMORPH. Finally, we illustrate the utility of the new proposed framework by evaluating evolutionary models fit, analyzing integration patterns, and reconstructing evolutionary trajectories for a high-dimensional 3-D dataset of brain shape in the New World monkeys. We find a clear support for an Early-burst model suggesting an early diversification of brain morphology during the ecological radiation of the clade. Penalized likelihood offers an efficient way to deal with high-dimensional multivariate comparative data.
Data from: Characterizing and comparing phylogenetic trait data from their normalized Laplacian spectrum
The dissection of the mode and tempo of phenotypic evolution is integral to our understanding of global biodiversity. Our ability to infer patterns of phenotypes across phylogenetic clades is essential to how we infer the macroevolutionary processes governing those patterns. Many methods are already available for fitting models of phenotypic evolution to data. However, there is currently no comprehensive non-parametric framework for characterising and comparing patterns of phenotypic evolution. Here we build on a recently introduced approach for using the phylogenetic spectral density profile to compare and characterize patterns of phylogenetic diversification, in order to provide a framework for non-parametric analysis of phylogenetic trait data. We show how to construct the spectral density profile of trait data on a phylogenetic tree from the normalized graph Laplacian. We demonstrate on simulated data the utility of the spectral density profile to successfully cluster phylogenetic trait data into meaningful groups and to characterise the phenotypic patterning within those groups. We furthermore demonstrate how the spectral density profile is a powerful tool for visualising phenotypic space across traits and for assessing whether distinct trait evolution models are distinguishable on a given empirical phylogeny. We illustrate the approach in two empirical datasets: a comprehensive dataset of traits involved in song, plumage and resource-use in tanagers, and a high-dimensional dataset of endocranial landmarks in New World monkeys. Considering the proliferation of morphometric and molecular data collected across the tree of life, we expect this approach will benefit big data analyses requiring a comprehensive and intuitive framework.
Data from: Interpreting the evolutionary regression: the interplay between observational and biological errors in phylogenetic comparative studies
Regressions of biological variables across species are rarely perfect. Usually there are residual deviations from the estimated model relationship, and such deviations commonly show a pattern of phylogenetic correlations indicating that they have biological causes. We discuss the origins and effects of phylogenetically correlated biological variation in regression studies. In particular, we discuss the interplay of biological deviations with deviations due to observational or measurement errors, which are also important in comparative studies based on estimated species means. We show how bias in estimated evolutionary regressions can arise from several sources, including phylogenetic inertia and either observational or biological error in the predictor variables. We show how all these biases can be estimated and corrected for in the presence of phylogenetic correlations. We present general formulas for incorporating measurement error in linear models with correlated data. We also show how alternative regression models, such as major-axis and reduced major-axis regression, which are often recommended when there is error in predictor variables, are strongly biased when there is biological variation in any part of the model. We argue that such methods should never be used to estimate evolutionary or allometric regression slopes.
Data from: Kakusan4 and Aminosan: two programs for comparing nonpartitioned, proportional, and separate models for combined molecular phylogenetic analyses of multilocus sequence data
Proportional and separate models able to apply different combination of substitution rate matrix and among-site rate variation model to each locus are frequently used in phylogenetic studies of multilocus data. However, the selection from among nonpartitioned (i.e., a common combination of models is applied to all-loci concatenated sequences), proportional, and separate models is usually based on the researcher's preference rather than on any information criteria. The present study describes two programs, "Kakusan4" (for DNA sequences) and "Aminosan" (for amino-acid sequences), that allow the selection of evolutionary models based on several types of information criteria. The programs can handle both multilocus and single-locus data, in addition to providing an easy-to-use wizard interface and a non-interactive command line interface. In the case of multilocus data, substitution rate matrices and among-site rate variation models are compared at each locus and at all-loci concatenated sequences, after which nonpartitioned, proportional, and separate models are compared based on information criteria. The programs also provide model configuration files for MrBayes, PAUP*, PHYML, RAxML, and Treefinder to support further phylogenetic analysis using a selected model. The best-fit models were found to differ depending on the data set. Furthermore, differences in the information criteria among nonpartitioned, proportional, and separate models were much larger than those among the nonpartitioned models. These findings suggest that selecting from nonpartitioned, proportional, and separate models results in a better phylogenetic tree. Kakusan4 and Aminosan are available at http://www.fifthdimension.jp/. They are licensed under GNU GPL Ver.2, and are able to run on Windows, MacOS X, and Linux.
Data from: Comparing the rates of speciation and extinction between phylogenetic trees
Over the past decade or so it has become increasingly popular to use reconstructed evolutionary trees to investigate questions about the rates of speciation and extinction. Although the methodology of this field has grown substantially in its sophistication in recent years, here I'll take a step back to present a very simple model that is designed to investigate the relatively straightforward question of whether the tempo of diversification (speciation and extinction) differs between two or more phylogenetic trees, without attempting to attribute a causal basis to this difference. It is a likelihood method, and I demonstrate that it generally shows type I error that is close to the nominal level. I also demonstrate that parameter estimates obtained with this approach are largely unbiased. Since this method can be used to compare trees of unknown relationship, it will be particularly well-suited to problems in which a difference in diversification rate between clades is suspected, but in which these clades are not particularly closely related. Since diversification methods can easily take into account an incomplete sampling fraction, but missing lineages are assumed to be missing at random, this method is also appropriate for cases in which we've hypothesized a difference in the process of diversification between two or more focal clades, but in which many un-sampled groups separate the few of interest. The method of this study is by no means an attempt to replace more sophisticated models in which, for instance, diversification depends on the state of an observed or unobserved discrete or continuous trait. Rather, my intention is to provide a complementary approach for circumstances in which a simpler hypothesis is warranted and of biological interest.
Data from: A unifying comparative phylogenetic framework including traits coevolving across interacting lineages
Models of phenotypic evolution fit to phylogenetic comparative data are widely used to make inferences regarding the tempo and mode of trait evolution. A wide range of models is already available for this type of analysis, and the field is still under active development. One of the most needed development concerns models that better account for the effect of within- and between-clade interspecific interactions on trait evolution, which can result from processes as diverse as competition, predation, parasitism, or mutualism. Here, we begin by developing a very general comparative phylogenetic framework for (multi)-trait evolution that can be applied to both ultrametric and nonultrametric trees. This framework not only encapsulates many previous models of continuous univariate and multivariate phenotypic evolution, but also paves the way for the consideration of a much broader series of models in which lineages coevolve, meaning that trait changes in one lineage are influenced by the value of traits in other, interacting lineages. Next, we provide a standard way for deriving the probabilistic distribution of traits at tip branches under our framework. We show that a multivariate normal distribution remains the expected distribution for a broad class of models accounting for interspecific interactions. Our derivations allow us to fit various models efficiently, and in particular greatly reduce the computation time needed to fit the recently proposed phenotype matching model. Finally, we illustrate the utility of our framework by developing a toy model for mutualistic coevolution. Our framework should foster a new era in the study of coevolution from comparative data.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.