Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.9.0
Dataset results
116 results for “Phylogenetics: methods”
Data from: Testing Phylogenetic Methods with Tree Congruence: Phylogenetic Analysis of Polymorphic Morphological Characters in Phrynosomatid Lizards
Congruence between trees from separately analyzed data sets is a powerful approach for assessing the performance of phylogenetic methods but has been applied primarily to the analysis of molecular data. In this study, different methods for treating polymorphic characters were compared using morphological data from phrynosomatid lizards. Clades were identified that are both traditionally recognized and supported by recent molecular analyses, and species were sampled from these clades to make three RknownS phylogenies of eight species each. The ability of different methods to estimate these "known" phylogenies with a finite sample of characters was tested. The phylogenetic methods included eight parsimony methods for coding polymorphism, three distance approaches (UPGMA, neighbor joining, and Fitch-Margoliash) applied to two genetic distance measures (Nei's and the modified Cavalli-Sforza and Edwards chord distance), and continuous maximum likelihood. The effects of excluding polymorphic characters and character weighting (a priori and successive) were also tested. Among the different parsimony approaches, the fixed-only method (excluding all polymorphic characters) performed relatively poorly, whereas the frequency method (including all polymorphic characters) performed relatively well. However, frequency-based distance methods consistently outperformed parsimony, especially with a small sample size (n= 1 individual per species). These results agree closely with those from recent simulation studies of polymorphic data and argue against the common practices of excluding polymorphic morphological characters, ignoring the frequencies of traits within species, and the exclusive use of parsimony to analyze morphological data.
Data from: A target enrichment method for gathering phylogenetic information from hundreds of loci: an example from the Compositae
Premise of the study: The Compositae (Asteraceae) are a large and diverse family of plants, and the most comprehensive phylogeny to date is a meta-tree based on 10 chloroplast loci that has several major unresolved nodes. We describe the development of an approach that enables the rapid sequencing of large numbers of orthologous nuclear loci to facilitate efficient phylogenomic analyses. Methods and Results: We designed a set of sequence capture probes that target conserved orthologous sequences in the Compositae. We also developed a bioinformatic and phylogenetic workflow for processing and analyzing the resulting data. Application of our approach to 15 species from across the Compositae resulted in the production of phylogenetically informative sequence data from 763 loci and the successful reconstruction of known phylogenetic relationships across the family. Conclusions: These methods should be of great use to members of the broader Compositae community, and the general approach should also be of use to researchers studying other families.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Fossil taxa are critical to inferences of historical diversity and the origins of modern biodiversity, but realizing their evolutionary significance is contingent on restoring fossil species to their correct position within the tree of life. For most fossil species, morphology is the only source of data for phylogenetic inference; this has traditionally been analysed using parsimony, the predominance of which is currently challenged by the development of probabilistic models that achieve greater phylogenetic accuracy. Here, based on simulated and empirical datasets, we explore the relative efficacy of competing phylogenetic methods in terms of clade support. We characterize clade support using bootstrapping for parsimony and Maximum Likelihood, and intrinsic Bayesian posterior probabilities, collapsing branches that exhibit less than 50% support. Ignoring node support, Bayesian inference is the most accurate method in estimating the tree used to simulate the data. After assessing clade support, Bayesian and Maximum Likelihood exhibit comparable levels of accuracy, and parsimony remains the least accurate method. However, Maximum Likelihood is less precise than Bayesian phylogeny estimation, and Bayesian inference recaptures more correct nodes with higher support compared to all other methods, including Maximum Likelihood. We assess the effects of these findings on empirical phylogenies. Our results indicate probabilistic methods should be favoured over parsimony.
Data from: A comparison of supermatrix and supertree methods for multilocus phylogenetics using organismal datasets
It has been proposed that supertree approaches should be applied to large multilocus sequence datasets to achieve computational tractability. Large datasets such as those derived from phylogenomics studies can be broken into many locus-specific tree searches and the resulting trees can be stitched together via a supertree method. Using simulated data, workers have reported that they can rapidly construct a supertree that is comparable to the results of heuristic tree search on the entire dataset. To test this assertion with organismal data, we compared tree length under the parsimony criterion and computational time for twenty multilocus datasets using supertree (SuperFine and SuperTriplets) and supermatrix (heuristic search in TNT) approaches. Tree length and computational times were compared among methods using the Wilcoxon matched-pairs signed rank test. Supermatrix searches produce significantly shorter trees than either supertree approach (SuperFine or SuperTriplets; p < 0.0002 in both cases). Moreover, the processing time of supermatrix search was significantly lower than SuperFine+locus-specific search (p < 0.01) but roughly equivalent to that of SuperTriplets+locus-specific search (p > 0.4, not significant). In conclusion, we show by using real rather than simulated data, that there is no basis, either in time tractability or tree length, for use of supertrees over heuristic tree search using a supermatrix for phylogenomics.
Data from: An improved hypergeometric probability method for identification of functionally linked proteins using phylogenetic profiles
Predicting functions of proteins and alternatively spliced isoforms encoded in a genome is one of the important applications of bioinformatics in the post-genome era. Due to the practical limitation of experimental characterization of all proteins encoded in a genome using biochemical studies, bioinformatics methods provide powerful tools for function annotation and prediction. These methods also help minimize the growing sequence-to-function gap. Phylogenetic profiling is a bioinformatics approach to identify the influence of a trait across species and can be employed to infer the evolutionary history of proteins encoded in genomes. Here we propose an improved phylogenetic profile-based method which considers the co-evolution of the reference genome to derive the basic similarity measure, the background phylogeny of target genomes for profile generation and assigning weights to target genomes. The ordering of genomes and the runs of consecutive matches between the proteins were used to define phylogenetic relationships in the approach. We used Escherichia coli K12 genome as the reference genome and its 4195 proteins were used in the current analysis. We compared our approach with two existing methods and our initial results show that the predictions have outperformed two of the existing approaches. In addition, we have validated our method using a targeted protein-protein interaction network derived from protein-protein interaction database STRING. Our preliminary results indicates that improvement in function prediction can be attained by using coevolution-based similarity measures and the runs on to the same scale instead of computing them in different scales. Our method can be applied at the whole-genome level for annotating hypothetical proteins from prokaryotic genomes.
Data from: Implementing and testing Bayesian and Maximum likelihood supertree methods in phylogenetics
Since their advent, supertrees have been increasingly used in large-scale evolutionary studies requiring a phylogenetic framework and substantial efforts have been devoted to developing a wide variety of supertree methods (SMs). Recent advances in supertree theory have allowed the implementation of maximum likelihood (ML) and Bayesian SMs, based on using an exponential distribution to model incongruence between input trees and the supertree. Such approaches are expected to have advantages over commonly used non-parametric SMs, e.g. matrix representation with parsimony (MRP). We investigated new implementations of ML and Bayesian SMs and compared these with some currently available alternative approaches. Comparisons include hypothetical examples previously used to investigate biases of SMs with respect to input tree shape and size, and empirical studies based either on trees harvested from the literature or on trees inferred from phylogenomic scale data. Our results provide no evidence of size or shape biases and demonstrate that the Bayesian method is a viable alternative to MRP and other non-parametric methods. Computation of input tree likelihoods allows the adoption of standard tests of tree topologies (e.g. the approximately unbiased test). The Bayesian approach is particularly useful in providing support values for supertree clades in the form of posterior probabilities.
FIGURE 1 in Tanaidacea from Brazil. II. A revision of the subfamily Hemikalliapseudinae (Kalliapseudidae; Tanaidacea; Crustacea) using phylogenetic methods
FIGURE 1. Strict consensus, bremer support values given adjacent to branches.
Data from: Effects of phylogenetic reconstruction method on the robustness of species delimitation using single-locus data
1. Coalescent-based species delimitation methods combine population genetic and phylogenetic theory to provide an objective means for delineating evolutionarily significant units of diversity. The Generalized Mixed Yule Coalescent (GMYC) and the Poisson Tree Process (PTP) are methods that use ultrametric (GMYC or PTP) or non-ultrametric (PTP) gene trees as input, intended for use mostly with single-locus data such as DNA barcodes. 2. Here we assess how robust the GMYC and PTP are to different phylogenetic reconstruction and branch smoothing methods. We reconstruct over 400 ultrametric trees using up to 30 different combinations of phylogenetic and smoothing methods and perform over 2,000 separate species delimitation analyses across 16 empirical datasets. We then assess how variable diversity estimates are, in terms of richness and identity, with respect to species delimitation, phylogenetic and smoothing methods. 3. The PTP method generally generates diversity estimates that are more robust to different phylogenetic methods. The GMYC is more sensitive, but provides consistent estimates for BEAST trees. The lower consistency of GMYC estimates is likely a result of differences among gene trees introduced by the smoothing step. Unresolved nodes (real anomalies or methodological artefacts) affect both GMYC and PTP estimates, but have a greater effect on GMYC estimates. Branch smoothing is a difficult step and perhaps an underappreciated source of bias that may be widespread among studies of diversity and diversification. 4. Nevertheless, careful choice of phylogenetic method does produce equivalent PTP and GMYC diversity estimates. We recommend simultaneous use of the PTP model with any model-based gene tree (e.g. RAxML) and GMYC approaches with BEAST trees for obtaining species hypotheses.
Data from: A novel Bayesian method for inferring and interpreting the dynamics of adaptive landscapes from phylogenetic comparative data
Our understanding of macroevolutionary patterns of adaptive evolution has greatly increased with the advent of large-scale phylogenetic comparative methods. Widely used Ornstein-Uhlenbeck (OU) models can describe an adaptive process of divergence and selection. However, inference of the dynamics of adaptive landscapes from comparative data is complicated by interpretational difficulties, lack of identifiability among parameter values and the common requirement that adaptive hypotheses must be assigned a priori. Here we develop a reversible-jump Bayesian method of fitting multi-optima OU models to phylogenetic comparative data that estimates the placement and magnitude of adaptive shifts directly from the data. We show how biologically informed hypotheses can be tested against this inferred posterior of shift locations using Bayes Factors to establish whether our a priori models adequately describe the dynamics of adaptive peak shifts. Furthermore, we show how the inclusion of informative priors can be used to restrict models to biologically realistic parameter space and test particular biological interpretations of evolutionary models. We argue that Bayesian model-fitting of OU models to comparative data provides a framework for integrating of multiple sources of biological data–such as microevolutionary estimates of selection parameters and paleontological timeseries–allowing inference of adaptive landscape dynamics with explicit, process-based biological interpretations.
Data from: Generalized Frequency Coding: A Method of Preparing Polymorphic Multistate Characters for Phylogenetic Analysis
A new method of coding polymorphic multistate characters for phylogenetic analysis is presented. By dividing such characters into subcharacters, their frequency distributions can be represented with discrete states. Differential weighting is employed to counter the effect of using multiple characters to represent one character. The new method, termed generalized frequency coding (GFC), is potentially superior to previously used methods in that it incorporates more information and can be applied to both qualitative and quantitative characters. The method was applied to a previously published data set that includes both types of polymorphic multistate characters, and performed well according to congruence with other studies and the g1 and nonparametric bootstrap statistics. The data set was also used to compare GFC to both gap-weighting and Manhattan distance step matrix coding. On these grounds and for philosophical reasons, GFC was found to be a better estimator of phylogeny.
Fig. 4 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 4. The proportion of different types of loci targeted by scarab and hydrophiloid probe sets.
Data from: Testing Phylogenetic Methods with Tree Congruence: Phylogenetic Analysis of Polymorphic Morphological Characters in Phrynosomatid Lizards
Open the record for dataset details and reuse information.
Data from: A comparison of supermatrix and supertree methods for multilocus phylogenetics using organismal datasets
Open the record for dataset details and reuse information.
Data from: Using novel phylogenetic methods to evaluate mammalian mtDNA, including amino acid-invariant sites-LogDet plus site stripping, to detect internal conflicts in the data, with special reference to the positions of hedgehog, armadillo, and elephant
Open the record for dataset details and reuse information.
Data from: Geneious! Simplified genome skimming methods for phylogenetic systematic studies: a case study in Oreocarya (Boraginaceae)
Open the record for dataset details and reuse information.
Data from: Generalized Frequency Coding: A Method of Preparing Polymorphic Multistate Characters for Phylogenetic Analysis
Open the record for dataset details and reuse information.
Data from: The efficacy of consensus tree methods for summarising phylogenetic relationships from a posterior sample of trees estimated from morphological data
Open the record for dataset details and reuse information.
Data from: The impact of phylogenetic dating method on interpreting trait evolution: a case study of Cretaceous–Palaeogene eutherian body-size evolution
Open the record for dataset details and reuse information.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Open the record for dataset details and reuse information.
Data from: A parametric method for assessing diversification rate variation in phylogenetic trees
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.