Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.9.0
Dataset results
116 results for “Phylogenetics: methods”
Fig. 2 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 2. Schematical overview of the workflow for the present study.Workflow proceeds from left to right and top to bottom.The tablets correspond to the three broader segments of the study: genomic resource generation, probe design, and in silico testing; within tablet boundaries can be found associated taxon sets, data, experimentation, and results. Arrows indicate the flow of data, associated results, and the location results can ultimately be found. Color corresponds to membership within a taxon set or probe set.The red boxes around Hydro 2.7kv1 and Scarab 3kv1 denote final optimized, tailored probe set design tailored for Hydrophiloidea and Scarabaeidae based on the results of this study. Length of 75CM box corresponds to alignment length. Abbreviations used: NCBI, National Center for Biotechnology Information; 75CM, 75% complete matrix; AMAS, alignment manipulation and summary statistics (Borowiec 2016); R-F, Robinson– Foulds distance (Robinson and Foulds 1981).
Data from: Assessing species boundaries and the phylogenetic position of the rare Szechwan Ratsnake, Euprepiophis perlacea (Serpentes: Colubridae), using coalescent-based methods
Open the record for dataset details and reuse information.
Data from: Multivariate phylogenetic comparative methods: evaluations, comparisons, and recommendations
Open the record for dataset details and reuse information.
Data from: A method for assessing phylogenetic least squares models for shape and other high-dimensional multivariate data
Open the record for dataset details and reuse information.
Data from: Phylogenetic comparative methods for evaluating the evolutionary history of function-valued traits
Open the record for dataset details and reuse information.
Data from: The impact of reconstruction methods, phylogenetic uncertainty and branch lengths on inference of chromosome number evolution in American daisies (Melampodium, Asteraceae)
Open the record for dataset details and reuse information.
Data from: Environmental conditions and biotic interactions acting together promote phylogenetic randomness in semi-arid plant communities: new methods help to avoid misleading conclusions
Open the record for dataset details and reuse information.
Data from: The evolutionary relationships and age of Homo naledi: an assessment using dated Bayesian phylogenetic methods
Open the record for dataset details and reuse information.
Data from: Predicting community structure in snakes on Eastern Nearctic islands using ecological neutral theory and phylogenetic methods
Open the record for dataset details and reuse information.
Data for: Theropod dinosaur diversity of the lower English Wealden: analysis of a tooth-based fauna from the Wadhurst Clay Formation (Lower Cretaceous: Valanginian) via phylogenetic, discriminant and machine learning methods
Open the record for dataset details and reuse information.
Data from: Rethinking phylogenetic comparative methods
Open the record for dataset details and reuse information.
Data from: A phylogenetic comparative method for evaluating trait coevolution across two phylogenies for sets of interacting species
Open the record for dataset details and reuse information.
A new method for integrating ecological niche modeling with phylogenetics to estimate ancestral distributions
Open the record for dataset details and reuse information.
Model trees and associated simulated nucleotide sequences for testing phylogenetic inference methods
<p>This repository contains 142 tar.gz archive files, each containing nucleotide sequence data that have been simulated using <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> for testing alignment-free phylogenetic inference methods. These datasets were generated by using the results (trees and model parameters) of 142 phylogenomic analyses of real-case data as model (available <a href="https://zenodo.org/record/4034261">here</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>Each archive contains the following files/directories:</p> <ul> <li><code>GTR.params.trees.tsv </code> a tab-delimited file summarizing the real-case GTR+Γ model parameters and the phylogenetic tree used to simulate the sequence dataset (gathered from <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>)</li> <li><code>tax.tsv </code> a tab-delimited file containing the initial (col 1) and simplified (col 2) taxon names</li> <li><code>model.nwk </code> a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing the initial model tree (gathered from <code>GTR.params.trees.tsv</code>) with simplified leaf names (following <code>tax.tsv</code>)</li> <li><code>control.txt </code> the <em>INDELible</em> input file used to simulate the evolution of a sequence along the tree in <code>model.nwk</code></li> <li><code>seq/ </code> a directory containing the simulated sequences (one FASTA file per leaf in the tree in <code>model.nwk</code>)</li> </ul> <p>___</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Data from: A penalized likelihood framework for high- dimensional phylogenetic comparative methods and an application to new-world monkeys brain evolution
Working with high-dimensional phylogenetic comparative datasets is challenging because likelihood-based multivariate methods suffer from low statistical performances as the number of traits p approaches the number of species n and because some computational complications occur when p exceeds n. Alternative phylogenetic comparative methods have recently been proposed to deal with the large p small n scenario but their use and performances are limited. Here we develop a penalized likelihood framework to deal with high-dimensional comparative datasets. We propose various penalizations and methods for selecting the intensity of the penalties. We apply this general framework to the estimation of parameters (the evolutionary trait covariance matrix and parameters of the evolutionary model) and model comparison for the high-dimensional multivariate Brownian (BM), Early-burst (EB), Ornstein-Uhlenbeck (OU) and Pagel's lambda models. We show using simulations that our penalized likelihood approach dramatically improves the estimation of evolutionary trait covariance matrices and model parameters when p approaches n, and allows for their accurate estimation when p equals or exceeds n. In addition, we show that penalized likelihood models can be efficiently compared using Generalized Information Criterion (GIC). We implement these methods, as well as the related estimation of ancestral states and the computation of phylogenetic PCA in the R package RPANDA and mvMORPH. Finally, we illustrate the utility of the new proposed framework by evaluating evolutionary models fit, analyzing integration patterns, and reconstructing evolutionary trajectories for a high-dimensional 3-D dataset of brain shape in the New World monkeys. We find a clear support for an Early-burst model suggesting an early diversification of brain morphology during the ecological radiation of the clade. Penalized likelihood offers an efficient way to deal with high-dimensional multivariate comparative data.
Data from: Phylogenetic tree estimation with and without alignment: new distance methods and benchmarking
Phylogenetic tree inference is a critical component of many systematic and evolutionary studies. The majority of these studies are based on the two-step process of multiple sequence alignment followed by tree inference, despite persistent evidence that the alignment step can lead to biased results. Here we present a two-part study that first presents PaHMM-Tree, a novel neighbour joining-based method that estimates pairwise distances without assuming a single alignment. We then use simulations to benchmark its performance against a wide-range of other phylogenetic tree inference methods, including the first comparison of alignment-free distance-based methods against more conventional tree estimation methods. Our new method for calculating pairwise distances based on statistical alignment provides distance estimates that are as accurate as those obtained using standard methods based on the true alignment. Pairwise distance estimates based on the two-step process tend to be substantially less accurate. This improved performance carries through to tree inference, where PaHMM-Tree provides more accurate tree estimates than all of the pairwise distance methods assessed. For close to moderately divergent sequence data we find that the two-step methods using statistical inference, where information from all sequences is included in the estimation procedure, tend to perform better than PaHMM-Tree, particularly full statistical alignment, which simultaneously estimates both the tree and the alignment. For deep divergences we find the alignment step becomes so prone to error that our distance-based PaHMM-Tree outperforms all other methods of tree inference. Finally, we find that the accuracy of alignment-free methods tends to decline faster than standard two-step methods in the presence of alignment uncertainty, and identify no conditions where alignment-free methods are equal to or more accurate than standard phylogenetic methods even in the presence of substantial alignment error.
Data from: A parametric method for assessing diversification rate variation in phylogenetic trees
Phylogenetic hypotheses are frequently used to examine variation in rates of diversification across the history of a group. Patterns of diversification-rate variation can be used to infer underlying ecological and evolutionary processes responsible for patterns of cladogenesis. Most existing methods examine rate variation through time. Methods for examining differences in diversification among groups are more limited. Here we present a new method, parametric rate comparison (PRC), that explicitly compares diversification rates among lineages in a tree using a variety of standard statistical distributions. PRC can identify subclades of the tree where diversification-rates are at variance with the remainder of the tree. A randomization test can be used to evaluate how often such variance would appear by chance alone. The method also allows for comparison of diversification-rate among a priori defined groups. Further, the application of the PRC method is not restricted to monophyletic groups. We examined the performance of PRC using simulated data which showed that PRC has acceptable false positive rates and statistical power to detect rate variation. We apply the PRC method to the well-studied radiation of North American Plethodon salamanders, and support the inference that the large-bodied P. glutinosus clade has a higher historical rate of diversification compared to other Plethodon salamanders.
Data from: The impact of phylogenetic dating method on interpreting trait evolution: a case study of Cretaceous–Palaeogene eutherian body-size evolution
The fossil record of the earliest Cenozoic contains the first large-bodied placental mammals. Several evolutionary models have been invoked to explain the transition from small to large body sizes, but methods for determining evolutionary mode of trait change depend on input from tree topology and divergence dates. Different dating methods may therefore affect inference of evolutionary model. Here, we fit models of body mass evolution onto dated phylogenies of Cretaceous and Palaeogene mammals, comparing the effect of dating method on interpretation of evolutionary model. Among traditional palaeontological dating approaches, an Ornstein–Uhlenbeck model with high alpha parameters is recovered as best-fitting when minimum-age dating is used, while branch-sharing methods are highly sensitive to topology. Release or release–radiate models are preferred when Bayesian fossilized birth–death method is used, but when using stochastic cal3 dating of trees, a model of increased evolutionary rate without a release in constraint at the Cretaceous–Palaeogene boundary has highest support. These results demonstrate unambiguously that choice of dating method is critical for interpretation of continuous trait evolution, and that care must therefore be taken to consider these effects in macroevolutionary studies.
Data from: The efficacy of consensus tree methods for summarising phylogenetic relationships from a posterior sample of trees estimated from morphological data
Consensus trees are required to summarise trees obtained through MCMC sampling of a posterior distribution, providing an overview of the distribution of estimated parameters such as topology, branch lengths and divergence times. Numerous consensus tree construction methods are available, each presenting a different interpretation of the tree sample. The rise of morphological clock and sampled-ancestor methods of divergence time estimation, in which times and topology are co-estimated, has increased the popularity of the maximum clade credibility (MCC) consensus tree method. The MCC method assumes that the sampled, fully resolved topology with the highest clade credibility contains an adequate summary of the most probable clades, with parameter estimates from compatible sampled trees used to obtain the marginal distributions of parameters such as clade ages and branch lengths. Using both simulated and empirical data, we demonstrate that MCC trees, and trees constructed using the similar maximum a posteriori (MAP) method, often include poorly supported and incorrect clades when summarising diffuse posterior samples of trees. We demonstrate that the paucity of information in morphological datasets contributes to the inability of MCC and MAP trees to present an accurate summary of the posterior distribution. Conversely, majority-rule consensus (MRC) trees report a lower proportion of incorrect nodes when summarising the same posterior samples of trees. Thus, we advocate the use of MRC trees, in place of MCC or MAP trees, in attempts to summarise the results of Bayesian phylogenetic analyses of morphological data.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.