Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.9.0
Dataset results
44 results for “species tree estimation”
Data from: Species tree estimation and the impact of gene loss following whole-genome duplication
Open the record for dataset details and reuse information.
Data from: To include or not to include: the impact of gene filtering on species tree estimation methods
With the increasing availability of whole genome data, many species trees are being constructed from hundreds to thousands of loci. Although concatenation analysis using maximum likelihood is a standard approach for estimating species trees, it does not account for gene tree heterogeneity, which can occur due to many biological processes, such as incomplete lineage sorting. Coalescent species tree estimation methods, many of which are statistically consistent in the presence of incomplete lineage sorting, include Bayesian methods that co-estimate the gene trees and the species tree, summary methods that compute the species tree by combining estimated gene trees, and site-based methods that infer the species tree from site patterns in the alignments of different loci. Due to concerns that poor quality loci will reduce the accuracy of estimated species trees, many recent phylogenomic studies have removed or filtered genes on the basis of phylogenetic signal and/or missing data prior to inferring species trees; little is known about the performance of species tree estimation methods when gene filtering is performed. We examine how incomplete lineage sorting, phylogenetic signal of individual loci, and missing data affect the absolute and the relative accuracy of species tree estimation methods and show how these properties affect methods' responses to gene filtering strategies. In particular, summary methods (ASTRAL-II, ASTRID, and MP-EST), a site-based coalescent method (SVDquartets within PAUP), and an unpartitioned concatenation analysis using maximum likelihood (RAxML) were evaluated on a heterogeneous collection of simulated multi-locus datasets, and the following trends were observed. Filtering genes based on gene tree estimation error improved the accuracy of the summary methods when levels of incomplete lineage sorting were low to moderate but did not benefit the summary methods under higher levels of incomplete lineage sorting, unless gene tree estimation error was also extremely high (a model condition with few replicates). Neither SVDquartets nor concatenation analysis using RAxML benefited from filtering genes on the basis of gene tree estimation error. Finally, filtering genes based on missing data was either neutral (i.e., did not impact accuracy) or else reduced the accuracy of all five methods. By providing insight into the consequences of gene filtering, we offer recommendations for estimating species tree in the presence of incomplete lineage sorting and reconcile seemingly conflicting observations made in prior studies regarding the impact of gene filtering.
Data from: Assessing the impacts of positive selection on coalescent-based species tree estimation and species delimitation.
The assumption of strictly neutral evolution is fundamental to the multispecies coalescent model and permits the derivation of gene tree distributions and coalescent times conditioned on a given species tree. In this study, we conduct computer simulations to explore the effects of violating this assumption in the form of species-specific positive selection when estimating species trees, species delimitations, and coalescent parameters under the model. We simulated datasets under an array of evolutionary scenarios that differ in both speciation parameters (i.e., divergence times, strength of selection) and experimental design (i.e., number of loci sampled) and incorporated species-specific positive selection occurring within branches of a species tree to identify the effects of selection on multispecies coalescent inferences. Our results highlight particular evolutionary scenarios and parameter combinations in which inferences may be more, or less, susceptible to the effects of positive selection. In some extreme cases, selection can decrease error in species delimitation and increase error in species tree estimation, yet these inferences appear to be largely robust to the effects of positive selection under many conditions likely to be encountered in empirical datasets.
Data from: Species tree estimation of North American chorus frogs (Hylidae: Pseudacris) with parallel tagged amplicon sequencing
The field of phylogenetics is changing rapidly with the application of high-throughput sequencing to non-model organisms. Cost-effective use of this technology for phylogenetic studies, which often include a relatively small portion of the genome but several taxa, requires strategies for genome partitioning and sequencing multiple individuals in parallel. In this study we estimated a multilocus phylogeny for the North American chorus frog genus Pseudacris using anonymous nuclear loci that were recently developed using a reduced representation library approach. We sequenced 27 nuclear loci and three mitochondrial loci for 44 individuals on 1/3 of an Illumina MiSeq run, obtaining 96.5% of the targeted amplicons at less than 20% of the cost of traditional Sanger sequencing. We found heterogeneity among gene trees, although four major clades (Trilling Frog, Fat Frog, crucifer, and West Coast) were consistently supported, and we resolved the relationships among these clades for the first time with strong support. We also found discordance between the mitochondrial and nuclear datasets that we attribute to mitochondrial introgression and a possible selective sweep. Bayesian concordance analysis in BUCKy and species tree analysis in *BEAST produced largely similar topologies, although we identify taxa that require additional investigation in order to clarify taxonomic and geographic range boundaries. Overall, we demonstrate the utility of a reduced representation library approach for marker development and parallel tagged sequencing on an Illumina MiSeq for phylogenetic studies of non-model organisms.
Data from: Accounting for uncertainty in gene tree estimation: summary-coalescent species tree inference in a challenging radiation of Australian lizards
Accurate gene tree inference is an important aspect of species tree estimation in a summary-coalescent framework. Yet, in empirical studies, inferred gene trees differ in accuracy due to stochastic variation in phylogenetic signal between targeted loci. Empiricists should, therefore, examine the consistency of species tree inference, while accounting for the observed heterogeneity in gene tree resolution of phylogenomic data sets. Here, we assess the impact of gene tree estimation error on summary-coalescent species tree inference by screening ${\sim}2000$ exonic loci based on gene tree resolution prior to phylogenetic inference. We focus on a phylogenetically challenging radiation of Australian lizards (genus Cryptoblepharus, Scincidae) and explore effects on topology and support. We identify a well-supported topology based on all loci and find that a relatively small number of high-resolution gene trees can be sufficient to converge on the same topology. Adding gene trees with decreasing resolution produced a generally consistent topology, and increased confidence for specific bipartitions that were poorly supported when using a small number of informative loci. This corroborates coalescent-based simulation studies that have highlighted the need for a large number of loci to confidently resolve challenging relationships and refutes the notion that low-resolution gene trees introduce phylogenetic noise. Further, our study also highlights the value of quantifying changes in nodal support across locus subsets of increasing size (but decreasing gene tree resolution). Such detailed analyses can reveal anomalous fluctuations in support at some nodes, suggesting the possibility of model violation. By characterizing the heterogeneity in phylogenetic signal among loci, we can account for uncertainty in gene tree estimation and assess its effect on the consistency of the species tree estimate. We suggest that the evaluation of gene tree resolution should be incorporated in the analysis of empirical phylogenomic data sets. This will ultimately increase our confidence in species tree estimation using summary-coalescent methods and enable us to exploit genomic data for phylogenetic inference.
FIGURE 18. Maximum Likelihood tree estimated from 1044 in Bythaelurus bachi n. sp., a new deep-water catshark (Carcharhiniformes, Scyliorhinidae) from the southwestern Indian Ocean, with a review of Bythaelurus species and a key to their identification
FIGURE 18. Maximum Likelihood tree estimated from 1044 aligned sites of the mitochondrial NADH2 gene using a General Time Reversible model and an accommodation for among site rate variation and Invariant sites (GTR+I+G model).
Data from: A hybrid phylogenetic–phylogenomic approach for species tree estimation in African Agama lizards with applications to biogeography, character evolution, and diversification
Africa is renowned for its biodiversity and endemicity, yet little is known about the factors shaping them across the continent. African Agama lizards (45 species) have a pan-continental distribution, making them an ideal model for investigating biogeography. Many species have evolved conspicuous sexually dimorphic traits, including extravagant breeding coloration in adult males, large adult male body sizes, and variability in social systems among colorful versus drab species. We present a comprehensive time-calibrated species tree for Agama, and their close relatives, using a hybrid phylogenetic-phylogenomic approach that combines traditional Sanger sequence data from five loci for 57 species (146 samples) with anchored phylogenomic data from 215 nuclear genes for 23 species. The Sanger data are analyzed using coalescent-based species tree inference using *BEAST, and the resulting posterior distribution of species trees is attenuated using the phylogenomic tree as a backbone constraint. The result is a time-calibrated species tree for Agama that includes 95% of all species, multiple samples for most species, strong support for the major clades, and strong support for most of the initial divergence events. Diversification within Agama began approximately 23 million years ago (Ma), and separate radiations in Southern, East, West, and Northern Africa have been diversifying for > 10 Myr. A suite of traits (morphological, coloration, and sociality) are tightly correlated and show a strong signal of high morphological disparity within clades, whereby the subsequent evolution of convergent phenotypes has accompanied diversification into new biogeographic areas.
Figure 7. Mitochondrial DNA gene tree estimated for Acanthocercus atricollis using a in Lifting the blue-headed veil - integrative taxonomy of the Acanthocercus atricollis species complex (Squamata: Agamidae)
Figure 7. Mitochondrial DNA gene tree estimated for Acanthocercus atricollis using a portion of the 16S gene. The support for branches from BI and ML are shown on each branch, respectively. The *BEAST species tree is shown in the top left with posterior probability values on branches.
Data from: Species tree estimation of North American chorus frogs (Hylidae: Pseudacris) with parallel tagged amplicon sequencing
Open the record for dataset details and reuse information.
Data from: A hybrid phylogenetic–phylogenomic approach for species tree estimation in African Agama lizards with applications to biogeography, character evolution, and diversification
Open the record for dataset details and reuse information.
Data from: Assessing the impacts of positive selection on coalescent-based species tree estimation and species delimitation.
Open the record for dataset details and reuse information.
Data from: Disentangling incomplete lineage sorting and introgression to refine species-tree estimates for Lake Tanganyika cichlid fishes
Open the record for dataset details and reuse information.
Data from: Accounting for uncertainty in gene tree estimation: summary-coalescent species tree inference in a challenging radiation of Australian lizards
Open the record for dataset details and reuse information.
Data from: To include or not to include: the impact of gene filtering on species tree estimation methods
Open the record for dataset details and reuse information.
Data from: Comparing species tree estimation with large anchored phylogenomic and small Sanger-sequenced molecular datasets: an empirical study on Malagasy pseudoxyrhophiine snakes
Open the record for dataset details and reuse information.
Data from: Probabilistic species tree distances: implementing the multispecies coalescent to compare species trees within the same model-based framework used to estimate them
Open the record for dataset details and reuse information.
Data from: Robustness to divergence time underestimation when inferring species trees from estimated gene trees
To infer species trees from gene trees estimated from phylogenomic data sets, tractable methods are needed that can handle dozens to hundreds of loci. We examine several computationally efficient approaches—MP-EST, STAR, STEAC, STELLS, and STEM—for inferring species trees from gene trees estimated using maximum likelihood (ML) and Bayesian approaches. Among the methods examined, we found that topology-based methods often performed better using ML gene trees and methods employing coalescent times typically performed better using Bayesian gene trees, with MP-EST, STAR, STEAC, and STELLS outperforming STEM under most conditions. We examine why the STEM tree (also called GLASS or Maximum Tree) is less accurate on estimated gene trees by comparing estimated and true coalescence times, performing species tree inference using simulations, and analyzing a great ape data set keeping track of false positive and false negative rates for inferred clades. We find that although true coalescence times are more ancient than speciation times under the multispecies coalescent model, estimated coalescence times are often more recent than speciation times. This underestimation can lead to increased bias and lack of resolution with increased sampling (either alleles or loci) when gene trees are estimated with ML. The problem appears to be less severe using Bayesian gene-tree estimates.
Fig. 3. Maximum likelihood gene trees estimated using PhyML. A in Morphological and Genetic Characterization of the First Species of Thalassodrilides (Annelida: Clitellata: Naididae: Limnodriloidinae) from Japan
Fig. 3. Maximum likelihood gene trees estimated using PhyML. A, COI; B, ITS. Numbers at branches denote aLRT branch support. Scale shows estimated numbers of nucleotide substitutions per site.
Data from: Evaluating summary methods for multi-locus species tree estimation in the presence of incomplete lineage sorting
Open the record for dataset details and reuse information.
Data from: Species tree estimation of diploid Helianthus (Asteraceae) using target enrichment
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.