Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “species-tree”
ASTRAL-Pro: Quartet-based species-tree inference despite paralogy
Open the record for dataset details and reuse information.
Data from: The effect of gene flow on coalescent-based species-tree inference
Most current methods for inferring species-level phylogenies under the coalescent model assume that no gene flow occurs following speciation. Several studies have examined the impact of gene flow (e.g., Eckert and Carstens (2008); Chung and Ane (2011); Leache et al. (2014); Solis-Lemus et al. (2016)) and of ancestral population structure (DeGeorgio and Rosenberg, 2016) on the performance of species-level phylogenetic inference, and analytic results have been proven for network models of gene flow (e.g., Solis-Lemus et al. (2016); Zhu et al. (2016)). However, there are few analytic results for a continuous model of gene flow following speciation, despite the development of mathematical tools that could facilitate such study (e.g., Hobolth et al. (2011); Andersen et al. (2014); Tian and Kubatko (2016)). In this paper, we consider a three-taxon isolation-with-migration model that allows gene flow between sister taxa for a brief period following speciation, as well as variation in the effective population sizes across the species tree. We derive the probabilities of each of the three gene tree topologies under this model, and show that for certain choices of the gene flow and effective population size parameters, anomalous gene trees (i.e., gene trees that are discordant with the species tree but that have higher probability than the gene tree concor- dant with the species tree) exist. We characterize the region of parameter space producing anomalous trees, and show that the probability of the gene tree that is concordant with the species tree can be arbitrarily small. We then show that there is theoretical support for using SVDQuartets with an outgroup to infer the rooted three-taxon species tree in a model of gene flow between sister taxa. We study the performance of SVDQuartets on simulated data and compare it to three other commonly-used methods for species tree inference, AS- TRAL, MP-EST, and concatenation. The simulations show that ASTRAL, MP-EST, and concatenation can be statistically inconsistent when gene flow is present, while SVDQuartets performs well, though large sample sizes may be required for certain parameter choices.
Data from: Disentangling incomplete lineage sorting and introgression to refine species-tree estimates for Lake Tanganyika cichlid fishes
Open the record for dataset details and reuse information.
Data from: The effect of gene flow on coalescent-based species-tree inference
Open the record for dataset details and reuse information.
Data from: Is recombination a problem for species-tree analyses?
As the field of phylogenetics transitions into phylogenomics it has spurred a shift in the general paradigm for data analysis whereby specific dataset attributes can now be considered in a model-based framework. Much effort has been put in to modeling the effects of nucleotide substitution within a genealogy (i.e., modeling mutations; Felsenstein 2005) and the sorting genes between lineages (i.e., coalescence; Knowles and Kubatko 2010), both of which are inherent properties of all species histories. Improvements in accuracy of species-tree estimation related to using models that account for these two sources of uncertainty have been well documented (Carstens and Knowles 2007; Edwards et al. 2007; Kubatko and Degnan 2007), but these are clearly not the only sources of gene tree species-tree discordance. As the field transitions towards using multilocus data, the role of other inherent dataset properties that may be sources of uncertainty, such as intralocus recombination needs to be examined and quantified. As all species-tree methods in current use rely on the simplifying assumption that recombination occurs between but not within loci, ignoring the presence of recombination represents a widespread and ubiquitous violation of species-tree models. Recombination within a locus may be a greater problem for species-tree methodologies than for concatenation because species-tree methods more accurately model the patterns arising from coalescent stochasticity. Concatenation assumes all genes share a common underlying tree, a model violated by most multilocus data (Knowles and Carstens 2007; Cranston 2010; Linnen 2010; Wiens et al. 2010). Relative to the presence of such a gross model violation, errors introduced by within-locus recombination are likely to be minor for datasets analyzed by concatenation. Coalescent methods for species-tree estimation explicitly model each gene tree separately, making the relative contribution of recombination to the uncertainty of the estimated species tree greater. Challenges posed by recombination may also be greatest for species-tree methods that rely on estimating parameters related to the scaled mutation rate (Liu 2008; Liu et al. 2008; Heled and Drummond 2010) because recombination-introduced heterogeneity may interfere with branch length estimates. Lastly, recombination may interact with other aspects of the speciation history, such as time to divergence, population size, and the length of time between speciation events, and sampling effort (McCormack et al. 2009). For example, recombination events occurring in recently speciated groups may not have accumulated a sufficient number of mutations for the effects to be problematic (see Fig. 1).
Data from: Effectiveness of phylogenomic data and coalescent species-tree methods for resolving difficult nodes in the phylogeny of advanced snakes (Serpentes: Caenophidia)
Next-generation genomic sequencing promises to quickly and cheaply resolve remaining contentious nodes in the Tree of Life, and facilitates species-tree estimation while taking into account stochastic genealogical discordance among loci. Recent methods for estimating species trees bypass full likelihood-based estimates of the multi-species coalescent, and approximate the true species-tree using simpler summary metrics. These methods converge on the true species-tree with sufficient genomic sampling, even in the anomaly zone. However, no studies have yet evaluated their efficacy on a large-scale phylogenomic dataset, and compared them to previous concatenation strategies. Here, we generate such a dataset for Caenophidian snakes, a group with >2500 species that contains several rapid radiations that were poorly resolved with fewer loci. We generate sequence data for 333 single-copy nuclear loci with ∼100% coverage (∼0% missing data) for 31 major lineages. We estimate phylogenies using neighbor joining, maximum parsimony, maximum likelihood, and three summary species-tree approaches (NJst, STAR, and MP-EST). All methods yield similar resolution and support for most nodes. However, not all methods support monophyly of Caenophidia, with Acrochordidae placed as the sister taxon to Pythonidae in some analyses. Thus, phylogenomic species-tree estimation may occasionally disagree with well-supported relationships from concatenated analyses of small numbers of nuclear or mitochondrial genes, a consideration for future studies. In contrast for at least two diverse, rapid radiations (Lamprophiidae and Colubridae), phylogenomic data and species-tree inference do little to improve resolution and support. Thus, certain nodes may lack strong signal, and larger datasets and more sophisticated analyses may still fail to resolve them.
Data from: Is recombination a problem for species-tree analyses?
Open the record for dataset details and reuse information.
Data from: Effectiveness of phylogenomic data and coalescent species-tree methods for resolving difficult nodes in the phylogeny of advanced snakes (Serpentes: Caenophidia)
Open the record for dataset details and reuse information.
Data from: When do species-tree and concatenated estimates disagree? An empirical analysis with higher-level scincid lizard phylogeny
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.