Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38
datasets available to search
ShareScore release 0.9.0
Dataset results
38 results for “Coalescent methods”
Figure 1 in Considering gene flow when using coalescent methods to delimit lineages of North American pitvipers of the genus Agkistrodon
Figure 1. Map of the USA showing the sampling localities and distribution of inferred lineages of Agkistrodon contortrix (A) and Agkistrodon piscivorus (B). The distribution of both species is indicated by a bold outline. The key to the right of each map indicates the population membership of individuals and also identifies admixture in individuals.
Figure 4 in Considering gene flow when using coalescent methods to delimit lineages of North American pitvipers of the genus Agkistrodon
Figure 4. Niche predictions for each species of the Agkistrodon contortrix complex with hybrids included in training the model (A), Agkistrodon contortrix complex with hybrids excluded (B), Agkistrodon piscivorus complex with hybrids included in training the model (C), Agkistrodon piscivorus complex with hybrids excluded (D), and East and Central lineages of Agkistrodon contortrix inferred by mtDNA (E). The distribution of colubriod samples obtained from the Herpnet database used to correct for sampling bias is shown in (F). Overlap between predicted suitable habitat is indicated by the darkest intermediate shade.
Figure 5 in Considering gene flow when using coalescent methods to delimit lineages of North American pitvipers of the genus Agkistrodon
Figure 5. The distribution of species within the Agkistrodon contortrix complex (A) and Agkistrodon piscivorus complex (B) inferred by Bayesian species delimitation. Putative hybrids zones are identified with cross-hatches. Distributions are adapted from Gloyd & Conant (1990) and the results from ecological niche modelling.
Data from: Assessing species boundaries and the phylogenetic position of the rare Szechwan Ratsnake, Euprepiophis perlacea (Serpentes: Colubridae), using coalescent-based methods
Open the record for dataset details and reuse information.
Data from: Species delimitation with ABC and other coalescent-based methods: a test of accuracy with simulations and an empirical example with lizards of the Liolaemus darwinii complex (Squamata: Liolaemidae)
Open the record for dataset details and reuse information.
Data from: Considering gene flow when using coalescent methods to delimit lineages of North American pitvipers of the genus Agkistrodon
Open the record for dataset details and reuse information.
A practical introduction to sequentially Markovian coalescent methods for estimating demographic history from genomic data
<p>A common goal of population genomics and molecular ecology is to reconstruct the demographic history of a species of interest. A pair of powerful tools based on the sequentially Markovian coalescent have been developed to infer past population sizes using genome sequences. These methods are most useful when sequences are available for only a limited number of genomes and when the aim is to study ancient demographic events. The results of these analyses can be difficult to interpret accurately, because doing so requires some understanding of their theoretical basis and of their sensitivity to confounding factors. In this practical review, we explain some of the key concepts underpinning the pairwise and multiple sequentially Markovian coalescent methods (PSMC and MSMC, respectively). We relate these concepts to the use and interpretation of these methods, and we explain how the choice of different parameter values by the user can affect the accuracy and precision of the inferences. Based on our survey of 100 PSMC studies and 30 MSMC studies, we describe how the two methods are used in practice. Readers of this article will become familiar with the principles, practice, and interpretation of the sequentially Markovian coalescent for inferring demographic history.</p>
Data from: Delimiting species using single-locus data and the Generalized Mixed Yule Coalescent approach: a revised method and evaluation on simulated data sets
DNA barcoding-type studies assemble single-locus data from large samples of individuals and species, and have provided new kinds of data for evolutionary surveys of diversity. An important goal of many such studies is to delimit evolutionarily significant species units, especially in biodiversity surveys from environmental DNA samples. The Generalized Mixed Yule Coalescent (GMYC) method is a likelihood method for delimiting species by fitting within- and between-species branching models to reconstructed gene trees. Although the method has been widely used, it has not previously been described in detail or evaluated fully against simulations of alternative scenarios of true patterns of population variation and divergence between species. Here, we present important reformulations to the GMYC method as originally specified, and demonstrate its robustness to a range of departures from its simplifying assumptions. The main factor affecting the accuracy of delimitation is the mean population size of species relative to divergence times between them. Other departures from the model assumptions, such as varying population sizes among species, alternative scenarios for speciation and extinction, and population growth or subdivision within species, have relatively smaller effects. Our simulations demonstrate that support measures derived from the likelihood function provide a robust indication of when the model performs well and when it leads to inaccurate delimitations. Finally, the so-called single-threshold version of the method outperforms the multiple-threshold version of the method on simulated data: we argue that this might represent a fundamental limit due to the nature of evidence used to delimit species in this approach. Together with other studies comparing its performance relative to other methods, our findings support the robustness of GMYC as a tool for delimiting species when only single-locus information is available.
Data from: Analysis of a rapid evolutionary radiation using ultraconserved elements (UCEs): Evidence for a bias in some multi-species coalescent methods
Rapid evolutionary radiations are expected to require large amounts of sequence data to resolve. To resolve these types of relationships many systematists believe that it will be necessary to collect data by next-generation sequencing (NGS) and use multispecies coalescent ("species tree") methods. Ultraconserved element (UCE) sequence capture is becoming a popular method to leverage the high throughput of NGS to address problems in vertebrate phylogenetics. Here we examine the performance of UCE data for gallopheasants (true pheasants and allies), a clade that underwent a rapid radiation 10–15 Ma. Relationships among gallopheasant genera have been difficult to establish. We used this rapid radiation to assess the performance of species tree methods, using ∼600 kilobases of DNA sequence data from ∼1500 UCEs. We also integrated information from traditional markers (nuclear intron data from 15 loci and three mitochondrial gene regions). Species tree methods exhibited troubling behavior. Two methods [Maximum Pseudolikelihood for Estimating Species Trees (MP-EST) and Accurate Species TRee ALgorithm (ASTRAL)] appeared to perform optimally when the set of input gene trees was limited to the most variable UCEs, though ASTRAL appeared to be more robust than MP-EST to input trees generated using less variable UCEs. In contrast, the rooted triplet consensus method implemented in Triplec performed better when the largest set of input gene trees was used. We also found that all three species tree methods exhibited a surprising degree of dependence on the program used to estimate input gene trees, suggesting that the details of likelihood calculations (e.g., numerical optimization) are important for loci with limited phylogenetic information. As an alternative to summary species tree methods we explored the performance of SuperMatrix Rooted Triple - Maximum Likelihood (SMRT-ML), a concatenation method that is consistent even when gene trees exhibit topological differences due to the multispecies coalescent. We found that SMRT-ML performed well for UCE data. Our results suggest that UCE data have excellent prospects for the resolution of difficult evolutionary radiations, though specific attention may need to be given to the details of the methods used to estimate species trees.
Data from: Detecting evolutionarily significant units above the species level using the Generalized Mixed Yule Coalescent method
1. There is renewed interest in inferring evolutionary history by modelling diversification rates using phylogenies. Understanding the performance of the methods used under different scenarios is essential for assessing empirical results. Recently we introduced a new approach for analysing broadscale diversity patterns, using the Generalized Mixed Yule Coalescent (GMYC) method to test for the existence of evolutionarily significant units above the species (higher ESUs). This approach focuses on identifying clades as well as estimating rates and we refer to it as clade-dependent. However, the ability of the GMYC to detect the phylogenetic signature of higher ESUs has not been fully explored, nor has it been placed in the context of other, clade-independent approaches. 2. We simulated >32,000 trees under two clade-independent models: constant-rate birth-death (CRBD) and variable-rate birth-death (VRBD), using parameter estimates from nine empirical trees and more general parameter values. The simulated trees were used to evaluate scenarios under which GMYC might incorrectly detect the presence of higher ESUs. 3. The GMYC null model was rejected at a high rate on CRBD-simulated trees. This would lead to spurious inference of higher ESUs. However, the support for the GMYC model was significantly greater in most of the empirical clades than expected under a CRBD process. Simulations with empirically derived parameter values could therefore be used to exclude CRBD as an explanation for diversification patterns. In contrast, a VRBD process could not be ruled out as an alternative explanation for the apparent signature of hESUs in the empirical clades, based on the GMYC method alone. Other metrics of tree shape, however, differed notably between the empirical and VRBD-simulated trees. These metrics could be used in future to distinguish clade-dependent and clade-independent models. 4. In conclusion, detection of higher ESUs using the GMYC is robust against some clade-independent models, as long as simulations are used to evaluate these alternatives, but not against others. The differences between clade-dependent and clade-independent processes are biologically interesting, but most current models focus on the latter. We advocate more research into clade-dependent models for broad diversity patterns.
Data from: Effectiveness of phylogenomic data and coalescent species-tree methods for resolving difficult nodes in the phylogeny of advanced snakes (Serpentes: Caenophidia)
Next-generation genomic sequencing promises to quickly and cheaply resolve remaining contentious nodes in the Tree of Life, and facilitates species-tree estimation while taking into account stochastic genealogical discordance among loci. Recent methods for estimating species trees bypass full likelihood-based estimates of the multi-species coalescent, and approximate the true species-tree using simpler summary metrics. These methods converge on the true species-tree with sufficient genomic sampling, even in the anomaly zone. However, no studies have yet evaluated their efficacy on a large-scale phylogenomic dataset, and compared them to previous concatenation strategies. Here, we generate such a dataset for Caenophidian snakes, a group with >2500 species that contains several rapid radiations that were poorly resolved with fewer loci. We generate sequence data for 333 single-copy nuclear loci with ∼100% coverage (∼0% missing data) for 31 major lineages. We estimate phylogenies using neighbor joining, maximum parsimony, maximum likelihood, and three summary species-tree approaches (NJst, STAR, and MP-EST). All methods yield similar resolution and support for most nodes. However, not all methods support monophyly of Caenophidia, with Acrochordidae placed as the sister taxon to Pythonidae in some analyses. Thus, phylogenomic species-tree estimation may occasionally disagree with well-supported relationships from concatenated analyses of small numbers of nuclear or mitochondrial genes, a consideration for future studies. In contrast for at least two diverse, rapid radiations (Lamprophiidae and Colubridae), phylogenomic data and species-tree inference do little to improve resolution and support. Thus, certain nodes may lack strong signal, and larger datasets and more sophisticated analyses may still fail to resolve them.
Data from: Delimiting species using single-locus data and the Generalized Mixed Yule Coalescent approach: a revised method and evaluation on simulated data sets
Open the record for dataset details and reuse information.
A practical introduction to sequentially Markovian coalescent methods for estimating demographic history from genomic data
Open the record for dataset details and reuse information.
Data from: Detecting evolutionarily significant units above the species level using the Generalized Mixed Yule Coalescent method
Open the record for dataset details and reuse information.
Data from: Simulated data for genomic selection and genome-wide association studies using a combination of coalescent and gene drop methods
Open the record for dataset details and reuse information.
Data from: Analysis of a rapid evolutionary radiation using ultraconserved elements (UCEs): Evidence for a bias in some multi-species coalescent methods
Open the record for dataset details and reuse information.
Data from: Coalescent versus concatenation methods and the placement of Amborella as sister to water lilies
Open the record for dataset details and reuse information.
Data from: Effectiveness of phylogenomic data and coalescent species-tree methods for resolving difficult nodes in the phylogeny of advanced snakes (Serpentes: Caenophidia)
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.