Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

51

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

51 results for “Multispecies coalescent”

Learn how ShareScore rates datasets ↗
dryad40/100

Supplementary material for: PhyloCoalSimulations: A simulator for network multispecies coalescent models, including a new extension for the inheritance of gene flow

<p>We consider the evolution of phylogenetic gene trees along phylogenetic species networks, according to the network multispecies coalescent process, and introduce a new network coalescent model with correlated inheritance of gene flow. This model generalizes two traditional versions of the network coalescent: with independent or common inheritance. At each reticulation, multiple lineages of a given locus are inherited from parental populations chosen at random, either independently across lineages, or with positive correlation according to a Dirichlet process. This process may account for locus-specific probabilities of inheritance, for example.</p> <p>We implemented the simulation of gene trees under these network coalescent models in the Julia package PhyloCoalSimulations, which depends on PhyloNetworks and its powerful network manipulation tools. Input species phylogenies can be read in extended Newick format, either in numbers of generations or in coalescent units. Simulated gene trees can be written in Newick format, and in a way that preserves information about their embedding within the species network. This embedding can be used for downstream purposes, such as to simulate species-specific processes like rate variation across species, or for other scenarios as illustrated in this note. This package should be useful for simulation studies and simulation-based inference methods. The software is available open source with documentation and a tutorial at <a href="https://github.com/cecileane/PhyloCoalSimulations.jl">https://github.com/cecileane/PhyloCoalSimulations.jl</a>.</p>

opencc-zeroMay 2023View details →
dryad40/100

Hierarchical heuristic species delimitation under the multispecies coalescent model with migration

<p>The multispecies coalescent (MSC) model accommodates genealogical fluctuations across the genome and provides a natural framework for comparative analysis of genomic sequence data to infer the history of species divergence and gene flow. Given a set of populations, hypotheses of species delimitation (and species phylogeny) may be formulated as instances of MSC models (e.g., MSC for one species versus MSC for two species) and compared using Bayesian model selection. This approach, implemented in the program bpp, has been found to be prone to over-splitting. Alternatively, heuristic criteria based on population parameters under the MSC model (such as population/species divergence times, population sizes, and migration rates) estimated from genomic sequence data may be used to delimit species. Here we extend the approach of species delimitation using the genealogical divergence index (𝑔𝑑𝑖) to develop hierarchical merge and split algorithms for heuristic species delimitation and implement them in a python pipeline called hhsd. Applied to data simulated under a model of isolation by distance, the approach was able to recover the correct species delimitation, whereas model comparison by bpp failed. Analyses of empirical datasets suggest that the procedure may be less prone to over-splitting. We discuss possible strategies for accommodating paraphyletic species in the procedure, as well as the challenges of species delimitation based on heuristic criteria.</p>

opencc-zeroSep 2023View details →
dryad40/100

Supplementary material for: PhyloCoalSimulations: A simulator for network multispecies coalescent models, including a new extension for the inheritance of gene flow

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: The multispecies coalescent model outperforms concatenation across diverse phylogenomic

Open the record for dataset details and reuse information.

publicJan 2020View details →
dryad40/100

Hierarchical heuristic species delimitation under the multispecies coalescent model with migration

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad40/100

Revisiting the multispecies coalescent model fit with an example from a complete molecular phylogeny of the Liolaemus wiegmannii species group (Squamata: Liolaemidae)

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Data from: Cryptic patterns of speciation in cryptic primates: microendemic mouse lemurs and the multispecies coalescent

Species delimitation is ever more critical for assessing biodiversity in threatened regions of the world, especially when undescribed lineages may be at risk from habitat loss. Mouse lemurs (Microcebus) are an example of a rapid radiation of morphologically cryptic species that are distributed throughout Madagascar in its rapidly vanishing forested habitats. Here, we focus on two pairs of sister lineages that occur in a region in northeastern Madagascar that shows high levels of microendemism. We revisit previous hypotheses of species diversity by filling geographic sampling gaps and by generating new genomic data for three named species, as well as an undescribed lineage previously identified to be of interest due to its highly divergent mtDNA. We analyzed RADseq data with multiple species delimitation methods based on the multispecies coalescent (MSC) model while accounting for introgression. Non-sister lineages occur sympatrically in two instances, despite an estimated divergence time of less than 1 Ma, thus suggesting rapid evolution of reproductive isolation in the mouse lemur clade. We note, however, that the divergence time estimates reported here are based on the MSC and calibrated with pedigree-based primate mutation rates. These dates are considerably more recent than previous analyses that used traditional relaxed clock methods and distant fossil calibrations. One pair of sister lineages passed all species delimitation tests while the other pair failed most, largely due to differences in Ne between the two pairs of lineages. Nevertheless, delimitation results were also supported by differences in levels of gene flow and patterns of isolation-by-distance between the two pairs. We conclude that MSC-based species delimitation methods are valuable tools for evaluating cryptic species, even though these methods can be strongly affected by variable Ne. We suggest that this result has general implications for species delimitation studies of other recently diverged lineages.

opencc-zeroJul 2020View details →
dryad36/100

Data from: The multilocus multispecies coalescent: a flexible new model of gene family evolution

<p>Incomplete lineage sorting (ILS), the interaction between coalescence and speciation, can generate incongruence between gene trees and species trees, as can gene duplication (D), transfer (T) and loss (L). These processes are usually modelled independently, but in reality, ILS can affect gene copy number polymorphism, i.e., interfere with DTL. This has been previously recognised, but not treated in a satisfactory way, mainly because DTL events are naturally modelled forward-in-time, while ILS is naturally modelled backwards-in-time with the coalescent. Here we consider the joint action of ILS and DTL on the gene tree/species tree problem in all its complexity. In particular, we show that the interaction between ILS and duplications/transfers (without losses) can result in patterns usually interpreted as resulting from gene loss, and that the realised rate of D, T and L becomes non-homogeneous in time when ILS is taken into account. We introduce algorithmic solutions to these problems. Our new model, the <em>multilocus multispecies coalescent</em> (MLMSC), which also accounts for any level of linkage between loci, generalises the multispecies coalescent model and offers a versatile, powerful framework for proper simulation and inference of gene family evolution.</p>

opencc-zeroAug 2020View details →
dryad36/100

Phase resolution of heterozygous sites in diploid genomes is important to phylogenomic analysis under the multispecies coalescent model

<p>Genome sequencing projects routinely generate haploid consensus sequences from diploid genomes, which are effectively chimeric sequences with the phase at heterozygous sites resolved at random. The impact of phasing errors on phylogenomic analyses under the multispecies coalescent (MSC) model is largely unknown. Here we conduct a computer simulation to evaluate the performance of four phase-resolution strategies (the true phase resolution, the diploid analytical integration algorithm which averages over all phase resolutions, computational phase resolution using the program PHASE, and random resolution) on estimation of the species tree and evolutionary parameters in analysis of multi-locus genomic data under the MSC model. We found that species tree estimation is robust to phasing errors when species divergences were much older than average coalescent times but may be affected by phasing errors when the species tree is shallow. Estimation of parameters under the MSC model with and without introgression is affected by phasing errors. In particular, random phase resolution causes serious overestimation of population sizes for modern species and biased estimation of cross-species introgression probability. In general the impact of phasing errors is greater when the mutation rate is higher, the data include more samples per species, and the species tree is shallower with recent divergences. Use of phased sequences inferred by the PHASE program produced small biases in parameter estimates. We analyze two real datasets, one of East Asian brown frogs and another of Rocky Mountains chipmunks, to demonstrate that heterozygote phase-resolution strategies have similar impacts on practical data analyses. We suggest that genome sequencing projects should produce unphased diploid genotype sequences if fully phased data are too challenging to generate, and avoid haploid consensus sequences, which have heterozygous sites phased at random. In case the analytical integration algorithm is computationally unfeasible, computational phasing prior to population genomic analyses is an acceptable alternative. </p>

opencc-zeroOct 2020View details →
dryad36/100

Data from: Cryptic patterns of speciation in cryptic primates: microendemic mouse lemurs and the multispecies coalescent

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad36/100

Bayesian inference under the multispecies coalescent with ancient DNA sequences

Open the record for dataset details and reuse information.

publicJul 2024View details →
dryad36/100

Data from: The multilocus multispecies coalescent: a flexible new model of gene family evolution

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad36/100

Phase resolution of heterozygous sites in diploid genomes is important to phylogenomic analysis under the multispecies coalescent model

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad32/100

Data from: A preliminary framework for DNA barcoding, incorporating the multispecies coalescent

The capacity to identify an unknown organism using the DNA sequence from a single gene has many applications. These include the development of biodiversity inventories (Janzen et al. 2005), forensics (Meiklejohn et al. 2011), biosecurity (Armstrong and Ball 2005), and the identification of cryptic species (Smith et al. 2006). The popularity and widespread use (Teletchea 2010) of the DNA barcoding approach (Hebert et al. 2003), despite broad misgivings (e.g., Smith 2005; Will et al. 2005; Rubinoff et al. 2006), attest to this. However, one major shortcoming to the standard barcoding approach is that it assumes that gene trees and species trees are synonymous, an assumption that is known not to hold in many cases (Pamilo and Nei 1988; Funk and Omland 2003). Biological processes that violate this assumption include incomplete lineage sorting and interspecific hybridization (Funk and Omland 2003). Indeed, simulation studies indicate that the concatenation approach (in which these two processes are ignored) can lead to statistically inconsistent estimation of the species tree (Kubatko and Degnan 2007). However, recent developments make a barcoding approach that utilizes a single locus outdated. The cost of sequencing multiple gene fragments is no longer inhibitory, but more importantly, a range of analytical approaches have been developed that account for incomplete lineage sorting (Degnan and Salter 2005; Edwards et al. 2007; Liu et al. 2008; Kubatko et al. 2009; Heled and Drummond 2010; Yang and Rannala 2010). These approaches incorporate coalescent theory into the analysis of species trees and species delimitation (Fujita et al. 2012) and are conveniently accessible as software programs (e.g., BEST, BPP, *BEAST, MrBayes v. 3.2, STEM, and COAL). Although the general mixed Yule coalescent (GMYC) approach has also been developed for species delimitation (Pons et al. 2006), we do not consider it further here. It operates quite differently to the approaches outlined above (i.e., BEST, BPP, *BEAST, MrBayes v. 3.2, STEM, and COAL). The GMYC approach seeks to identify the shift in the rate of lineage branching that should be evident when interspecific evolutionary processes switch to population-level processes (Pons et al. 2006). Both empirical (Esselstyn et al. 2012) and simulation studies (Esselstyn et al. 2012; Fujisawa and Barraclough 2013) report that it performs poorly when effective population sizes and speciation rates are high, but within biologically relevant ranges. Ideally, a "next-generation" barcoding approach would (1) identify a minimal set of barcoding genes (perhaps specific to certain lineages), (2) generate a large and cladistically divergent database for comparisons, and (3) identify species using species delimitation approaches that incorporate the multispecies coalescent. The first two of these conditions are straightforward and require only discussion (requirement 1) and resources (requirement 2). However, the third requirement is much more problematic. Some of the recently developed approaches for species delimitation could not be used alone; for example, BPP requires a user-specified guide tree (Yang and Rannala 2010). All of the recently developed approaches are computationally intensive (Degnan and Rosenberg 2009), with many having practical limitations on the number of individuals that can be compared. By contrast, the current barcoding approach is able to compare enormous numbers of sequences in a very short time, primarily because the approach is analytically simple; a single sequence is compared with all sequences in the database by calculating all possible pairwise K2P distances. As long as exemplars exist within the database that have K2P distances below some predetermined threshold (usually 4%), the species is considered identified. The speed of analysis is due primarily to the use of distance-based measures. The purpose of this article is to initiate the development of a framework for "next-gen barcoding": one that incorporates the multispecies coalescent, but does so by comparing multiple gene sequences from an unknown taxon with a database of sequences.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Impact of model violations on the inference of species boundaries under the multispecies coalescent

The use of genetic data for identifying species-level lineages across the tree of life has received increasing attention in the field of systematics over the past decade. The multispecies coalescent model provides a framework for understanding the process of lineage divergence, and has become widely adopted for delimiting species. However, because these studies lack an explicit assessment of model fit, in many cases, the accuracy of the inferred species boundaries are unknown. This is concerning given the large amount of empirical data and theory that highlight the complexity of the speciation process. Here, we seek to fill this gap by using simulation to characterize the sensitivity of inference under the multispecies coalescent to several violations of model assumptions thought to be common in empirical data. We also assess the fit of the multispecies coalescent model to empirical data in the context of species delimitation. Our results show substantial variation in model fit across datasets. Posterior predictive tests find the poorest model performance in datasets that were hypothesized to be impacted by model violations. We also show that while the inferences assuming the multispecies coalescent are robust to minor model violations, such inferences can be biased under some biologically plausible scenarios. Taken together, these results suggest that researchers can identify individual datasets in which species delimitation under the multispecies coalescent is likely to be problematic, thereby highlighting the cases where additional lines of evidence to identify species boundaries are particularly important to collect. Our study supports a growing body of work highlighting the importance of model checking in phylogenetics, and the usefulness of tailoring tests of model fit to assess the reliability of particular inferences.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Sky island diversification meets the multispecies coalescent – divergence in the spruce-fir moss spider (Microhexura montivaga, Araneae, Mygalomorphae) on the highest peaks of southern Appalachia

Microhexura montivaga is a miniature tarantula-like spider endemic to the highest peaks of the southern Appalachian mountains and is known only from six allopatric, highly disjunct montane populations. Because of severe declines in spruce-fir forest in the late 20th century, M. montivaga was formally listed as a US federally endangered species in 1995. Using DNA sequence data from one mitochondrial and seven nuclear genes, patterns of multigenic genetic divergence were assessed for six montane populations. Independent mitochondrial and nuclear discovery analyses reveal obvious genetic fragmentation both within and among montane populations, with five to seven primary genetic lineages recovered. Multispecies coalescent validation analyses [guide tree and unguided Bayesian Phylogenetics and Phylogeography (BPP), Bayes factor delimitation (BFD)] using nuclear-only data congruently recover six or seven distinct lineages; BFD analyses using combined nuclear plus mitochondrial data favour seven or eight lineages. In stark contrast to this clear genetic fragmentation, a survey of secondary sexual features for available males indicates morphological conservatism across montane populations. While it is certainly possible that morphologically cryptic speciation has occurred in this taxon, this system may alternatively represent a case where extreme population genetic structuring (but not speciation) leads to an oversplitting of lineage diversity by multispecies coalescent methods. Our results have clear conservation implications for this federally endangered taxon and illustrate a methodological issue expected to become more common as genomic-scale data sets are gathered for taxa found in naturally fragmented habitats.

opencc-zeroDec 2014View details →
dryad32/100

StarBeast3: Adaptive parallelised Bayesian inference under the multispecies coalescent

<p><span><span>As genomic sequence data becomes increasingly available, inferring the phylogeny of the species as that of concatenated genomic data can be enticing. However, this approach makes for a biased estimator of branch lengths and substitution rates and an inconsistent estimator of tree topology. Bayesian multispecies coalescent methods address these issues. This is achieved by constraining a set of gene trees within a species tree and jointly inferring both under a Bayesian framework. However, this approach comes at the cost of increased computational demand. Here, we introduce StarBeast3 -- a software package for efficient. Bayesian inference under the multispecies coalescent model via Markov chain Monte Carlo. We gain efficiency by introducing cutting-edge proposal kernels and adaptive operators, and StarBeast3 is particularly efficient when a relaxed clock model is applied. Furthermore, gene tree inference is parallelised, allowing the software to scale with the size of the problem. We validated our software and benchmarked its performance using three real and two synthetic datasets. Our results indicate that StarBeast3 is up to one-and-a-half orders of magnitude faster than StarBeast2, and therefore more than two orders faster than *BEAST, depending on the dataset and on the parameter, and can achieve convergence on large datasets with hundreds of genes. StarBeast3 is open-source and is easy to set up with a friendly graphical user interface.</span></span></p>

opencc-zeroMay 2022View details →
zenodo32/100

Figure 11 in Fantastic beasts and how to delimit them: an integrative approach using multispecies coalescent methods reveals two new, endemic Dugesia species (Platyhelminthes: Tricladida) from Corsica and Sardinia

Figure 11. Dugesia hoidi: A, holotype RMNH.VER.21056.1, photomicrograph showing the penial fold (pf) in sagiưal section; B, paratype RMNH.VER.21056.2, photomicrograph showing the penis papilla (pp) and the penial fold (pf) in transverse section.

opennotspecifiedNov 2023View details →
zenodo32/100

Figure 10. Dugesia hoidi. Holotype RMNH.VER.21056.1 in Fantastic beasts and how to delimit them: an integrative approach using multispecies coalescent methods reveals two new, endemic Dugesia species (Platyhelminthes: Tricladida) from Corsica and Sardinia

Figure 10. Dugesia hoidi. Holotype RMNH.VER.21056.1: A, sagiưal reconstruction of the male copulatory apparatus (anterior to the right); B, sagiưal reconstruction of the penial fold and female copulatory apparatus; C, photomicrograph of sagiưal section, showing penis bulb (pb) with the seminal vesicle (sv), right (rvd) and the less (lvd) vas deferens, penis papilla (pp) with the pointed diaphragm (d), and the ejaculatory duct (ed).

opennotspecifiedNov 2023View details →
zenodo32/100

Figure 7. Dugesia benazzii s.s., CGAS Pla 25.1 in Fantastic beasts and how to delimit them: an integrative approach using multispecies coalescent methods reveals two new, endemic Dugesia species (Platyhelminthes: Tricladida) from Corsica and Sardinia

Figure 7. Dugesia benazzii s.s., CGAS Pla 25.1: A, sagiưal reconstruction of the male copulatory apparatus (anterior to the right); B, sagiưal reconstruction of the fold and female copulatory apparatus; C, photomicrograph of sagiưal section, showing the penis bulb (pb), penis papilla (pp) with conical, pointed diaphragm (d), ejaculatory duct (ed), penial fold (pf), and 'angled' bursal canal (abc).

opennotspecifiedNov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record