Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “multispecies modeling”
Using interview surveys and multispecies occupancy models to inform vertebrate conservation
<p>The Excel workbook WGAllSpeciesCov.xlsx contains detection histories of 30 species of vertebrates in 395 sites in the Western Ghats, India obtained from interviews with field staff of the Forest Department, people from local communities and formally trained people. The workbook also contains site level covariates or determinants of species occurrence. The associated README text file contains metadata to describe the data in the xlsx workbook. Complete R code to read in the data and run the model is provided in the R file: FP_MSp_SS_SpR_AllTraitGr.R and the JAGS code for the multispecies occupancy model is provided in the text file FP_MSp_SS_SpR_AllTraitGr_Exp_NoFP.txt</p>
Supplementary material for: PhyloCoalSimulations: A simulator for network multispecies coalescent models, including a new extension for the inheritance of gene flow
<p>We consider the evolution of phylogenetic gene trees along phylogenetic species networks, according to the network multispecies coalescent process, and introduce a new network coalescent model with correlated inheritance of gene flow. This model generalizes two traditional versions of the network coalescent: with independent or common inheritance. At each reticulation, multiple lineages of a given locus are inherited from parental populations chosen at random, either independently across lineages, or with positive correlation according to a Dirichlet process. This process may account for locus-specific probabilities of inheritance, for example.</p> <p>We implemented the simulation of gene trees under these network coalescent models in the Julia package PhyloCoalSimulations, which depends on PhyloNetworks and its powerful network manipulation tools. Input species phylogenies can be read in extended Newick format, either in numbers of generations or in coalescent units. Simulated gene trees can be written in Newick format, and in a way that preserves information about their embedding within the species network. This embedding can be used for downstream purposes, such as to simulate species-specific processes like rate variation across species, or for other scenarios as illustrated in this note. This package should be useful for simulation studies and simulation-based inference methods. The software is available open source with documentation and a tutorial at <a href="https://github.com/cecileane/PhyloCoalSimulations.jl">https://github.com/cecileane/PhyloCoalSimulations.jl</a>.</p>
Hierarchical heuristic species delimitation under the multispecies coalescent model with migration
<p>The multispecies coalescent (MSC) model accommodates genealogical fluctuations across the genome and provides a natural framework for comparative analysis of genomic sequence data to infer the history of species divergence and gene flow. Given a set of populations, hypotheses of species delimitation (and species phylogeny) may be formulated as instances of MSC models (e.g., MSC for one species versus MSC for two species) and compared using Bayesian model selection. This approach, implemented in the program bpp, has been found to be prone to over-splitting. Alternatively, heuristic criteria based on population parameters under the MSC model (such as population/species divergence times, population sizes, and migration rates) estimated from genomic sequence data may be used to delimit species. Here we extend the approach of species delimitation using the genealogical divergence index (𝑔𝑑𝑖) to develop hierarchical merge and split algorithms for heuristic species delimitation and implement them in a python pipeline called hhsd. Applied to data simulated under a model of isolation by distance, the approach was able to recover the correct species delimitation, whereas model comparison by bpp failed. Analyses of empirical datasets suggest that the procedure may be less prone to over-splitting. We discuss possible strategies for accommodating paraphyletic species in the procedure, as well as the challenges of species delimitation based on heuristic criteria.</p>
Supplementary material for: PhyloCoalSimulations: A simulator for network multispecies coalescent models, including a new extension for the inheritance of gene flow
Open the record for dataset details and reuse information.
Data from: The multispecies coalescent model outperforms concatenation across diverse phylogenomic
Open the record for dataset details and reuse information.
Hierarchical heuristic species delimitation under the multispecies coalescent model with migration
Open the record for dataset details and reuse information.
Revisiting the multispecies coalescent model fit with an example from a complete molecular phylogeny of the Liolaemus wiegmannii species group (Squamata: Liolaemidae)
Open the record for dataset details and reuse information.
Data from: The multilocus multispecies coalescent: a flexible new model of gene family evolution
<p>Incomplete lineage sorting (ILS), the interaction between coalescence and speciation, can generate incongruence between gene trees and species trees, as can gene duplication (D), transfer (T) and loss (L). These processes are usually modelled independently, but in reality, ILS can affect gene copy number polymorphism, i.e., interfere with DTL. This has been previously recognised, but not treated in a satisfactory way, mainly because DTL events are naturally modelled forward-in-time, while ILS is naturally modelled backwards-in-time with the coalescent. Here we consider the joint action of ILS and DTL on the gene tree/species tree problem in all its complexity. In particular, we show that the interaction between ILS and duplications/transfers (without losses) can result in patterns usually interpreted as resulting from gene loss, and that the realised rate of D, T and L becomes non-homogeneous in time when ILS is taken into account. We introduce algorithmic solutions to these problems. Our new model, the <em>multilocus multispecies coalescent</em> (MLMSC), which also accounts for any level of linkage between loci, generalises the multispecies coalescent model and offers a versatile, powerful framework for proper simulation and inference of gene family evolution.</p>
Phase resolution of heterozygous sites in diploid genomes is important to phylogenomic analysis under the multispecies coalescent model
<p>Genome sequencing projects routinely generate haploid consensus sequences from diploid genomes, which are effectively chimeric sequences with the phase at heterozygous sites resolved at random. The impact of phasing errors on phylogenomic analyses under the multispecies coalescent (MSC) model is largely unknown. Here we conduct a computer simulation to evaluate the performance of four phase-resolution strategies (the true phase resolution, the diploid analytical integration algorithm which averages over all phase resolutions, computational phase resolution using the program PHASE, and random resolution) on estimation of the species tree and evolutionary parameters in analysis of multi-locus genomic data under the MSC model. We found that species tree estimation is robust to phasing errors when species divergences were much older than average coalescent times but may be affected by phasing errors when the species tree is shallow. Estimation of parameters under the MSC model with and without introgression is affected by phasing errors. In particular, random phase resolution causes serious overestimation of population sizes for modern species and biased estimation of cross-species introgression probability. In general the impact of phasing errors is greater when the mutation rate is higher, the data include more samples per species, and the species tree is shallower with recent divergences. Use of phased sequences inferred by the PHASE program produced small biases in parameter estimates. We analyze two real datasets, one of East Asian brown frogs and another of Rocky Mountains chipmunks, to demonstrate that heterozygote phase-resolution strategies have similar impacts on practical data analyses. We suggest that genome sequencing projects should produce unphased diploid genotype sequences if fully phased data are too challenging to generate, and avoid haploid consensus sequences, which have heterozygous sites phased at random. In case the analytical integration algorithm is computationally unfeasible, computational phasing prior to population genomic analyses is an acceptable alternative. </p>
Multispecies modelling reveals potential for habitat restoration to re-establish boreal vertebrate community dynamics
<p>1. The restoration of habitats degraded by industrial disturbance is essential for achieving conservation objectives in disturbed landscapes. In boreal ecosystems, disturbances from seismic exploration lines and other linear features have adversely affected biodiversity, most notably leading to declines in threatened woodland caribou. Large-scale restoration of disturbed habitats is needed, yet empirical assessments of restoration effectiveness on wildlife communities remain rare.</p> <p>2. We used 73 camera trap deployments from 2015-2019 and joint species distribution models to investigate how habitat use by the larger vertebrate community (>0.2 kg) responded to variation in key seismic line characteristics (line-of-sight, width, density and mounding) following restoration treatments in a landscape disturbed by oil and gas development in northeastern Alberta.</p> <p>3. The proportion of variation explained by line characteristics was low in comparison to habitat type and season, suggesting short-term responses to restoration treatments were relatively weak. However, we predicted that lines with characteristics consistent with restored conditions would support altered community composition, with reduced use by wolf and coyote, thereby indicating that line restoration will result in reduced contact rates between caribou and these key predators.</p> <p>4. Our analysis provides a framework to assess and predict wildlife community responses to emerging restoration efforts. With the growing importance of habitat restoration for caribou and other vertebrate species, we recommend longer-term monitoring combined with landscape-scale comparisons of different restoration approaches to more fully understand and direct these critical conservation investments. Only by combining rigorous multispecies monitoring with large-scale restoration will we effectively conserve biodiversity within rapidly changing environments.</p>
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
<p>Although environmental DNA (eDNA) metabarcoding has become widely applied to gauge ecosystems in a noninvasive and cost-efficient manner, false negatives can occur due to various factors in its inherent multistage workflow. It is therefore essential to deal with this kind of species detection errors in eDNA metabarcoding to achieve accurate assessment of species distribution and diversity. To address this issue, we proposed a variant of the multispecies site occupancy model for eDNA metabarcoding studies and applied it to an eDNA metabarcoding dataset of freshwater fish communities collected in the Kasumigaura watershed in Japan.</p> <ul> </ul>
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
Open the record for dataset details and reuse information.
Multispecies modelling reveals potential for habitat restoration to re-establish boreal vertebrate community dynamics
Open the record for dataset details and reuse information.
Data from: The multilocus multispecies coalescent: a flexible new model of gene family evolution
Open the record for dataset details and reuse information.
Phase resolution of heterozygous sites in diploid genomes is important to phylogenomic analysis under the multispecies coalescent model
Open the record for dataset details and reuse information.
Data from: Impact of model violations on the inference of species boundaries under the multispecies coalescent
The use of genetic data for identifying species-level lineages across the tree of life has received increasing attention in the field of systematics over the past decade. The multispecies coalescent model provides a framework for understanding the process of lineage divergence, and has become widely adopted for delimiting species. However, because these studies lack an explicit assessment of model fit, in many cases, the accuracy of the inferred species boundaries are unknown. This is concerning given the large amount of empirical data and theory that highlight the complexity of the speciation process. Here, we seek to fill this gap by using simulation to characterize the sensitivity of inference under the multispecies coalescent to several violations of model assumptions thought to be common in empirical data. We also assess the fit of the multispecies coalescent model to empirical data in the context of species delimitation. Our results show substantial variation in model fit across datasets. Posterior predictive tests find the poorest model performance in datasets that were hypothesized to be impacted by model violations. We also show that while the inferences assuming the multispecies coalescent are robust to minor model violations, such inferences can be biased under some biologically plausible scenarios. Taken together, these results suggest that researchers can identify individual datasets in which species delimitation under the multispecies coalescent is likely to be problematic, thereby highlighting the cases where additional lines of evidence to identify species boundaries are particularly important to collect. Our study supports a growing body of work highlighting the importance of model checking in phylogenetics, and the usefulness of tailoring tests of model fit to assess the reliability of particular inferences.
Figure 6 in Bayesian Poisson tree processes and multispecies coalescent models shed new light on the diversification of Nawab butterflies in the Solomon Islands (Nymphalidae, Charaxinae, Polyura)
Figure 6. Habitus of Polyura bicolor and Polyura epigenes across the Solomon Islands. Pictures of the upperside (left half) and underside (right half) of the two sexes. All pictures taken by Bernard Turlin.
Figure 4 in Bayesian Poisson tree processes and multispecies coalescent models shed new light on the diversification of Nawab butterflies in the Solomon Islands (Nymphalidae, Charaxinae, Polyura)
Figure 4. Nuclear haplotype network and maximum-likelihood phylogenetic reconstruction. The haplotype network reconstructed using the concatenated RPS5 and wingless alignments is presented on the left. The phylogenetic hypothesis inferred using the same data set in RA×ML is presented at the top right. A table highlighting the nucleotide substitutions found between Polyura epigenes bicolor and Polyura epigenes* is presented at the bottom right.
Figure 3 in Bayesian Poisson tree processes and multispecies coalescent models shed new light on the diversification of Nawab butterflies in the Solomon Islands (Nymphalidae, Charaxinae, Polyura)
Figure 3. Comparison of species delimitation results. Graph showing the result of each analysis of species delimitation. The posterior probability of either Polyura bicolor or Polyura epigenes (including Polyura bicolor) as a valid species are shown.
Figure 2 in Bayesian Poisson tree processes and multispecies coalescent models shed new light on the diversification of Nawab butterflies in the Solomon Islands (Nymphalidae, Charaxinae, Polyura)
Figure 2. Molecular phylogeny and molecular species delimitation results. Bayesian phylogeny as recovered from MrBayes analyses. The nodal supports of both Bayesian-inference (BI) and maximum-likelihood (ML) analyses are shown, with colour coding as indicated in the figure. The cloudogram of the *BEAST analysis showing all posterior species trees is presented at the bottom left. Denser regions indicate a robust support for the inferred relationship.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.