Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
199
datasets available to search
ShareScore release 0.9.0
Dataset results
199 results for “base tree”
................................................................................................................................................. Fig. 5. Phylogenetic tree based on 16S rRNA sequences (a) and gltA sequences (b) showing the position of strains R1T, R3, R4 and R6 in relation to the known Bartonella spp. The tree was rooted by using Brucella abortus (a) and Sinorhizobium meliloti (b) as the outgroup. in Bartonella schoenbuchii sp. nov., isolated from the blood of wild roe deer.
................................................................................................................................................. Fig. 5. Phylogenetic tree based on 16S rRNA sequences (a) and gltA sequences (b) showing the position of strains R1T, R3, R4 and R6 in relation to the known Bartonella spp. The tree was rooted by using Brucella abortus (a) and Sinorhizobium meliloti (b) as the outgroup.
Data from: Accuracy and precision of species trees: effects of locus, individual, and base-pair sampling on inference of species trees of the Liolaemus darwinii group (Squamata, Liolaemidae)
Molecular phylogenetics has entered a new era in which species trees are estimated from a collection of gene trees using methods that accommodate their heterogeneity and discordance with the species tree. Empirical evaluation of species trees is necessary to assess the performance (i.e., accuracy and precision) of these methods with real data, which consist of gene genealogies likely shaped by different historical and demographic processes. We analyzed 20 loci for 16 species of the South American lizards of the Liolaemus darwinii species group and reconstructed a species tree with *BEAST, then compared the performance of this method under different sampling strategies of loci, individuals, and sequence lengths. We found an increase in the accuracy and precision of species trees with the number of loci, but for any number of loci, accuracy decreased when using only one individual per species or 25% of the full sequence length. In addition, locus 'informativeness' was an important factor in the accuracy/precision of species trees when using a few loci, but it became increasingly irrelevant with additional loci. Our empirical results combined with previous simulation studies suggest that there is an optimal range of sampling effort of loci, individuals, and sequence lengths for a given speciation history and information content of the data. Future studies should be directed towards further assessment of other factors that can impact performance of species trees, including gene flow, data 'informativeness', tree shape, missing data, and uncertain species boundaries.
Data from: The effect of gene flow on coalescent-based species-tree inference
Most current methods for inferring species-level phylogenies under the coalescent model assume that no gene flow occurs following speciation. Several studies have examined the impact of gene flow (e.g., Eckert and Carstens (2008); Chung and Ane (2011); Leache et al. (2014); Solis-Lemus et al. (2016)) and of ancestral population structure (DeGeorgio and Rosenberg, 2016) on the performance of species-level phylogenetic inference, and analytic results have been proven for network models of gene flow (e.g., Solis-Lemus et al. (2016); Zhu et al. (2016)). However, there are few analytic results for a continuous model of gene flow following speciation, despite the development of mathematical tools that could facilitate such study (e.g., Hobolth et al. (2011); Andersen et al. (2014); Tian and Kubatko (2016)). In this paper, we consider a three-taxon isolation-with-migration model that allows gene flow between sister taxa for a brief period following speciation, as well as variation in the effective population sizes across the species tree. We derive the probabilities of each of the three gene tree topologies under this model, and show that for certain choices of the gene flow and effective population size parameters, anomalous gene trees (i.e., gene trees that are discordant with the species tree but that have higher probability than the gene tree concor- dant with the species tree) exist. We characterize the region of parameter space producing anomalous trees, and show that the probability of the gene tree that is concordant with the species tree can be arbitrarily small. We then show that there is theoretical support for using SVDQuartets with an outgroup to infer the rooted three-taxon species tree in a model of gene flow between sister taxa. We study the performance of SVDQuartets on simulated data and compare it to three other commonly-used methods for species tree inference, AS- TRAL, MP-EST, and concatenation. The simulations show that ASTRAL, MP-EST, and concatenation can be statistically inconsistent when gene flow is present, while SVDQuartets performs well, though large sample sizes may be required for certain parameter choices.
Data from: Assessing the impacts of positive selection on coalescent-based species tree estimation and species delimitation.
The assumption of strictly neutral evolution is fundamental to the multispecies coalescent model and permits the derivation of gene tree distributions and coalescent times conditioned on a given species tree. In this study, we conduct computer simulations to explore the effects of violating this assumption in the form of species-specific positive selection when estimating species trees, species delimitations, and coalescent parameters under the model. We simulated datasets under an array of evolutionary scenarios that differ in both speciation parameters (i.e., divergence times, strength of selection) and experimental design (i.e., number of loci sampled) and incorporated species-specific positive selection occurring within branches of a species tree to identify the effects of selection on multispecies coalescent inferences. Our results highlight particular evolutionary scenarios and parameter combinations in which inferences may be more, or less, susceptible to the effects of positive selection. In some extreme cases, selection can decrease error in species delimitation and increase error in species tree estimation, yet these inferences appear to be largely robust to the effects of positive selection under many conditions likely to be encountered in empirical datasets.
Data from: Towards a common methodology for developing logistic tree mortality models based on ring-width data
Tree mortality is a key process shaping forest dynamics. Thus, there is a growing need for indicators of the likelihood of tree death. During the last decades, an increasing number of tree-ring based studies have aimed to derive growth–mortality functions, mostly using logistic models. The results of these studies, however, are difficult to compare and synthesize due to the diversity of approaches used for the sampling strategy (number and characteristics of alive and death observations), the type of explanatory growth variables included (level, trend, etc.), and the length of the time window (number of years preceding the alive/death observation) that maximized the discrimination ability of each growth variable. We assess the implications of key methodological decisions when developing tree-ring based growth–mortality relationships using logistic mixed-effects regression models. As examples, we use published tree-ring datasets from Abies alba (13 different sites), Nothofagus dombeyi (one site), and Quercus petraea (one site). Our approach is based on a constant sampling size and aims at (1) assessing the dependency of growth–mortality relationships on the statistical sampling scheme used, (2) determining the type of explanatory growth variables that should be considered, and (3) identifying the best length of the time window used to calculate them. The performance of tree-ring-based mortality models was reasonably high for all three species (area under the receiving operator characteristics curve, AUC > 0.7). Growth level variables were the most important predictors of mortality probability for two species (A. alba, N. dombeyi), while growth-trend variables need to be considered for Q. petraea. In addition, the length of the time window used to calculate each growth variable was highly uncertain and depended on the sampling scheme, as some growth–mortality relationships varied with tree age. The present study accounts for the main sampling-related biases to determine reliable species-specific growth–mortality relationships. Our results highlight the importance of using a sampling strategy that is consistent with the research question. Moving towards a common methodology for developing reliable growth–mortality relationships is an important step towards improving our understanding of tree mortality across species and its representation in dynamic vegetation models.
Data from: Phylogenomic analyses resolve an ancient trichotomy at the base of Ischyropsalidoidea (Arachnida, Opiliones) despite high levels of gene tree conflict and unequal minority resolution frequencies
Phylogenetic resolution of ancient rapid radiations has remained problematic despite major advances in statistical approaches and DNA sequencing technologies. Here we report on a combined phylogenetic approach utilizing transcriptome data in conjunction with Sanger sequence data to investigate a tandem of ancient divergences in the harvestmen superfamily Ischyropsalidoidea (Arachnida, Opiliones, Dyspnoi). We rely on Sanger sequences to resolve nodes within and between closely related genera, and use RNA-seq data from a subset of taxa to resolve a short and ancient internal branch. We use several analytical approaches to explore this succession of ancient diversification events, including concatenated and coalescent-based analyses and maximum likelihood gene trees for each locus. We evaluate the robustness of phylogenetic inferences using a randomized locus sub-sampling approach, and find congruence across these methods despite considerable incongruence across gene trees. Incongruent gene trees are not recovered in frequencies expected from a simple multispecies coalescent model, and we reject incomplete lineage sorting as the sole contributor to gene tree conflict. Using these approaches we attain robust support for higher-level phylogenetic relationships within Ischyropsalidoidea.
Data from: The indicator side of tree microhabitats: a multi-taxon approach based on bats, birds and saproxylic beetles
1. National and international forest biodiversity assessments largely rely on indirect indicators, based on elements of forest structure that are used as surrogates for species diversity. These proxies are reputedly easier and cheaper to assess than biodiversity. Tree microhabitats – tree-borne singularities such as cavities, conks of fungi or bark characteristics – have gained attention as potential forest biodiversity indicators. However, as with most biodiversity indicators, there is a lack of scientific evidence documenting their quantitative link with the biodiversity they are supposed to assess. 2. We explored the link between microhabitat indices and the richness and abundance of three taxonomic groups: bats, birds, and saproxylic beetles. Using a nation-wide multi-taxon sampling design in France, we compared 213 plots located inside and outside strict forest reserves. We hypothesized that the positive effect setting aside forest reserves has on biodiversity conservation is indirectly due to an increase in the proportion of large structural elements (e.g. living trees, standing and lying deadwood). These, in turn, are likely to favour the quantity and diversity of microhabitats. We analysed the relationship between the abundance and species richness of different groups and guilds (e.g. red-listed species, forest specialists, cavity dwellers) and microhabitat density and diversity. We then used confirmatory structural equation models to assess the direct and indirect effects of management abandonment, large structural elements and microhabitats on the biodiversity of the target species. 3. For several groups of birds and bats, the indirect effect of management abandonment and large structural elements on biodiversity was mediated by microhabitats. However, the magnitude of the link between microhabitat indices and biodiversity was moderate. In particular, saproxylic beetles' biodiversity was poorly explained by microhabitats, large structural elements or management abandonment. 4. Synthesis and applications: Tree microhabitats may serve as indicators for bats and birds, but they are not a universal biodiversity indicator. Rather, compared to large structural elements, they most likely have a complementary role to biodiversity. In terms of forest management and conservation, preserving diversity of microhabitats at the local scale benefits several groups of both bats and birds.
Data from: Congruent species delimitation of two controversial gold-thread nanmu tree species based on morphological and restriction site-associated DNA sequencing data
Species delimitation is fundamental to conservation and sustainable use of economically important forest tree species. However, the delimitation of two highly valued gold-thread nanmu species (Phoebe bournei and P. zhennan) has been confusing and debated. To address this problem, we integrated morphology and restriction site-associated DNA sequencing (RADseq) to define their species boundaries. We obtained highly consistent results from both data sets, supporting two distinct lineages corresponding to P. bournei and P. zhennan. In Phoebe bournei, higher order leaf venation is more prominent, petioles are thicker and leaf apex angle is narrower, compared to P. zhennan. Both data sets also showed that putative P. bournei localities from north-eastern Guizhou were P. zhennan. The two species have different distributions and only overlap in the Wuling Mountains. Phoebe bournei occurs mainly in Central Fujian, southern Jiangxi, the Nanling Mountains and the Wuling Mountains, whereas P. zhennan is found in the adjoining eastern regions of the Qionglai Mountains, the Southern Sichuan Hills and the Wuling Mountains. The improved delimitation of P. bournei and P. zhennan and clarification of their ranges provide a better guidance for conservation and sustainable utilization of these tree species.
FIGURE 6. Fast distance based analysis tree for 16s ribosomal RNA gene. Note total genetic uniformity among 28 in Billions and billions sold: Pet-feeder crickets (Orthoptera: Gryllidae), commercial cricket farms, an epizootic densovirus, and government regulations make for a potential disaster
FIGURE 6. Fast distance based analysis tree for 16s ribosomal RNA gene. Note total genetic uniformity among 28 individuals of G. locorojo from eight "localities" on three continents. See Appendix A for specimen source data.
FIGURE 13. Phylogenetic tree inferred from a in Validation of the taxon Ixodes aragaoi Fonseca (Acari: Ixodidae) based on morphological and molecular data
FIGURE 13. Phylogenetic tree inferred from a partial sequence (435 characters, 113 parsimony informative) of the 16S rRNA mitochondrial gene of 12 tick species of the Ixodes ricinus complex, using I. nipponensis as outgroup. Numbers at nodes are the support values for the major branches (bootstrap) derived from 500 replicates for maximum parsimony. Number within brackets are GenBank accession numbers.
FIGURE 7. Maximum Likelihood best tree for 33 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 7. Maximum Likelihood best tree for 33 specimens of Devadatta and one outgroup taxon from the combined COI+16S+ITS+28S data set. Bootstrap support values below 100 are superimposed on the tree. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 5. 28S gene tree for 33 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 5. 28S gene tree for 33 specimens of Devadatta and one outgroup taxon, from Bayesian Inference analysis. Posterior probability values are shown (as percentages) if less than 100%. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 3. 16S gene tree for 33 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 3. 16S gene tree for 33 specimens of Devadatta and one outgroup taxon, from Bayesian Inference analysis. Posterior probability values are shown (as percentages) if less than 100%. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 4. ITS gene tree for 33 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 4. ITS gene tree for 33 specimens of Devadatta and one outgroup taxon, from Bayesian Inference analysis. Posterior probability values are shown (as percentages) if less than 100%. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 2. COI gene tree for the 33 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 2. COI gene tree for the 33 specimens of Devadatta also used in the analysis with other markers and one outgroup taxon. Posterior probability values are shown (as percentages) if less than 100%. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 1. COI gene tree for 73 in Revision of the genus Devadatta Kirby, 1890 in Borneo based on molecular and morphological methods, with descriptions of four new species (Odonata: Zygoptera: Devadattidae)
FIGURE 1. COI gene tree for 73 specimens of Devadatta and one outgroup taxon, from Bayesian Inference analysis. Posterior probability values are shown (as percentages) if less than 100%. RMNH collection codes are shown for each specimen, with the RMNH.INS. prefix omitted for clarity.
FIGURE 4. Tree from Bayesian inference for Bungarus candidus, B in A new color pattern of the Bungarus candidus complex (Squamata: Elapidae) from Vietnam based on morphological and molecular data
FIGURE 4. Tree from Bayesian inference for Bungarus candidus, B. magnimaculatus, and outgroup taxa. Values (%) at internal nodes are Bayesian posterior probabilities and bootstrap values from maximum parsimony and maximum likelihood, respectively.
FIGURE 5. Circular maximum parsimony phylogenetic tree with all sequenced recognised Thai Aleiodes species with a in A turbo-taxonomic study of Thai Aleiodes (Aleiodes) and Aleiodes (Arcaleiodes) (Hymenoptera: Braconidae: Rogadinae) based largely on COI barcoded specimens, with rapid descriptions of 179 new species
FIGURE 5. Circular maximum parsimony phylogenetic tree with all sequenced recognised Thai Aleiodes species with a number of named, primarily Palaearctic taxa included. Species groups that are characterizable morphologically and discussed are indicated in different colours. The tree is rooted using Heterogamus species.
FIGURE. Bayesian tree based on nuclear (ITS) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches. in Hedysarum sunhangii (Fabaceae, Hedysareae), a new species from Pamir-Alay (Babatag Ridge - Uzbekistan)
FIGURE. Bayesian tree based on nuclear (ITS) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches.
FIGURE. Bayesian tree based on combined plastid (matK, trnL-trnF) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches in Hedysarum sunhangii (Fabaceae, Hedysareae), a new species from Pamir-Alay (Babatag Ridge - Uzbekistan)
FIGURE. Bayesian tree based on combined plastid (matK, trnL-trnF) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.