Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
28
datasets available to search
ShareScore release 0.7.1
Dataset results
28 results for “supertree”
Spectral Cluster Supertree: Analysis Data
<p>Contains all datasets used in the Spectral Cluster Supertree paper. The datasets are composed of a set of rooted model trees, and rooted source trees to predict them. Please cite the appropriate papers, depending on which of the datasets you use.</p> <p>The <code>birth_death</code> folder contains our own dataset generated for our paper (where the generation process is explained), it aims to mimic what may be seen through divide and conquer algorithms for phylogenetic reconstruction. Parameters used to simulate an alignment were simulated under parameters estimated from a sequence alignment of 3 bacterial species (Kaehler et al., 2015) - see <code>alignment</code> folder.</p> <p>The <code>SMIDGenOutgrouped</code> folder contains both the SMIDGenOG (Fleischauer and Böcker, 2016) and SMIDGenOG-5500 dataset (Fleischauer and Böcker, 2017).</p> <p>The <code>SuperTriplets</code> folder contains the SuperTriplets dataset (Ranwez et al, 2010).</p> <p> </p>
Fig. 2 in A species-level supertree for stylophoran echinoderms
Fig. 2. Strict consensus supertree for stylophorans. The lower thecal surfaces of representative taxa are not drawn to the same scale (modified from Lefebvre and Vizcaino 1999; Lefebvre 2000b, 2001; and Martí Mus 2002). See text for details.
Fig. 1 in A species-level supertree for stylophoran echinoderms
Fig. 1. Basic anatomical organization of stylophorans. A. The cornute Flabellicarpus rushtoni (Lower Ordovician; England) in ventral (A1) and dorsal (A2) views (modified from Martí Mus 2002). B. The mitrate Rhenocystis latipedunculata (Lower Devonian; Germany) in ventral (B1) and dorsal (B2) views (modified from Ruta and Bartels 1998). Scale bars 1 cm.
Fig. 1 in Stylophoran supertrees revisited
Fig. 1. Morphology of stylophorans; all reconstructions in lower aspect. A. The cornute Cothurnocystis elizae (Upper Ordovician, Scotland); redrawn and modified from Ubaghs (1967). B. The primitive stylophoran Ceratocystis perneri (Middle Cambrian, Bohemia); redrawn from Ubaghs (1967). C. The mitrate Chinianocarpos thorali (Lower Ordovician, Montagne Noire, France); redrawn from Jefferies (1986). Scale bars 5 mm.
Fig. 3. 70 in Stylophoran supertrees revisited
Fig. 3. 70% majority−rule consensus supertree B for stylophorans. See text for details.
Fig. 2. 70 in Stylophoran supertrees revisited
Fig. 2. 70% majority−rule consensus supertree A for stylophorans. See text for details.
Data from: Supertrees based on the subtree prune-and-regraft distance
Supertree methods reconcile a set of phylogenetic trees into a single structure that is often interpreted as a branching history of species. A key challenge is combining conflicting evolutionary histories that are due to artifacts of phylogenetic reconstruction and phenomena such as lateral gene transfer (LGT). Although they often work well in practice, existing supertree approaches use optimality criteria that do not reflect underlying processes, have known biases and may be unduly influenced by LGT. We present the first method to construct supertrees by using the subtree prune-and-regraft (SPR) distance as an optimality criterion. Although calculating the rooted SPR distance between a pair of trees is NP-hard, our new maximum agreement forest-based methods can reconcile trees with hundreds of taxa and > 50 transfers in fractions of a second, which enables repeated calculations during the course of an iterative search. Our approach can accommodate trees in which uncertain relationships have been collapsed to multifurcating nodes. Using a series of simulated benchmark datasets, we show that SPR supertrees are more similar to correct species histories under plausible rates of LGT than supertrees based on parsimony or Robinson-Foulds distance criteria. We successfully constructed an SPR supertree from a phylogenomic dataset of 40,631 gene trees that covered 244 genomes representing several major bacterial phyla. Our SPR-based approach also allowed direct inference of highways of gene transfer between bacterial classes and genera; a small number of these highways connect genera in different phyla and can highlight specific genes implicated in long-distance LGT.
Data from: A comparison of supermatrix and supertree methods for multilocus phylogenetics using organismal datasets
It has been proposed that supertree approaches should be applied to large multilocus sequence datasets to achieve computational tractability. Large datasets such as those derived from phylogenomics studies can be broken into many locus-specific tree searches and the resulting trees can be stitched together via a supertree method. Using simulated data, workers have reported that they can rapidly construct a supertree that is comparable to the results of heuristic tree search on the entire dataset. To test this assertion with organismal data, we compared tree length under the parsimony criterion and computational time for twenty multilocus datasets using supertree (SuperFine and SuperTriplets) and supermatrix (heuristic search in TNT) approaches. Tree length and computational times were compared among methods using the Wilcoxon matched-pairs signed rank test. Supermatrix searches produce significantly shorter trees than either supertree approach (SuperFine or SuperTriplets; p < 0.0002 in both cases). Moreover, the processing time of supermatrix search was significantly lower than SuperFine+locus-specific search (p < 0.01) but roughly equivalent to that of SuperTriplets+locus-specific search (p > 0.4, not significant). In conclusion, we show by using real rather than simulated data, that there is no basis, either in time tractability or tree length, for use of supertrees over heuristic tree search using a supermatrix for phylogenomics.
Data from: Implementing and testing Bayesian and Maximum likelihood supertree methods in phylogenetics
Since their advent, supertrees have been increasingly used in large-scale evolutionary studies requiring a phylogenetic framework and substantial efforts have been devoted to developing a wide variety of supertree methods (SMs). Recent advances in supertree theory have allowed the implementation of maximum likelihood (ML) and Bayesian SMs, based on using an exponential distribution to model incongruence between input trees and the supertree. Such approaches are expected to have advantages over commonly used non-parametric SMs, e.g. matrix representation with parsimony (MRP). We investigated new implementations of ML and Bayesian SMs and compared these with some currently available alternative approaches. Comparisons include hypothetical examples previously used to investigate biases of SMs with respect to input tree shape and size, and empirical studies based either on trees harvested from the literature or on trees inferred from phylogenomic scale data. Our results provide no evidence of size or shape biases and demonstrate that the Bayesian method is a viable alternative to MRP and other non-parametric methods. Computation of input tree likelihoods allows the adoption of standard tests of tree topologies (e.g. the approximately unbiased test). The Bayesian approach is particularly useful in providing support values for supertree clades in the form of posterior probabilities.
Data from: SuperFine: fast and accurate supertree estimation
Many research groups are estimating trees containing anywhere from a few thousand to hundreds of thousands of species, towards the eventual goal of the estimation of a Tree of Life, containing perhaps as many as several million leaves. These phylogenetic estimations present enormous computational challenges, and current computational methods are likely to fail to run even on datasets in the low end of this range. One approach to estimate a large species tree is to use phylogenetic estimation methods (such as maximum likelihood) on a supermatrix produced by concatenating multiple sequence alignments for a collection of markers; however, the most accurate of these phylogenetic estimation methods are extremely computationally intensive for datasets with more than a few thousand sequences. Supertree methods, which assemble phylogenetic trees from a collection of trees on subsets of the taxa, are important tools for phylogeny estimation where phylogenetic analyses based upon maximum likelihood are infeasible. In this paper, we introduce SuperFine, a meta-method that utilizes a novel two-step procedure in order to improve the accuracy and scalability of supertree methods. Our study, using both simulated and empirical data, shows that SuperFine-boosted supertree methods produce more accurate trees than standard supertree methods, and run quickly on very datasets with thousands of sequences. Furthermore, SuperFine-boosted MRP (Matrix Representation with Parsimony, the most well known supertree method) approaches the accuracy of maximum likelihood methods on supermatrix datasets under realistic conditions.
Supplementary material 2 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
The machine-readable text from the screenshot.
Supplementary material 1 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
A one-per-line UT8-encoded plain-text list of URLs of the 5816 source PDFs used in this research
Figure 6 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 6 - The consensus supertree produced from an analysis of 924 source trees from the journal IJSEM.
Figure 1 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 1 - Overall workflow; from content acquisition to stripping figure images out of the PDF, to image filtering, image analysis and reconversion back into re-usable, machine-readable phylogenetic data.
Figure 4 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 4 - Output from image analysis of the input tree image in figure 1. All taxa and relationships are correctly reproduced, with branch lengths also preserved with high fidelity. (Note that the vertical ordering of the tips is not meaningful and is arbitrarily created by the display software.)
Figure 5 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 5 - Screenshot of exemplar machine-readable NeXML formatted data output from our automated analysis of the figure image from figure 1 of Park et al. 2008. Note that the genus, species, strain, and Genbank Accession numbers are semantically distinguished where detected. Heuristic post-OCR autocorrection processes are also noted where these have been applied (e.g. the conversion of a letter 'Z' to the number '2' in many Genbank Accession numbers). A machine-readable version of this file is supplied as supplementary material (Suppl. material 2).
Figure 3 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 3 - Number of leaves (terminal taxa) in each of 1614 source tree images (blue) and number of leaves recovered-from each image (orange). The modal number of taxa recovered per image was 12, the median was 13, and the mean was 13.96. The modal number of taxa not recovered from the trees was 2, the median was 5 and the mean was 7.15. The image mining process is lossy since most output tree files did not recover all of the taxa from the source image.
Figure 2 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 2 - A typical source input tree raster image (figure 1 from Park et al. 2008). Note the low resolution image quality. As this computer-generated ilustration follows predefined rules and conventions for the visual display of phylogenetic trees, we do not believe that it qualifies as a copyrightable work in itself (see Egloff et al. 2017 for more).
Figure 7 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 7 - Comparison between our supertree (left) and the NCBI Taxonomy reference tree (right): This example section of the supertree corresponds to taxa mostly from Rhodospirillaceae with the exception of rogue taxa indicated with a red asterisk. This section is related to the NCBI taxonomy reference tree on the right, containing those Rhodospirillaceae species leaves included in the supertree analysis (27). Nine taxa out of the 27 Rhodospirillaceae included were reconstructed elsewhere in our supertree (not shown). This is representative of the phylogenetic placement errors found throughout the supertree: individual rogue taxa, as well as misplaced clades of related taxa.
Figure 8 from: Mounce R, Murray-Rust P, Wills M (2017) A machine-compiled microbial supertree from figure-mining thousands of papers. Research Ideas and Outcomes 3: e13589. https://doi.org/10.3897/rio.3.e13589
Figure 8 - A visual exploration of taxon overlap of the 924 source trees used in this supertree analysis using the Supertree Toolkit 2 (Hill and Davis 2014). This demonstrates that there is not connectivity between all of the source trees we used in our supertree analysis.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.