Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
501
datasets available to search
ShareScore release 0.7.1
Dataset results
501 results for “Phylogenetic tree”
Data from: Multiple genotypes of Phelipanche ramosa indicate repeated introductions to the Americas: Sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Data from: PlaceMyFossils: An integrative approach to analyse and visualize the phylogenetic placement of fossils using backbone trees
Open the record for dataset details and reuse information.
Data from: Functional and phylogenetic dimensions of tree biodiversity reveal unique geographic patterns
Open the record for dataset details and reuse information.
Machine learning can be as good as maximum likelihood when reconstructing phylogenetic trees and determining the best evolutionary model on four taxon alignments
Open the record for dataset details and reuse information.
Empirical data for: Extending phylogenetic regression models for comparing within-species patterns across the Tree of Life
Open the record for dataset details and reuse information.
Data alignment and phylogenetic trees from Phylogeny, ecology, morphological evolution, and reclassification of the diatom orders Surirellales and Rhopalodiales
<p>Alignment and tree files from Ruck et al. 2016:</p> <p>Phylogeny, ecology, morphological evolution, and reclassification of the diatom orders Surirellales and Rhopalodiales</p>
Data for: Foliar spectra accurately distinguish most temperate tree species and show strong phylogenetic signal.
<p>Gene supermatrix and partitions used for Blanchard, F., Bruneau, A., Laliberté, E. (2023). Foliar spectra accurately distinguish most temperate tree species and show strong phylogenetic signal. <i>Am.J.Bot</i>., [Submitted]. See text for more information.</p><p>All leaf spectral and trait data can be found at https://data.caboscience.org/leaf/</p>
Data from: Phylogenetic conflict between species tree and maternally inherited gene trees in a clade of Emberiza buntings (Aves: Emberizidae)
<p>Different genomic regions may reflect conflicting phylogenetic topologies on account of incomplete lineage sorting and/or gene flow. Genomic data are necessary to reconstruct the true species tree and explore potential causes of phylogenetic conflict. Here, we investigate the phylogenetic relationships of four <em>Emberiza</em> species (Aves: Emberizidae) and discuss the potential causes of the observed mitochondrial non-monophyly of <em>Emberiza godlewskii</em> (Godlewski's bunting) using phylogenomic analyses based on whole genome resequencing data from 41 birds. Phylogenetic analyses based on both the whole mitochondrial genome and ~39 kilobases from the non-recombining W chromosome reveal that the northern and southern populations of <em>E. godlewskii</em> are each sister to <em>E. cioides</em> and <em>E. cia</em>, respectively. In contrast, phylogenetic analysis based on genome-wide data support the monophyly of <em>E. godlewskii</em> with the following tree topology: (((<em>E. godlewskii</em>, <em>E. cia</em>), <em>E. cioides</em>), <em>E. jankowskii</em>).<em> </em>Using D-statistics, we detected multiple gene flow events among different lineages, indicating pervasive introgressive hybridization within this clade. Introgression from an unsampled lineage that is sister to <em>E. cioides</em> or introgression from an unsampled mitochondrial + W chromosomal lineage of <em>E. cioides</em> into northern <em>E. godlewskii </em>may explain the phylogenetic conflict between the species tree estimated from genome-wide data and mtDNA/W trees. These results underscore the importance of using genomic data for phylogenetic reconstruction and species delimitation.</p>
The effect of neighbor species' phylogenetic and trait difference on tree growth in subtropical forests
<p>To comprehensively understand ecological dynamics within a forest ecosystem, it is vital to explore how surrounding trees influence the growth of individual trees in a community. Biotic interactions have a significant potential impacts on individual tree growth and their effects can be evaluated through trait and phylogenetic-based approaches. This study investigates the relative importance of biotic interactions on tree growth by examining several metrics and considering three classes of intrinsic growth rates among the focal individuals: slower, intermediate and faster-growing trees. The metrics include hierarchical and absolute trait differences of focal trees to neighbors, neighborhood crowding index, phylogenetic distance, and trait community metrics. Our results indicated that the phylogenetic distance between the focal tree and its neighbors positively impacted the growth of all classes, whereas different traits have distinct effects on slower and faster-growing trees. Specific leaf area (SLA) and leaf area (LA) showed hierarchical importance to tree growth. Trees surrounded by neighbors with higher SLA and LA than themselves grow better, particularly for slower-growing trees. Higher levels of wood density difference between the focal trees and their neighbors positively impacts slower and faster-growing trees, while height difference negatively impacts faster-growing trees. We conclude that the interactions between trees are mediated by their ecological differences, but the performance and responses to surrounding competitors vary along with their growth class within a community. This study has revealed that the tree's intrinsic growth rate mediates the effect of traits and phylogeny of surrounding trees on individual tree growth.</p>
Alignments and tree files from: Phylogenetic relationships within tribe Hibisceae (Malvaceae) reveal complex patterns of polyphyly in Hibiscus and Pavonia
<p>The diverse and spectacular Hibisceae tribe comprises over 750 species. No studies, however, have broadly sampled across the dozens of genera in the tribe, leading to uncertainty in the relationships among genera. The non-monophyly of the genus <em>Hibiscus </em>is infamous and challenging, whereas the monophyly of most other genera in the tribe has yet to be assessed, including the large genus <em>Pavonia</em>. Here we significantly increase taxon sampling in the most complete phylogenetic study of the tribe to date. We assess monophyly of most currently recognized genera in the tribe and include three and thirteen newly sampled sections of <em>Hibiscus </em>and <em>Pavonia</em>, respectively. We also include five rarely sampled genera and 137 species previously unsampled. Our phylogenetic trees demonstrate that <em>Hibiscus</em>, as traditionally defined, encompasses at least 20 additional genera. The status of <em>Pavonia </em>emerges as comparable in complexity to <em>Hibiscus</em>. We offer clarity in the phylogenetic placement of several taxa of uncertain affinity (e.g., <em>Helicteropsis, Hibiscadelphus, Jumelleanthus, and Wercklea)</em>. We also identify two new clades and elevate them to the generic rank with the recognition of two, new monotypic genera: 1) <em>Blanchardia </em>M.M.Hanes & R.L.Barrett is a surprising Caribbean lineage that is sister to the entire tribe, and 2) <em>Astrohibiscus </em>McLay & R.L.Barrett represents former members of <em>Hibiscus caesius</em> s.l. <em>Cravenia </em>McLay & R.L.Barrett is also described as a new genus for the <em>Hibiscus panduriformis</em> clade which is allied to <em>Abelmoschus</em>. Finally, we introduce a new classification for the tribe and clarify the boundaries of <em>Hibiscus </em>and <em>Pavonia</em>.</p>
Data from: Geography and ecology shape the phylogenetic composition of Amazonian tree communities
<p><strong>Aim:</strong> Amazonia hosts more tree species, from numerous evolutionary lineages both young and ancient, than any other biogeographic region. Previous studies have shown that tree lineages colonised multiple edaphic environments and dispersed widely across Amazonia, leading to a hypothesis, which we test, that lineages should not be strongly associated with either geographic regions or edaphic forest types.</p> <p><strong>Location:</strong> Amazonia.</p> <p><strong>Taxon:</strong> Angiosperms (Magnoliids; Monocots; Eudicots).</p> <p><strong>Methods:</strong> Data for the abundance of 5,082 tree species in 1,989 plots were combined with a mega-phylogeny. We applied evolutionary ordination to assess how phylogenetic composition varies across Amazonia. We used variation partitioning and Moran's eigenvector maps (MEM) to test and quantify the separate and joint contributions of spatial and environmental variables to explain the phylogenetic composition of plots. We tested the indicator value of lineages for geographic regions and edaphic forest types and mapped associations onto the phylogeny.</p> <p><strong>Results:</strong> In the terra firme and várzea forest types, phylogenetic composition varies by geographic region, but the igapó and white-sand forest types retain a unique evolutionary signature regardless of region. Overall, we find that soil chemistry, climate, and topography explain 24% of the variation in phylogenetic composition, with 79% of that variation being spatially structured (R <sup>2</sup> = 19% overall for combined spatial/environmental effects). Phylogenetic composition also shows substantial spatial patterns not related to the environmental variables we quantified (R <sup>2</sup> = 28%). A greater number of lineages were significant indicators of geographic regions than forest types.</p> <p><strong>Main conclusions:</strong> Numerous tree lineages, including some ancient ones (>66 Ma), show strong associations with geographic regions and edaphic forest types of Amazonia. This shows that specialization on specific edaphic environments has played a long-standing role in the evolutionary assembly of Amazonian forests. Furthermore, many lineages, even those that have dispersed across Amazonia, dominate within a specific region, likely because of phylogenetically conserved niches for environmental conditions that are prevalent within regions. </p>
Phylogenetic trees and morphological traits dataset of Geonoma palm genus
<p>The phylogenetic trees correspond to the <em>Geonoma</em> genus and Geonomateae tribe phylogenies based on a concatenation analysis of 20 small informative genes and the csv file corresponds to the mean value of 24 morphological traits for 68 <em>Geonoma</em> species based on Henderson (2012) dataset.</p>
The thermal niche and phylogenetic assembly of evergreen tree metacommunities in a mid-to-upper tropical montane zone
<p>Frost and freezing temperatures have posed an obstacle to tropical woody evergreen plants over evolutionary timescales. Thus, along tropical elevation gradients, frost may influence woody plant community structure by filtering out lowland tropical clades, and allowing extra-tropical lineages to establish at higher elevations.</p> <p>Here we assess the extent to which frost influences the taxonomic and phylogenetic structure of naturally-patchy evergreen forests (locally known as <em>shola</em>) along a mid-upper montane elevation gradient in the Western Ghats, India. Specifically, we examine the role of large-scale macroclimate and factors affecting local microclimates, including <em>shola</em> patch size and distance from <em>shola</em> edge, in driving <em>shola </em>metacommunity structure. We find that the <em>shola </em>metacommunity shows phylogenetic overdispersion with elevation, with greater representation of extra-tropical lineages above 2000m, and marked turnover in taxonomic composition of <em>shola </em>woody communities near the frost-affected forest edge above 2000m, from those below 2000m. Both minimum winter temperature and patch size were equally important in determining metacommunity structure, with plots inside very large <em>sholas </em>dominated by older tropical lineages, with many endemics. Phylogenetic overdispersion in the upper montane <em>shola </em>metacommunity thus resulted from tropical lineages persisting in the interiors of large closed frost-free <em>sholas</em>, where their regeneration niche has been preserved over time.</p>
Amino acid sequences of RWP-RK domain containing proteins used for the construction of phylogenetic tree shown in Fig. 1
<p><span>The RWP-RK protein family is a group</span><span> of transcription factors containing </span><span>the RWP-RK DNA-binding domain. The RWP-RK DNA-binding domain is an ancient motif that emerged before the establishment of the Viridiplantae (green plants), which consist of green algae and land plants. This domain is mostly absent in other kingdoms but widely distributed in Viridiplantae. In green algae, a liverwort, and several angiosperms, RWP-RK proteins play essential roles in nitrogen responses and sexual reproduction-associated processes, which</span><span> </span><span>are seemingly unrelated phenomena but possible interdependent processes</span><span> </span><span>in autotrophs. Consistent with</span><span> related but diversified roles of the RWP-RK proteins in these organisms, the RWP-RK protein family appears to have expanded intensively, but independently, in the algal and land plant lineages. Therefore, bryophyte RWP-RK proteins occupy a unique position in the evolutionary process of establishing the RWP-RK protein family. In this review, we summarize current knowledge about the RWP-RK protein family in the Viridiplantae, and discuss the significance of bryophyte RWP-RK proteins in clarifying the relationship between diversification in the RWP-RK protein family and </span><span>procurement</span><span> of sophisticated mechanisms for adaptation to the terrestrial environment.</span></p>
Robust analysis of phylogenetic tree space
<p>Phylogenetic analyses often produce large numbers of trees. Mapping trees' distribution in "tree space" can illuminate the behavior and performance of search strategies, reveal distinct clusters of optimal trees, and expose differences between different data sources or phylogenetic methods—but the high-dimensional spaces defined by metric distances are necessarily distorted when represented in fewer dimensions. Here, I explore the consequences of this transformation in phylogenetic search results from 128 morphological data sets, using stratigraphic congruence—a complementary aspect of tree similarity—to evaluate the utility of low-dimensional mappings. I find that phylogenetic similarities between cladograms are most accurately depicted in tree spaces derived from information-theoretic tree distances or the quartet distance. Robinson–Foulds tree spaces exhibit prominent distortions and often fail to group trees according to phylogenetic similarity, whereas the strong influence of tree shape on the Kendall–Colijn distance makes its tree space unsuitable for many purposes. Distances mapped into two or even three dimensions often display little correspondence with true distances, which can lead to profound misrepresentation of clustering structure. Without explicit testing, one cannot be confident that a tree space mapping faithfully represents the true distribution of trees, nor that visually evident structure is valid. My recommendations for tree space validation and visualization are implemented in a new graphical user interface in the "TreeDist" R package.</p>
Distribution dataset of the 140 Chinese mountain floras and dated phylogenetic tree
<p>We compiled<strong> </strong>checklists of species of angiosperm for 140 Chinese mountain flora from previously published, comprehensive species checklists, white papers, and research papers. From our initial checklists, we excluded all nonnative species, and we reconciled taxonomy to the Leipzig Catalogue of Vascular Plants (LCVP), with infraspecific taxa combined under their respective species. Following taxonomic reconciliation and categorization within higher ranks, our dataset comprised a total of 17,576 species in 2,585 genera belonging to 251 families and 56 orders.</p> <p>In addition, six dated phylogenetic trees by the V.PhyloMaker2 approach are provided.</p> <p> </p>
Supplementary Material 1: Phylogenetic tree from Unraveling an unknown diversity of archaeal and bacterial tetraether membrane lipid producers in a euxinic marine system
<p>Phylogenetic tree (Black Sea MAGs)</p>
MAST: Phylogenetic inference with mixtures across sites and trees (revised)
<p>Hundreds or thousands of loci are now routinely used in modern phylogenomic studies. Concatenation approaches to tree inference assume that there is a single topology for the entire dataset, but different loci may have different evolutionary histories due to incomplete lineage sorting, introgression, and/or horizontal gene transfer; even single loci may not be treelike due to recombination. To overcome this shortcoming, we introduce an implementation of a multi-tree mixture model that we call MAST. This model extends a prior implementation by Boussau et al. (2009) by allowing users to estimate the weight of each of a set of pre-specified bifurcating trees in a single alignment. The MAST model allows each tree to have its own weight, topology, branch lengths, substitution model, nucleotide or amino acid frequencies, and model of rate heterogeneity across sites. We implemented the MAST model in a maximum-likelihood framework in the popular phylogenetic software, IQ-TREE. Simulations show that we can accurately recover the true model parameters, including branch lengths and tree weights for a given set of tree topologies, under a wide range of biologically realistic scenarios. We also show that we can use standard statistical inference approaches to reject a single-tree model when data are simulated under multiple trees (and vice versa). We applied the MAST model to multiple primate datasets and found that it can recover the signal of incomplete lineage sorting in the Great Apes, as well as the asymmetry in minor trees caused by introgression among several macaque species. When applied to a dataset of four Platyrrhine species for which standard concatenated maximum likelihood and gene tree approaches disagree, we observe that MAST gives the highest weight (i.e. the largest proportion of sites) to the tree also supported by gene tree approaches. These results suggest that the MAST model is able to analyse a concatenated alignment using maximum likelihood while avoiding some of the biases that come with assuming there is only a single tree. We discuss how the MAST model can be extended in the future.</p>
Supplementary Data: MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation
<p>Supplementary Data<br> MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation<br> Submitted to BMC Evolutionary Biology</p> <p>This record contains PANDIT based dataset and TreeBASE dataset (Nguyen et al. 2015) which are analyzed by different bootstrap methods in the study "MPBoot: Fast phylogenetic maximum parsimony tree inference and bootstrap approximation". The PANDIT based dataset (compressed in file data_pandit.tar.gz) is used to benchmark the accuracy of bootstrap estimates. The TreeBASE dataset (compressed in file data_treebase.tar.gz) is used to benchmark computing times and capability of finding the best-known MP scores. </p> <p>After being uncompressed, the PANDIT based dataset comprises two subdirectories corresponding to the simulated DNA and AA MSAs. They were generated by Seq-Gen (Rambaut and Grass 1997), where the model parameters and true tree were inferred from the original MSAs downloaded from the PANDIT database (Whelan et al. 2006).</p> <ul> <li>Inside "dna" subdirectory, there are 6,207 numbered directories corresponding to 6,207 DNA MSAs. Note that the numbering of these directories is not consecutive because we excluded MSAs where TNT or PAUP* runs did not finish. In each numbered directory N, there are three files: (1) data.N contains the simulated MSA in PHYLIP format; (2) model.N contains the best-fit model detected from the corresponding original MSA; (3) tree.N contains the tree (in Newick format) inferred from the corresponding original MSA. tree.N and model.N are used by Seq-Gen to simulate the MSA in data.N.</li> <li>The "aa" subdirectory is organized similarly for 6,165 AA MSAs.</li> </ul> <p>After being uncompressed, the TreeBASE dataset comprises 115 files corresponding to 115 MSAs. There are:</p> <ul> <li>70 DNA MSAs in PHYLIP format. These files follow the naming scheme dna_[number of sequences]_[number of sites].phy.</li> <li>45 protein MSAs in PHYLIP format. These files follow the naming scheme prot_[number of sequences]_[number of sites].phy.<br> </li> </ul>
Fig. 1. Phylogenetic relationships of species of Eusurbus and Zentamyia. Tree generated from morpho- logical phylogenetic analysis, unambiguous apomorphies mapped on branches, black circles indicate non- homoplasious changes.
Fig. 1. Phylogenetic relationships of species of Eusurbus and Zentamyia. Tree generated from morpho- logical phylogenetic analysis, unambiguous apomorphies mapped on branches, black circles indicate non- homoplasious changes.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.