Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,344
datasets available to search
ShareScore release 0.9.0
Dataset results
1,344 results for “: phylogenomics”
Sequence-capture phylogenomics of true spiders reveals convergent evolution of respiratory systems
<p>The common ancestor of spiders likely used silk to line burrows or make simple webs, with specialized spinning organs and aerial webs originating with the evolution of the megadiverse "true spiders" (Araneomorphae). The base of the araneomorph tree also concentrates the greatest number of changes in respiratory structures, a character system whose evolution is still poorly understood, and that might be related to the evolution of silk glands. Emphasizing a dense sampling of multiple araneomorph lineages where tracheal systems likely originated, we gathered genomic-scale data and reconstructed a phylogeny of true spiders. This robust phylogenomic framework was used to conduct maximum likelihood and Bayesian character evolution analyses for respiratory systems, silk glands, and aerial webs, based on a combination of original and published data. Our results indicate that in true spiders, posterior book lungs were transformed into morphologically similar tracheal systems six times independently, after the evolution of novel silk gland systems and the origin of aerial webs. From these comparative data we put forth a novel hypothesis that early-diverging web building spiders were faced with new energetic demands for spinning, which prompted the evolution of similar tracheal systems via convergence; we also propose tests of predictions derived from this hypothesis.</p>
Data from: Extracting phylogenetic signal from phylogenomic data: higher-level relationships of the nightbirds (Strisores)
A well-resolved phylogeny would facilitate study of adaptation to nocturnality in the avian superorder Strisores, a group that includes both nocturnal and diurnal lineages. Based on previous estimates, it could be hypothesized that there were multiple independent origins of nocturnality in this group. In order to refine the Strisores phylogeny, we generated genome-scale datasets of 2,289 – 4,243 ultra-conserved elements for 23 taxa representing all major living lineages in the group. Among the considerations for using genome-scale, molecular sequence data in phylogenomic analysis are issues related to GC content, GC variance and their effects on model selection. In this study, we employed a variety of analytical techniques to empirically investigate those issues in our data, as well as biases and errors resulting from alignment trimming, taxon selection and matrix completeness. Extensive analyses revealed conflict within the data, especially in regard to variation in GC content, that would not have been detected with more cursory study. Our results indicate that readily available models of molecular evolution are insufficient to encapsulate all phenomena present in genome-scale matrices, and that this problem may be at the root of many current issues in phylogenomic analysis. The analytical methods employed in this study are relevant to phylogenomic analysis of any large, heterogeneous matrix. In conclusion, we present a strongly supported estimate of the Strisores tree and discuss potential evolutionary pathways of nocturnality in this clade.
Phylogenomic analysis sheds light on the evolutionary pathways towards acoustic communication in Orthoptera
<p>Acoustic communication is enabled by the evolution of specialised hearing and sound producing organs. In this study, we performed a large-scale macroevolutionary study to understand how both hearing and sound production evolved and affected diversification in the insect order Orthoptera, which includes many familiar singing insects, such as crickets, katydids, and grasshoppers. Using phylogenomic data, we firmly establish phylogenetic relationships among the major lineages and divergence time estimates within Orthoptera, as well as the lineage-specific and dynamic patterns of evolution for hearing and sound producing organs. In the suborder Ensifera, we infer that forewing-based stridulation and tibial tympanal ears co-evolved, but in the suborder Caelifera, abdominal tympanal ears first evolved in a non-sexual context, and later co-opted for sexual signalling when sound producing organs evolved. However, we find little evidence that the evolution of hearing and sound producing organs increased diversification rates in those lineages with known acoustic communication.</p>
Phylogenomic data reveal widespred introgression across the range of an alpine and arctic specialist
<p>Understanding how gene flow affects population divergence and speciation remains challenging. Differentiating one evolutionary process from another can be difficult because multiple processes can produce similar patterns, and more than one process can occur simultaneously. While simple population models produce predictable results, how these processes balance in taxa with patchy distributions and complicated natural histories is less certain. These types of populations might be highly connected through migration (gene flow), but can experience stronger effects of genetic drift and inbreeding, or localized selection. While different signals can be difficult to separate, the application of high throughput sequence data can provide the resolution necessary to distinguish many of these processes. We present whole genome sequence data for an avian species group with an alpine and arctic tundra distribution to examine the role that different population genetic processes have played in their evolutionary history. Rosy-finches inhabit high elevation mountaintop sky islands and high-latitude island and continental tundra. They exhibit extensive plumage variation coupled with low levels of genetic variation. Additionally, the number of species within the complex is debated, making them excellent for studying the forces involved in the process of diversification, as well as an important species group in which to investigate species boundaries. Total genomic variation suggests a broadly continuous pattern of allele frequency changes across the mainland taxa of this group in North America. However, phylogenomic analyses recover multiple distinct, well supported, groups that coincide with previously described morphological variation and current species-level taxonomy. Tests of introgression using D-statistics and approximate Bayesian computation reveal significant levels of introgression between multiple North American taxa. These results provide insight into the balance between divergent and homogenizing population genetic processes, and highlight remaining challenges in interpreting conflict between different types of analytical approaches with whole genome sequence data.</p>
Phylogenomics, biogeography and taxonomic revision of New Guinean pythons (Pythonidae, Leiopython) harvested for international trade
<p>The large and enigmatic New Guinean pythons in the genus <i>Leiopython</i> are harvested from the wild to supply the international trade in pets. Six species are currently recognized (<i>albertisii</i>, <i>biakensis</i>, <i>fredparkeri</i>, <i>huonensis</i>, <i>meridionalis</i>, <i>montanus</i>) but the taxonomy of this group has been controversial. We combined analysis of 421 nuclear loci and complete mitochondrial genomes with morphological data to construct a detailed phylogeny of this group, understand their biogeographic patterns and establish the systematic diversity of this genus. Our molecular genetic data support two major clades, corresponding to <i>L. albertisii</i> and <i>L. meridionalis</i>, but offer no support for the other four species. Our morphological data also only support two species. We therefore recognize <i>L. albertisii</i> and <i>L. meridionalis</i> as valid species and place <i>L. biakensis, L. fredparkeri, L. huonensis </i>and<i> L. montanus </i>into synonymy. We found that <i>L. albertisii</i>and <i>L. meridionalis</i> are sympatric in western New Guinea; an atypical pattern compared to other Papuan species complexes in which the distributions of sister taxa are partitioned to the north and south of the island's central mountain range. For the purpose of conservation management, overestimation of species diversity within <i>Leiopython</i> has resulted in the unnecessary allocation of resources that could have been expended elsewhere. We strongly caution against revising the taxonomy of geographically widespread species groups when little or no molecular genetic data and only small morphological samples are available.</p>
Phylogenomic resolution of sea spider diversification through integration of multiple data classes
<p><span><span><span><span><span><span><span><span><span><span><span>Despite significant advances in invertebrate phylogenomics over the past decade, the higher-level phylogeny of Pycnogonida (sea spiders) remains elusive. Due to the inaccessibility of some small-bodied lineages, few phylogenetic studies have sampled all sea spider families. Previous efforts based on a handful of genes have yielded unstable tree topologies. Here, we inferred the relationships of 89 sea spider species using targeted capture of the mitochondrial genome, 56 conserved exons, 101 ultraconserved elements, and three nuclear ribosomal genes. We inferred molecular divergence times by integrating morphological data for fossil species to calibrate 15 nodes in the arthropod tree of life. This integration of data classes resolved the basal topology of sea spiders with high support. The enigmatic family Austrodecidae was resolved as the sister group to the remaining Pycnogonida and the small-bodied family Rhynchothoracidae as the sister group of the robust-bodied family Pycnogonidae. Molecular divergence time estimation recovered a basal divergence of crown group sea spiders in the Ordovician. Comparison of diversification dynamics with other marine invertebrate taxa that originated in the Paleozoic suggests that sea spiders and some crustacean groups exhibit resilience to mass extinction episodes, relative to mollusk and echinoderm lineages. </span></span></span></span></span></span></span></span></span></span></span></p>
Mimicry and mitonuclear discordance in nudibranchs: new insights from exon capture phylogenomics
<p>Phylogenetic inference and species delimitation can be challenging in taxonomic groups that have recently radiated and where introgression produces conflicting gene trees, especially when species delimitation has traditionally relied on mitochondrial data and colour pattern. <i>Chromodoris</i>, a genus of colourful and toxic nudibranch in the Indo-Pacific, has been shown to have extraordinary cryptic diversity and mimicry, and has recently radiated, ultimately complicating species delimitation. In these cases, additional genome-wide data can help improve phylogenetic resolution and provide important insights about evolutionary history. Here, we employ a transcriptome-based exon capture approach to resolve <i>Chromodoris</i> phylogeny with data from 2,925 exons and 1,630 genes, derived from 15 nudibranch transcriptomes. We show that some previously identified mimics instead show mitonuclear discordance, likely deriving from introgression or mitochondrial capture, but we confirm one 'pure' mimic in Western Australia. Sister-species relationships and species-level entities were recovered with high support in both concatenated Maximum Likelihood (ML) and summary coalescent phylogenies, but the ML topologies were highly variable while the coalescent topologies were consistent across datasets. Our work also demonstrates the broad phylogenetic utility of 149 genes that were previously identified from eupulmonate gastropods. This study is one of the first to i) demonstrate the efficacy of exon capture for recovering relationships among recently radiated invertebrate taxa, ii) employ genome-wide nuclear markers to test mimicry hypotheses in nudibranchs and iii) provide evidence for introgression and mitochondrial capture in nudibranchs.</p>
Supplementary Information for: A cautionary note on the use of genotype callers in phylogenomics
<p>Next-generation-sequencing genotype callers are commonly used in studies to call variants from newly-sequenced species. However, due to the current availability of genomic resources, it is still common practice to use only one reference genome for a given genus, or even one reference for an entire clade of a higher taxon. The problem with traditional genotype callers, such as the one from GATK, is that they are optimized for variant calling at the population level. However, when these callers are used at the phylogenetic level, the consequences for downstream analyses can be substantial. Here, we performed simulations to compare the performance between the genotype callers of GATK and ATLAS, and present their differences at various phylogenetic scales. We show that the genotype caller of GATK substantially underestimates the number of variants at the phylogenetic level, but not at the population level. We also found that the accuracy of heterozygote calls declines with increasing distance to the reference genome. We quantified this decline, and found that it is very sharp in GATK, while ATLAS maintains a high accuracy even at moderately-divergent species from the reference. We further suggest that efforts should be taken towards acquiring more reference genomes per species, before pursuing high-scale phylogenomic studies.</p>
Phylogenomics of the Andean tetraploid clade of the American Amaryllidaceae (subfamily Amaryllidoideae): unlocking a polyploid generic radiation abetted by continental geodynamics
<p>The second large clade of the endemic American Amaryllidaceae subfam. Amaryllidoideae constitutes the tetraploid-derived (<i>n</i> = 23) Andean-centered tribes, most of which have 46 chromosomes. Despite progress in resolving phylogenetic relationships of the group with nrDNA, certain subclades were poorly resolved or weakly supported in those previous studies. Sequence capture using anchored hybrid enrichment was employed across 95 species of the clade along with five outgroups and generated sequences of 524 nuclear genes and a partial plastome. Maximum likelihood phylogenetic analyses were conducted on concatenated supermatrices, and coalescent species tree analyses were run on the gene trees, followed by hybridization network, age diversification and biogeographic analyses. The four tribes Clinantheae, Eucharideae, Eustephieae (the first branch), and Hymenocallideae (sister to <i>Clinanthus</i>) are resolved in all analyses with 100% support. Nuclear gene supermatrix and species tree results were largely in concordance; however cytonuclear discordance was evident. Hybridization network analysis identified significant reticulation in <i>Clinanthus</i>, <i>Hymenocallis</i>, <i>Stenomesson</i> and the subclade of Eucharideae comprising <i>Eucharis</i>, <i>Caliphruria</i>, and <i>Urceolina</i>. Our data support a previous treatment of the latter as a single genus, <i>Urceolina</i>, with the addition of <i>Eucrosia</i> <i>dodsonii</i>. Biogeographic analysis and penalized likelihood age estimation suggests an origin in the central Andean region (north and central Peru) for the complex in the mid-Oligocene, with more dispersals than vicariances in its history, but no extinctions. The Eucharideae experienced a sudden lineage radiation ca. 10 Mya. We tie much of the divergences in the Andean-centered lineages to the rise of the Andes, directly and indirectly, and suggest that the Amotape-Huancabamba Zone functioned as both a corrider (dispersal) and a barrier to migration (vicariance). Several taxonomic changes are made. This is the largest DNA sequence data set to be applied within Amaryllidaceae to date.The second large clade of the endemic American Amaryllidaceae subfam. Amaryllidoideae constitutes the tetraploid-derived (<i>n</i> = 23) Andean-centered tribes, most of which have 46 chromosomes. Despite progress in resolving phylogenetic relationships of the group with nrDNA, certain subclades were poorly resolved or weakly supported in those previous studies. Sequence capture using anchored hybrid enrichment was employed across 95 species of the clade along with five outgroups and generated sequences of 524 nuclear genes and a partial plastome. Maximum likelihood phylogenetic analyses were conducted on concatenated supermatrices, and coalescent species tree analyses were run on the gene trees, followed by hybridization network, age diversification and biogeographic analyses. The four tribes Clinantheae, Eucharideae, Eustephieae (the first branch), and Hymenocallideae (sister to <i>Clinanthus</i>) are resolved in all analyses with 100% support. Nuclear gene supermatrix and species tree results were largely in concordance; however cytonuclear discordance was evident. Hybridization network analysis identified significant reticulation in <i>Clinanthus</i>, <i>Hymenocallis</i>, <i>Stenomesson</i> and the subclade of Eucharideae comprising <i>Eucharis</i>, <i>Caliphruria</i>, and <i>Urceolina</i>. Our data support a previous treatment of the latter as a single genus, <i>Urceolina</i>, with the addition of <i>Eucrosia</i> <i>dodsonii</i>. Biogeographic analysis and penalized likelihood age estimation suggests an origin in the central Andean region (north and central Peru) for the complex in the mid-Oligocene, with more dispersals than vicariances in its history, but no extinctions. The Eucharideae experienced a sudden lineage radiation ca. 10 Mya. We tie much of the divergences in the Andean-centered lineages to the rise of the Andes, directly and indirectly, and suggest that the Amotape-Huancabamba Zone functioned as both a corrider (dispersal) and a barrier to migration (vicariance). Several taxonomic changes are made. This is the largest DNA sequence data set to be applied within Amaryllidaceae to date.</p>
Data from: Schneider et al. (2020). Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa. Taxon.
<p>DNA sequence alignments with all loci concatenated (e.g., "Alignment_concatenated_xxx_dataset") or for each locus separated (see folders "Gene_alignments_xxx_dataset"). The individual gene alignments are identified by their locus number. Numbers in the sequence headers of the fasta files correspond to the Lab IDs of specimens (see related publication for detailed voucher information). These alignments were used for the phylogenetic analyses in the related publication. The bait set contains the probe sequences used for the targeted enrichment of nuclear loci of Ochnaceae.</p>
Phylogenomics of Perityleae (Compositae) provides new insights into morphological and chromosomal evolution of the rock daisies
<p>Rock daisies (Perityleae; Compositae) are a diverse clade of seven genera and ca. 84 minimum-rank taxa that mostly occur as narrow endemics on sheer rock-cliffs throughout the southwest U.S. and northern Mexico. Taxonomy of Perityleae has traditionally been based on morphology and cytogenetics. To test taxonomic hypotheses and utility of characters emphasized in past treatments, we present the first densely sampled molecular phylogenies of Perityleae and reconstruct trait and chromosome evolution. We inferred phylogenetic trees from whole chloroplast genomes, nuclear ribosomal cistrons, and hundreds of low-copy nuclear genes using genome skimming and target-capture. Discordance between sources of molecular data suggests a underappreciated history of hybridization in Perityleae. Phylogenies support the monophyly of subtribe Peritylinae, a distinctive group possessing a four-lobed disc corolla; however, all of the phylogenetic trees generated in this study reject the monophyly of the most species-rich genus, Perityle as well as its sections Perityle sect. Perityle, Perityle sect. Laphamia, and Perityle sect. Pappothrix. Using reversible jump MCMC, our results suggest that morphological characters traditionally used to classify members of Perityleae have evolved multiple times within the group. A base chromosome number of x=18 gave rise to lower base numbers in subtribe Peritylinae (x=12, 13, 16, 17, and 19) by descending dysploidization. Most taxa constitute a monophyletic lineage with a base chromosome number of x=17, with many polyploidization events. These results demonstrate the advantages and obstacles to next-generation sequencing approaches in synantherology while laying the foundation for taxonomic revision and comparative study of the evolutionary ecology of Perityleae</p>
Mitochondrial genes from eighteen angiosperms fill sampling gaps for phylogenomic inferences of the early diversification of flowering plants
<p class="Default"><span>The early diversification of angiosperms is a rapid yet complicated process and thus it renders the phylogenetic analyses of early-diverging angiosperms much difficulty. Plastid and nuclear phylogenomic studies have raised several controversial hypotheses regarding the angiosperm phylogeny, whereas mitochondrial genomes have been largely ignored. In this study, we newly sequenced mitochondrial genomes from 18 angiosperms to fill the sampling gaps in magnoliids, Austrobaileyales, Chloranthales, Ceratophyllales, and early-diverging lineages of eudicots and monocots. A data matrix of 38 mitochondrial genes from 107 taxa was assembled to address this question. Although conflicting phylogenies were recovered in this study from different datasets and analytical methods, congruence was achieved regarding the deep relationships of several major angiosperm lineages: Chloranthales always groups with Ceratophyllales, Austrobaileyales is sister to mesangiosperms, and a previously unplaced clade—Dilleniales—is consistently resolved as a sister to superasterids. Substitutional saturation, GC compositional heterogeneity, and codon-usage bias are suggested as common reasons for the noisy signals that impact phylogenetic reconstructions, and angiosperm mitochondrial genes seem to hardly suffer from these factors. In addition, the 3<sup>rd</sup> codon positions of the mitochondrial genes contained more phylogenetic signals than the 1<sup>st</sup> and 2<sup>nd</sup> codon positions, which might be responsible for the incongruent results recovered among different datasets. Due to the rapid radiation process, the relationships among these early lineages are not well resolved. Nevertheless, this study based on mitochondrial genomes provides additional evidence and alternative hypotheses for the early evolution and diversification of angiosperms.</span></p>
Data from: Chloroplast phylogenomics resolves key relationships in ferns
Studies on chloroplast genomes of ferns and lycophytes are relatively few in comparison with those on seed plants. Although a basic phylogenetic framework of extant ferns is available, relationships among a few key nodes remain unresolved or poorly supported. The primary objective of this study is to explore the phylogenetic utility of large chloroplast gene data in resolving difficult deep nodes in ferns. We sequenced the chloroplast genomes from Cyrtomium devexiscapulae (eupolypod I) and Woodwardia unigemmata (eupolypod II), and constructed the phylogeny of ferns based on both 48 genes and 64 genes. The trees based on 48 genes and 64 genes are identical in topology, differing only in support values for four nodes, three of which showed higher support values for the 48-gene dataset. Equisetum was resolved as the sister to the clade of Psilotales-Ophioglossales, and the clade of Equisetales-Psilotales-Ophioglossales was sister to the clade of the leptosporangiate and marattioid ferns. The sister-relationship between the tree fern clade and polypods was supported by 82%, and 100%, bootstrap values in the 64-gene and the 48-gene trees, respectively. In polypod ferns, Pteridaceae was sister to the clade of Dennstaedtiaceae and eupolypods with a high support value, and the relationship of Dennstaedtiaceae-eupolypods was strongly supported. With recent parallel advances in the phylogenetics of ferns using nuclear data, chloroplast phylogenomics shows great potential in providing a framework for testing the impact of reticulate evolution in the early evolution of ferns.
Data from: A phylogenomic approach based on PCR target enrichment and high throughput sequencing: resolving the diversity within the South American species of Bartsia l. (Orobanchaceae)
Advances in high-throughput sequencing (HTS) have allowed researchers to obtain large amounts of biological sequence information at speeds and costs unimaginable only a decade ago. Phylogenetics, and the study of evolution in general, is quickly migrating towards using HTS to generate larger and more complex molecular datasets. In this paper, we present a method that utilizes microfluidic PCR and HTS to generate large amounts of sequence data suitable for phylogenetic analyses. The approach uses the Fluidigm Access Array System (Fluidigm, San Francisco, CA, USA) and two sets of PCR primers to simultaneously amplify 48 target regions across 48 samples, incorporating sample-specific barcodes and HTS adapters (2,304 unique amplicons per Access Array). The final product is a pooled set of amplicons ready to be sequenced, and thus, there is no need to construct separate, costly genomic libraries for each sample. Further, we present a bioinformatics pipeline to process the raw HTS reads to either generate consensus sequences (with or without ambiguities) for every locus in every sample or—more importantly—recover the separate alleles from heterozygous target regions in each sample. This is important because it adds allelic information that is well suited for coalescent-based phylogenetic analyses that are becoming very common in conservation and evolutionary biology. To test our approach and bioinformatics pipeline, we sequenced 576 samples across 96 target regions belonging to the South American clade of the genus Bartsia L. in the plant family Orobanchaceae. After sequencing cleanup and alignment, the experiment resulted in ~25,300bp across 486 samples for a set of 48 primer pairs targeting the plastome, and ~13,500bp for 363 samples for a set of primers targeting regions in the nuclear genome. Finally, we constructed a combined concatenated matrix from all 96 primer combinations, resulting in a combined aligned length of ~40,500bp for 349 samples.
Data from: How should genes and taxa be sampled for phylogenomic analyses with missing data? An empirical study in iguanian lizards
Targeted sequence capture is becoming a widespread tool for generating large phylogenomic data sets to address difficult phylogenetic problems. However, this methodology often generates data sets in which increasing the number of taxa and loci increases amounts of missing data. Thus, a fundamental (but still unresolved) question is whether sampling should be designed to maximize sampling of taxa or genes, or to minimize the inclusion of missing data cells. Here, we explore this question for an ancient, rapid radiation of lizards, the pleurodont iguanians. Pleurodonts include many well-known clades (e.g., anoles, basilisks, iguanas, and spiny lizards) but relationships among families have proven difficult to resolve strongly and consistently using traditional sequencing approaches. We generated up to 4921 ultraconserved elements with sampling strategies including 16, 29, and 44 taxa, from 1179 to approximately 2.4 million characters per matrix and approximately 30% to 60% total missing data. We then compared mean branch support for interfamilial relationships under these 15 different sampling strategies for both concatenated (maximum likelihood) and species tree (NJst) approaches (after showing that mean branch support appears to be related to accuracy). We found that both approaches had the highest support when including loci with up to 50% missing taxa (matrices with ∼40–55% missing data overall). Thus, our results show that simply excluding all missing data may be highly problematic as the primary guiding principle for the inclusion or exclusion of taxa and genes. The optimal strategy was somewhat different for each approach, a pattern that has not been shown previously. For concatenated analyses, branch support was maximized when including many taxa (44) but fewer characters (1.1 million). For species-tree analyses, branch support was maximized with minimal taxon sampling (16) but many loci (4789 of 4921). We also show that the choice of these sampling strategies can be critically important for phylogenomic analyses, since some strategies lead to demonstrably incorrect inferences (using the same method) that have strong statistical support. Our preferred estimate provides strong support for most interfamilial relationships in this important but phylogenetically challenging group.
Data from: Phylogenomic analysis supports the ancestral presence of LPS-outer membranes in the Firmicutes
One of the major unanswered questions in evolutionary biology is when and how the transition between diderm (two membranes) and monoderm (one membrane) cell envelopes occurred in Bacteria. The Negativicutes and the Halanaerobiales belong to the classically monoderm Firmicutes, but possess outer membranes with lipopolysaccharide (LPS-OM). Here, we show that they form two phylogenetically distinct lineages, each close to different monoderm relatives. In contrast, their core LPS biosynthesis enzymes were inherited vertically, as in the majority of bacterial phyla. Finally, annotation of key OM systems in the Halanaerobiales and the Negativicutes shows a puzzling combination of monoderm and diderm features. Together, these results support the hypothesis that the LPS-OMs of Negativicutes and Halanaerobiales are remnants of an ancient diderm cell envelope that was present in the ancestor of the Firmicutes, and that the monoderm phenotype in this phylum is a derived character that arose multiple times independently through OM loss.
Data from: SimPhy: phylogenomic simulation of gene, locus and species trees
We present a fast and flexible software package—SimPhy—for the simulation of multiple gene families evolving under incomplete lineage sorting, gene duplication and loss, horizontal gene transfer—all three potentially leading to species tree/gene tree discordance—and gene conversion. SimPhy implements a hierarchical phylogenetic model in which the evolution of species, locus, and gene trees is governed by global and local parameters (e.g., genome-wide, species-specific, locus-specific), that can be fixed or be sampled from a priori statistical distributions. SimPhy also incorporates comprehensive models of substitution rate variation among lineages (uncorrelated relaxed clocks) and the capability of simulating partitioned nucleotide, codon, and protein multilocus sequence alignments under a plethora of substitution models using the program INDELible. We validate SimPhy's output using theoretical expectations and other programs, and show that it scales extremely well with complex models and/or large trees, being an order of magnitude faster than the most similar program (DLCoal-Sim). In addition, we demonstrate how SimPhy can be useful to understand interactions among different evolutionary processes, conducting a simulation study to characterize the systematic overestimation of the duplication time when using standard reconciliation methods. SimPhy is available at https://github.com/adamallo/SimPhy, where users can find the source code, precompiled executables, a detailed manual and example cases.
Data from: Phylogenomic analysis of glycogen branching and debranching enzymatic duo
Branched polymers of glucose are universally used for energy storage in cells, taking the form of glycogen in animals, fungi, Bacteria, and Archaea, and of amylopectin in plants. Some enzymes involved in glycogen and amylopectin metabolism are similarly conserved in all forms of life, but some, interestingly, are not. In this paper we focus on phylogeny of glycogen branching and debranching enzymes, respectively involved in introducing and removing of the alpha(1->6) bonds in glucose polymers, bonds that provide the unique branching structure to glucose polymers. Results: We performed a large scale phylogenomic analysis of branching and debranching enzymes in over 400 completely sequenced genomes, including more than 200 from eukaryotes. We show that branching and debranching enzymes can be found in all kingdoms of life, including all major groups of eukaryotes, and thus were likely to have been present in the last universal common ancestor (LUCA) but have been lost in seemingly random fashion in numerous single celled eukaryotes. We also show how animal branching and debranching enzymes evolved from their LUCA ancestors by acquiring additional domains. Furthermore, we show that enzymes commonly perceived as orthologous, such as human branching enzyme GBE1 and E. coli branching enzyme glgB, are in fact related by a gene duplication and consequently paralogous. Conclusions: Despite its well known association with animal liver cells and plant starch, energy storage in the form of branched glucose polymers is clearly an ancient process and has been present in the earliest cells. The evolution of enzymes enabling this form of energy storage is more complex than previously thought and illustrates the need for explicit phylogenomic analysis in the study of even seemingly "simple" metabolic enzymes. Patterns of conservation and divergence in the evolution of the glycogen/starch branching and debranching enzymes have interesting biomedical connotations, as mutations in these enzymes lead to a variety of inheritable diseases in humans and other mammals.
Data from: Order-level fern plastome phylogenomics: new insights from Hymenophyllales
PREMISE OF THE STUDY: Filmy ferns (Hymenophyllales) are a highly specialized lineage, having mesophyll one cell layer thick and inhabiting particularly shaded and humid environments. The phylogenetic placement of Hymenophyllales has been inconclusive, and while over 87 whole fern plastomes have been published, none was from Hymenophyllales. To better understand the evolutionary history of filmy ferns, we sequenced the first complete plastome for this order. METHODS: We compiled a plastome phylogenomic dataset encompassing all eleven fern orders, and reconstructed phylogenies using different data types (nucleotides, codons, and amino acids) and partition schemes (codon positions and loci). To infer the evolution of fern plastome organization, we coded plastomic features, including inversions, inverted repeat boundary shifts, gene losses, and tRNA anticodon sequences as characters, and reconstructed the ancestral states for these characters. KEY RESULTS: We discovered a suite of novel, Hymenophyllales-specific plastome structures that likely resulted from repeated expansions and contractions of the inverted repeat regions. Our phylogenetic analyses reveal that Hymenophyllales is highly supported as either sister to Gleicheniales or to Gleicheniales + the remaining non-Osmundales leptosporangiates, depending on the data type and partition scheme. CONCLUSIONS: Although our analyses could not confidently resolve the phylogenetic position of Hymenophyalles, the results here highlight the danger of drawing conclusions from "all-in" phylogenomic dataset without exploring potential inconsistencies in the data. Finally, our first order-level reconstruction of fern plastome structural evolution provides a useful framework for future plastome research.
Data from: Phylogenomics resolves the timing and pattern of insect evolution
Insects are the most speciose group of animals, but the phylogenetic relationships of many major lineages remain unresolved. We inferred the phylogeny of insects from 1478 protein-coding genes. Phylogenomic analyses of nucleotide and amino acid sequences, with site-specific nucleotide or domain-specific amino acid substitution models, produced statistically robust and congruent results resolving previously controversial phylogenetic relations hips. We dated the origin of insects to the Early Ordovician [~479 million years ago (Ma)], of insect flight to the Early Devonian (~406 Ma), of major extant lineages to the Mississippian (~345 Ma), and the major diversification of holometabolous insects to the Early Cretaceous. Our phylogenomic study provides a comprehensive reliable scaffold for future comparative analyses of evolutionary innovations among insects.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.