Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,344
datasets available to search
ShareScore release 0.9.0
Dataset results
1,344 results for “: phylogenomics”
Data from: Ancestral gene flow and parallel organellar genome capture result in extreme phylogenomic discord in a lineage of angiosperms
While hybridization has recently received a resurgence of attention from systematists and evolutionary biologists, there remains a dearth of case studies on ancient, diversified hybrid lineages-clades of organisms that originated through reticulation. Studies on these groups are valuable in that they would speak to the long-term phylogenetic success of lineages following gene flow between species. We present a phylogenomic view of Heuchera, long known for frequent hybridization, incorporating all three independent genomes: targeted nuclear (~400,000 bp), plastid (~160,000 bp), and mitochondrial (~470,000 bp) data. We analyze these data using multiple concatenation and coalescence strategies. The nuclear phylogeny is consistent with previous work and with morphology, confidently suggesting a monophyletic Heuchera. By contrast, analyses of both organellar genomes recover a grossly polyphyletic Heuchera,consisting of three primary clades with relationships extensively rearranged within these as well. A minority of nuclear loci also exhibit phylogenetic discord; yet these topologies remarkably never resemble the pattern of organellar loci and largely present low levels of discord inter alia. Two independent estimates of the coalescent branch length of the ancestor of Heuchera using nuclear data suggest rare or nonexistent incomplete lineage sorting with related clades, inconsistent with the observed gross polyphyly of organellar genomes (confirmed by simulation of gene trees under the coalescent). These observations, in combination with previous work, strongly suggest hybridization as the cause of this phylogenetic discord.
Data from: Phylogenomics from whole genome sequences using aTRAM
Novel sequencing technologies are rapidly expanding the size of data sets that can be applied to phylogenetic studies. Currently the most commonly used phylogenomic approaches involve some form of genome reduction. While these approaches make assembling phylogenomic data sets more economical for organisms with large genomes, they reduce the genomic coverage and thereby the long-term utility of the data. Currently, for organisms with moderate to small genomes (<1000 Mbp) it is feasible to sequence the entire genome at modest coverage (10−30×). Computational challenges for handling these large data sets can be alleviated by assembling targeted reads, rather than assembling the entire genome, to produce a phylogenomic data matrix. Here we demonstrate the use of automated Target Restricted Assembly Method (aTRAM) to assemble 1107 single-copy ortholog genes from whole genome sequencing of sucking lice (Anoplura) and out-groups. We developed a pipeline to extract exon sequences from the aTRAM assemblies by annotating them with respect to the original target protein. We aligned these protein sequences with the inferred amino acids and then performed phylogenetic analyses on both the concatenated matrix of genes and on each gene separately in a coalescent analysis. Finally, we tested the limits of successful assembly in aTRAM by assembling 100 genes from close- to distantly related taxa at high to low levels of coverage.
Data from: Methodological congruence in phylogenomic analyses with morphological support for teiid lizards (Sauria: Teiidae)
A well-known issue in phylogenetics is discordance among gene trees, species trees, morphology, and other data types. Gene-tree discordance is often caused by incomplete lineage sorting, lateral gene transfer, and gene duplication. Multispecies-coalescent methods can account for incomplete lineage sorting and are believed by many to be more accurate than concatenation. However, simulation studies and empirical data have demonstrated that concatenation and species tree methods often recover similar topologies. We use three popular methods of phylogenetic reconstruction (one concatenation, two species tree) to evaluate relationships within Teiidae. These lizards are distributed across the United States to Argentina and the West Indies, and their classification has been controversial due to incomplete sampling and the discordance among various character types (chromosomes, DNA, musculature, osteology, etc.) used to reconstruct phylogenetic relationships. Recent morphological and molecular analyses of the group resurrected three genera and created five new genera to resolve non-monophyly in three historically ill-defined genera: Ameiva, Cnemidophorus, and Tupinambis. Here, we assess the phylogenetic relationships of the Teiidae using "next-generation" anchored-phylogenomics sequencing. Our final alignment includes 316 loci (488,656 bp DNA) for 244 individuals (56 species of teiids, representing all currently recognized genera) and all three methods (ExaML, MP-EST, and ASTRAL-II) recovered essentially identical topologies. Our results are basically in agreement with recent results from morphology and smaller molecular datasets, showing support for monophyly of the eight new genera. Interestingly, even with hundreds of loci, the relationships among some genera in Tupinambinae remain ambiguous (i.e. low nodal support for the position of Salvator and Dracaena).
Data from: Phylogenomic analyses support traditional relationships within Cnidaria
Cnidaria, the sister group to Bilateria, is the most diverse group of animals in terms of morphology, lifecycles, ecology, and development. How this diversity originated and evolved is not well understood because phylogenetic relationships among major cnidarian lineages are unclear, and recent studies present contrasting phylogenetic hypotheses. Here, we use transcriptome data from 15 newly-sequenced species in combination with 26 publicly available genomes and transcriptomes to assess phylogenetic relationships among major cnidarian lineages. Phylogenetic analyses using different partition schemes and models of molecular evolution, as well as topology tests for alternative phylogenetic relationships, support the monophyly of Medusozoa, Anthozoa, Octocorallia, Hydrozoa, and a clade consisting of Staurozoa, Cubozoa, and Scyphozoa. Support for the monophyly of Hexacorallia is weak due to the equivocal position of Ceriantharia. Taken together, these results further resolve deep cnidarian relationships, largely support traditional phylogenetic views on relationships, and provide a historical framework for studying the evolutionary processes involved in one of the most ancient animal radiations.
Data from: Phylogenomic analyses of Echinodermata support the sister groups of Asterozoa and Echinozoa
Echinoderms (sea urchins, sea stars, brittle stars, sea lilies and sea cucumbers) are a group of diverse organisms, second in number within deuterostome species to only the chordates. Echinoderms serve as excellent model systems for developmental biology due to their diverse developmental mechanisms, tractable laboratory use, and close phylogenetic distance to chordates. In addition, echinoderms are very well represented in the fossil record, including some larval features, making echinoderms a valuable system for studying evolutionary development. The internal relationships of Echinodermata have not been consistently supported across phylogenetic analyses, however, and this has hindered the study of other aspects of their biology. In order to test echinoderm phylogenetic relationships, we sequenced 23 de novo transcriptomes from all five clades of echinoderms. Using multiple phylogenetic methods at a variety of sampling depths we have constructed a well-supported phylogenetic tree of Echinodermata, including support for the sister groups of Asterozoa (sea stars and brittle stars) and Echinozoa (sea urchins and sea cucumbers). These results will help inform developmental and evolutionary studies specifically in echinoderms and deuterostomes in general.
Data from: Integrating phylogenomic and population genomic patterns in avian lice provides a more complete picture of parasite evolution
Parasite diversity accounts for most of the biodiversity on earth, and is shaped by many processes (e.g. cospeciation, host-switching). To identify the effects of the processes that shape parasite diversity, it is ideal to incorporate both deep (phylogenetic) and shallow (population) perspectives. To this end, we developed a novel workflow to obtain phylogenetic and population genetic data from whole genome sequences of body lice parasitizing New World ground-doves. Phylogenies from these data showed consistent, highly resolved species-level relationships for the lice. By comparing the louse and ground-dove phylogenies, we found that over long-term evolutionary scales their phylogenies were largely congruent. Many louse lineages (both species and populations) also demonstrated high host-specificity, suggesting ground-dove divergence is a primary driver of their parasites' diversity. However, the few louse taxa that are generalists are structured according to biogeography at the population level. This suggests dispersal among sympatric hosts has some effect on body louse diversity, but over deeper time scales the parasites eventually sort according to host species. Overall, our results demonstrate that multiple factors explain the patterns of diversity in this group of parasites, and that the effects of these factors can vary over different evolutionary scales. The integrative approach we employed was crucial for uncovering these patterns, and should be broadly applicable to other studies.
Data from: Phylogenomics provides new insight into evolutionary relationships and genealogical discordance in the reef-building coral genus Acropora
Understanding the genetic basis of reproductive isolation is a long-standing goal of speciation research. In recently diverged populations, genealogical discordance may reveal genes and genomic regions that contribute to the speciation process. Previous work has shown that conspecific colonies of Acropora that spawn in different seasons (spring and autumn) are associated with highly diverged lineages of the phylogenetic marker PaxC. Here, we used 10 034 single-nucleotide polymorphisms to generate a genome-wide phylogeny and compared it with gene genealogies from the PaxC intron and the mtDNA Control Region in 20 species of Acropora, including three species with spring- and autumn-spawning cohorts. The PaxC phylogeny separated conspecific autumn and spring spawners into different genetic clusters in all three species; however, this pattern was not supported in two of the three species at the genome level, suggesting a selective connection between PaxC and reproductive timing in Acropora corals. This genome-wide phylogeny provides an improved foundation for resolving phylogenetic relationships in Acropora and, combined with PaxC, provides a fascinating platform for future research into regions of the genome that influence reproductive isolation and speciation in corals.
Data from: Phylogenomic analyses of Sabal (Arecaceae) species relationships using targeted sequence capture
With the increasing availability of high-throughput sequencing, phylogenetic analyses are no longer constrained by the limited availability of a few loci. Here, we describe a sequence capture methodology, which we used to collect data for analyses of diversification within Sabal (Arecaceae), a palm genus native to the south-eastern USA, Caribbean, Bermuda and Central America. RNA probes were developed and used to enrich DNA samples for putatively low copy nuclear genes and the plastomes for all Sabal species and two outgroup species. Sequence data were generated on an Illumina MiSeq sequencer and target sequences were assembled using custom workflows. Both coalescence and supermatrix analyses of 133 nuclear genes were used to estimate species trees relationships. Plastid genomes were also analysed, yielding generally poor resolution with regard to species relationships. Species relationships described in both nuclear gene and plastome sequences largely reflect the biogeography of the group and, to a lesser extent, previous morphology-based hypotheses. Beyond the biological implications, this research validates a high-throughput methodology for generating a large number of genes for coalescence-based phylogenetic analyses in plant lineages.
Data from: Phylogenomics and the evolution of hemipteroid insects
Hemipteroid insects (Paraneoptera), with over 10% of all known insect diversity, are a major component of terrestrial and aquatic ecosystems. Previous phylogenetic analyses have not consistently resolved the relationships among major hemipteroid lineages. We provide maximum likelihood-based phylogenomic analyses of a taxonomically comprehensive dataset comprising sequences of 2,395 single-copy, protein-coding genes for 193 samples of hemipteroid insects and outgroups. These analyses yield a well-supported phylogeny for hemipteroid insects. Monophyly of each of the three hemipteroid orders (Psocodea, Thysanoptera, and Hemiptera) is strongly supported, as are most relationships among suborders and families. Thysanoptera (thrips) is strongly supported as sister to Hemiptera. However, as in a recent large-scale analysis sampling all insect orders, trees from our data matrices support Psocodea (bark lice and parasitic lice) as the sister group to the holometabolous insects (those with complete metamorphosis). In contrast, four-cluster likelihood mapping of these data does not support this result. A molecular dating analysis using 23 fossil calibration points suggests hemipteroid insects began diversifying before the Carboniferous, over 365 million years ago. We also explore implications for understanding the timing of diversification, the evolution of morphological traits, and the evolution of mitochondrial genome organization. These results provide a phylogenetic framework for future studies of the group.
Data from: Phylogenomics uncovers confidence and conflict in the rapid radiation of Australo-Papuan rodents
The estimation of robust and accurate measures of branch support has proven challenging in the era of phylogenomics. In datasets of potentially millions of sites, bootstrap support for bifurcating relationships around very short internal branches can be inappropriately inflated. Such over-estimation of branch support may be particularly problematic in rapid radiations, where phylogenetic signal is low and incomplete lineage sorting severe. Here, we explore this issue by comparing various branch support estimates under both concatenated and coalescent frameworks, in the recent radiation Australo-Papuan murine rodents (Muridae: Hydromyini). Using nucleotide sequence data from 1245 independent loci and several phylogenomic inference methods, we unequivocally resolve the majority of genus-level relationships within Hydromyini. However, at four nodes we recover inconsistency in branch support estimates both within and among concatenated and coalescent approaches. In most cases, concatenated likelihood approaches using standard fast bootstrap algorithms did not detect any uncertainty at these four nodes, regardless of partitioning strategy. However, we found this could be overcome with two-stage resampling, i.e. across genes and sites within genes (using -bsam GENESITE in IQtree). In addition, low confidence at recalcitrant nodes was recovered using UFBoot2, a recent revision to the bootstrap protocol in IQtree, but this depended on partitioning strategy. Summary coalescent approaches also failed to detect uncertainty under some circumstances. For each of four recalcitrant nodes, an equivalent (or close to equivalent) number of genes were in strong support (> 75% bootstrap) of both the primary and at least one alternative topological hypothesis, suggesting notable phylogenetic conflict among loci not detected using some standard branch support metrics. Recent debate has focused on the appropriateness of concatenated versus multi-genealogical approaches to resolving species relationships, but less so on accurately estimating uncertainty in large datasets. Our results demonstrate the importance of employing multiple approaches when assessing confidence, and highlight the need for greater attention to the development of robust measures of uncertainty in the era of phylogenomics.
Data from: Analysis of phylogenomic tree space resolves relationships among marsupial families
A fundamental challenge in resolving evolutionary relationships across the Tree of Life is to account for heterogeneity in the evolutionary signal across loci. Studies of marsupial mammals have demonstrated that this heterogeneity can be substantial, leaving considerable uncertainty in the evolutionary timescale and relationships within the group. Using simulations and a new phylogenomic data set comprising nucleotide sequences of 1550 loci from 18 of the 22 extant marsupial families, we demonstrate the power of a method for identifying clusters of loci that support different phylogenetic trees. We find two distinct clusters of loci, each providing an estimate of the species tree that matches previously proposed resolutions of the marsupial phylogeny. We also identify a well supported placement for the enigmatic marsupial moles (Notoryctes) that contradicts previous molecular estimates but is consistent with morphological evidence. The pattern of gene-tree variation across tree-space is characterized by changes in information content, GC content, substitution-model adequacy, and signatures of purifying selection in the data. In a simulation study, we show that incomplete lineage sorting can explain the division of loci into the two tree-topology clusters, as found in our phylogenomic analysis of marsupials. We also demonstrate the potential benefits of minimizing uncertainty from phylogenetic conflict for molecular dating. Our analyses reveal that Australasian marsupials appeared in the early Paleocene, whereas the diversification of present-day families occurred primarily during the late Eocene and early Oligocene. Our methods provide an intuitive framework for improving the accuracy and precision of phylogenetic inference and molecular dating using genome-scale data.
Data from: Phylogenomics of an extra-Antarctic notothenioid radiation reveals a previously unrecognized lineage and diffuse species boundaries
Background: The impressive adaptive radiation of notothenioid fishes in Antarctic waters is generally thought to have been facilitated by an evolutionary key innovation, antifreeze glycoproteins, permitting the rapid evolution of more than 120 species subsequent to the Antarctic glaciation. By way of contrast, the second-most species-rich notothenioid genus, Patagonotothen, which is nested within the Antarctic clade of Notothenioidei, is almost exclusively found in the non-Antarctic waters of Patagonia. While the drivers of the diversification of Patagonotothen are currently unknown, they are unlikely to be related to antifreeze glycoproteins, given that water temperatures in Patagonia are well above freezing point. Here we performed a phylogenetic analysis based on genome-wide single nucleotide polymorphisms (SNPs) derived from restriction site-associated DNA sequencing (RADseq) in a total of twelve Patagonotothen species. Results: We present a well-supported, time-calibrated phylogenetic hypothesis including closely and distantly related outgroups, confirming the monophyly of the genus Patagonotothen with an origin approximately 3 million year ago and the paraphyly of both the sister genus Lepidonotothen and the family Notothenidae. Our phylogenomic and population genetic analyses highlight a previously unrecognized linage and provide evidence for shared genetic variation between some closely related species. We also provide a mitochondrial phylogeny showing mitonuclear discordance. Conclusions: Based on a combination of phylogenomic and population genomic approaches, we provide evidence for the existence of a new, potentially cryptic, Patagonotothen species, and demonstrate that genetic boundaries between some closely related species are diffuse, likely due to recent introgression and/or incomplete linage sorting. The detected mitonuclear discordance highlights the limitations of relying on a single locus for species barcoding. In addition, our time calibrated phylogenetic hypothesis shows that the early burst of diversification roughly coincides with the onset of the intensification of Quaternary glacial cycles and that the rate of species accumulation may have been stepwise rather than constant. Our phylogenetic framework not only advances our understanding of the origin of a high-latitude marine radiation, but also provides the basis for the study of the ecology and life history of the genus Patagonotothen, as well as for their conservation and commercial management.
Data from: Phylogenomic analyses reveal extensive gene flow within the magic flowers (Achimenes)
Premise of the study: The Neotropical Gesneriaceae is a lineage known for its colorful and diverse flowers, as well as an extensive history of intra- and intergeneric hybridization, particularly among Achimenes (the magic flowers) and other members of the subtribe Gloxiniinae. Despite numerous studies seeking to elucidate the evolutionary relationships of these lineages, relatively few have sought to infer specific patterns of gene flow despite evidence of widespread hybridization. Methods: To explore the utility of phylogenomic data for reassessing phylogenetic relationships and inferring patterns of gene flow among species of Achimenes, we sequenced 12 transcriptomes. We used a variety of methods to infer the species tree, examine gene tree discordance, and infer patterns of gene flow. Key results: Phylogenomic analyses resolve clade relationships at the crown of the lineage with high support. In contrast to previous analyses, we recovered strong support for several new relationships despite a significant amount of gene tree discordance. We present evidence for at least two introgression events between two species pairs that share pollinators and suggest that the species status of A. admirabilis be reexamined. Conclusions: Our study demonstrates the utility of transcriptome data for phylogenomic analyses and inferring patterns of gene flow despite gene tree discordance. Moreover, these data provide another example of prevalent interspecific gene flow among Neotropical plants that share pollinators.
Data from: Phylogenomic analyses of deep gastropod relationships reject Orthogastropoda
Gastropods are a highly diverse clade of molluscs that includes many familiar animals, such as limpets, snails, slugs and sea slugs. It is one of the most abundant groups of animals in the sea and the only molluscan lineage that has successfully colonized land. Yet the relationships among and within its constituent clades have remained in flux for over a century of morphological, anatomical and molecular study. Here, we re-evaluate gastropod phylogenetic relationships by collecting new transcriptome data for 40 species and analysing them in combination with publicly available genomes and transcriptomes. Our datasets include all five main gastropod clades: Patellogastropoda, Vetigastropoda, Neritimorpha, Caenogastropoda and Heterobranchia. We use two different methods to assign orthology, subsample each of these matrices into three increasingly dense subsets, and analyse all six of these supermatrices with two different models of molecular evolution. All 12 analyses yield the same unrooted network connecting the five major gastropod lineages. This reduces deep gastropod phylogeny to three alternative rooting hypotheses. These results reject the prevalent hypothesis of gastropod phylogeny, Orthogastropoda. Our dated tree is congruent with a possible end-Permian recovery of some gastropod clades, namely Caenogastropoda and some Heterobranchia subclades.
Data from: Information dropout patterns in restriction site associated DNA phylogenomics and a comparison with multilocus Sanger data in a species-rich moth genus
A rapid shift from traditional Sanger sequencing-based molecular methods to the phylogenomic approach with large numbers of loci is underway. Among phylogenomic methods, RAD (Restriction site Associated DNA) sequencing approaches have gained much attention as they enable rapid generation of up to thousands of loci randomly scattered across the genome and are suitable for non-model species. RAD data sets however suffer from large amounts of missing data and rapid locus dropout along with decreasing relatedness among taxa. The relationship between locus dropout and the amount of phylogenetic information retained in the data has remained largely un-investigated. Similarly, phylogenetic hypotheses based on RAD have rarely been compared with phylogenetic hypotheses based on multilocus Sanger sequencing, even less so using exactly the same species and specimens. We compared the Sanger-based phylogenetic hypothesis (8 loci; 6,172 bp) of 32 species of the diverse moth genus Eupithecia (Lepidoptera, Geometridae) to that based on double-digest RAD sequencing (3,256 loci; 726,658 bp). We observed that topologies were largely congruent, with some notable exceptions that we discuss. The locus dropout effect was strong. We demonstrate that number of loci is not a precise measure of phylogenetic information since the number of single-nucleotide polymorphisms (SNPs) may remain low at very shallow phylogenetic levels despite large numbers of loci. As we hypothesize, the number of SNPs and parsimony informative SNPs (PIS) is low at shallow phylogenetic levels, peaks at intermediate levels and, thereafter, declines again at the deepest levels as a result of decay of available loci. Similarly, we demonstrate with empirical data that the locus dropout affects the type of loci retained, the loci found in many species tending to show lower interspecific distances than those shared among fewer species. We also examine the effects of the numbers of loci, SNPs and PIS on nodal bootstrap support, but could not demonstrate with our data our expectation of a positive correlation between them. We conclude that RAD methods provide a powerful tool for phylogenomics at an intermediate phylogenetic level as indicated by its broad congruence with an eight-gene Sanger data set in a genus of moths. When assessing the quality of the data for phylogenetic inference, the focus should be on the distribution and number of SNPs and PIS rather than on loci.
Data from: Phylogenomic analyses confirm a novel invasive North American Corbicula (Bivalvia: Cyrenidae) lineage
The genus Corbicula consists of estuarine or freshwater clams native to temperate/tropical regions of Asia, Africa, and Australia that collectively encompass both sexual species and clonal (androgenetic) lineages. The latter have become globally invasive in freshwater systems and they represent some of the most successful aquatic invasive lineages. Previous studies have documented four invasive clonal lineages, Forms A, B, C, and Rlc, with varying known distributions. Form A (R in Europe) occurs globally, Form B is found solely in North America, mainly the western United States, Form C (S in Europe) occurs both in European watersheds and in South America, and Rlc is known from Europe. A putative fifth invasive morph, Form D, was recently described in the New World from the Illinois River (Great Lakes watershed), where it occurs in sympatry with Forms A and B. An initial study showed Form D to be conchologically distinct: possessing rust-colored rays and white nacre with purple teeth. However, its genetic distinctiveness using standard molecular markers (mitochondrial cytochrome c oxidase subunit I and nuclear ribosomal 28S RNA) was ambiguous. To resolve this issue, we performed a phylogenomic analysis using 1,699-30,027 nuclear genomic loci collected via the next generation double digested restriction-site associated DNA sequencing method. Our results confirmed Form D to be a distinct invasive New World lineage with a population genomic profile consistent with clonality. A majority (7/9) of the phylogenomic analyses recovered the four New World invasive Corbicula lineages (Forms A, B, C, and D) as members of a clonal clade, sister to the non-clonal Lake Biwa (Japan) endemic, C. sandai. The age of the clonal clade was estimated at 1.49 million years (my; ± 0.401– 2.955 my) whereas the estimated ages of the four invasive lineage crown clades ranged from 0.27-0.44 my. We recovered no evidence of nuclear genomic admixture among the four invasive lineages in our study populations. In contrast, 2/6 C. sandai individuals displayed partial nuclear genomic Structure assignments with multiple invasive clonal lineages. These results provide new insights into the origin and maintenance of clonality in this complex system.
Data from: Phylogenomic incongruence, hypothesis testing, and taxonomic sampling: the monophyly of characiform fishes
Phylogenomic studies using genome‐wide datasets are quickly becoming the state of the art for systematics and comparative studies, but in many cases, they result in strongly supported incongruent results. The extent to which this conflict is real depends on different sources of error potentially affecting big datasets (assembly, stochastic, and systematic error). Here, we apply a recently developed methodology (GGI or gene genealogy interrogation) and data curation to new and published datasets with more than 1000 exons, 500 ultraconserved element (UCE) loci, and transcriptomic sequences that support incongruent hypotheses. The contentious non‐monophyly of the order Characiformes proposed by two studies is shown to be a spurious outcome induced by sample contamination in the transcriptomic dataset and an ambiguous result due to poor taxonomic sampling in the UCE dataset. By exploring the effects of number of taxa and loci used for analysis, we show that the power of GGI to discriminate among competing hypotheses is diminished by limited taxonomic sampling, but not equally sensitive to gene sampling. Taken together, our results reinforce the notion that merely increasing the number of genetic loci for a few representative taxa is not a robust strategy to advance phylogenetic knowledge of recalcitrant groups. We leverage the expanded exon capture dataset generated here for Characiformes (206 species in 23 out of 24 families) to produce a comprehensive phylogeny and a revised classification of the order.
Data from: Phylogenomic analysis of the Chilean clade of Liolaemus lizards (Squamata: Liolaemidae) based on sequence capture data
The genus Liolaemus is one of the most ecologically diverse and species-rich genera of lizards worldwide. It currently includes more than 250 recognized species, which have been subject to many ecological and evolutionary studies. Nevertheless, Liolaemus lizards have a complex taxonomic history, mainly due to the incongruence between morphological and genetic data, incomplete taxon sampling, incomplete lineage sorting and hybridization. In addition, as many species have restricted and remote distributions, this has hampered their examination and inclusion in molecular systematic studies. The aims of this study are to infer a robust phylogeny for a subsample of lizards representing the Chilean clade (subgenus Liolaemus sensu stricto), and to test the monophyly of several of the major species groups. We use a phylogenomic approach, targeting 541 ultra-conserved elements (UCEs) and 44 protein-coding genes for 16 taxa. We conduct a comparison of phylogenetic analyses using maximum-likelihood and several species tree inference methods. The UCEs provide stronger support for phylogenetic relationships compared to the protein-coding genes; however, the UCEs outnumber the protein-coding genes by 10-fold. On average, the protein-coding genes contain over twice the number of informative sites. Based on our phylogenomic analyses, all the groups sampled are polyphyletic. Liolaemus tenuis tenuis is difficult to place in the phylogeny, because only a few loci (nine) were recovered for this species. Topologies or support values did not change dramatically upon exclusion of L. t. tenuis from analyses, suggesting that missing data did not had a significant impact on phylogenetic inference in this data set. The phylogenomic analyses provide strong support for sister group relationships between L. fuscus, L. monticola, L. nigroviridis and L. nitidus, and L. platei and L. velosoi. Despite our limited taxon sampling, we have provided a reliable starting hypothesis for the relationships among many major groups of the Chilean clade of Liolaemus that will help future work aimed at resolving the Liolaemus phylogeny.
Data from: Integrating phylogenomic and morphological data to assess candidate species-delimitation models in brown and red-bellied snakes (Storeria)
Systematics at the species level is still marked by theoretical and empirical tensions amongst the desires to identify geographical lineages, delimit species, and estimate their relationships. These goals are often confounded because each relies, at least to some extent, on the others being known. However, recently developed methods can simultaneously address all three. Furthermore, next-generation genomic sequencing allows us to generate large-scale molecular data sets to examine variation within species at a fine scale. Finally, a renaissance in morphological species validation allows us to integrate historical species definitions with coalescent models for species delimitation. Here, we investigate the applicability of these methods in an empirical case, in the Nearctic snake genus Storeria. Integrating trait data into species delimitation reduces the number of species delimited from molecular data alone. Whereas molecular data support eight distinct species-level lineages, including morphological data reduces this to four. The taxa Storeria dekayi, Storeria occipitomaculata, Storeria storerioides, and Storeria victa are considered distinct, monotypic species, with no subspecies recognized. We highlight the need for careful assessment of species delimitation, combining both computational genetic methods as well as traditional character-based descriptions. It is now possible to identify phylogeographical lineages, delimit species using molecular and morphological data, and estimate their relationships in a single coherent set of analyses. Moving forward, this will allow for more rapid and objective assessments of cryptic diversity at the species level.
Data from: Resolving rapid radiations within angiosperm families using anchored phylogenomics
Despite the promise that molecular data would provide a seemingly unlimited source of independent characters, many plant phylogenetic studies are still based on only two regions, the plastid genome and nuclear ribosomal DNA (nrDNA). Their popularity can be explained by high copy numbers and universal PCR primers that make their sequences easily amplified and converted into parallel datasets. Unfortunately, their utility is limited by linked loci and limited characters resulting in low confidence in the accuracy of phylogenetic estimates, especially when rapid radiations occur. In another contribution on anchored phylogenomics in angiosperms, we presented flowering plant-specific anchored enrichment probes for hundreds of conserved nuclear genes and demonstrated their use at the level of all angiosperms. In this contribution, we focus on a common problem in phylogenetic reconstructions below the family level: weak or unresolved backbone due to rapid radiations (≤10 million years) followed by long divergence, using the Cariceae-Dulichieae-Scirpeae clade (CDS, Cyperaceae) as a test case. By comparing our nuclear matrix of 461 genes to a typical Sanger-sequence dataset consisting of a few plastid genes (matK, ndhF) and an nrDNA marker (ETS), we demonstrate that our nuclear data is fully compatible with the Sanger dataset and resolves short backbone internodes with high support in both concatenated and coalescence-based analyses. In addition, we show that nuclear gene tree incongruence is inversely proportional to phylogenetic information content, indicating that incongruence is mostly due to gene tree estimation error. This suggests that large numbers of conserved nuclear loci could produce more accurate trees than sampling rapidly evolving regions prone to saturation and long-branch attraction. The robust phylogenetic estimates obtained here, and high congruence with previous morphological and molecular analyses, are strong evidence for a complete tribal revision of CDS. The anchored hybrid enrichment probes used in this study should be similarly effective in other flowering plant groups.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.