Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,344

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,344 results for “: phylogenomics”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Modeling site heterogeneity with posterior mean site frequency profiles accelerates accurate phylogenomic estimation

Proteins have distinct structural and functional constraints at different sites that lead to site-specific preferences for particular amino acid residues as the sequences evolve. Heterogeneity in the amino acid substitution process between sites is not modeled by commonly used empirical amino acid exchange matrices. Such model misspecification can lead to artefacts in phylogenetic estimation such as long-branch attraction. Although sophisticated site-heterogeneous mixture models have been developed to address this problem in both Bayesian and maximum likelihood (ML) frameworks, their formidable computational time and memory usage severely limits their use in large phylogenomic analyses. Here we propose a posterior mean site frequency (PMSF) method as a rapid and efficient approximation to full empirical profile mixture models for ML analysis. The PMSF approach assigns a conditional mean amino acid frequency profile to each site calculated based on a mixture model fitted to the data using a preliminary guide tree. These PMSF profiles can then be used for in-depth tree-searching in place of the full mixture model. Compared with widely used empirical mixture models with k classes, our implementation of PMSF in IQ-TREE (http://www.iqtree.org) speeds up the computation by approximately k /1.5-fold and requires a small fraction of the RAM. Furthermore, this speedup allows, for the first time, full nonparametric bootstrap analyses to be conducted under complex site-heterogeneous models on large concatenated data matrices. Our simulations and empirical data analyses demonstrate that PMSF can effectively ameliorate long-branch attraction artefacts. In some empirical and simulation settings PMSF provided more accurate estimates of phylogenies than the mixture models from which they derive.

opencc-zeroDec 2016View details →
dryad28/100

Ordered phylogenomic subsampling enables diagnosis of systematic errors in the placement of the enigmatic arachnid order Palpigradi

<p><span><span><span><span><span><span><span><span><span><span><span>The miniaturized arachnid order Palpigradi has ambiguous phylogenetic affinities, due to its odd combination of plesiomorphic and derived morphological traits. This lineage has never been sampled in phylogenomic datasets because of its small body size and fragility of most species, a sampling gap of immediate concern to recent disputes over arachnid monophyly. To redress this gap, we sampled a population of the cave-inhabiting species <i>Eukoenenia spelaea</i> from Slovakia and inferred its placement in the phylogeny of Chelicerata using dense phylogenomic matrices of up to 1450 loci, drawn from high-quality transcriptomic libraries and complete genomes. The complete matrix included exemplars of all extant orders of Chelicerata. Analyses of the complete matrix recovered palpigrades as the sister group of the long-branch order Parasitiformes (ticks) with high support. However, sequential deletion of long-branch taxa revealed that the position of palpigrades is prone to topological instability. Phylogenomic subsampling approaches that maximized taxon or dataset completeness recovered palpigrades as the sister group of camel spiders (Solifugae), with modest support. While this relationship is congruent with the location and architecture of the coxal glands, a long-forgotten character system that opens in the pedipalpal segments only in palpigrades and solifuges, we show that nodal support values in concatenated supermatrices can mask high levels of underlying topological conflict in the placement of the enigmatic Palpigradi. </span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroNov 2019View details →
dryad28/100

Data from: Phylogenomic analysis of transcriptome data elucidates co-occurrence of a paleopolyploid event and the origin of bimodal karyotypes in Agavoideae (Asparagaceae)

PREMISE OF THE STUDY: The stability of the bimodal karyotype found in Agave and closely related species has long interested botanists. The origin of the bimodal karyotype has been attributed to allopolyploidy, but this hypothesis has not been tested. Next Generation transcriptome sequence data were used to test whether a paleopolyploid event occurred on the same branch of the Agavoideae phylogenetic tree as the origin of the Yucca-Agave bimodal karyotype. METHODS: Illumina RNAseq data were generated for phylogenetically strategic species in Agavoideae. Paleopolyploidy was inferred in analyses of frequency plots for synonymous substitutions per synonymous site (Ks) between Hosta, Agave and Chlorophytum paralogous and orthologous gene pairs. Phylogenies of gene families including paralogous genes for these species and outgroup species were estimated in order to place inferred paleopolyploid events on a species tree. KEY RESULTS: Ks frequency plots suggested paleopolyploid events in the history of the genera Agave, Hosta and Chlorophytum. Phylogenetic analyses of gene families estimated from transcriptome data revealed two polyploid events: one predating the last common ancestor of Agave and Hosta and one within the lineage leading to Chlorophytum. CONCLUSIONS: We found that allopolyoidy and the origin of the Yucca-Agave bimodal karyotype co-occur on the same lineage consistent with the hypothesis that the bimodal karyotype is a consequence of allopolyploidy. We discuss this and alternative mechanisms for the formation of the Yucca-Agave bimodal karyotype. More generally, we illustrate how the use of next generation sequencing technology is a cost-efficient means for assessing genome evolution in non-model species.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Microevolutionary processes generate phylogenomic discordance at ancient divergences

Stochastic population processes may cause differences between species histories and gene histories. These processes are assumed to only influence the most recent divergences in the tree of life; however, there may be under-appreciated potential for microevolutionary processes to impact deep divergences. I used multispecies coalescent models to determine the impact of stochastic processes on deep phylogenomic histories. Here I show phylogenomic discordance between gene histories and species histories is expected at deep divergences for many eukaryotic taxa, and the probability of discordance increases with population size, generation time, and the number of species in the tree. Five eukaryotic clades (angiosperms, birds, harpaline beetles, mammals, and nymphalid butterflies) demonstrate significant discordance potential at divergences over 50 million years old, and this discordance potential is independent of the age of divergence. These findings demonstrate population processes acting over very short timescales will leave a lasting impact on genomic histories, even for divergence events occurring tens to hundreds of millions of years ago.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Comparison of target-capture and restriction-site associated DNA sequencing for phylogenomics: a test in cardinalid tanagers (Aves, genus: Piranga)

Restriction-site associated DNA sequencing (RAD-seq) and target capture of specific genomic regions, such as ultraconserved elements (UCEs), are emerging as two of the most popular methods for phylogenomics using reduced-representation genomic datasets. These two methods were designed to target different evolutionary timescales: RAD-seq was designed for population-genomic level questions and UCEs for deeper phylogenetics. The utility of both datasets to infer phylogenies across a variety of taxonomic levels has not been adequately compared within the same taxonomic system. Additionally, the effects of uninformative gene trees on species tree analyses (for target capture data) have not been explored. Here, we utilize RAD-seq and UCE data to infer a phylogeny of the bird genus Piranga. The group has a range of divergence dates (0.5 my – 6 my), contains eleven recognized species, and lacks a resolved phylogeny. We compared two species tree methods for the RAD-seq data and six species tree methods for the UCE data. Additionally, in the UCE data, we analyzed a complete matrix as well as datasets with only highly informative loci. A complete matrix of 189 UCE loci with ten or more parsimony informative (PI) sites, and an ~80% complete matrix of 1128 PI SNPs (from RAD-seq) yield the same fully resolved phylogeny of Piranga. We inferred non-monophyletic relationships of P. lutea individuals, with all other a priori species identified as monophyletic. Finally, we found that species tree analyses that included predominantly uninformative gene trees provided strong support for different topologies, with consistent phylogenetic results when limiting species tree analyses to highly informative loci or only using less informative loci with concatenation or methods meant for SNPs alone.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Phylogenomics and analysis of shared genes suggest a single transition to mutualism in Wolbachia of nematodes

Wolbachia, endosymbiotic bacteria of the order Rickettsiales, are widespread in arthropods but also present in nematodes. In arthropods, A and B supergroup Wolbachia are generally associated with distortion of host reproduction. In filarial nematodes, including some human parasites, multiple lines of experimental evidence indicate that C and D supergroup Wolbachia are essential for the survival of the host, and here the symbiotic relationship is considered mutualistic. The origin of this mutualistic endosymbiosis is of interest for both basic and applied reasons: How does a parasite become a mutualist? Could intervention in the mutualism aid in treatment of human disease? Correct rooting and high-quality resolution of Wolbachia relationships are required to resolve this question. However, because of the large genetic distance between Wolbachia and the nearest outgroups, and the limited number of genomes so far available for large-scale analyses, current phylogenies do not provide robust answers. We therefore sequenced the genome of the D supergroup Wolbachia endosymbiont of Litomosoides sigmodontis, revisited the selection of loci for phylogenomic analyses, and performed a phylogenomic analysis including available complete genomes (from isolates in supergroups A, B, C and D). Using ninety orthologous genes with reliable phylogenetic signals, we obtained a robust phylogenetic reconstruction, including a highly supported root to the Wolbachia phylogeny between an (A+B) clade and a (C+D) clade. While we currently lack data from several Wolbachia supergroups, notably F, our analysis supports a model wherein the putatively mutualist endosymbiotic relationship between Wolbachia and nematodes originated from a single transition event.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Bayes factors unmask highly variable information content, bias, and extreme influence in phylogenomic analyses

As the application of genomic data in phylogenetics has become routine, a number of cases have arisen where alternative data sets strongly support conflicting conclusions. This sensitivity to analytical decisions has prevented firm resolution of some of the most recalcitrant nodes in the tree of life. To better understand the causes and nature of this sensitivity, we analyzed several phylogenomic data sets using an alternative measure of topological support (the Bayes factor) that both demonstrates and averts several limitations of more frequently employed support measures (such as Markov chain Monte Carlo estimates of posterior probabilities). Bayes factors reveal important, previously hidden, differences across six "phylogenomic" data sets collected to resolve the phylogenetic placement of turtles within Amniota. These data sets vary substantially in their support for well-established amniote relationships, particularly in the proportion of genes that contain extreme amounts of information as well as the proportion that strongly reject these uncontroversial relationships. All six data sets contain little information to resolve the phylogenetic placement of turtles relative to other amniotes. Bayes factors also reveal that a very small number of extremely influential genes (less than 1% of genes in a data set) can fundamentally change significant phylogenetic conclusions. In one example, these genes are shown to contain previously unrecognized paralogs. This study demonstrates both that the resolution of difficult phylogenomic problems remains sensitive to seemingly minor analysis details and that Bayes factors are a valuable tool for identifying and solving these challenges.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Chloroplast phylogenomic analyses resolve deep-level relationships of an intractable bamboo tribe Arundinarieae (Poaceae)

The temperate woody bamboos constitute a distinct tribe Arundinarieae (Poaceae: Bambusoideae) with high species diversity. Estimating phylogenetic relationships among the 11 major lineages of Arundinarieae has been particularly difficult, owing to a possible rapid radiation and the extremely low rate of sequence divergence. Here we explore the use of chloroplast genome sequencing for phylogenetic inference. We sampled 25 species (22 temperate bamboos and 3 outgroups) for the complete genome representing 8 major lineages of Arundinarieae in an attempt to resolve backbone relationships. Phylogenetic analyses of coding versus noncoding sequences, and of different regions of the genome (large and small single-copy, and inverted repeats regions) yielded no well-supported contradicting topologies but potential incongruence was found between the coding and noncoding sequences. The use of various data partitioning schemes in analysis of the complete sequences resulted in nearly identical topologies and node support values, although the partitioning schemes were decisively different from each other as to the fit to the data. Our full genomic data set substantially increased resolution along the backbone and provided strong support for most relationships despite the very short internodes and long branches in the tree. The inferred relationships were also robust to potential confounding factors (e.g., long-branch attraction) and received support from independent indels in the genome. We then added taxa from the three Arundinarieae lineages that were not included in the full-genome data set; each of these were sampled for &gt;50% genome sequences. The resulting trees not only corroborated the reconstructed deep-level relationships but also largely resolved the phylogenetic placements of these three additional lineages. Furthermore, adding 129 additional taxa sampled for only 8 chloroplast loci to the combined data set yielded almost identical relationships, albeit with low support values. We believe that the inferred phylogeny is robust to taxon sampling. Having resolved the deep-level relationships of Arundinarieae, we illuminate how chloroplast phylogenomics can be used for elucidating difficult phylogeny at low taxonomic levels in intractable plant groups.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Organellar phylogenomics inform systematics in the green algal family Hydrodictyaceae (Chlorophyceae) and provide clues to the complex evolutionary history of plastid genomes in the green algal Tree of Life.

Premise of the study: Phylogenomic analyses across the green algae are resolving relationships at the class, order and family levels, and highlighting dynamic patterns of evolution in organellar genomes. Here we present a within-family phylogenomic study to resolve genera and species relationships in the family Hydrodictyaceae (Chlorophyceae), for which poor resolution in previous phylogenetic studies, along with divergent morphological traits, have precluded taxonomic revisions. Methods: Complete plastome sequences and mitochondrial protein-coding gene sequences were acquired from representatives of the Hydrodictyaceae using Next-Generation sequencing methods. Plastomes were characterized and gene order and content were compared with plastomes spanning the Sphaeropleales. Single-gene and concatenated-gene phylogenetic analyses of plastid and mitochondrial genes were performed. Key results: The Hydrodictyaceae contain the largest sphaeroplealean plastomes thus far fully sequenced. Conservation of plastome gene order within Hydrodictyaceae is striking compared with more dynamic patterns revealed across Sphaeropleales. Phylogenetic analyses resolve Hydrodictyon sister to a monophyletic Pediastrum, though the morphologically distinct P. angulosum and P. duplex continue to be polyphyletic. Analyses of plastid data supported the neochloridacean genus Chlorotetraëdron as sister to Hydrodictyaceae, while conflicting signal was found in the mitochondrial data. Conclusions: A phylogenomic approach resolved within-family relationships not obtainable with previous phylogenetic analyses. Denser taxon sampling across Sphaeropleales is necessary to capture patterns in plastome evolution, and further taxa and studies are needed to fully resolve sister lineage to Hydrodictyaceae and polyphyly of Pediastrum angulosum and P. duplex.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Phylogenomic systematics of Ostariophysan fishes: ultraconserved elements support the surprising non-monophyly of Characiformes

Ostariophysi is a superorder of bony fishes including more than 10,300 species in 1,100 genera and 70 families. This superorder is traditionally divided into five major groups (orders): Gonorynchiformes (milkfishes and sandfishes), Cypriniformes (carps and minnows), Characiformes (tetras and their allies), Siluriformes (catfishes), and Gymnotiformes (electric knifefishes). Unambiguous resolution of the relationships among these lineages remains elusive, with previous molecular and morphological analyses failing to produce a consensus phylogeny. In this study, we use over 350 ultraconserved element (UCEs) loci comprising five million base pairs collected across thirty-five representative ostariophysan species to compile one of the most data-rich phylogenies of fishes to date. We use these data to infer higher-level (interordinal) relationships among ostariophysan fishes, focusing on the monophyly of the Characiformes— one the most contentiously debated groups in fish systematics. As with most previous molecular studies, we recover a non-monophyletic Characiformes with the two monophyletic suborders, Citharinoidei and Characoidei, more closely related to other ostariophysan clades than to each other. We also explore incongruence between results from different UCE datasets, issues of orthology, and the use of morphological characters in combination with our molecular data.

opencc-zeroDec 2016View details →
dryad28/100

Data from: An evaluation of transcriptome-based exon capture for frog phylogenomics across multiple scales of divergence (Class: Amphibia, Order: Anura)

Custom sequence capture experiments are becoming an efficient approach for gathering large sets of orthologous markers in nonmodel organisms. Transcriptome-based exon capture utilizes transcript sequences to design capture probes, typically using a reference genome to identify intron–exon boundaries to exclude shorter exons (&lt;200 bp). Here, we test directly using transcript sequences for probe design, which are often composed of multiple exons of varying lengths. Using 1260 orthologous transcripts, we conducted sequence captures across multiple phylogenetic scales for frogs, including outgroups ~100 Myr divergent from the ingroup. We recovered a large phylogenomic data set consisting of sequence alignments for 1047 of the 1260 transcriptome-based loci (~561 000 bp) and a large quantity of highly variable regions flanking the exons in transcripts (~70 000 bp), the latter improving substantially by only including ingroup species (~797 000 bp). We recovered both shorter (&lt;100 bp) and longer exons (&gt;200 bp), with no major reduction in coverage towards the ends of exons. We observed significant differences in the performance of blocking oligos for target enrichment and nontarget depletion during captures, and differences in PCR duplication rates resulting from the number of individuals pooled for capture reactions. We explicitly tested the effects of phylogenetic distance on capture sensitivity, specificity, and missing data, and provide a baseline estimate of expectations for these metrics based on a priori knowledge of nuclear pairwise differences among samples. We provide recommendations for transcriptome-based exon capture design based on our results, cost estimates and offer multiple pipelines for data assembly and analysis.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Phylogenomics resolves evolutionary relationships among ants, bees, and wasps

Eusocial behavior has arisen in few animal groups, most notably in the aculeate Hymenoptera, a clade comprising ants, bees, and stinging wasps. Phylogeny is crucial to understanding the evolution of the salient features of these insects, including eusociality. Yet the phylogenetic relationships among the major lineages of aculeate Hymenoptera remain contentious. We address this problem here by generating and analyzing genomic data for a representative series of taxa. We obtain a single well-resolved and strongly supported tree, robust to multiple methods of phylogenetic inference. Apoidea (spheciform wasps and bees) and ants are sister groups, a novel finding that contradicts earlier views that ants are closer to ectoparasitoid wasps. Vespid wasps (paper wasps, yellow jackets, and relatives) are sister to all other aculeates except chrysidoids. Thus, all eusocial species of Hymenoptera are contained within two major groups, characterized by transport of larval provisions and nest construction, likely prerequisites for the evolution of eusociality. These two lineages are interpolated among three other clades of wasps whose species are predominantly ectoparasitoids on concealed hosts, the inferred ancestral condition for aculeates. This phylogeny provides a new framework for exploring the evolution of nesting, feeding, and social behavior within the stinging Hymenoptera.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Phylogenomic evidence for ancient hybridization in the genomes of living cats (Felidae)

Interspecies hybridization has been recently recognized as potentially common in wild animals, but the extent to which it shapes modern genomes is still poorly understood. Distinguishing historical hybridization events from other processes leading to phylogenetic discordance among different markers requires a well-resolved species tree that considers all modes of inheritance, and overcomes systematic problems due to rapid lineage diversification by sampling large genomic character sets. Here we assessed genome-wide phylogenetic variation across a diverse mammalian family, Felidae (cats). We combined genotypes from a genome-wide SNP array with additional autosomal, X- and Y-linked variants to sample ~150 kilobases of nuclear sequence, in addition to complete mitochondrial genomes generated using light-coverage Illumina sequencing. We present the first robust felid timetree that accounts for unique maternal, paternal, and biparental evolutionary histories. Signatures of phylogenetic discordance were abundant in the genomes of modern cats, in many cases indicating hybridization as the most likely cause. Comparison of big cat whole-genome sequences revealed a substantial reduction of X-linked divergence times across several large recombination coldspots, which were highly enriched for signatures of selection-driven post-divergence hybridization between the ancestors of the snow leopard and lion lineages. These results highlight the mosaic origin of modern felid genomes and the influence of sex chromosomes and sex-biased dispersal in post-speciation gene flow. A complete resolution of the Tree of Life will require comprehensive genomic sampling of biparental and sex-limited genetic variation to identify and control for phylogenetic conflict caused by ancient admixture and sex-biased differences in genomic transmission.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Using phylogenomic data to explore the effects of relaxed clocks and calibration strategies on divergence time estimation: primates as a test case

Primates have long been a test case for the development of phylogenetic methods for divergence time estimation. Despite a large number of studies, however, the timing of origination of crown Primates relative to the K-Pg boundary and the timing of diversification of the main crown groups remain controversial. Here we analysed a dataset of 372 taxa (367 Primates and 5 outgroups, 3.4 million aligned base pairs) that includes nine primate genomes. We systematically explore the effect of different interpretations of fossil calibrations and molecular clock models on primate divergence time estimates. We find that even small differences in the construction of fossil calibrations can have a noticeable impact on estimated divergence times, especially for the oldest nodes in the tree. Notably, choice of molecular rate model (auto-correlated or independently distributed rates) has an especially strong effect on estimated times, with the independent rates model producing considerably more ancient age estimates for the deeper nodes in the phylogeny. We implement thermodynamic integration, combined with Gaussian quadrature, in the program MCMCTree, and use it to calculate Bayes factors for clock models. Bayesian model selection indicates that the auto-correlated rates model fits the primate data substantially better, and we conclude that time estimates under this model should be preferred. We show that for eight core nodes in the phylogeny, uncertainty in time estimates is close to the theoretical limit imposed by fossil uncertainties. Thus, these estimates are unlikely to be improved by collecting additional molecular sequence data. All analyses place the origin of Primates close to the K-Pg boundary, either in the Cretaceous or straddling the boundary into the Palaeogene.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Deep phylogenomics of a tandem-repeat galectin regulating appendicular skeletal pattern formation

Background: A multiscale network of two galectins Galectin-1 (Gal-1) and Galectin-8 (Gal-8) patterns the avian limb skeleton. Among vertebrates with paired appendages, chondrichthyan fins typically have one or more cartilage plates and many repeating parallel endoskeletal elements, actinopterygian fins have more varied patterns of nodules, bars and plates, while tetrapod limbs exhibit tandem arrays of few, proximodistally increasing numbers of elements. We applied a comparative genomic and protein evolution approach to understand the origin of the galectin patterning network. Having previously observed a phylogenetic constraint on Gal-1 structure across vertebrates, we asked whether evolutionary changes of Gal-8 could have critically contributed to the origin of the tetrapod pattern. Results: Translocations, duplications, and losses of Gal-8 genes in Actinopterygii established them in different genomic locations from those that the Sarcopterygii (including the tetrapods) share with chondrichthyans. The sarcopterygian Gal-8 genes acquired a potentially regulatory non-coding motif and underwent purifying selection. The actinopterygian Gal-8 genes, in contrast, did not acquire the non-coding motif and underwent positive selection. Conclusion: These observations interpreted through the lens of a reaction-diffusion-adhesion model based on avian experimental findings can account for the distinct endoskeletal patterns of cartilaginous, ray-finned, and lobe-finned fishes, and the stereotypical limb skeletons of tetrapods.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Chloroplast phylogenomic data from the green algal order Sphaeropleales (Chlorophyceae, Chlorophyta) reveal complex patterns of sequence evolution

Chloroplast sequence data are widely used to infer phylogenies of plants and algae. With the increasing availability of complete chloroplast genome sequences, the opportunity arises to resolve ancient divergences that were heretofore problematic. On the flip side, properly analyzing large multi-gene data sets can be a major challenge, as these data may be riddled with systematic biases and conflicting signals. Our study contributes new data from nine complete and four fragmentary chloroplast genome sequences across the green algal order Sphaeropleales. Our phylogenetic analyses of a 56-gene data set show that analyzing these data on a nucleotide level yields a well-supported phylogeny – yet one that is quite different from a corresponding amino acid analysis. We offer some possible explanations for this conflict through a range of analyses of modified data sets. In addition, we characterize the newly sequenced genomes in terms of their structure and content, thereby further contributing to the knowledge of chloroplast genome evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Phylogenomic mining of the mints reveals multiple mechanisms contributing to the evolution of chemical diversity in Lamiaceae

The evolution of chemical complexity has been a major driver of plant diversification, with novel compounds serving as key innovations. The species-rich mint family (Lamiaceae) produces an enormous variety of compounds that act as attractants and defense molecules in nature and are used widely by humans as flavor additives, fragrances, and anti-herbivory agents. To elucidate the mechanisms by which such diversity evolved, we combined leaf transcriptome data from 48 Lamiaceae species and four outgroups with a robust phylogeny and chemical analyses of three terpenoid classes (monoterpenes, sesquiterpenes, iridoids) that share and compete for precursors. Our integrated chemical-genomic-phylogenetic approach revealed that: 1) gene family expansion rather than increased enzyme promiscuity of terpene synthases is correlated with mono- and sesqui-terpene diversity; 2) differential expression of core genes within the iridoid biosynthetic pathway is associated with iridoid presence/absence; 3) generally, production of iridoids and canonical monoterpenes appeared to be inversely correlated; and 4) iridoid biosynthesis was significantly associated with expression of geraniol synthase, which diverts metabolic flux away from canonical monoterpenes, suggesting that competition for common precursors can be a central control point in specialized metabolism. These results suggest that multiple mechanisms contributed to the evolution of chemodiversity in this economically important family.

opencc-zeroDec 2017View details →
dryad28/100

Data from: A phylogenomic resolution of the sea urchin tree of life

Background: Echinoidea is a clade of marine animals including sea urchins, heart urchins, sand dollars and sea biscuits. Found in benthic habitats across all latitudes, echinoids are key components of marine communities such as coral reefs and kelp forests. A little over 1,000 species inhabit the oceans today, a diversity that traces its roots back at least to the Permian. Although much effort has been devoted to elucidating the echinoid tree of life using a variety of morphological data, molecular attempts have relied on only a handful of genes. Both of these approaches have had limited success at resolving the deepest nodes of the tree, and their disagreement over the positions of a number of clades remains unresolved. Results: We performed de novo sequencing and assembly of 17 transcriptomes to complement available genomic resources of sea urchins and produce the first phylogenomic analysis of the clade. Multiple methods of probabilistic inference recovered identical topologies, with virtually all nodes showing maximum support. In contrast, the coalescent-based method ASTRAL-II resolved one node differently, a result apparently driven by gene tree error induced by evolutionary rate heterogeneity. Regardless of the method employed, our phylogenetic structure deviates from the currently accepted classification of echinoids, with neither Acroechinoidea (all euechinoids except echinothurioids), nor Clypeasteroida (sand dollars and sea biscuits) being monophyletic as currently defined. We show that phylogenetic signal for novel resolutions of these lineages is strong and distributed throughout the genome, and fail to recover systematic biases as drivers of our results. Conclusions: Our investigation substantially augments the molecular resources available for sea urchins, providing the first transcriptomes for many of its main lineages. Using this expanded genomic dataset, we resolve the position of several clades in agreement with early molecular analyses but in disagreement with morphological data. Our efforts settle multiple phylogenetic uncertainties, including the position of the enigmatic deep-sea echinothurioids and the identity of the sister clade to sand dollars. We offer a detailed assessment of evolutionary scenarios that could reconcile our findings with morphological evidence, opening up new lines of research into the development and evolutionary history of this ancient clade.

opencc-zeroDec 2017View details →
dryad28/100

Data from: A phylogenomic perspective on the radiation of ray-finned fishes based upon targeted sequencing of ultraconserved elements (UCEs)

Ray-finned fishes constitute the dominant radiation of vertebrates with over 32,000 species. Although molecular phylogenetics has begun to disentangle major evolutionary relationships within this vast section of the Tree of Life, there is no widely available approach for efficiently collecting phylogenomic data within fishes, leaving much of the enormous potential of massively parallel sequencing technologies for resolving major radiations in ray-finned fishes unrealized. Here, we provide a genomic perspective on longstanding questions regarding the diversification of major groups of ray-finned fishes through targeted enrichment of ultraconserved nuclear DNA elements (UCEs) and their flanking sequence. Our workflow efficiently and economically generates data sets that are orders of magnitude larger than those produced by traditional approaches and is well-suited to working with museum specimens. Analysis of the UCE data set recovers a well-supported phylogeny at both shallow and deep time-scales that supports a monophyletic relationship between Amia and Lepisosteus (Holostei) and reveals elopomorphs and then osteoglossomorphs to be the earliest diverging teleost lineages. Our approach additionally reveals that sequence capture of UCE regions and their flanking sequence offers enormous potential for resolving phylogenetic relationships within ray-finned fishes.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Anchored hybrid enrichment for massively high-throughput phylogenomics

The field of phylogenetics is on the cusp of a major revolution, enabled by new methods of data collection that leverage both genomic resources and recent advances in DNA sequencing. Previous phylogenetic work has required labor-intensive marker development coupled with single-locus PCR and DNA sequencing on a clade-by-clade and marker-by-marker basis. Here, we present a new, cost-efficient, and rapid approach to obtaining data from hundreds of genes for potentially hundreds of individuals for deep and shallow phylogenetic studies. Specifically, we designed probes for target enrichment of &gt;500 loci in highly-conserved anchor regions of vertebrate genomes (flanked by less conserved regions) from five model species and tested enrichment efficiency in non-model species up to 254 million years divergent from the nearest model. We found that hybrid enrichment using conserved probes (anchored enrichment) can recover a large number of unlinked loci that are useful at a diversity of phylogenetic timescales. This new approach has the potential to not only expedite resolution of deep-scale portions of the Tree of Life but also to greatly accelerate resolution of the large number of shallow clades that remain unresolved. The combination of low cost (~1% of the cost of traditional Sanger sequencing and ~3.5% of the cost of high-throughput amplicon sequencing for projects on the scale of 500 loci x 100 individuals) and rapid data collection (~2 weeks of laboratory time) are expected to make this approach tractable even for researchers working on systems with limited or non-existent genomic resources.

opencc-zeroDec 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record