Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,344

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,344 results for “: phylogenomics”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Vitis phylogenomics: hybridization intensities from a SNP array outperform genotype calls

Understanding relationships among species is a fundamental goal of evolutionary biology. Single nucleotide polymorphisms (SNPs) identified through next generation sequencing and related technologies enable phylogeny reconstruction by providing unprecedented numbers of characters for analysis. One approach to SNP-based phylogeny reconstruction is to identify SNPs in a subset of individuals, and then to compile SNPs on an array that can be used to genotype additional samples at hundreds or thousands of sites simultaneously. Although powerful and efficient, this method is subject to ascertainment bias because applying variation discovered in a representative subset to a larger sample favors identification of SNPs with high minor allele frequencies and introduces bias against rare alleles. Here, we demonstrate that the use of hybridization intensity data, rather than genotype calls, reduces the effects of ascertainment bias. Whereas traditional SNP calls assess known variants based on diversity housed in the discovery panel, hybridization intensity data survey variation in the broader sample pool, regardless of whether those variants are present in the initial SNP discovery process. We apply SNP genotype and hybridization intensity data derived from the Vitis9kSNP array developed for grape to show the effects of ascertainment bias and to reconstruct evolutionary relationships among Vitis species. We demonstrate that phylogenies constructed using hybridization intensities suffer less from the distorting effects of ascertainment bias, and are thus more accurate than phylogenies based on genotype calls. Moreover, we reconstruct the phylogeny of the genus Vitis using hybridization data, show that North American subgenus Vitis species are monophyletic, and resolve several previously poorly known relationships among North American species. This study builds on earlier work that applied the Vitis9kSNP array to evolutionary questions within Vitis vinifera and has general implications for addressing ascertainment bias in array-enabled phylogeny reconstruction.

opencc-zeroDec 2012View details →
dryad28/100

Data from: The impact of anchored phylogenomics and taxon sampling on phylogenetic inference in narrow-mouthed frogs (Anura, Microhylidae)

Despite considerable progress in unravelling the phylogenetic relationships of microhylid frogs, relationships among subfamilies remain largely unstable and many genera are not demonstrably monophyletic. Here, we used five alternative combinations of DNA sequence data (ranging from seven loci for 48 taxa to up to 73 loci for as many as 142 taxa) generated using the anchored phylogenomics sequencing method (66 loci, derived from conserved genome regions, for 48 taxa) and Sanger sequencing (seven loci for up to 142 taxa) to tackle this problem. We assess the effects of character sampling, taxon sampling, analytical methods and assumptions in phylogenetic inference of microhylid frogs. The phylogeny of microhylids shows high susceptibility to different analytical methods and datasets used for the analyses. Clades inferred from maximum-likelihood are generally more stable across datasets than those inferred from parsimony. Parsimony trees inferred within a tree-alignment framework are generally better resolved and better supported than those inferred within a similarity-alignment framework, even under the same cost matrix (equally weighted) and same treatment of gaps (as a fifth nucleotide state). We discuss potential causes for these differences in resolution and clade stability among discovery operations. We also highlight the problem that commonly used algorithms for model-based analyses do not explicitly model insertion and deletion events (i.e. gaps are treated as missing data). Our results corroborate the monophyly of Microhylidae and most currently recognized subfamilies but fail to provide support for relationships among subfamilies. Several taxonomic updates are provided, including naming of two new subfamilies, both monotypic.

opencc-zeroDec 2014View details →
dryad28/100

Data from: When outgroups fail; phylogenomics of rooting the emerging pathogen, Coxiella burnetii

Rooting phylogenies is critical for understanding evolution, yet the importance, intricacies and difficulties of rooting are often overlooked. For rooting, polymorphic characters among the group of interest (ingroup) must be compared to those of a relative (outgroup) that diverged before the last common ancestor (LCA) of the ingroup. Problems arise if an outgroup does not exist, is unknown, or is so distant that few characters are shared, in which case duplicated genes originating before the LCA can be used as proxy outgroups to root diverse phylogenies. Here, we describe a genome-wide expansion of this technique that can be used to solve problems at the other end of the evolutionary scale: where ingroup individuals are all very closely related to each other, but the next closest relative is very distant. We used shared orthologous single nucleotide polymorphisms (SNPs) from 10 whole genome sequences of Coxiella burnetii, the causative agent of Q fever in humans, to create a robust, but unrooted phylogeny. To maximize the number of characters informative about the rooting, we searched entire genomes for polymorphic duplicated regions where orthologs of each paralog could be identified so that the paralogs could be used to root the tree. Recent radiations, such as those of emerging pathogens, often pose rooting challenges due to a lack of ingroup variation and large genomic differences with known outgroups. Using a phylogenomic approach, we created a robust, rooted phylogeny for C. burnetii.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Golden orbweavers ignore biological rules: phylogenomic and comparative analyses unravel a complex evolution of sexual size dimorphism

Instances of sexual size dimorphism (SSD) provide the context for rigorous tests of biological rules of size evolution, such as Cope's Rule (phyletic size increase), Rensch's Rule (allometric patterns of male and female size), as well as male and female body size optima. In certain spider groups, such as the golden orbweavers (Nephilidae), extreme female-biased SSD (eSSD, female:male body length ≥ 2) is the norm. Nephilid genera construct webs of exaggerated proportions, which can be aerial, arboricolous, or intermediate (hybrid). First, we established the backbone phylogeny of Nephilidae using 367 Anchored Hybrid Enrichment (AHE) markers, then combined these data with classical markers for a reference species-level phylogeny. Second, we used the phylogeny to test Cope and Rensch's Rules, sex specific size optima, and the coevolution of web size, type, and features with female and male body size and their ratio, SSD. Male, but not female, size increases significantly over time, and refutes Cope's Rule. Allometric analyses reject the converse, Rensch's Rule. Male and female body sizes are uncorrelated. Female size evolution is random, but males evolve towards an optimum size (3.2-4.9 mm). Overall, female body size correlates positively with absolute web size. However, intermediate sized females build the largest webs (of the hybrid type), giant female Nephila and Trichonephila build smaller webs (of the aerial type), and the smallest females build the smallest webs (of the arboricolous type). We propose taxonomic changes based on the criteria of clade age, monophyly and exclusivity, classification information content, and diagnosability. Spider families, as currently defined, tend to be between 37-98 million years old, and Nephilidae is estimated at 133 Ma (97 - 146), thus deserving family status. We therefore resurrect the family Nephilidae Simon 1894 that contains Clitaetra Simon 1889, the Cretaceous Geratonephila Poinar & Buckley 2012, Herennia Thorell 1877, Indoetra Kuntner 2006, new rank, Nephila Leach 1815, Nephilengys L. Koch 1872, Nephilingis Kuntner 2013, Palaeonephila Wunderlich 2004 from Tertiary Baltic amber, and Trichonephila Dahl 1911, new rank. We propose the new clade Orbipurae to contain Araneidae Clerck 1757, Phonognathidae Simon 1894, new rank, and Nephilidae. Nephilid female gigantism is a phylogenetically-ancient phenotype (over 100 Ma), as is eSSD, though their magnitudes vary by lineage.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Phylogenomics resolves a spider backbone phylogeny and rejects a prevailing paradigm for orb web evolution

Spiders represent an ancient predatory lineage known for their extraordinary biomaterials, including venoms and silks. These adaptations make spiders key arthropod predators in most terrestrial ecosystems. Despite ecological, biomedical, and biomaterial importance, relationships among major spider lineages remain unresolved or poorly supported. Current working hypotheses for a spider "backbone" phylogeny are largely based on morphological evidence, as most molecular markers currently employed are generally inadequate for resolving deeper-level relationships. We present here a phylogenomic analysis of spiders including taxa representing all major spider lineages. Our robust phylogenetic hypothesis recovers some fundamental and uncontroversial spider clades, but rejects the prevailing paradigm of a monophyletic Orbiculariae, the most diverse lineage, containing orb-weaving spiders. Based on our results, the orb web either evolved much earlier than previously hypothesized and is ancestral for a majority of spiders or else it has multiple independent origins, as hypothesized by precladistic authors. Cribellate deinopoid orb weavers that use mechanically adhesive silk are more closely related to a diverse clade of mostly webless spiders than to the araneoid orb-weaving spiders that use adhesive droplet silks. The fundamental shift in our understanding of spider phylogeny proposed here has broad implications for interpreting the evolution of spiders, their remarkable biomaterials, and a key extended phenotype—the spider web.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Marker development for phylogenomics: the case of Orobanchaceae, a plant family with contrasting nutritional modes

Phylogenomic approaches, employing next-generation sequencing (NGS) techniques, have revolutionized systematic and evolutionary biology. Target enrichment is an efficient and cost-effective method in phylogenomics and is becoming increasingly popular. Depending on availability and quality of reference data as well as on biological features of the study system, (semi-)automated identification of suitable markers will require specific bioinformatic pipelines. Here, we established a highly flexible bioinformatic pipeline, BaitsFinder, to identify putative orthologous single copy genes (SCGs) and to construct bait sequences in a single workflow. Additionally, this pipeline has been constructed to be able to cope with challenging data sets, such as the nutritionally heterogeneous plant family Orobanchaceae. To this end, we used transcriptome data of differing quality available for four Orobanchaceae species and, as reference, SCG data from monkeyflower (Erythranthe guttata, syn. Mimulus g.; 1,915 genes) and tomato (Solanum lycopersicum; 391 genes). Depending on whether gaps were permitted in initial blast searches of the four Orobanchaceae species against the reference, our pipeline identified 1,307 and 981 SCGs with average length of 994 bp and 775 bp, respectively. Automated bait sequence construction (using 2× tiling) resulted in 38,170 and 21,856 bait sequences, respectively. In comparison to the recently published MarkerMiner 1.0 pipeline BaitsFinder identified about 1.6 times as many SCGs (of at least 900 bp length). Skipping steps specific to analyses of Orobanchaceae, BaitsFinder was successfully used in a group of non-parasitic plants (three Asteraceae species and, as reference, SCG data from Arabidopsis thaliana based on previously compiled SCGs). Thus, BaitsFinder is expected to be broadly applicable in groups, where only transcriptomes or partial genome data of differing quality are available.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Phylogenomic resolution of the Hemichordate and Echinoderm clade

Ambulacraria, comprising Hemichordata and Echinodermata, is closely related to Chordata, making it integral to understanding chordate origins and polarizing chordate molecular and morphological characters. Unfortunately, relationships within Hemichordata and Echinodermata have remained unresolved, compromising our ability to extrapolate findings from the most closely related molecular and developmental models outside of Chordata (e.g., the acorn worms Saccoglossus kowalevskii and Ptychodera flava and the sea urchin Strongylocentrotus purpuratus). To resolve long-standing phylogenetic issues within Ambulacraria, we sequenced transcriptomes for 14 hemichordates as well as 8 echinoderms and complemented these with existing data for a total of 33 ambulacrarian operational taxonomic units (OTUs). Examination of leaf stability values revealed rhabdopleurid pterobranchs and the enteropneust Stereobalanus canadensis were unstable in placement; therefore, analyses were also run without these taxa. Analyses of 185 genes resulted in reciprocal monophyly of Enteropneusta and Pterobranchia, placed the deep-sea family Torquaratoridae within Ptychoderidae, and confirmed the position of ophiuroid brittle stars as sister to asteroid sea stars (the Asterozoa hypothesis). These results are consistent with earlier perspectives concerning plesiomorphies of Ambulacraria, including pharyngeal gill slits, a single axocoel, and paired hydrocoels and somatocoels. The resolved ambulacrarian phylogeny will help clarify the early evolution of chordate characteristics and has implications for our understanding of major fossil groups, including graptolites and somasteroideans.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Taxon-rich phylogenomic analyses resolve the eukaryotic tree of life and reveal the power of subsampling by sites

Most eukaryotic lineages are microbial, and many have only recently been sampled for phylogenetic studies or remain in the 'dark area' of the tree of life where there are no molecular data. To assess relationships among eukaryotic lineages, we perform a taxon-rich phylogenomic analysis including 232 eukaryotes selected to maximize taxonomic diversity and up to 1554 genes chosen as vertically inherited based on their broad distribution among eukaryotes. We also include sequences from 486 bacteria and 84 archaea to assess the impact of endosymbiotic gene transfer (EGT) from plastids and to detect contamination. Overall, our analyses are consistent with other less taxon-rich estimates of the eukaryotic tree of life and we recover strong support for five major clades: Amoebozoa, Excavata (without the genus Malawimonas), Opisthokonta, Archaeplastida and SAR (Stramenopila, Alveolata and Rhizaria). Our analyses also highlight the existence of 'orphan' lineages, lineages that lack robust placement in the eukaryotic tree of life and indicate the possibility of as yet undiscovered diversity. In analyses including bacteria and archaea, we find that ~10% of the 1554 genes, which we choose because they are found in four or five of the five major eukaryotic clades and hence may be more likely to be inherited vertically, appear to have been acquired from cyanobacteria through EGT in photosynthetic lineages. Removing these EGT genes places the green algae as sister to the glaucophytes instead of the red algae, suggesting that unknowingly including of genes of plastid origin, and combining them with genes of nuclear origin, may mislead phylogenetic estimates. Finally, the large size of our dataset allows comparative analyses of subsets of data; alignments built from randomly sampled sites provide greater support, particularly for deep relationships, than do equivalent sized datasets built from randomly sampled genes.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Untangling the early diversification of eukaryotes: a phylogenomic study of the evolutionary origins of Centrohelida, Haptophyta, and Cryptista

Assembling the global eukaryotic tree of life has long been a major effort of Biology. In recent years, pushed by the new availability of genome-scale data for microbial eukaryotes, it has become possible to revisit many evolutionary enigmas. However, some of the most ancient nodes, which are essential for inferring a stable tree, have remained highly controversial. Among other reasons, the lack of adequate genomic datasets for key taxa has prevented the robust reconstruction of early diversification events. In this context, the centrohelid heliozoans are particularly relevant for reconstructing the tree of eukaryotes because they represent one of the last substantial groups that was missing large and diverse genomic data. Here, we filled this gap by sequencing high-quality transcriptomes for four centrohelid lineages, each corresponding to a different family. Combining these new data with a broad eukaryotic sampling, we produced a gene-rich taxon-rich phylogenomic dataset that enabled us to refine the structure of the tree. Specifically, we show that (i) centrohelids relate to haptophytes, confirming Haptista; (ii) Haptista relates to SAR; (iii) Cryptista share strong affinity with Archaeplastida; and (iv) Haptista + SAR is sister to Cryptista + Archaeplastida. The implications of this topology are discussed in the broader context of plastid evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: The relative importance of modeling site pattern heterogeneity versus partition-wise heterotachy in phylogenomic inference

Large taxa-rich genome-scale data sets are often necessary for resolving ancient phylogenetic relationships. But accurate phylogenetic inference requires that they are analyzed with realistic models that account for the heterogeneity in substitution patterns amongst the sites, genes and lineages. Two kinds of adjustments are frequently used: models that account for heterogeneity in amino acid frequencies at sites in proteins, and partitioned models that accommodate the heterogeneity in rates (branch lengths) among different proteins in different lineages (protein-wise heterotachy). Although partitioned and site-heterogeneous models are both widely used in isolation, their relative importance to the inference of correct phylogenies has not been carefully evaluated. We conducted several empirical analyses and a large set of simulations to compare the relative performances of partitioned models, site-heterogeneous models and combined partitioned site heterogeneous models. In general, site-homogeneous models (partitioned or not) performed worse than site heterogeneous, except in simulations with extreme protein-wise heterotachy. Furthermore, simulations using empirically-derived realistic parameter settings showed a marked long-branch attraction (LBA) problem for analyses employing protein-wise partitioning even when the generating model included partitioning. This LBA problem results from a small sample bias compounded over many single protein alignments. In some cases, this problem was ameliorated by clustering similarly-evolving proteins together into larger partitions using the PartitionFinder method. Similar results were obtained under simulations with larger numbers of taxa or heterogeneity in simulating topologies over genes. For an empirical Microsporidia test data set, all but one tested site-heterogeneous models (with or without partitioning) obtain the correct Microsporidia+Fungi grouping, whereas site-homogenous models (with or without partitioning) did not. The single exception was the fully partitioned site-heterogeneous analysis that succumbed to the compounded small sample LBA bias. In general unless protein-wise heterotachy effects are extreme, it is more important to model site-heterogeneity than protein-wise heterotachy in phylogenomic analyses. Complete protein-wise partitioning should be avoided as it can lead to a serious LBA bias. In cases of extreme protein-wise heterotachy, approaches that cluster similarly-evolving proteins together and coupled with site-heterogeneous models work well for phylogenetic estimation.

opencc-zeroDec 2018View details →
dryad28/100

Data from: A phylogenomic analysis of the role and timing of molecular adaptation in the aquatic transition of cetartiodactyl mammals

Recent studies have reported multiple cases of molecular adaptation in cetaceans related to their aquatic abilities. However, none of these has included the hippopotamus, precluding an understanding of whether molecular adaptations in cetaceans occurred before or after they split from their semi-aquatic sister taxa. Here, we obtained new transcriptomes from the hippopotamus and humpback whale, and analysed these together with available data from eight other cetaceans. We identified more than 11 000 orthologous genes and compiled a genome-wide dataset of 6845 coding DNA sequences among 23 mammals, to our knowledge the largest phylogenomic dataset to date for cetaceans. We found positive selection in nine genes on the branch leading to the common ancestor of hippopotamus and whales, and 461 genes in cetaceans compared to 64 in hippopotamus. Functional annotation revealed adaptations in diverse processes, including lipid metabolism, hypoxia, muscle and brain function. By combining these findings with data on protein–protein interactions, we found evidence suggesting clustering among gene products relating to nervous and muscular systems in cetaceans. We found little support for shared ancestral adaptations in the two taxa; most molecular adaptations in extant cetaceans occurred after their split with hippopotamids.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Anchored phylogenomics illuminates the skipper butterfly tree of life

Butterflies (Papilionoidea) are perhaps the most charismatic insect lineage, yet phylogenetic relationships among them remain incompletely studied and controversial. We sequenced nearly 400 loci using Anchored Hybrid Enrichment and sampled all tribes and more than 120 genera of skippers (Hesperiidae), one of the most species-rich and poorly studied butterfly families. Maximum-likelihood, parsimony and coalescent multi-species methods all converged on a novel, robust phylogenetic hypothesis for skippers. Different optimality criteria and methodologies recovered almost identical phylogenetic trees with strong nodal support at nearly all taxonomic levels. Our results support Coeliadinae as the sister group to the remaining skippers, the monotypic Euschemoninae as sister group to all other subfamilies but Coeliadinae, and the monophyly of Eudaminae plus Pyrginae. Within Pyrginae, Celaenorrhinini and Tagiadini are sister groups, the Neotropical firetips, Pyrrhopygini, are sister to all other tribes but Celaenorrhinini and Tagiadini. Achlyodini is recovered as the sister group to Carcharodini, and Erynnini as sister group to Pyrgini. Within Hesperiinae, there is strong support for the monophyly of Aeromachini plus remaining Hesperiinae. The giant skippers (Agathymus and Megathymus) once classified as a single subfamily, are recovered as monophyletic with strong support, but are deeply nested within grass skippers (Hesperiinae). These results enhance understanding of the evolution of one of the most species-rich butterfly families.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Phylogenomic support for evolutionary relationships of New World direct-developing frogs (Anura: Terraranae)

Phylogenomic approaches have proven able to resolve difficult branches in the tree of life. New World direct-developing frogs (Terraranae) represent a large evolutionary radiation in which interrelationships at key points in the phylogeny have not been adequately determined, affecting evolutionary, biogeographic, and taxonomic interpretations. We employed anchored hybrid enrichment to generate a data set containing 389 loci and >600,000 nucleotide positions for 30 terraranan and several outgroup frog species encompassing all major lineages in the clade. Concatenated maximum likelihood and coalescent species-tree approaches recover nearly identical topologies with strong support for nearly all relationships in the tree. These results are similar to previous phylogenetic results but provide additional resolution at short internodes. Among taxa whose placement varied in previous analyses, Ceuthomantis is shown to be the sister taxon to all other terraranans, rather than deeply embedded within the radiation, and Strabomantidae is monophyletic rather than paraphyletic with respect to Craugastoridae. We present an updated taxonomy to reflect these results, and describe a new subfamily for the genus Hypodactylus.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Resources for phylogenomic analyses of Australian terrestrial vertebrates

High-throughput sequencing methods promise to improve our ability to infer the evolutionary histories of lineages and to delimit species. These are exciting prospects for the study of Australian vertebrates, a group comprised of many globally unique lineages with a long history of isolation. The evolutionary relationships within many of these lineages have been difficult to resolve with small numbers of loci, and we now know that many lineages also exhibit substantial cryptic diversity. Here, we present a set of phylogenetically diverse transcriptome resources to enable exon-based sequence capture studies of Australian vertebrates, including transcriptome sequences for four species of birds, four frogs, seven lizards and seven mammals. We also use exon data from the marsupial transcriptomes we generated to examine an approach for choosing a moderate number (dozens or hundreds) of phylogenetically informative exons based on a single transcriptome sequence, and a relatively distant reference genome.

opencc-zeroDec 2015View details →
dryad28/100

Data from: A phylogenomic approach to vertebrate phylogeny supports a turtle-archosaur affinity and a possible paraphyletic Lissamphibia

In resolving the vertebrate tree of life, two fundamental questions remain: 1) what is the phylogenetic position of turtles within amniotes, and 2) what are the relationships between the three major lissamphibian (extant amphibian) groups? These relationships have historically been difficult to resolve, with five different hypotheses proposed for turtle placement, and four proposed branching patterns within Lissamphibia. We compiled a large cDNA/EST dataset for vertebrates (75 genes for 129 taxa) to address these outstanding questions. Gene-specific phylogenetic analyses revealed a great deal of variation in preferred topology, resulting in topologically ambiguous conclusions from the combined dataset. Due to consistent preferences for the same divergent topologies across genes, we suspected systematic phylogenetic error as a cause of some variation. Accordingly, we developed and tested a novel statistical method that identifies sites that have a high probability of containing biased signal for a specific phylogenetic relationship. After removing putatively biased sites, support emerged for a sister relationship between turtles and either crocodilians or archosaurs, as well as for a caecilian-salamander sister relationship within Lissamphibia, with Lissamphibia potentially paraphyletic.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Step-wise evolution of complex chemical defenses in millipedes: a phylogenomic approach

With fossil representatives from the Silurian capable of respiring atmospheric oxygen, millipedes are among the oldest terrestrial animals, and likely the first to acquire diverse and complex chemical defenses against predators. Exploring the origin of complex adaptive traits is critical for understanding the evolution of Earth's biological complexity, and chemical defense evolution serves as an ideal study system. The classic explanation for the evolution of complexity is by gradual increase from simple to complex, passing through intermediate "stepping stone" states. Here we present the first phylogenetic-based study of the evolution of complex chemical defenses in millipedes by generating the largest genomic-based phylogenetic dataset ever assembled for the group. Our phylogenomic results demonstrate that chemical complexity shows a clear pattern of escalation through time. New pathways are added in a stepwise pattern, leading to greater chemical complexity, independently in a number of derived lineages. This complexity gradually increased through time, leading to the advent of three distantly related chemically complex evolutionary lineages, each uniquely characteristic of each of the respective millipede groups.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Effectiveness of phylogenomic data and coalescent species-tree methods for resolving difficult nodes in the phylogeny of advanced snakes (Serpentes: Caenophidia)

Next-generation genomic sequencing promises to quickly and cheaply resolve remaining contentious nodes in the Tree of Life, and facilitates species-tree estimation while taking into account stochastic genealogical discordance among loci. Recent methods for estimating species trees bypass full likelihood-based estimates of the multi-species coalescent, and approximate the true species-tree using simpler summary metrics. These methods converge on the true species-tree with sufficient genomic sampling, even in the anomaly zone. However, no studies have yet evaluated their efficacy on a large-scale phylogenomic dataset, and compared them to previous concatenation strategies. Here, we generate such a dataset for Caenophidian snakes, a group with >2500 species that contains several rapid radiations that were poorly resolved with fewer loci. We generate sequence data for 333 single-copy nuclear loci with ∼100% coverage (∼0% missing data) for 31 major lineages. We estimate phylogenies using neighbor joining, maximum parsimony, maximum likelihood, and three summary species-tree approaches (NJst, STAR, and MP-EST). All methods yield similar resolution and support for most nodes. However, not all methods support monophyly of Caenophidia, with Acrochordidae placed as the sister taxon to Pythonidae in some analyses. Thus, phylogenomic species-tree estimation may occasionally disagree with well-supported relationships from concatenated analyses of small numbers of nuclear or mitochondrial genes, a consideration for future studies. In contrast for at least two diverse, rapid radiations (Lamprophiidae and Colubridae), phylogenomic data and species-tree inference do little to improve resolution and support. Thus, certain nodes may lack strong signal, and larger datasets and more sophisticated analyses may still fail to resolve them.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Tetraconatan phylogeny with special focus on Malacostraca and Branchiopoda—Highlighting the strength of taxon-specific matrices in phylogenomics

Understanding the evolution of Tetraconata or Pancrustacea —the clade that includes crustaceans and insects—requires a well-resolved hypothesis regarding the relationships within and among its constituent taxa. Herein, we assembled a taxon-rich phylogenomic data set focusing on crustacean lineages based solely on genomes and new-generation Illumina-generated transcriptomes, including 89 representatives of Tetraconata. This constitutes the first phylogenomic study specifically addressing internal relationships of Malacostraca (with 26 species included) and Branchiopoda (36 species). Seven matrices comprising 81 to 684 orthogroups and 17,690 to 242,530 amino acid positions were assembled and analysed under five different analytical approaches. To maximize gene occupancy and to improve resolution, taxon-specific matrices were designed for Malacostraca and Branchiopoda. Key tetraconatan taxa (i.e., Oligostraca, Multicrustacea, Branchiopoda, Malacostraca, Thecostraca, Copepoda, Hexapoda) were monophyletic and well supported. Within Branchiopoda Phyllopoda, Diplostraca, Cladoceromorpha and Cladocera were monophyletic. Within Malacostraca the clades Eumalacostraca, Decapoda and Reptantia were well supported. Recovery of Caridoida or Peracarida was highly depending on the analysis for the complete matrix but were consistently recovered monophyletic in the malacostracan-specific matrix. From such examples, we demonstrate that taxon-specific matrices and particular evolutionary models and analytical methods, namely CAT-GTR and Dayhoff recoding, outperform other approaches in resolving certain recalcitrant nodes in phylogenomic analyses.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Phylogenomic analyses of more than 4000 nuclear loci resolve the origin of snakes among lizard families

Squamate reptiles (lizards and snakes) are the most diverse group of terrestrial vertebrates, with more than 10 000 species. Despite considerable effort to resolve relationships among major squamates clades, some branches have remained difficult. Among the most vexing has been the placement of snakes among lizard families, with most studies yielding only weak support for the position of snakes. Furthermore, the placement of iguanian lizards has remained controversial. Here we used targeted sequence capture to obtain data from 4178 nuclear loci from ultraconserved elements from 32 squamate taxa (and five outgroups) including representatives of all major squamate groups. Using both concatenated and species-tree methods, we recover strong support for a sister relationship between iguanian and anguimorph lizards, with snakes strongly supported as the sister group of these two clades. These analyses strongly resolve the difficult placement of snakes within squamates and show overwhelming support for the contentious position of iguanians. More generally, we provide a strongly supported hypothesis of higher-level relationships in the most species-rich tetrapod clade using coalescent-based species-tree methods and approximately 100 times more loci than previous estimates.

opencc-zeroDec 2016View details →
dryad28/100

Data from: The evolution of peafowl and other taxa with ocelli (eyespots): a phylogenomic approach

The most striking feature of peafowl (Pavo) is the males' elaborate train, which exhibits ocelli (ornamental eyespots) that are under sexual selection. Two additional genera within the Phasianidae (Polyplectron and Argusianus) exhibit ocelli, but the appearance and location of these ornamental eyespots exhibit substantial variation among these genera, raising the question of whether ocelli are homologous. Within Polyplectron, ocelli are ancestral, suggesting ocelli may have evolved even earlier, prior to the divergence among genera. However, it remains unclear whether Pavo, Polyplectron and Argusianus form a monophyletic clade in which ocelli evolved once. We estimated the phylogeny of the ocellated species using sequences from 1966 ultraconserved elements (UCEs) and three mitochondrial regions. The three ocellated genera did form a strongly supported clade, but each ocellated genus was sister to at least one genus without ocelli. Indeed, Polyplectron and Galloperdix, a genus not previously suggested to be related to any ocellated taxon, were sister genera. The close relationship between taxa with and without ocelli suggests multiple gains or losses. Independent gains, possibly reflecting a pre-existing bias for eye-like structures among females and/or the existence of a simple mutational pathway for the origin of ocelli, appears to be the most likely explanation.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record