Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,344
datasets available to search
ShareScore release 0.9.0
Dataset results
1,344 results for “phylogenomics”
Fig. 5 in Phylogenomic Species Delimitation, Taxonomy, and 'Bird Guide' Identification for the Neotropical Ant Genus Rasopone (Hymenoptera: Formicidae)
Fig. 5. Male of Rasopone mesoamericana sp. nov. (Nicaragua, CASENT0627722). Scale bars are 0.5 mm for face view, 1.0 mm for dorsal and lateral views.
Fig. 3. Phylogenetic relationships among a in Phylogenomic Species Delimitation, Taxonomy, and 'Bird Guide' Identification for the Neotropical Ant Genus Rasopone (Hymenoptera: Formicidae)
Fig. 3. Phylogenetic relationships among a curated set of COI barcode sequences for Rasopone. Black samples were sequenced for UCEs. Red samples were downloaded from the BOLD database.The tree was inferred using IQ-TREE with the data partitioned by codon position. Black circles on nodes indicate high support, which we define as ≥95% ultrafast bootstrap support and ≥95% SH-like branch support.Terminal names match taxonomic changes proposed in paper and provide useful sample identifiers (e.g., extraction codes [EX#] or BOLD process IDs).A complete, unpruned COI tree is available in Supp Fig. S1 (online only).
Phylogenomics indicates Amazonia as the major source of Neotropical swarm-founding social wasp diversity
The Neotropical realm harbors unparalleled species richness and hence has challenged biologists to explain the cause of its high biotic diversity. Empirical studies to shed light on the processes underlying biological diversification in the Neotropics are focused mainly on vertebrates and plants, with little attention to the hyperdiverse insect fauna. Here, we use phylogenomic data from ultraconserved element (UCE) loci to reconstruct for the first time the evolutionary history of Neotropical swarm-founding social wasps (Hymenoptera, Vespidae, Epiponini). Using maximum likelihood, Bayesian, and species tree approaches we recovered a highly resolved phylogeny for epiponine wasps. Additionally, we estimated divergence dates, diversification rates, and the biogeographic history for these insects in order to test whether the group followed a "museum" (speciation events occurred gradually over many millions of years) or "cradle" (lineages evolved rapidly over a short time period) model of diversification. The origin of many genera and all sampled extant Epiponini species occurred during the Miocene and Plio-Pleistocene. Moreover, we detected no major shifts in the estimated diversification rate during the evolutionary history of Epiponini, suggesting a relatively gradual accumulation of lineages with low extinction rates. Several lines of evidence suggest that the Amazonian region played a major role in the evolution of Epiponini wasps. This spatio-temporal diversification pattern, most likely concurrent with climatic and landscape changes in the Neotropics during the Miocene and Pliocene, establishes the Amazonian region as the major source of Neotropical swarm-founding social wasp diversity.
Phylogenomic analyses of 142 prokaryotic genera
<p>This repository contains 142 tar archive files, each corresponding to a prokaryotic genus. Each archive contains the following files/directories:</p> <ul> <li><code>accn.tax.tsv </code> a tab-delimited file containing the assembly accession (col 1) and an associated taxon name (col 2) for each selected genome (one per line)</li> <li><code>gff/ </code> a directory containing a gzip-compressed <a href="https://m.ensembl.org/info/website/upload/gff3.html">GFF3</a> file (as returned by the annotation tool <a href="https://github.com/tseemann/prokka"><em>Prokka</em></a>) for each genome specified in <code>accn.tax.tsv</code></li> <li><code>cds.fna.gz </code> a gzip-compressed FASTA file containing all the coding sequences (at the codon level) from the GFF3 files in the directory <code>gff/</code></li> <li><code>msa/ </code> a directory containing multiple amino acid and codon sequence alignments (compressed FASTA files with extensions .afa.gz and .afc.gz, respectively) for each cluster of at least four homologous sequences (as determined by the pipeline <a href="https://github.com/sanger-pathogens/Roary"><em>Roary</em></a> from the GFF3 files in <code>gff/</code>)</li> <li><code>supermatrix.fasta.gz </code> a gzip-compressed FASTA file obtained by concatenating all the multiple codon sequence alignments in the directory <code>msa/</code></li> <li><code>tree.nwk </code> a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing a maximum-likelihood (ML) phylogenetic tree inferred from the file <code>supermatrix.fasta.gz</code> using <a href="http://www.iqtree.org/"><em>IQ-TREE</em></a></li> <li><code>iqtree.txt </code> a txt file summarizing the ML estimates of the GTR+Γ evolutionary model parameters, as returned by <em>IQ-TREE</em> when inferring the phylogenetic tree <code>tree.nwk</code></li> </ul> <p> </p> <p>A summary of the 142 phylogenomic analyses can be found in the tab-delimited file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1"><code>GTR.params.tree.tsv</code></a>. Each line corresponds to one genus and contains the 12 following fields:<br> <code>[1] </code> genus name,<br> <code>[2-5] </code> frequencies of T, C, A, G, respectively,<br> <code>[6-10]</code> C-T, A-T, G-T, A-C, C-G rate parameters, respectively (normalized such that A-G rate = 1),<br> <code>[11] </code> Γ shape parameter alpha,<br> <code>[12] </code> Newick-formatted phylogenetic tree.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Phylogenomics reveals an extensive history of genome duplication in diatoms (Bacillariophyta)
<p>Abstract</p> <p>Premise of the Study</p> <p>Diatoms are one of the most species‐rich lineages of microbial eukaryotes. Similarities in clade age, species richness, and primary productivity motivate comparisons to angiosperms, whose genomes have been inordinately shaped by whole‐genome duplication (WGD). WGDs have been linked to speciation, increased rates of lineage diversification, and identified as a principal driver of angiosperm evolution. We synthesized a large but scattered body of evidence that suggests polyploidy may be common in diatoms as well.</p> <p>Methods</p> <p>We used gene counts, gene trees, and distributions of synonymous divergence to carry out a phylogenomic analysis of WGD across a diverse set of 37 diatom species.</p> <p>Key Results</p> <p>Several methods identified WGDs of varying age across diatoms. Determining the occurrence, exact number, and placement of events was greatly impacted by uncertainty in gene trees. WGDs inferred from synonymous divergence of paralogs varied depending on how redundancy in transcriptomes was assessed, gene families were assembled, and synonymous distances (Ks) were calculated. Our results highlighted a need for systematic evaluation of key methodological aspects of Ks‐based approaches to WGD inference. Gene tree reconciliations supported allopolyploidy as the predominant mode of polyploid formation, with strong evidence for ancient allopolyploid events in the thalassiosiroid and pennate diatom clades.</p> <p>Conclusions</p> <p>Our results suggest that WGD has played a major role in the evolution of diatom genomes. We outline challenges in reconstructing paleopolyploid events in diatoms that, together with these results, offer a framework for understanding the impact of genome duplication in a group that likely harbors substantial genomic diversity.</p>
Do alignment and trimming methods matter for phylogenomic (UCE) analyses?
Alignment is a crucial issue in molecular phylogenetics because different alignment methods can potentially yield very different topologies for individual genes. But it is unclear if the choice of alignment methods remains important in phylogenomic analyses, which incorporate data from dozens, hundreds, or thousands of genes. For example, problematic biases in alignment might be multiplied across many loci, whereas alignment errors in individual genes might become irrelevant. The issue of alignment trimming (i.e. removing poorly aligned regions or missing data from individual genes) is also poorly explored. Here, we test the impact of 12 different combinations of alignment and trimming methods on phylogenomic analyses. We compare these methods using published phylogenomic data from ultraconserved elements (UCEs) from squamate reptiles (lizards and snakes), birds, and tetrapods. We compare the properties of alignments generated by different alignment and trimming methods (e.g., length, informative sites, missing data). We also test whether these datasets can recover well-established clades when analyzed with concatenated (RAxML) and species-tree methods (ASTRAL-III), using the full data (~5,000 loci) and subsampled datasets (10% and 1% of loci). We show that different alignment and trimming methods can significantly impact various aspects of phylogenomic datasets (e.g. length, informative sites). However, these different methods generally had little impact on the recovery and support values for well-established clades, even across very different numbers of loci. Nevertheless, our results suggest several "best practices" for alignment and trimming. Intriguingly, the choice of phylogenetic methods impacted the results most strongly, with concatenated analyses recovering significantly more well-established clades (with stronger support) than the species-tree analyses.
New genetic markers for Sapotaceae phylogenomics: more than 600 nuclear genes applicable from family to population levels
<p>Some tropical plant families, such as the Sapotaceae, have a complex taxonomy, which can be resolved using Next Generation Sequencing (NGS). For most groups however, methodological protocols are still missing. Here we identified 531 monocopy genes and 227 Short tandem repeats (STR) markers and tested them on Sapotaceae using target capture and NGS. The probes were designed using two genome skimming samples from<em>Capurodendron delphinense</em> and <em>Bemangidia lowryi</em>, both from the Tseboneae tribe, as well as the published <em>Manilkara zapota</em> transcriptome from the Sapotoideae tribe. We combined our probes with 261 additional ones previously published and designed for the entire angiosperm group. On a total of 792 low-copy genes, 638 showed no signs of paralogy and were used to build a phylogeny of the family with 231 individuals from all main lineages. A highly supported topology was obtained at high taxonomic ranks but also at the species level. This phylogeny revealed the existence of more than 20 putative new species. Single nucleotide polymorphisms (SNPs) extracted from the 638 genes were able to distinguish lineages within a species complex and to highlight geographical structuration. STR were recovered efficiently for the species used as reference (<em>C. delphinense</em>) but the recovery rate decreased dramatically with the phylogenetic distance to the focal species. All together, the new loci will help reaching a sound taxonomic understanding of the family Sapotaceae for which many circumscriptions and relationships are still debated, at the species, genus and tribe levels.</p>
Data from: Phylogenomic insights into the evolution of stinging wasps and the origins of ants and bees
The stinging wasps (Hymenoptera: Aculeata) are an extremely diverse lineage of hymenopteran insects, encompassing over 70,000 described species and a diversity of life history traits, including ectoparasitism, cleptoparasitism, predation, pollen feeding (bees [Anthophila] and Masarinae) and eusociality (social vespid wasps, ants, and some bees) [1]. The most well-studied lineages of Aculeata are the ants, which are ecologically dominant in most terrestrial ecosystems [2], and the bees, the most important lineage of angiosperm-pollinating insects [3]. Establishing the phylogenetic affinities of ants and bees helps us understand and reconstruct patterns of social evolution as well as fully appreciate the biological implications of the switch from carnivory to pollen feeding (pollenivory). Despite recent advancements in aculeate phylogeny [4–11], considerable uncertainty remains regarding higher level relationships within Aculeata, including the phylogenetic affinities of ants and bees [5–7]. We used ultraconserved element (UCE) phylogenomics [7,12] to resolve relationships among stinging wasp families, gathering sequence data from > 800 UCE loci and 187 samples, including 30 out of 31 aculeate families. We analyzed the 187-taxon data set using multiple analytical approaches, and we evaluated several alternative taxon sets. We also tested alternative hypotheses for the phylogenetic positions of ants and bees. Our results present a highly supported phylogeny of the stinging wasps. Most importantly, we find unequivocal evidence that ants are the sister group to bees+apoid wasps (Apoidea) and that bees are nested within a paraphyletic Crabronidae. We also demonstrate that taxon choice can fundamentally impact tree topology and clade support in phylogenomic inference.
Data from: A phylogenomic approach to clarifying the relationship of Mesodinium within the Ciliophora: a case study in the complexity of mixed-species transcriptome analyses
<p>Recent high-throughput sequencing endeavors have yielded multi-gene/protein phylogenies that confidently resolve several inter- and intra-class relationships within the phylum Ciliophora. We leverage the massive sequencing efforts from the Marine Microbial Eukaryote Transcriptome Sequencing Project, other SRA submissions, and available genome data with our own sequencing efforts to determine the phylogenetic position of <i>Mesodinium</i> and to generate the most taxonomically-rich phylogenomic ciliate tree to date. Regardless of the data mining strategy, the multi-protein dataset, or the molecular models of evolution employed, we consistently recovered the same well-supported relationships among ciliate classes, confirming many of the higher-level relationships previously identified. <i>Mesodinium</i> always formed a monophyletic group with members of the Litostomatea, with mixotrophic species of <i>Mesodinium</i> – <i>M. rubrum</i>, <i>M. major</i>, and <i>M. chamaeleon</i> - being more closely related to each other than to the heterotrophic member, <i>M. pulex</i>. The well-supported position of <i>Mesodinium</i> as sister to other litostomes contrasts with previous molecular analyses including those from phylogenomic studies that exploited the same transcriptomic databases. These topological discrepancies illustrate the need for caution when mining mixed-species transcriptomes and indicate that identifying ciliate sequences among prey contamination - particularly for <i>Mesodinium</i> species where expression from stolen prey nuclei appears to dominate – requires thorough and iterative vetting with phylogenies that incorporate sequences from a large outgroup of prey.</p>
Data from: A phylogenomic approach to clarifying the relationship of Mesodinium within the Ciliophora: a case study in the complexity of mixed-species transcriptome analyses
<p>Recent high-throughput sequencing endeavors have yielded multi-gene/protein phylogenies that confidently resolve several inter- and intra-class relationships within the phylum Ciliophora. We leverage the massive sequencing efforts from the Marine Microbial Eukaryote Transcriptome Sequencing Project, other SRA submissions, and available genome data with our own sequencing efforts to determine the phylogenetic position of <i>Mesodinium</i> and to generate the most taxonomically-rich phylogenomic ciliate tree to date. Regardless of the data mining strategy, the multi-protein dataset, or the molecular models of evolution employed, we consistently recovered the same well-supported relationships among ciliate classes, confirming many of the higher-level relationships previously identified. <i>Mesodinium</i> always formed a monophyletic group with members of the Litostomatea, with mixotrophic species of <i>Mesodinium</i> – <i>M. rubrum</i>, <i>M. major</i>, and <i>M. chamaeleon</i> - being more closely related to each other than to the heterotrophic member, <i>M. pulex</i>. The well-supported position of <i>Mesodinium</i> as sister to other litostomes contrasts with previous molecular analyses including those from phylogenomic studies that exploited the same transcriptomic databases. These topological discrepancies illustrate the need for caution when mining mixed-species transcriptomes and indicate that identifying ciliate sequences among prey contamination - particularly for <i>Mesodinium</i> species where expression from stolen prey nuclei appears to dominate – requires thorough and iterative vetting with phylogenies that incorporate sequences from a large outgroup of prey.</p>
Phylogenomic analysis of cell-surface receptors and downstream signaling components in the plant lineage
<p>Here we identified cell-surface receptors and downstream signaling components from the genomes of 350 plant species. </p><p>Zip file contains:</p><p>Folder 'seqeunces for downstream signaling components' - FASTA and TREE files of the identified downstream signaling components.</p><p>Folder 'sequences for cell-surface receptors' - FASTA and TREE files of the identified cell-surface receptors.</p><p>Folder 'Specific analysis' - Contains specific analysis for the identified cell-surface receptors.</p><p>Subfolder 'ID analysis' - Contains information on ID clusters and motifs analysis in IDs.</p><p>Subfolder 'LRR motif gap analysis' - Contains information on small (10-29 aa) and large (30-90) gaps between LRR motifs in RLPs and RLKs.</p><p>Subfolder 'LRR-RLK & LRR-RLP phylogenetic analysis' - Contains FASTA and TREE files of the specific domain/regions (C3, C3-F, eJM-TM-cJM, and all) in LRR-RLPs and LRR-RLKs. This subfolder also contains the specific amino acid, charge and motif analysis in this region (see C3F-end features.xlsx).</p><p>Protein counts per species file - Contains the total number of each protein family/subfamily in each of the 350 species.</p><p>simpleToFullNames (translator file)- Translator file for the original ID of each gene. </p>
Code and sequence data pertaining to: A phylogenomic perspective on interspecific competition
<p>Evolutionary processes may have substantial impacts on community assembly, but evidence for phylogenetic relatedness as a determinant of interspecific interaction strength remains mixed. In this perspective, we consider a possible role for discordance between gene trees and species trees in the interpretation of phylogenetic signal in studies of community ecology. Modern genomic data show that the evolutionary histories of many taxa are better described by a patchwork of histories that vary along the genome rather than a single species tree. If a subset of genomic loci harbor trait-related genetic variation, then the phylogeny at these loci may be more informative of interspecific trait differences than the genome background. We develop a simple method to detect loci harboring phylogenetic signal and demonstrate its application through a proof of principle analysis of Penicillium genomes and pairwise interaction strength. Our results show that phylogenetic signal that may be masked genome-wide could be detectable using phylogenomic techniques and may provide a window into the genetic basis for interspecific interactions.</p>
Data from: Phylogenomics of American pika (Ochotona princeps) lineage diversification
<p>Quaternary climate oscillations have profoundly influenced current species distributions. For many montane species, these fluctuations were a prominent driver in species range shifts, often resulting in intraspecific diversification, as has been the case for American pikas (<em>Ochotona princeps</em>). Range shifts and population declines in this thermally-sensitive lagomorph have been linked to historical and contemporary environmental changes across its western North American range, with previous research reconstructing five mitochondrial DNA lineages. Here, we paired genome-wide data (25,244 SNPs) with range-wide sampling to re-examine the number and distribution of intra-specific lineages, and investigate patterns of within- and among-lineage divergence and diversity. Our results provide genomic evidence of <em>O. princeps</em> monophyly, reconstructing six distinct lineages that underwent multiple rounds of divergence (0.809-2.81mya), including a new Central Rocky Mountain lineage. We further found evidence for population differentiation across multiple spatial scales, and reconstructed levels of standing variation comparable to those found in other small mammals. Overall, our findings demonstrate the influence of past glacial cycles on <em>O. princeps</em> lineage diversification, suggest that current subspecific taxonomy may need to be revisited, and provide an important framework for investigations of American pika adaptive potential in the face of anthropogenic climate change.</p>
Data from: Revised evolutionary and taxonomic synthesis for parrots (order: Psittaciformes) guided by phylogenomic analysis
<p>Parrots (Order: Psittaciformes) are a diverse clade that are easily distinguishable from other birds. Despite the clear characters that define the Psittaciformes (hooked bills, zygodactylous feet, and plumage that is often predominantly green or red), relative morphological uniformity among parrots has made taxonomic classification a fraught endeavor for over a century. Parrot systematics were propelled forward when DNA sequencing data shed insights into higher- and species-level relationships. However, despite these significant advances, major gaps in taxon sampling and uncertainty in relationships remained due to inferring phylogenetic relationships with short fragments of DNA. Recent work using genome-wide molecular markers with nearly complete parrot species-level sampling has brought clarity to many of the remaining outstanding questions on taxonomic relationships. Here, we build on this work by including four additional species to present a taxonomic revision of Psittaciformes better aligned with its evolutionary tree. We infer maximum likelihood and time-calibrated phylogenies for parrots, present accounts for 106 genera, compare how our findings relate to previous work, and highlight future areas of research. The family-group nomenclature we propose reflects deep evolutionary divergences with diagnosable synapomorphies that are commensurate across comparable ranks in psittaciform clades. We erect three new family-group names at the rank of tribe (Brotogerini Smith, Thom and Joseph, 2024; Neophemini Schodde, Smith, Thom and Joseph, 2024; Bolbopsittacini Smith, Thom and Joseph, 2024). We elevate one tribe to subfamily rank for the cacatuid genus <em>Probosciger</em> and we restrict usage of the recently introduced tribe Touitini to its type-genus <em>Touit</em>. At shallower taxonomic scales, recognition of more rather than fewer genera addresses issues of paraphyly or high discordance in morphological and genomic characters at those levels. We support many reinstatements of older generic names advocated in recent decades and we further reinstate five valid, available generic names not widely used in recent literature if at all (<em>Licmetis</em>, <em>Gymnopsittacus</em>, <em>Clarkona</em>, <em>Suavipsitta</em>, <em>Cardeos</em>). We advocate the retention of <em>Vini</em> Lesson, 1833 over <em>Coriphilus</em> Wagler, 1832 based on preliminary examination showing substantially more frequent usage of the former. We redraw generic limits in some other cases (e.g., <em>Bolborhynchus</em> parrotlets and allies) and this includes recognizing fewer genera than recently proposed for the <em>Psittacula</em> <em>sensu lato</em> ringneck parakeets. Our revised classification of parrots addresses many longstanding taxonomic questions including those that have arisen through the acquisition of genetic data. It provides context for the temporal origins of psittaciform clades and the taxonomic and phenotypic diversification throughout their evolutionary history. We hope that it will be a benchmark guiding further taxonomic study as well as for downstream analyses in many other fields.</p>
F I G U R E 2 in Mitochondrial phylogenomics of the Australian scribbly gum moth Ogmograptis (Lepidoptera: Bucculatricidae) and an examination of deep-level relationships within Lepidoptera
F I G U R E 2 Phylogeny of the Lepidoptera inferred from mitochondrial genomes, Part B—Apoditrysia. Topology and branch lengths are from the ML-PCG12-R analysis with nodal supports, maximum likelihood (ML) bootstraps (BS) and Bayesian inference (BI) posterior probabilities (PP) mapped for all eight analyses. Nodal supports depict the range of values: <70%/0.9, 70%–89%/0.9–0.94, 90%–99%/0.95–0.99 and 100%/1.0 (see key). Branch lengths are equal to expected substitutions/site. The complete trees for each analysis including precise nodal support values are included in Figures S7–S14.
F I G U R E 1 in Mitochondrial phylogenomics of the Australian scribbly gum moth Ogmograptis (Lepidoptera: Bucculatricidae) and an examination of deep-level relationships within Lepidoptera
F I G U R E 1 Phylogeny of the Lepidoptera inferred from mitochondrial genomes, Part A—non-Apoditrysia. Topology and branch lengths are from the ML-PCG12-R analysis with nodal supports, maximum likelihood (ML) bootstraps (BS) and Bayesian inference (BI) posterior probabilities (PP) mapped for all eight analyses. Nodal supports depict the range of values: <70%/0.9, 70%–89%/0.9–0.94, 90%–99%/0.95–0.99 and 100%/1.0 (see key). Branch lengths are equal to expected substitutions/site. The complete trees for each analysis including precise nodal support values are included in Figures S7–S14.
Figure 1 in Construction of a phylogenetic matrix: Scripts and guidelines for phylogenomics
Figure 1. Flowchart of constructing a phylogenetic matrix for phylogenomics. The custom scripts used in each step are marked as italic. Dashed boxes indicate that these strategies of each step choose only one or more suitable strategies.
Supplementary materials for Phylogenomics and genetic analysis of solvent-producing Clostridium species
<p>The genus <em>Clostridium</em> is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 <em>C. beijerenckiI, </em>57 <em>C. saccharobutylicum</em>, 4 <em>C. saccharoperbutylacetonicum</em>, 5 <em>C. butyricum</em>, 7 <em>C. acetobutylicum</em>, and 3 <em>C. tetanomorphum </em>genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or <em>Escherichia coli </em>for enhanced chemical production. Recently, gene sequences from this collection were used to engineer <em>Clostridium autoethanogenum</em>, a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production<em>, </em>as well as butanol, butanoic acid, hexanol and hexanoic acid production. </p>
Reference genome choice and filtering thresholds jointly influence phylogenomic analyses
<p>Molecular phylogenies are a cornerstone of modern comparative biology and are commonly employed to investigate a range of biological phenomena, such as diversification rates, patterns in trait evolution, biogeography, and community assembly. Recent work has demonstrated that significant biases may be introduced into downstream phylogenetic analyses from processing genomic data; however, it remains unclear whether there are interactions among bioinformatic parameters or biases introduced through the choice of reference genome for sequence alignment and variant-calling. We address these knowledge gaps by employing a combination of simulated and empirical data sets to investigate to what extent the choice of reference genome in upstream bioinformatic processing of genomic data influences phylogenetic inference, as well as the way that reference genome choice interacts with bioinformatic filtering choices and phylogenetic inference method. We demonstrate that more stringent minor allele filters bias inferred trees away from the true species tree topology, and that these biased trees tend to be more imbalanced and have a higher center of gravity than the true trees. We find the greatest topological accuracy when filtering sites for minor allele count > 3–4 in our 51-taxa data sets, while tree center of gravity was closest to the true value when filtering for sites with minor allele count > 1-2. In contrast, filtering for missing data increased accuracy in the inferred topologies; however, this effect was small in comparison to the effect of minor allele filters and may be undesirable due to a subsequent mutation spectrum distortion. The bias introduced by these filters differs based on the reference genome used in short read alignment, providing further support that choosing a reference genome for alignment is an important bioinformatic decision with implications for downstream analyses. These results demonstrate that attributes of the study system and dataset (and their interaction) add important nuance for how best to assemble and filter short read genomic data for phylogenetic inference.</p>
Data for phylogenomic analysis of chelicerate gene family evolution
<p>We used phylogenomics to investigate patterns of gene family evolution across ticks and other chelicerates, which include a diverse array of parasites. We used phylogenetic profiling and trait-association tests to predict gene families that may enable parasitic species to feed on hosts undetected for prolonged periods (>1 day). This release accompanies the pub, “<a href="https://doi.org/10.57844/arcadia-4e3b-bbea">Comparative phylogenomic analysis of Chelicerates points to gene families associated with long-term suppression of host detection</a>." Please see the pub for more information.</p> <ul> <li>chelicerata-v1-10062023.zip contains the outputs from NovelTree that are needed as inputs for phylogenetic profiling.</li> <li>annotated.zip contains gene annotations used to do orthogroup filtering.</li> <li>tx2gene.tsv has presence/absence of expression for each Amblyomma americanum transcript. </li> <li>chelicerate_proteome_preprocessing_outputs.zip contains the outputs of chelicerate protein data curation.</li> <li>chelicerata-v1-parameterfile.json & chelicerata-v1-samplesheet.csv were inputs for setting up the initial NovelTree run.</li> <li>2024-06-24-all-chelicerate-noveltree-proteins.fasta has the full set of chelicerate protein sequences.</li> <li>summary_of_noveltree_results.zip contains summary figures from the outputs of the NovelTree run.</li> <li>chelicerate-samples.tsv is the sample sheet used in proteome curation upstream of NovelTree.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.