Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
43
datasets available to search
ShareScore release 0.9.0
Dataset results
43 results for “lineage tree”
Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages
<p><strong>Abstract</strong></p> <p>Advances in phylogenetic methods and high-throughput sequencing have allowed the reconstruction of deep phylogenetic relationships in the evolutionary history of eukaryotes. Yet, the root of the eukaryotic tree of life remains elusive. The most ‘popular’ (i.e. in textbooks and reviews) hypothesis for the root is between Unikonta (Opisthokonta + Amoebozoa) and Bikonta (all other eukaryotes), which emerged from analyses of a single gene fusion and a limited sampling of eukaryotic lineages. Subsequent highly-cited studies based on concatenation of genes supported this hypothesis with some variations or proposed a root within the Excavata. However, concatenation of genes neither considers phylogenetically-informative events (i.e. gene duplications and losses), nor provides an estimate of the root. A more recent study using gene tree-species tree reconciliation methods suggested the root lies between Opisthokonta and all other eukaryotes, but only including 59 taxa and 20 genes. Here we apply a gene tree – species tree reconciliation approach to a gene-rich and taxon-rich dataset (i.e. 2,786 gene families from two sets of ~158 diverse eukaryotic lineages) to assess the root, and we iterate each analysis 100 times to quantify tree space uncertainty. Our results estimate a root between Fungi and all other eukaryotes, or between Opisthokonta and all other eukaryotes, and reject alternative popular roots from the literature. Based on further analysis of genome size we propose Opisthokonta + others as the most likely root. Finding the root of the eukaryotic tree of life is critical for the field of comparative biology as it allows us to understand the timing and mode of evolution of characters across the evolutionary history of eukaryotes.</p> <p>Methods<br> Here we provide the alignments, gene trees, inputs, and outputs from our project entitled "Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages". Sequences and alignments were produced using the phylogenomic pipeline PhyloToL, which contains a taxon- and gene-rich database (including eukaryotes, archaea, and bacteria). These data were then used for 1) assessing the root of the eukaryotes and 2) for comparison with other previously published hypotheses. In both cases, we used the species tree - gene tree reconciliation tool iGTP. We also did a comparison of hypotheses using the likelihood-based tool SpeciesRax</p> <p><strong>iGTP</strong></p> <p>Input data</p> <p>These data are divided into four datasets based on taxa selection. For dataset SEL+, taxa were selected based on their taxonomy; for RAN+, taxa were selected randomly among the major eukaryotic clades Opisthokonta, Amoebozoa, Archaeplastida, Excavata, SAR, and some orphan lineages. Datasets SEL- and RAN- are the same as SEL+ and RAN+, but exclude microsporidians in order to account for and avoid long branch attraction due to microsporidians fast-evolutionary rates. We chose the gene families that contain at least 25 taxa representing at least four of the five major eukaryotic clades. Additionally, at least 2 of the major clades had to contain at least 2 minor clades (e.g. Glaucophytes and Rhodophyta are minor clades in the major clade Archaeplastida). In a pilot analysis, we produced an alignment and a phylogenetic tree for each gene family using the default settings of a previous version of PhyloToL (GUIDANCE V1.3.1 sequence cutoff = 0.3 and column cutoff = 0.4; RAxML quick tree with model PROTGAMMALG and no bootstraps). Then, we kept the gene families that are exclusive of eukaryotes or the ones in which eukaryotes were monophyletic. From a total of 3,002 gene families that met our criteria, 2786 passed the initial steps of PhyloToL when including only the data from the dataset SEL+. These 2,786 gene families were used for further analyses with all datasets.</p> <p>MSAs were produced with PhyloToL (GUIDANCE V2.02 sequence cutoff = 0.3, column cutoff = 0.4, number of iterations = 5; Sela, et al. 2015). The default parameters of PhyloToL include up to five iterations of GUIDANCE V2.02 with 10 bootstraps and MAFFT V7 with algorithm E-INS-i for less than 200 sequences or “auto” option if more than 200 sequences, and maxiterate = 1000. Instead, here we run up to five iterations of GUIDANCE with 20 bootstraps and the simple MAFFT algorithm FFT-NS-2. Then, we perform an additional GUIDANCE run with 100 bootstraps and the default MAFFT parameters for PhyloToL.</p> <p>Gene trees were inferred with RAxML v.8.2.4 with 10 ML searches for best-ML tree (option "-# 10"), using the rapid hill-climbing algorithm (option "-f d") and no bootstrap replicates. The protein evolution model used was evaluated during the gene tree inference (option "-m PROTCATAUTO") by testing all models available in RAxML (e.g. JTT, LG, WAG, etc) with optimization of substitution rates and of site-specific evolutionary rates which were categorized into four distinct rate categories for greater computational efficiency.</p> <p>We ran 100 repetitions of iGTP analyses per dataset. But, given the complexity of the datasets and the heuristic nature of some key steps of the iGTP algorithm (e.g. gene tree rooting and initial starting species tree generation), in a preliminary analysis, we faced two systematic challenges with iGTP as the inferred species tree was affected by: 1) the order of the leaves in the input unrooted gene tree Newick strings (i.e. the input trees were treated as rooted even though we specified that they were not); and 2) the input gene order in the 100 replicates. Therefore, we randomly shuffled the order of the leaves in the unrooted gene trees (keeping the same topology), and randomly shuffled the order of the input gene trees in each of the 100 replicates per dataset. Here we provide the 100 input files generated for those iGTP analyses. </p> <p>Output data</p> <p>Here we also share the data generated after two analyses: 1) root assessment and 2) hypothesis testing. For the former, we allowed iGTP to calculate the more parsimonious root given our input files. For the latter, we allowed iGTP to calculate the reconciliation cost of the gene trees given the input files and constraints in the species trees to reflect previously published root hypotheses. The constraints are explained in the README file.</p> <p>SpeciesRax</p> <p>Since we removed LGT and contamination from our dataset using a series of filters, we applied the model UndatedDL instead of UndatedDTL, which implies that we only took into consideration duplications and losses and ignored the transferences. Then, the command used for SpeciesRax was...</p> <p>./generax --families forGeneRax/famFile --species-tree forGeneRax/spsTreeAn.newick --strategy SKIP --rec-model UndatedDL --per-family-rates --prefix An_DS1 --si-strategy EVAL </p> <p>input </p> <p>Here, we are sharing all the necessary files to run SpeciesRax, including the mapping file (famFile), the species trees (spsTrees; the best iGTP constrained species trees per hypothesis), and their underlying gene trees (trees_r)</p> <p>output</p> <p>Output folder from SpeciesRax, which includes log files, events (duplications, losses) counts, and statistics (i.e., reconciliation likelihood values)</p> <p>Note: <br> * As in the iGTP analyses, for the SpeciesRax files, the words Op, Fu, Di, Un, An, refer to the five root hypotheses compared: Opisthokonta-others, Fungi-others, Discoba-others, Unikonta-Bikonta, and (Ancyromonadida + Metamonada)-others, respectively. <br> * For the SpeciesRax analyses, the Un word refers to Ut (Cavalier-Smith 2003) <br> </p>
Fig. 1. Bayesian phylogenetic tree constructed using partial cytochrome b in Unexpected absence of exo-erythrocytic merogony during high gametocytaemia in two species of Haemoproteus (Haemosporida: Haemoproteidae), including description of Haemoproteus angustus n. sp. (lineage hCWT7) and a report of previously unknown residual bodies during in vitro gametogenesis
Fig. 1. Bayesian phylogenetic tree constructed using partial cytochrome b sequences of 61 lineages of Haemoproteus, 4 lineages of Plasmodium, and Leucocytozoon sp. lSISKIN2 as outgroup. Posterior probabilities higher than 0.8 are indicated close to the respective nodes. Red font indicates the parasite lineage described in this publication. Vertical bars (A–D) show groups of closely related lineages, which complete development and produce gametocytes only in non-passerines (A, D), both non-passerines and passerines (B), and only passerines (C). Blue font indicates Haemoproteus species, which develop in non-passerine avian hosts, which are indicated by symbols (● – Psittaciformes; ∎ - Coraciiformes; ▴ - Strigiformes; ◆ - Anseriformes; ★ - Charadriiformes; ♥ - Pelecaniformes; ⋄ - Piciformes; ⊠ - Sphenisciformes; Ω - Musophagiformes; § - Trochiliformes; Ψ – Falconiformes; Σ – Columbiformes; Φ - Galliformes). Lineage names were provided (according to MalAvi database), followed by parasite species names and sequence GenBank accession numbers.
Fig. 2. Bayesian tree reconstructed using 478 in Detection of haemosporidian parasites in wild and domestic birds in northern and central provinces of Iran: Introduction of new lineages and hosts
Fig. 2. Bayesian tree reconstructed using 478-bp mitochondrial cytb gene for avian blood parasites lineages. The amplified sequences in the current study are highlighted in bold. Posterior probability support of>0.8 is displayed for each branch. Schematic tree is summarized in section A and separated clade for each genus is given in sections of B (Plasmodium), C (Haemoproteus), and D (Leucocytozoon).
Fig. 2. Majority rule tree from the 73 in The European Early Cretaceous cryptodiran turtle Chitracephalus dumonii and the diversity of a poorly known lineage of turtles
Fig. 2. Majority rule tree from the 73 most parsimonious trees produced by the cladistic analysis of Chitracephalus dumonii using the modified data set of Joyce (2007) proposed in Pérez−García et al. (2012). Retention index (RI) = 0.872 and consistency index (CI) = 0.567. Values refer to percentages under 100% obtained in the majority rule analysis; those with values below 50% are collapsed. Letters refer to the nodes mentioned in the text.
Data from: Species tree branch length estimation despite incomplete lineage sorting, duplication, and loss
Open the record for dataset details and reuse information.
The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves (metazoa data)
<p>This dataset is associated to the following publication: <strong>Macé, B.</strong>, Mouillot, D., Dalongeville, A., Bruno, M., Deter, J., Varenne, A., Gudefin, A., Boissery, P., & Manel, S. (<strong>2024</strong>). The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves. <em>Molecular Ecology</em>, e17373. <a href="https://doi.org/10.1111/mec.17373">https://doi.org/10.1111/mec.17373</a></p> <p>It contains the data obtained with the <strong>metazoa</strong> marker:</p> <ul> <li><em>fastq</em> files are the raw NGS eDNA sequencing outputs</li> <li><em>dat</em> file records the adapters names and oligos used for sequencing</li> </ul> <p>Metadata associated to each eDNA sample are also provided.</p> <p> </p> <p><strong>Methods</strong></p> <blockquote> <p>eDNA extractions were performed in a BSL-2 lab dedicated for eDNA samples following the protocol described in Polanco Fernández et al. (2021). Four PCR amplifications were conducted with different assays covering the whole tree of life. The teleo primer pair (Valentini et al., 2016) targets a 12S mitochondrial DNA marker from teleosts and elasmobranchs; the metazoa primer pair (Kelly et al., 2016) targets a 16S mitochondrial DNA marker from metazoans; the euka2 primer pair (Guardiola et al., 2015) targets a marker from eukaryotes located on the V7 region of the 18S ribosomal RNA; and the bact2 primer pair (Taberlet et al., 2018) targets a marker from prokaryotes located on the V4 region of the 16S ribosomal RNA. The idea of this experimental design is to give a holistic overview of communities, with a nested hierarchy euka2-metazoa-teleo to obtain a finer taxonomic resolution over animal communities, and particularly fish. Twelve PCR replicates per sample were run, with negative extractions and PCR positive and negative controls analyzed in parallel. Unique tags were used for each PCR replicate amplified with the teleo primers only, allowing to differentiate them in the bioinformatic analysis (see after). NGS library preparation and MiSeq paired-end sequencing (2 × 150 bp) were performed at DNA Gensee (Le Bourget-du-Lac, France).</p> </blockquote> <p> </p> <p><strong>References</strong></p> <p>Guardiola, M., Uriz, M. J., Taberlet, P., Coissac, E., Wangensteen, O. S., & Turon, X. (2015). Deep-Sea, Deep-Sequencing: Metabarcoding Extracellular DNA from Sediments of Marine Canyons. <em>PLOS ONE</em>, <em>10</em>(10), e0139633. https://doi.org/10.1371/journal.pone.0139633</p> <p>Kelly, R. P., O’Donnell, J. L., Lowell, N. C., Shelton, A. O., Samhouri, J. F., Hennessey, S. M., Feist, B. E., & Williams, G. D. (2016). Genetic signatures of ecological diversity along an urbanization gradient. <em>PeerJ</em>, <em>4</em>, e2444. https://doi.org/10.7717/peerj.2444</p> <p>Polanco Fernández, A., Marques, V., Fopp, F., Juhel, J.-B., Borrero-Pérez, G. H., Cheutin, M.-C., Dejean, T., González Corredor, J. D., Acosta-Chaparro, A., Hocdé, R., Eme, D., Maire, E., Spescha, M., Valentini, A., Manel, S., Mouillot, D., Albouy, C., & Pellissier, L. (2021). Comparing environmental DNA metabarcoding and underwater visual census to monitor tropical reef fishes. <em>Environmental DNA</em>, <em>3</em>(1), 142–156. https://doi.org/10.1002/edn3.140</p> <p>Taberlet, P., Bonin, A., Zinger, L., & Coissac, E. (2018). <em>Environmental DNA: For Biodiversity Research and Monitoring</em>. Oxford University Press.</p> <p>Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., Bellemain, E., Besnard, A., Coissac, E., Boyer, F., Gaboriaud, C., Jean, P., Poulet, N., Roset, N., Copp, G. H., Geniez, P., Pont, D., Argillier, C., Baudoin, J.-M., … Dejean, T. (2016). Next-generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. <em>Molecular Ecology</em>, <em>25</em>(4), 929–942. https://doi.org/10.1111/mec.13428</p>
The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves (bact2 data)
<p>This dataset is associated to the following publication: <strong>Macé, B.</strong>, Mouillot, D., Dalongeville, A., Bruno, M., Deter, J., Varenne, A., Gudefin, A., Boissery, P., & Manel, S. (<strong>2024</strong>). The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves. <em>Molecular Ecology</em>, e17373. <a href="https://doi.org/10.1111/mec.17373">https://doi.org/10.1111/mec.17373</a></p> <p>It contains the data obtained with the <strong>bact2</strong> marker:</p> <ul> <li><em>fastq</em> files are the raw NGS eDNA sequencing outputs</li> <li><em>dat</em> file records the adapters names and oligos used for sequencing</li> </ul> <p>Metadata associated to each eDNA sample are also provided.</p> <p> </p> <p><strong>Methods</strong></p> <blockquote> <p>eDNA extractions were performed in a BSL-2 lab dedicated for eDNA samples following the protocol described in Polanco Fernández et al. (2021). Four PCR amplifications were conducted with different assays covering the whole tree of life. The teleo primer pair (Valentini et al., 2016) targets a 12S mitochondrial DNA marker from teleosts and elasmobranchs; the metazoa primer pair (Kelly et al., 2016) targets a 16S mitochondrial DNA marker from metazoans; the euka2 primer pair (Guardiola et al., 2015) targets a marker from eukaryotes located on the V7 region of the 18S ribosomal RNA; and the bact2 primer pair (Taberlet et al., 2018) targets a marker from prokaryotes located on the V4 region of the 16S ribosomal RNA. The idea of this experimental design is to give a holistic overview of communities, with a nested hierarchy euka2-metazoa-teleo to obtain a finer taxonomic resolution over animal communities, and particularly fish. Twelve PCR replicates per sample were run, with negative extractions and PCR positive and negative controls analyzed in parallel. Unique tags were used for each PCR replicate amplified with the teleo primers only, allowing to differentiate them in the bioinformatic analysis (see after). NGS library preparation and MiSeq paired-end sequencing (2 × 150 bp) were performed at DNA Gensee (Le Bourget-du-Lac, France).</p> </blockquote> <p> </p> <p><strong>References</strong></p> <p>Guardiola, M., Uriz, M. J., Taberlet, P., Coissac, E., Wangensteen, O. S., & Turon, X. (2015). Deep-Sea, Deep-Sequencing: Metabarcoding Extracellular DNA from Sediments of Marine Canyons. <em>PLOS ONE</em>, <em>10</em>(10), e0139633. https://doi.org/10.1371/journal.pone.0139633</p> <p>Kelly, R. P., O’Donnell, J. L., Lowell, N. C., Shelton, A. O., Samhouri, J. F., Hennessey, S. M., Feist, B. E., & Williams, G. D. (2016). Genetic signatures of ecological diversity along an urbanization gradient. <em>PeerJ</em>, <em>4</em>, e2444. https://doi.org/10.7717/peerj.2444</p> <p>Polanco Fernández, A., Marques, V., Fopp, F., Juhel, J.-B., Borrero-Pérez, G. H., Cheutin, M.-C., Dejean, T., González Corredor, J. D., Acosta-Chaparro, A., Hocdé, R., Eme, D., Maire, E., Spescha, M., Valentini, A., Manel, S., Mouillot, D., Albouy, C., & Pellissier, L. (2021). Comparing environmental DNA metabarcoding and underwater visual census to monitor tropical reef fishes. <em>Environmental DNA</em>, <em>3</em>(1), 142–156. https://doi.org/10.1002/edn3.140</p> <p>Taberlet, P., Bonin, A., Zinger, L., & Coissac, E. (2018). <em>Environmental DNA: For Biodiversity Research and Monitoring</em>. Oxford University Press.</p> <p>Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., Bellemain, E., Besnard, A., Coissac, E., Boyer, F., Gaboriaud, C., Jean, P., Poulet, N., Roset, N., Copp, G. H., Geniez, P., Pont, D., Argillier, C., Baudoin, J.-M., … Dejean, T. (2016). Next-generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. <em>Molecular Ecology</em>, <em>25</em>(4), 929–942. https://doi.org/10.1111/mec.13428</p>
The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves (euka2 data)
<p>This dataset is associated to the following publication: <strong>Macé, B.</strong>, Mouillot, D., Dalongeville, A., Bruno, M., Deter, J., Varenne, A., Gudefin, A., Boissery, P., & Manel, S. (<strong>2024</strong>). The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves. <em>Molecular Ecology</em>, e17373. <a href="https://doi.org/10.1111/mec.17373">https://doi.org/10.1111/mec.17373</a></p> <p>It contains the data obtained with the <strong>euka2</strong> marker:</p> <ul> <li><em>fastq</em> files are the raw NGS eDNA sequencing outputs</li> <li><em>dat</em> file records the adapters names and oligos used for sequencing</li> </ul> <p>Metadata associated to each eDNA sample are also provided.</p> <p> </p> <p><strong>Methods</strong></p> <blockquote> <p>eDNA extractions were performed in a BSL-2 lab dedicated for eDNA samples following the protocol described in Polanco Fernández et al. (2021). Four PCR amplifications were conducted with different assays covering the whole tree of life. The teleo primer pair (Valentini et al., 2016) targets a 12S mitochondrial DNA marker from teleosts and elasmobranchs; the metazoa primer pair (Kelly et al., 2016) targets a 16S mitochondrial DNA marker from metazoans; the euka2 primer pair (Guardiola et al., 2015) targets a marker from eukaryotes located on the V7 region of the 18S ribosomal RNA; and the bact2 primer pair (Taberlet et al., 2018) targets a marker from prokaryotes located on the V4 region of the 16S ribosomal RNA. The idea of this experimental design is to give a holistic overview of communities, with a nested hierarchy euka2-metazoa-teleo to obtain a finer taxonomic resolution over animal communities, and particularly fish. Twelve PCR replicates per sample were run, with negative extractions and PCR positive and negative controls analyzed in parallel. Unique tags were used for each PCR replicate amplified with the teleo primers only, allowing to differentiate them in the bioinformatic analysis (see after). NGS library preparation and MiSeq paired-end sequencing (2 × 150 bp) were performed at DNA Gensee (Le Bourget-du-Lac, France).</p> </blockquote> <p> </p> <p><strong>References</strong></p> <p>Guardiola, M., Uriz, M. J., Taberlet, P., Coissac, E., Wangensteen, O. S., & Turon, X. (2015). Deep-Sea, Deep-Sequencing: Metabarcoding Extracellular DNA from Sediments of Marine Canyons. <em>PLOS ONE</em>, <em>10</em>(10), e0139633. https://doi.org/10.1371/journal.pone.0139633</p> <p>Kelly, R. P., O’Donnell, J. L., Lowell, N. C., Shelton, A. O., Samhouri, J. F., Hennessey, S. M., Feist, B. E., & Williams, G. D. (2016). Genetic signatures of ecological diversity along an urbanization gradient. <em>PeerJ</em>, <em>4</em>, e2444. https://doi.org/10.7717/peerj.2444</p> <p>Polanco Fernández, A., Marques, V., Fopp, F., Juhel, J.-B., Borrero-Pérez, G. H., Cheutin, M.-C., Dejean, T., González Corredor, J. D., Acosta-Chaparro, A., Hocdé, R., Eme, D., Maire, E., Spescha, M., Valentini, A., Manel, S., Mouillot, D., Albouy, C., & Pellissier, L. (2021). Comparing environmental DNA metabarcoding and underwater visual census to monitor tropical reef fishes. <em>Environmental DNA</em>, <em>3</em>(1), 142–156. https://doi.org/10.1002/edn3.140</p> <p>Taberlet, P., Bonin, A., Zinger, L., & Coissac, E. (2018). <em>Environmental DNA: For Biodiversity Research and Monitoring</em>. Oxford University Press.</p> <p>Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., Bellemain, E., Besnard, A., Coissac, E., Boyer, F., Gaboriaud, C., Jean, P., Poulet, N., Roset, N., Copp, G. H., Geniez, P., Pont, D., Argillier, C., Baudoin, J.-M., … Dejean, T. (2016). Next-generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. <em>Molecular Ecology</em>, <em>25</em>(4), 929–942. https://doi.org/10.1111/mec.13428</p>
The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves (teleo data)
<p>This dataset is associated to the following publication: <strong>Macé, B.</strong>, Mouillot, D., Dalongeville, A., Bruno, M., Deter, J., Varenne, A., Gudefin, A., Boissery, P., & Manel, S. (<strong>2024</strong>). The Tree of Life eDNA metabarcoding reveals a similar taxonomic richness but dissimilar evolutionary lineages between seaports and marine reserves. <em>Molecular Ecology</em>, e17373. <a href="https://doi.org/10.1111/mec.17373">https://doi.org/10.1111/mec.17373</a></p> <p>It contains the data obtained with the <strong>teleo</strong> marker:</p> <ul> <li><em>fastq</em> files are the raw NGS eDNA sequencing outputs</li> <li><em>dat</em> file records the adapters names and oligos used for sequencing</li> </ul> <p>Metadata associated to each eDNA sample are also provided.</p> <p> </p> <p><strong>Methods</strong></p> <blockquote> <p>eDNA extractions were performed in a BSL-2 lab dedicated for eDNA samples following the protocol described in Polanco Fernández et al. (2021). Four PCR amplifications were conducted with different assays covering the whole tree of life. The teleo primer pair (Valentini et al., 2016) targets a 12S mitochondrial DNA marker from teleosts and elasmobranchs; the metazoa primer pair (Kelly et al., 2016) targets a 16S mitochondrial DNA marker from metazoans; the euka2 primer pair (Guardiola et al., 2015) targets a marker from eukaryotes located on the V7 region of the 18S ribosomal RNA; and the bact2 primer pair (Taberlet et al., 2018) targets a marker from prokaryotes located on the V4 region of the 16S ribosomal RNA. The idea of this experimental design is to give a holistic overview of communities, with a nested hierarchy euka2-metazoa-teleo to obtain a finer taxonomic resolution over animal communities, and particularly fish. Twelve PCR replicates per sample were run, with negative extractions and PCR positive and negative controls analyzed in parallel. Unique tags were used for each PCR replicate amplified with the teleo primers only, allowing to differentiate them in the bioinformatic analysis (see after). NGS library preparation and MiSeq paired-end sequencing (2 × 150 bp) were performed at DNA Gensee (Le Bourget-du-Lac, France).</p> </blockquote> <p> </p> <p><strong>References</strong></p> <p>Guardiola, M., Uriz, M. J., Taberlet, P., Coissac, E., Wangensteen, O. S., & Turon, X. (2015). Deep-Sea, Deep-Sequencing: Metabarcoding Extracellular DNA from Sediments of Marine Canyons. <em>PLOS ONE</em>, <em>10</em>(10), e0139633. https://doi.org/10.1371/journal.pone.0139633</p> <p>Kelly, R. P., O’Donnell, J. L., Lowell, N. C., Shelton, A. O., Samhouri, J. F., Hennessey, S. M., Feist, B. E., & Williams, G. D. (2016). Genetic signatures of ecological diversity along an urbanization gradient. <em>PeerJ</em>, <em>4</em>, e2444. https://doi.org/10.7717/peerj.2444</p> <p>Polanco Fernández, A., Marques, V., Fopp, F., Juhel, J.-B., Borrero-Pérez, G. H., Cheutin, M.-C., Dejean, T., González Corredor, J. D., Acosta-Chaparro, A., Hocdé, R., Eme, D., Maire, E., Spescha, M., Valentini, A., Manel, S., Mouillot, D., Albouy, C., & Pellissier, L. (2021). Comparing environmental DNA metabarcoding and underwater visual census to monitor tropical reef fishes. <em>Environmental DNA</em>, <em>3</em>(1), 142–156. https://doi.org/10.1002/edn3.140</p> <p>Taberlet, P., Bonin, A., Zinger, L., & Coissac, E. (2018). <em>Environmental DNA: For Biodiversity Research and Monitoring</em>. Oxford University Press.</p> <p>Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., Bellemain, E., Besnard, A., Coissac, E., Boyer, F., Gaboriaud, C., Jean, P., Poulet, N., Roset, N., Copp, G. H., Geniez, P., Pont, D., Argillier, C., Baudoin, J.-M., … Dejean, T. (2016). Next-generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. <em>Molecular Ecology</em>, <em>25</em>(4), 929–942. https://doi.org/10.1111/mec.13428</p> <p> </p>
A new tree-based methodological framework to infer the evolutionary history of Mesopolyploid lineages: An application to the Brassiceae tribe (Brassicaceae)
<p>Whole genome duplication events are notably widespread in plants and this poses particular challenges for phylogenetic inference in allopolyploid lineages, i.e. lineages that result from the merging of two or more diverged genomes after interspecific hybridization. The nuclear genomes resulting from allopolyploidization contain homologous gene copies from different evolutionary origins called homoeologs, whose orthologs must be sorted out in order to reconstruct the evolutionary history of polyploid clades. In this study, we propose a methodological approach to resolve the phylogeny of allopolyploid clades focusing on mesopolyploid genomes, which experienced some level of genome reshuffling and gene fractionation across their subgenomes. To illustrate our methodological framework, we applied it to a clade belonging to the model Brassicaceae plant family, the Brassiceae tribe, that experienced a mesohexaploidy event. The dataset analysed consists of both publically available genomic sequences and new transcriptomic data according to taxa. The present methodology requires a well-annotated reference genome, for which the identification of the parental subgenome fragments has been performed (e.g. Brassica rapa and Brassica oleracea). Focusing on fully retained genes (i.e., genes for which all homoeologous gene copies inherited from the parental lineages are still present in the reference genome), the method constructs multilabelled gene trees that allow subsequent assignment of each gene copy to its diploid parental lineage. Once the orthologous copies are identified, genes from the same parental origin are concatenated and tree-building methods are used to reconstruct the species tree. This method allows resolving the phylogenetic relationships (i) among extant species within a mesopolyploid clade, (ii) among the parental lineages of a mesopolyploid lineage, and (iii) between the parental lineages and closely related extant species. We report here the first well-resolved nuclear-based phylogeny of the Brassiceae tribe.</p>
Phylogenomic analyses of 2,786 genes in 158 lineages support a root of the eukaryotic tree of life between opisthokonts and all other lineages
<p>Advances in phylogenetic methods and high-throughput sequencing have allowed the reconstruction of deep phylogenetic relationships in the evolutionary history of eukaryotes. Yet, the root of the eukaryotic tree of life remains elusive. The most 'popular' (i.e. in textbooks and reviews) hypothesis for the root is between Unikonta (Opisthokonta + Amoebozoa) and Bikonta (all other eukaryotes), which emerged from analyses of a single gene fusion and a limited sampling of eukaryotic lineages. Subsequent highly-cited studies based on concatenation of genes supported this hypothesis with some variations or proposed a root within the Excavata. However, concatenation of genes neither considers phylogenetically-informative events (i.e. gene duplications and losses) nor provides an estimate of the root. A more recent study using gene tree-species tree reconciliation methods suggested the root lies between Opisthokonta and all other eukaryotes, but only including 59 taxa and 20 genes. Here we apply a gene tree – species tree reconciliation approach to a gene-rich and taxon-rich dataset (i.e. 2,786 gene families from two sets of ~158 diverse eukaryotic lineages) to assess the root, and we iterate each analysis 100 times to quantify tree space uncertainty. Our results estimate a root between Fungi and all other eukaryotes, or between Opisthokonta and all other eukaryotes, and reject alternative popular roots from the literature. Based on further analysis of genome size, we propose Opisthokonta + others as the most likely root. Finding the root of the eukaryotic tree of life is critical for the field of comparative biology as it allows us to understand the timing and mode of evolution of characters across the evolutionary history of eukaryotes.</p>
Data for Isotype-aware Inference of B cell Clonal Lineage Trees from Single-cell Sequencing Data
<p>This is the accompanying data to the manuscript titled<em> Isotype-aware Inference of B cell Clonal Lineage Trees from Single-cell Sequencing Data</em>. To reproduce the TRIBAL output please use this <a href="https://doi.org/10.5281/zenodo.12741290">code repository</a> as the arguments and codebase may have changed since release. </p>
Data from: A 4-lineage statistical suite to evaluate the support of large-scale retrotransposon insertion data to reconstruct evolutionary trees
Open the record for dataset details and reuse information.
A new tree-based methodological framework to infer the evolutionary history of Mesopolyploid lineages: An application to the Brassiceae tribe (Brassicaceae)
Open the record for dataset details and reuse information.
Nuclear phylogenomic analyses of asterids conflict with plastome trees and support novel relationships among major lineages
Open the record for dataset details and reuse information.
Phylogenomic analyses of 2,786 genes in 158 lineages support a root of the eukaryotic tree of life between opisthokonts and all other lineages
Open the record for dataset details and reuse information.
The perfect storm: Gene tree estimation error, incomplete lineage sorting, and ancient gene flow explain the most recalcitrant ancient angiosperm clade, Malpighiales
<p>The genomic revolution offers renewed hope of resolving rapid radiations in the Tree of Life. The development of the multispecies coalescent (MSC) model and improved gene tree estimation methods can better accommodate gene tree heterogeneity caused by incomplete lineage sorting (ILS) and gene tree estimation error stemming from the short internal branches. However, the relative influence of these factors in species tree inference is not well understood. Using anchored hybrid enrichment, we generated a data set including 423 single-copy loci from 64 taxa representing 39 families to infer the species tree of the flowering plant order Malpighiales. This order includes nine of the top ten most unstable nodes in angiosperms, which have been hypothesized to arise from the rapid radiation during the Cretaceous. Here, we show that coalescent-based methods do not resolve the backbone of Malpighiales and concatenation methods yield inconsistent estimations, providing evidence that gene tree heterogeneity is high in this clade. Despite high levels of ILS and gene tree estimation error, our simulations demonstrate that these two factors alone are insufficient to explain the lack of resolution in this order. To explore this further, we examined triplet frequencies among empirical gene trees and discovered some of them deviated significantly from those attributed to ILS and estimation error, suggesting gene flow as an additional and previously unappreciated phenomenon promoting gene tree variation in Malpighiales. Finally, we applied a novel method to quantify the relative contribution of these three primary sources of gene tree heterogeneity and demonstrated that ILS, gene tree estimation error, and gene flow contributed to 15%, 52%, and 32% of the variation, respectively. Together, our results suggest that a perfect storm of factors likely influence this lack of resolution, and further indicate that recalcitrant phylogenetic relationships like the backbone of Malpighiales may be better represented as phylogenetic networks. Thus, reducing such groups solely to existing models that adhere strictly to bifurcating trees greatly oversimplifies reality, and obscures our ability to more clearly discern the process of evolution.</p>
Data from: Nuclear and chloroplast DNA phylogeography reveals Pleistocene divergence and subsequent secondary contact of two genetic lineages of the tropical rainforest tree species Shorea leprosula (Dipterocarpaceae) in Southeast Asia
Tropical rainforests in Southeast Asia have been affected by climatic fluctuations during past glacial eras. To examine how the accompanying changes in land areas and temperature have affected the genetic properties of rainforest trees in the region, we investigated the phylogeographic patterns of a widespread dipterocarp species, Shorea leprosula. Two types of DNA markers were used: expressed sequence tag-based simple sequence repeats (EST-SSRs) and chloroplast DNA (cpDNA) sequence variations. Both sets of markers revealed clear genetic differentiation between populations in Borneo and those in the Malay Peninsula and Sumatra (Malay/Sumatra). However, in the southwestern part of Borneo genetic admixture of the lineages was observed in the two marker types. Coalescent simulation based on cpDNA sequence variation suggested that the two lineages arose 0.28 to 0.09 million years before present, and that following their divergence migration from Malay/Sumatra to Borneo strongly exceeded migration in the opposite direction. We conclude that the genetic structure of S. leprosula was largely formed during the middle Pleistocene and was subsequently modified by eastward migration across the subaerially exposed Sunda Shelf.
Figure 1. A, Bayesian time tree for Hemiphyllodactylus with 95 in Repeated evolution of sympatric, palaeoendemic species in closely related, co-distributed lineages of Hemiphyllodactylus Bleeker, 1860 (Squamata: Gekkonidae) across a sky-island archipelago in Peninsular Malaysia
Figure 1. A, Bayesian time tree for Hemiphyllodactylus with 95% highest posterior density (95% HPD) intervals for major nodes represented by purple bars. Black circles at nodes are posterior probabilities ≥ 0.95; grey circles at nodes are posterior probabilities <0.95. B, Bayesian time tree for the Hemiphyllodactylus harterti group. C, Distribution of the H. harterti group in Peninsular Malaysia.
Figure 6. Maximum credibility Bayesian tree obtained from a in Morphology and Bayesian tip-dating recover deep Cretaceous-age divergences among major chrysidid lineages (Hymenoptera: Chrysididae)
Figure 6. Maximum credibility Bayesian tree obtained from a combined analysis of a partitioned dataset of 300 morphological characters employing a relaxed morphological clock model with tip-dating. Character-state transformations (ChN, character; ChS, character state) are colour-coded with tagma as represented in the diagram on the bottom, and are depicted as solid (unequivocal changes) or empty (reversed or multiple changes) charts; mesosomal characters derived from legs and wings are depicted as rectangle and hexagon portrayals respectively; slow changes (DelTran optimization) are indicated by '+'. Fossil taxa are indicated by daggers. Phylogenetic relationships among Chrysididae (Clade 1) are shown in detail in Figures 7–12.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.