Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “eukaryotic tree of life”

Learn how ShareScore rates datasets ↗
zenodo44/100

The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life. Data file underpinning Figure 2A and Figure 2B

<div>These datasheets accompany the article "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life" in Frontiers in Science</div> <div>This file contains data processed from Catalog of Life on 31 December 2023. The catalog was downloaded and post-processed to</div> <div>remove prokaryotic taxa</div> <div>remove extinct and fossil taxa</div> <div>remove taxon names that were listed as junior synonyms</div> <div>remove taxon names listed as "invalid"</div> <div>Total living, valid eukaryotic genera 167,085</div> <div>The taxa were sorted by the nomenclatorial Code under which they were declared (to avoid namespace clashes)</div> <div>International Code for Algae, Fungi and Plants https://www.iapt-taxon.org/nomen/main.php</div> <div>Algal, Fungal, Plant code genera 31,076</div> <div>International Code of Zoological Nomenclature https://www.iczn.org/the-code/the-code-online/</div> <div>Zoological code genera 136,009</div> <div>The Code-sorted taxa were aggregated by the generic portion of their names, and two plots were generated:</div> <div>a plot aggregating the cumulative number of species in genera sorted by species number (Figure 2A)</div> <div>a plot illustrating the distribution of the size of genera (Figure 2B)</div> <div>This data file gives access to these processed data for</div> <div>Figure 2 A Data</div> <div>Figure 2 B Data</div> <div>The original data including the intermediate calculations of values, and the plotted graphs, are available as a GoogleDoc at https://docs.google.com/spreadsheets/d/1V-bTtWjIRasC3AgID0jGlyToKqI-H9h1aPeSxUpNrjk/edit?usp=sharing</div>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages

<p><strong>Abstract</strong></p> <p>Advances in phylogenetic methods and high-throughput sequencing have allowed the reconstruction of deep phylogenetic relationships in the evolutionary history of eukaryotes. Yet, the root of the eukaryotic tree of life remains elusive. The most &lsquo;popular&rsquo; (i.e. in textbooks and reviews) hypothesis for the root is between Unikonta (Opisthokonta + Amoebozoa) and Bikonta (all other eukaryotes), which emerged from analyses of a single gene fusion and a limited sampling of eukaryotic lineages. Subsequent highly-cited studies based on concatenation of genes supported this hypothesis with some variations or proposed a root within the Excavata. However, concatenation of genes neither considers phylogenetically-informative events (i.e. gene duplications and losses), nor provides an estimate of the root. A more recent study using gene tree-species tree reconciliation methods suggested the root lies between Opisthokonta and all other eukaryotes, but only including 59 taxa and 20 genes. Here we apply a gene tree &ndash; species tree reconciliation approach to a gene-rich and taxon-rich dataset (i.e. 2,786 gene families from two sets of ~158 diverse eukaryotic lineages) to assess the root, and we iterate each analysis 100 times to quantify tree space uncertainty. Our results estimate a root between Fungi and all other eukaryotes, or between Opisthokonta and all other eukaryotes, and reject alternative popular roots from the literature. Based on further analysis of genome size we propose Opisthokonta + others as the most likely root. Finding the root of the eukaryotic tree of life is critical for the field of comparative biology as it allows us to understand the timing and mode of evolution of characters across the evolutionary history of eukaryotes.</p> <p>Methods<br> Here we provide the alignments, gene trees, inputs, and outputs from our project entitled &quot;Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages&quot;. Sequences and alignments were produced using the phylogenomic pipeline PhyloToL, which contains a taxon- and gene-rich database (including eukaryotes, archaea, and bacteria). These data were then used for 1) assessing the root of the eukaryotes and 2) for comparison with other previously published hypotheses. In both cases, we used the species tree - gene tree reconciliation tool iGTP. We also did a comparison of hypotheses using the likelihood-based tool SpeciesRax</p> <p><strong>iGTP</strong></p> <p>Input data</p> <p>These data are divided into four datasets based on taxa selection. For dataset SEL+, taxa were selected based on their taxonomy; for RAN+, taxa were selected randomly among the major eukaryotic clades Opisthokonta, Amoebozoa, Archaeplastida, Excavata, SAR, and some orphan lineages. Datasets SEL- and RAN- are the same as SEL+ and RAN+, but exclude microsporidians in order to account for and avoid long branch attraction due to microsporidians fast-evolutionary rates. We chose the gene families that contain at least 25 taxa representing at least four of the five major eukaryotic clades. Additionally, at least 2 of the major clades had to contain at least 2 minor clades (e.g. Glaucophytes and Rhodophyta are minor clades in the major clade Archaeplastida). In a pilot analysis, we produced an alignment and a phylogenetic tree for each gene family using the default settings of a previous version of PhyloToL (GUIDANCE V1.3.1 sequence cutoff = 0.3 and column cutoff = 0.4; RAxML quick tree with model PROTGAMMALG and no bootstraps). Then, we kept the gene families that are exclusive of eukaryotes or the ones in which eukaryotes were monophyletic. From a total of 3,002 gene families that met our criteria, 2786 passed the initial steps of PhyloToL when including only the data from the dataset SEL+. These 2,786 gene families were used for further analyses with all datasets.</p> <p>MSAs were produced with PhyloToL (GUIDANCE V2.02 sequence cutoff = 0.3, column cutoff = 0.4, number of iterations = 5; Sela, et al. 2015). The default parameters of PhyloToL include up to five iterations of GUIDANCE V2.02 with 10 bootstraps and MAFFT V7 with algorithm E-INS-i for less than 200 sequences or &ldquo;auto&rdquo; option if more than 200 sequences, and maxiterate = 1000. Instead, here we run up to five iterations of GUIDANCE with 20 bootstraps and the simple MAFFT algorithm FFT-NS-2. Then, we perform an additional GUIDANCE run with 100 bootstraps and the default MAFFT parameters for PhyloToL.</p> <p>Gene trees were inferred with RAxML v.8.2.4 with 10 ML searches for best-ML tree (option &quot;-# 10&quot;), using the rapid hill-climbing algorithm (option &quot;-f d&quot;) and no bootstrap replicates. The protein evolution model used was evaluated during the gene tree inference (option &quot;-m PROTCATAUTO&quot;) by testing all models available in RAxML (e.g. JTT, LG, WAG, etc) with optimization of substitution rates and of site-specific evolutionary rates which were categorized into four distinct rate categories for greater computational efficiency.</p> <p>We ran 100 repetitions of iGTP analyses per dataset. But, given the complexity of the datasets and the heuristic nature of some key steps of the iGTP algorithm (e.g. gene tree rooting and initial starting species tree generation), in a preliminary analysis, we faced two systematic challenges with iGTP as the inferred species tree was affected by: 1) the order of the leaves in the input unrooted gene tree Newick strings (i.e. the input trees were treated as rooted even though we specified that they were not); and 2) the input gene order in the 100 replicates. Therefore, we randomly shuffled the order of the leaves in the unrooted gene trees (keeping the same topology), and randomly shuffled the order of the input gene trees in each of the 100 replicates per dataset. Here we provide the 100 input files generated for those iGTP analyses.&nbsp;</p> <p>Output data</p> <p>Here we also share the data generated after two analyses: 1) root assessment and 2) hypothesis testing. For the former, we allowed iGTP to calculate the more parsimonious root given our input files. For the latter, we allowed iGTP to calculate the reconciliation cost of the gene trees given the input files and constraints in the species trees to reflect previously published root hypotheses. The constraints are explained in the README file.</p> <p>SpeciesRax</p> <p>Since we removed LGT and contamination from our dataset using a series of filters, we applied the model UndatedDL instead of UndatedDTL, which implies that we only took into consideration duplications and losses and ignored the transferences. Then, the command used for SpeciesRax was...</p> <p>./generax --families forGeneRax/famFile --species-tree forGeneRax/spsTreeAn.newick --strategy SKIP --rec-model UndatedDL --per-family-rates --prefix An_DS1 --si-strategy EVAL&nbsp;</p> <p>input&nbsp;</p> <p>Here, we are sharing all the necessary files to run SpeciesRax, including the mapping file (famFile), the species trees (spsTrees; the best iGTP constrained species trees per hypothesis), and their underlying gene trees (trees_r)</p> <p>output</p> <p>Output folder from SpeciesRax, which includes log files, events (duplications, losses) counts, and statistics (i.e., reconciliation likelihood values)</p> <p>Note:&nbsp;<br> * As in the iGTP analyses, for the SpeciesRax files, the words Op, Fu, Di, Un, An, refer to the five root hypotheses compared: Opisthokonta-others, Fungi-others, Discoba-others, Unikonta-Bikonta, and (Ancyromonadida + Metamonada)-others, respectively.&nbsp;<br> * For the SpeciesRax analyses, the Un word refers to Ut (Cavalier-Smith 2003)&nbsp;<br> &nbsp;</p>

opencc-by-4.0Feb 2021View details →
dryad36/100

Phylogenomic analyses of 2,786 genes in 158 lineages support a root of the eukaryotic tree of life between opisthokonts and all other lineages

<p>Advances in phylogenetic methods and high-throughput sequencing have allowed the reconstruction of deep phylogenetic relationships in the evolutionary history of eukaryotes. Yet, the root of the eukaryotic tree of life remains elusive. The most 'popular' (i.e. in textbooks and reviews) hypothesis for the root is between Unikonta (Opisthokonta + Amoebozoa) and Bikonta (all other eukaryotes), which emerged from analyses of a single gene fusion and a limited sampling of eukaryotic lineages. Subsequent highly-cited studies based on concatenation of genes supported this hypothesis with some variations or proposed a root within the Excavata. However, concatenation of genes neither considers phylogenetically-informative events (i.e. gene duplications and losses) nor provides an estimate of the root. A more recent study using gene tree-species tree reconciliation methods suggested the root lies between Opisthokonta and all other eukaryotes, but only including 59 taxa and 20 genes. Here we apply a gene tree – species tree reconciliation approach to a gene-rich and taxon-rich dataset (i.e. 2,786 gene families from two sets of ~158 diverse eukaryotic lineages) to assess the root, and we iterate each analysis 100 times to quantify tree space uncertainty. Our results estimate a root between Fungi and all other eukaryotes, or between Opisthokonta and all other eukaryotes, and reject alternative popular roots from the literature. Based on further analysis of genome size, we propose Opisthokonta + others as the most likely root. Finding the root of the eukaryotic tree of life is critical for the field of comparative biology as it allows us to understand the timing and mode of evolution of characters across the evolutionary history of eukaryotes.</p>

opencc-zeroFeb 2021View details →
zenodo36/100

Earth BioGenome Project: Authors and author contributions for publication "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life"

<p>The zipped file cntains three tab-separated datasets (worksheets). These worksheets list</p> <p>The contributing authors for the manuscript "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life" and the roles of these authors in the manuscript.</p> <p>The funding sources for these authors</p> <p>A list of the authors contributing to the "EBP Community of Scientists" collective authorship for the same manuscript.</p>

opencc-by-4.0Sep 2024View details →
dryad36/100

Truly ubiquitous CRESS DNA viruses scattered across the eukaryotic tree of life

<p>Until recently, most viruses detected and characterized were of economic significance, associated with agricultural and medical diseases. This was certainly true for the eukaryote-infecting circular Rep (replication-associated protein)-encoding single-stranded DNA (CRESS DNA) viruses, which were thought to be a relatively small group of viruses. With the explosion of metagenomic sequencing over the past decade and increasing use of rolling-circle replication for sequence amplification, scientists have identified and annotated copious numbers of novel CRESS DNA viruses – many without known hosts but which have been found in association with eukaryotes. Similar advances in cellular genomics have revealed that many eukaryotes have endogenous sequences homologous to viral Reps, which not only provide "fossil records" to reconstruct the evolutionary history of CRESS DNA viruses but also reveal potential host species for viruses known by their sequences alone. The Rep protein is a conserved protein that all CRESS DNA viruses use to assist rolling circle replication that is known to be endogenized in a few eukaryotic species (notably tobacco and water yam). A systematic search for endogenous Rep-like sequences in GenBank's non-redundant eukaryotic database was performed using tBLASTn. We utilized relaxed search criteria for the capture of integrated Rep sequence within eukaryotic genomes, identifying 93 unique species with an endogenized fragment of Rep in their nuclear (78 species), plasmid (1 species), mitochondrial (6 species) or chloroplast (8 species) genomes. These species come from 19 different phyla, scattered across the eukaryotic tree of life. Exogenous and endogenous CRESS DNA viral Rep tree topology suggested potential hosts for one family of uncharacterized viruses and supports a primarily fungal host range for genomoviruses.</p>

opencc-zeroSep 2021View details →
dryad36/100

Truly ubiquitous CRESS DNA viruses scattered across the eukaryotic tree of life

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad36/100

Phylogenomic analyses of 2,786 genes in 158 lineages support a root of the eukaryotic tree of life between opisthokonts and all other lineages

Open the record for dataset details and reuse information.

publicAug 2022View details →
zenodo32/100

Author information for the publication "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life"

<p>This Excel format file includes three sheets:</p> <p>1: CREDIT information for the named authors of the publication "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life"</p> <p>2: Funding information for the named authors of the publication "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life"</p> <p>3: A listing of the individuals included in the collective authorship "The EBP Community of Scientists" in the publication "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life"</p>

opencc-by-4.0Sep 2024View details →
dryad28/100

Data from: Taxon-rich phylogenomic analyses resolve the eukaryotic tree of life and reveal the power of subsampling by sites

Most eukaryotic lineages are microbial, and many have only recently been sampled for phylogenetic studies or remain in the 'dark area' of the tree of life where there are no molecular data. To assess relationships among eukaryotic lineages, we perform a taxon-rich phylogenomic analysis including 232 eukaryotes selected to maximize taxonomic diversity and up to 1554 genes chosen as vertically inherited based on their broad distribution among eukaryotes. We also include sequences from 486 bacteria and 84 archaea to assess the impact of endosymbiotic gene transfer (EGT) from plastids and to detect contamination. Overall, our analyses are consistent with other less taxon-rich estimates of the eukaryotic tree of life and we recover strong support for five major clades: Amoebozoa, Excavata (without the genus Malawimonas), Opisthokonta, Archaeplastida and SAR (Stramenopila, Alveolata and Rhizaria). Our analyses also highlight the existence of 'orphan' lineages, lineages that lack robust placement in the eukaryotic tree of life and indicate the possibility of as yet undiscovered diversity. In analyses including bacteria and archaea, we find that ~10% of the 1554 genes, which we choose because they are found in four or five of the five major eukaryotic clades and hence may be more likely to be inherited vertically, appear to have been acquired from cyanobacteria through EGT in photosynthetic lineages. Removing these EGT genes places the green algae as sister to the glaucophytes instead of the red algae, suggesting that unknowingly including of genes of plastid origin, and combining them with genes of nuclear origin, may mislead phylogenetic estimates. Finally, the large size of our dataset allows comparative analyses of subsets of data; alignments built from randomly sampled sites provide greater support, particularly for deep relationships, than do equivalent sized datasets built from randomly sampled genes.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Turning the crown upside down: gene tree parsimony roots the eukaryotic tree of life

Open the record for dataset details and reuse information.

publicFeb 2012View details →
dryad28/100

Data from: Taxon-rich phylogenomic analyses resolve the eukaryotic tree of life and reveal the power of subsampling by sites

Open the record for dataset details and reuse information.

publicDec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record