Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

410

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

410 results for “eukaryotic”

Learn how ShareScore rates datasets ↗
dryad32/100

Genomic analysis finds no evidence of canonical eukaryotic DNA processing complexes in a free-living protist

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad32/100

Data from: Systematic evaluation of horizontal gene transfer between eukaryotes and viruses

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad32/100

Data from: The evolution of protein-coding gene structure in eukaryotes

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo28/100

Extended Data Table 3 in Isolation of an archaeon at the prokaryote eukaryote interface

Extended Data Table 3 | Growth of MK-D1 after incubation of 120 days with a range of substrates

opennotspecifiedJan 2020View details →
zenodo28/100

Sequences of microbial eukaryotic genes obtained from the metagenome of the Mariana Trench

<p>nonredundant_eukaryotic_genes.faa contains the non-redundant amino acid sequences of predicted eukaryotic genes in all samples by MetaEuk.</p> <p>nonredundant_eukaryotic_genes.fna contains the non-redundant nucleotide sequences of predicted eukaryotic genes in all samples by MetaEuk.</p> <p>tax_for_per_gene.txt contains the taxonomic information of per non-redundant gene.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Limited host-specificity of eukaryotic virome in Hymenoptera

<p>Additional datafiles for Bee_Euvir repository.</p>

opencc-by-4.0Aug 2020View details →
dryad28/100

Data from: Recent events dominate interdomain lateral gene transfers between prokaryotes and eukaryotes and, with the exception of endosymbiotic gene transfers, few ancient transfer events persist

While there is compelling evidence for the impact of endosymbiotic gene transfer (EGT; transfer from either mitochondrion or chloroplast to the nucleus) on genome evolution in eukaryotes, the role of interdomain transfer from bacteria and/or archaea (i.e. prokaryotes) is less clear. Lateral gene transfers (LGTs) have been argued to be potential sources of phylogenetic information, particularly for reconstructing deep nodes that are difficult to recover with traditional phylogenetic methods. We sought to identify interdomain LGTs by using a phylogenomic pipeline that generated 13 465 single gene trees and included up to 487 eukaryotes, 303 bacteria and 118 archaea. Our goals include searching for LGTs that unite major eukaryotic clades, and describing the relative contributions of LGT and EGT across the eukaryotic tree of life. Given the difficulties in interpreting single gene trees that aim to capture the approximately 1.8 billion years of eukaryotic evolution, we focus on presence–absence data to identify interdomain transfer events. Specifically, we identify 1138 genes found only in prokaryotes and representatives of three or fewer major clades of eukaryotes (e.g. Amoebozoa, Archaeplastida, Excavata, Opisthokonta, SAR and orphan lineages). The majority of these genes have phylogenetic patterns that are consistent with recent interdomain LGTs and, with the notable exception of EGTs involving photosynthetic eukaryotes, we detect few ancient interdomain LGTs. These analyses suggest that LGTs have probably occurred throughout the history of eukaryotes, but that ancient events are not maintained unless they are associated with endosymbiotic gene transfer among photosynthetic lineages.

opencc-zeroDec 2014View details →
dryad28/100

Data from: A potential case of reinforcement in a facultatively sexual unicellular eukaryote

The origin of a new species requires a mechanism to prevent divergent populations from interbreeding. In the classic allopatric model, divided populations evolve independently and accumulate genetic differences. If contact is restored, hybrids suffer reduced fitness and selection may favor traits that prevent mistakes in mating, a process known as reinforcement. This decisive but transient phase is challenging to document and has been reported mostly in macroorganisms. Very little is known about the processes through which new microbial species originate. In particular, it is unclear whether microbial eukaryotes, many of which can reproduce sexually during complex life cycles, speciate in much the same way as do well-studied plants and animals. Using individual cellular mate choice trials, we investigated the mating behavior of sympatric and allopatric woodland populations of the yeast Saccharomyces paradoxus. We find evidence consistent with reinforcement, potentially representing an example of microbial speciation in progress.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Hemimastigophora is a novel supra-kingdom-level lineage of Eukaryotes

Almost all eukaryote life forms have now been placed within one of five to eight supra-kingdom-level groups using molecular phylogenetics1,2,3,4. The 'phylum' Hemimastigophora is probably the most distinctive morphologically defined lineage that still awaits such a phylogenetic assignment. First observed in the nineteenth century, hemimastigotes are free-living predatory protists with two rows of flagella and a unique cell architecture5,6,7; to our knowledge, no molecular sequence data or cultures are currently available for this group. Here we report phylogenomic analyses based on high-coverage, cultivation-independent transcriptomics that place Hemimastigophora outside of all established eukaryote supergroups. They instead comprise an independent supra-kingdom-level lineage that most likely forms a sister clade to the 'Diaphoretickes' half of eukaryote diversity (that is, the 'stramenopiles, alveolates and Rhizaria' supergroup (Sar), Archaeplastida and Cryptista, as well as other major groups). The previous ranking of Hemimastigophora as a phylum understates the evolutionary distinctiveness of this group, which has considerable importance for investigations into the deep-level evolutionary history of eukaryotic life—ranging from understanding the origins of fundamental cell systems to placing the root of the tree. We have also established the first culture of a hemimastigote (Hemimastix kukwesjijk sp. nov.), which will facilitate future genomic and cell-biological investigations into eukaryote evolution and the last eukaryotic common ancestor.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Meta‐analysis of chromosome‐scale crossover rate variation in eukaryotes and its significance to evolutionary genomics

Understanding the distribution of crossovers along chromosomes is crucial to evolutionary genomics because the crossover rate determines how strongly a genome region is influenced by natural selection. Nevertheless, generalities in the chromosome-scale distribution of crossovers have not been investigated formally. We fill this gap by synthesizing joint information on genetic and physical maps across 62 animal, plant, and fungal species. Our quantitative analysis reveals a strong and taxonomically wide-spread reduction of the crossover rate in the center of chromosomes relative to their peripheries. We demonstrate that this pattern is poorly explained by the position of the centromere, but find that the magnitude of the relative reduction in the crossover rate in chromosome centers increases with chromosome length. That is, long chromosomes often display a dramatically low crossover rate in their center whereas short chromosomes exhibit a relatively homogeneous crossover rate. This observation is compatible with a model in which crossovers are initiated from the chromosome tips, an idea with preliminary support from mechanistic investigations of meiotic recombination. Consequently, we show that organisms achieve a higher genome-wide crossover rate by evolving smaller chromosomes. Summarizing theory and providing empirical examples, we finally highlight that taxonomically wide-spread and systematic heterogeneity in crossover rate along chromosomes generates predictable broad-scale trends in genetic diversity and population differentiation by modifying the impact of natural selection among regions within a genome. We conclude by emphasizing that chromosome-scale heterogeneity in crossover rate should urgently be incorporated into analytical tools in evolutionary genomics, and in the interpretation of emerging patterns.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Do saline taxa evolve faster? comparing relative rates of molecular evolution between freshwater and marine eukaryotes

The major branches of life diversified in the marine realm, and numerous taxa have since transitioned between marine and freshwaters. Previous studies have demonstrated higher rates of molecular evolution in crustaceans inhabiting continental saline habitats as compared with freshwaters, but it is unclear whether this trend is pervasive or whether it applies to the marine environment. We employ the phylogenetic comparative method to investigate relative molecular evolutionary rates between 148 pairs of marine or continental saline vs. freshwater lineages representing disparate eukaryote groups, including bony fish, elasmobranchs, cetaceans, crustaceans, mollusks, annelids, algae, and other eukaryotes, using available protein-coding and non-coding genes. Overall, we observed no consistent pattern in nucleotide substitution rates linked to habitat across all genes and taxa. However, we observed some trends of higher evolutionary rates within protein-coding genes in freshwater taxa—the comparisons mainly involving bony fish—compared with their marine relatives. The results suggest no systematic differences in substitution rate between marine and freshwater organisms.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Potential and pitfalls of eukaryotic metagenome skimming: A test case for lichens

Whole genome shotgun sequencing of multi species communities using only a single library layout is commonly used to assess taxonomic and functional diversity of microbial assemblages. Here we investigate to what extent such metagenome skimming approaches are applicable for in-depth genomic characterizations of eukaryotic communities, e.g. lichens. We address how to best assemble a particular eukaryotic metagenome skimming data, what pitfalls can occur, and what genome quality can be expected from this data. To facilitate a project specific benchmarking, we introduce the concept of twin sets, simulated data resembling the outcome of a particular metagenome sequencing study. We show that the quality of genome reconstructions depends essentially on assembler choice. Individual tools, including the metagenome assemblers Omega and MetaVelvet, are surprisingly sensitive to low and uneven coverages. In combination with the routine of assembly parameter choice to optimize the assembly N50 size, these tools can preclude an entire genome from the assembly. In contrast, MIRA, an all-purpose overlap assembler, and SPAdes, a multi-sized de Bruijn graph assembler, facilitate a comprehensive view on the individual genomes across a wide range of coverage ratios. Testing assemblers on a real-world metagenome skimming data from the lichen Lasallia pustulata demonstrates the applicability of twin sets for guiding method selection. Furthermore, it reveals that the assembly outcome for the photobiont Trebouxia sp. falls behind the a-priori expectation given the simulations. Although the underlying reasons remain still unclear this highlights that further studies on this organism require special attention during sequence data generation and downstream analysis.

opencc-zeroDec 2014View details →
dryad28/100

Data from: A single Tim translocase in the mitosomes of Giardia intestinalis illustrates convergence of protein import machines in anaerobic eukaryotes

Mitochondria have evolved diverse forms across eukaryotic diversity in adaptation to anaerobiosis. Mitosomes are the simplest and the least well-studied type of anaerobic mitochondria. Transport of proteins via TIM complexes, composed of three proteins of the Tim17 protein family (Tim17/22/23), is one of the key unifying aspects of mitochondrial and mitochondria-derived organelles. However, multiple experimental and bioinformatic attempts have so far failed to identify the nature of TIM in mitosomes of the anaerobic metamonad protist, Giardia intestinalis, one of the few experimental models for mitosome biology. Here, we present the identification of a single G. intestinalis Tim17 protein (GiTim17), made possible only by the implementation of a metamonad-specific hidden Markov model. While very divergent in primary sequence and in predicted membrane topology, experimental data suggest that GiTim17 is an inner membrane mitosomal protein, forming a disulphide-linked dimer. We suggest that the peculiar GiTim17 sequence reflects adaptation to the unusual, detergent resistant, inner mitosomal membrane. Specific pull-down experiments indicate interaction of GiTim17 with mitosomal Tim44, the tethering component of the import motor complex. Analysis of TIM complexes across eukaryote diversity indicates that a "single Tim" translocase is a convergent adaptation of mitosomes in anaerobic protists, with Tim22 and Tim17 (but not Tim23), providing the protein backbone.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Breakdown of phylogenetic signal: a survey of microsatellite densities in 454 shotgun sequences from 154 non model eukaryote species

Microsatellites are ubiquitous in Eukaryotic genomes. A more complete understanding of their origin and spread can be gained from a comparison of their distribution within a phylogenetic context. Although information for model species is accumulating rapidly, it is insufficient due to a lack of species depth, thus intragroup variation is necessarily ignored. As such, apparent differences between groups may be overinflated and generalizations cannot be inferred until an analysis of the variation that exists within groups has been conducted. In this study, we examined microsatellite coverage and motif patterns from 454 shotgun sequences of 154 Eukaryote species from eight distantly related phyla (Cnidaria, Arthropoda, Onychophora, Bryozoa, Mollusca, Echinodermata, Chordata and Streptophyta) to test if a consistent phylogenetic pattern emerges from the microsatellite composition of these species. It is clear from our results that data from model species provide incomplete information regarding the existing microsatellite variability within the Eukaryotes. A very strong heterogeneity of microsatellite composition was found within most phyla, classes and even orders. Autocorrelation analyses indicated that while microsatellite contents of species within clades more recent than 200 Mya tend to be similar, the autocorrelation breaks down and becomes negative or non-significant with increasing divergence time. Therefore, the age of the taxon seems to be a primary factor in degrading the phylogenetic pattern present among related groups. The most recent classes or orders of Chordates still retain the pattern of their common ancestor. However, within older groups, such as classes of Arthropods, the phylogenetic pattern has been scrambled by the long independent evolution of the lineages.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Complex phylogeographic patterns in the freshwater alga Synura provide new insights on ubiquity versus endemism in microbial eukaryotes

The global distribution, abundance, and diversity of microscopic freshwater algae demonstrate an ability to overcome significant barriers such as dry land and oceans by exploiting a range of biotic and abiotic colonization vectors. If these vectors are considered unlimited and colonization occurs in proportion to population size, then globally ubiquitous distributions are predicted to arise. This model contrasts with observations that many freshwater microalgal taxa possess true biogeographies. Here, using a concatenated multi-gene dataset, we study the phylogeography of the freshwater heterokont alga Synura petersenii sensu lato. Our results suggest that this Synura morphotaxon contains both cosmopolitan and regionally endemic cryptic species, co-occurring in some cases, and masked by a common ultrastructural morphology. Phylogenies based on both proteins (seven protein-coding plastid and mitochondrial genes) and DNA (nine genes including ITS and 18S rDNA) reveal pronounced biogeographic delineations within phylotypes of this cryptic species complex, while retaining one clade that is globally distributed. Relaxed molecular clock calculations, constrained by fossil records, suggest that the genus Synura is considerably older than currently proposed. The availability of tectonically-relevant geological time (10^7-10^8 years) has enabled the development of the observed, complex biogeographic patterns. Our comprehensive analysis of freshwater algal biogeography suggests that neither ubiquity nor endemism wholly explain global patterns of microbial eukaryote distribution, and that processes of dispersal remain poorly understood.

opencc-zeroDec 2009View details →
dryad28/100

Data from: Taxon-rich phylogenomic analyses resolve the eukaryotic tree of life and reveal the power of subsampling by sites

Most eukaryotic lineages are microbial, and many have only recently been sampled for phylogenetic studies or remain in the 'dark area' of the tree of life where there are no molecular data. To assess relationships among eukaryotic lineages, we perform a taxon-rich phylogenomic analysis including 232 eukaryotes selected to maximize taxonomic diversity and up to 1554 genes chosen as vertically inherited based on their broad distribution among eukaryotes. We also include sequences from 486 bacteria and 84 archaea to assess the impact of endosymbiotic gene transfer (EGT) from plastids and to detect contamination. Overall, our analyses are consistent with other less taxon-rich estimates of the eukaryotic tree of life and we recover strong support for five major clades: Amoebozoa, Excavata (without the genus Malawimonas), Opisthokonta, Archaeplastida and SAR (Stramenopila, Alveolata and Rhizaria). Our analyses also highlight the existence of 'orphan' lineages, lineages that lack robust placement in the eukaryotic tree of life and indicate the possibility of as yet undiscovered diversity. In analyses including bacteria and archaea, we find that ~10% of the 1554 genes, which we choose because they are found in four or five of the five major eukaryotic clades and hence may be more likely to be inherited vertically, appear to have been acquired from cyanobacteria through EGT in photosynthetic lineages. Removing these EGT genes places the green algae as sister to the glaucophytes instead of the red algae, suggesting that unknowingly including of genes of plastid origin, and combining them with genes of nuclear origin, may mislead phylogenetic estimates. Finally, the large size of our dataset allows comparative analyses of subsets of data; alignments built from randomly sampled sites provide greater support, particularly for deep relationships, than do equivalent sized datasets built from randomly sampled genes.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Untangling the early diversification of eukaryotes: a phylogenomic study of the evolutionary origins of Centrohelida, Haptophyta, and Cryptista

Assembling the global eukaryotic tree of life has long been a major effort of Biology. In recent years, pushed by the new availability of genome-scale data for microbial eukaryotes, it has become possible to revisit many evolutionary enigmas. However, some of the most ancient nodes, which are essential for inferring a stable tree, have remained highly controversial. Among other reasons, the lack of adequate genomic datasets for key taxa has prevented the robust reconstruction of early diversification events. In this context, the centrohelid heliozoans are particularly relevant for reconstructing the tree of eukaryotes because they represent one of the last substantial groups that was missing large and diverse genomic data. Here, we filled this gap by sequencing high-quality transcriptomes for four centrohelid lineages, each corresponding to a different family. Combining these new data with a broad eukaryotic sampling, we produced a gene-rich taxon-rich phylogenomic dataset that enabled us to refine the structure of the tree. Specifically, we show that (i) centrohelids relate to haptophytes, confirming Haptista; (ii) Haptista relates to SAR; (iii) Cryptista share strong affinity with Archaeplastida; and (iv) Haptista + SAR is sister to Cryptista + Archaeplastida. The implications of this topology are discussed in the broader context of plastid evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Expression of eukaryotic-like protein in the microbiome of sponges

Eukaryotic-like proteins (ELPs) are classes of proteins that are found in prokaryotes, but have a likely evolutionary origin in eukaryotes. ELPs have been postulated to mediate host-microbiome interactions. Recent work has discovered that prokaryotic symbionts of sponges contain abundant and diverse genes for ELPs, which could modulate interactions with their filter-feeding and phagocytic host. However, the extent to which these ELP genes are actually used and expressed by the symbionts is poorly understood. Here we use metatranscriptomics to investigate ELP expression in the microbiomes of three different sponges (Cymbastella concentrica, Scopalina sp. and Tedania anhelens). We developed a workflow with optimized rRNA removal and in silico subtraction of host sequences to obtain a reliable symbiont metatranscriptome. This showed that between 1.3 and 2.3% of all symbiont transcripts contain genes for ELPs. Two classes of ELPs (cadherin and tetratricopetide repeats) were abundantly expressed by in the C. concentrica and Scopalina sp. microbiomes, while ankyrin repeat ELPs were predominant in the T. anhelens metatranscriptome. Comparison to non-ELP containing transcripts indicated a constitutive expression of ELPs across a range of bacterial and archaeal symbionts. Expressed ELPs also contained domains involved in protein secretion and/or were co-expressed with proteins involved in extra-cellular transport. This suggests these ELPs are likely exported, which could allow for direct interaction with the sponge. Our study shows that ELP genes in sponge symbionts represent actively expressed functions that could mediate molecular interaction between symbiosis partners.

opencc-zeroDec 2015View details →
zenodo28/100

Data from: Identification of prokaryotic and eukaryotic virus-derived sequences in virome using deep learning

<h4>This repository contains the data and Docker image to reproduce the results of our paper: <strong>identification of prokaryotic and eukaryotic virus-derived sequences in virome using deep learning</strong></h4> <p>Authors: Hengchuang Yin, Shufang Wu, Jie Tan, Qian Guo, Mo Li, Jinyuan Guo, Yaqi Wang, Xiaoqing Jiang, and Huaiqiu Zhu*</p> <p><strong>This work has been accepted by GigaScience.&nbsp;</strong></p> <p><strong>Hengchuang Yin, Shufang Wu, Jie Tan, Qian Guo, Mo Li, Jinyuan Guo, Yaqi Wang, Xiaoqing Jiang, and Huaiqiu Zhu. "IPEV: Identification of Prokaryotic and Eukaryotic Virus-Derived Sequences in Virome Using Deep Learning." GigaScience 13 (2024): giae018.&nbsp;<a href="https://doi.org/10.1093/gigascience/giae018" rel="nofollow">https://doi.org/10.1093/gigascience/giae018</a>.</strong></p> <div>&nbsp;</div> <p>&nbsp;</p> <p><strong>Background: </strong>The virome obtained through virus-like particle enrichment contains a mixture of prokaryotic and eukaryotic virus-derived fragments. Accurate identification and classification of these elements are crucial to understanding their roles and functions in microbial communities. However, the rapid mutation rates of viral genomes pose challenges in developing high-performance tools for classification, potentially limiting downstream analyses.</p> <p><strong>Findings: </strong>We present IPEV, a novel method to distinguish prokaryotic and eukaryotic viruses in viromes, with a 2D convolutional neural network combining trinucleotide pair relative distance and frequency. Cross-validation assessments of IPEV demonstrate its state-of-the-art precision, significantly improving the F1-score by approximately 22% on an independent test set compared to existing methods when query viruses share less than 30% sequence similarity with known viruses.&nbsp;Furthermore, IPEV outperforms other methods in accuracy on marine and gut virome samples based on annotations by sequence alignments. IPEV reduces runtime by at most 1,225 times compared to existing methods under the same computing configuration. We also utilized IPEV to analyze longitudinal samples and found that the gut virome exhibits a higher degree of temporal stability than previously observed in persistent personal viromes, providing novel insights into the resilience of the gut virome in individuals.&nbsp;</p> <p><strong>Conclusions:&nbsp;</strong>IPEV is a high-performance, user-friendly tool that assists biologists in identifying and classifying prokaryotic and eukaryotic viruses within viromes. The tool is available at&nbsp;https://github.com/basehc/IPEV.</p> <p>&nbsp;</p> <p><strong>5_fold_cross_validation.zip:</strong> Dataset of cross-validation of IPEV</p> <p><strong>Eukaryotic_virus_CV_Dataset-1.csv:</strong> GI, and accession ID for the cross-validation Dataset-1 (eukaryotic virus)</p> <p><strong>Prokaryotic_virus_CV_Dataset-1.csv:</strong> GI, and accession ID for the cross-validation Dataset-1 (prokaryotic virus)</p> <p><strong>Test_Prokaryotic_virus_Dataset-1.fasta: </strong>An independent test set of IPEV (prokaryotic&nbsp;virus)</p> <p><strong>Test_Eukaryotic_virus_Dataset-1.fasta: </strong>An independent test set of IPEV (eukaryotic&nbsp;virus)</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Dataset_sequencing_error.zip: </strong>Simulated dataset with sequencing errors</p> <p><strong>Cap_enzyme_sequence.fasta: </strong>Accession IDs of Receptor Binding Proteins (RBPs) in phages collected by our article</p> <p><strong>Dataset_runtime_evaluation.zip: </strong>Dataset for evaluating the runtime of IPEV</p> <p><strong>Receptor_binding_protein_accession_id:</strong>&nbsp;Accession IDs of Receptor Binding Proteins (RBPs) in phages collected by our article</p> <p>&nbsp;</p> <p><strong>archaea_ID.txt</strong> Accession ID information for the reference archaea dataset</p> <p><strong>bacteria_ID.txt</strong> Accession ID information for the reference bacterial dataset</p> <p><strong>marine_virome_id.csv:</strong> Ocean virome data information used in our paper</p> <p><strong>gut_virome.csv</strong>:Gur virome data information used in our paper</p> <p><strong>fungi.txt: </strong>Negative sequence information used to train, validate, and test the model in the decontamination function</p> <p><strong>bacteria.txt: </strong>Negative sequence information used to train, validate, and test the model in the decontamination function</p> <p>&nbsp;</p> <p>&nbsp;</p> <h4><strong>Reproduce the results of our paper from a Docker image.</strong></h4> <p>&nbsp;</p> <p>We also provide a Docker image file that does not require any environment configuration. You can reproduce the results of our paper (e.g., train and test our IPEV model) in a Docker image.</p> <p>Pull the<a href="https://hub.docker.com/r/dryinhc/ipev_v1"> <em>dryinhc/ipev_v1</em></a> image from Docker Hub. Open a terminal window and run the following command:</p> <p><em>docker pull dryinhc/ipev_v1</em></p> <p>This will download the image to your local machine.</p> <p>Run the <em>dryinhc/ipev_v1</em> image. In the same terminal window, run the following command:</p> <p><em>docker run -it --rm dryinhc/ipev_v1</em></p> <p>This will start a container based on the image and run the IPEV tool.</p> <p>And you can run cd train or cd other file folders in the container.</p> <p>To exit the container, press <em>Ctrl+D</em> or type&nbsp;<em>exit</em>.</p> <p>It contains 4 directories, namely 5 fold cross validation, independent set, marine virome, and gut virome. The 5-fold cross-validation directory holds the scripts required for implementing the 5-fold cross-validation method. The independent set directory contains scripts necessary for working with an independent set. Lastly, the marine virome and gut virome directories store scripts for analyzing real datasets.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>We hereby confirm that the dataset associated with the research described in this work is made available to the public under the Creative Commons Zero (CC0) license.</p> <p>&nbsp;</p> <p>&nbsp;</p> <h4><strong>Contact&nbsp;</strong></h4> <p>&nbsp;</p> <p>If you have any questions, please don't hesitate to ask me: yinhengchuang@pku.edu.cn or hqzhu@pku.edu.cn</p>

openother-pdNov 2023View details →
dryad28/100

Ubiquity and evolution of structural maintenance of chromosomes (SMC) proteins in eukaryotes

<p>Structural maintenance of chromosomes (SMC) protein complexes are common in Bacteria, Archaea, and Eukaryota. SMC proteins, together with the proteins related to SMC (SMC-related proteins), constitute a superfamily of ATPases. Bacteria/Archaea and Eukaryotes are distinctive from one another in terms of the repertory of SMC proteins. A single type of SMC protein is dimerized in the bacterial and archaeal complexes, whereas eukaryotes possess six distinct SMC subfamilies (SMC1-6), constituting three heterodimeric complexes, namely cohesin, condensin, and SMC5/6 complex. Thus, to bridge the homodimeric SMC complexes in Bacteria and Archaea to the heterodimeric SMC complexes in Eukaryota, we need to invoke multiple duplications of an SMC gene followed by functional divergence. However, to our knowledge, the evolution of the SMC proteins in Eukaryota had not been examined for more than a decade. In this study, we reexamined the ubiquity of SMC1-6 in phylogenetically diverse eukaryotes that cover the major eukaryotic taxonomic groups recognized to date and provide two novel insights into the SMC evolution in eukaryotes. First, multiple secondary losses of SMC5 and SMC6 occurred in the eukaryotic evolution. Second, the SMC proteins constituting cohesin and condensin (i.e., SMC1-4), and SMC5 and SMC6 were derived from closely related but distinct ancestral proteins. Based on the above-mentioned findings, we discuss how SMC1-6 have diverged from the archaeal homologs.</p>

opencc-zeroDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record