Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

117

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

117 results for “archaea”

Learn how ShareScore rates datasets ↗
edi64/100

Seasonal Distribution of Ammonia-Oxidizing Archaea and Ammonia-Oxidation Rates in the South Atlantic Bight from April to November 2014

Previous work in nearshore waters of the Georgia USA coast has demonstrated mid-summer peaks in the abundance of Thaumarchaeota (blooms with 100 to 1,000-fold increases) accompanied by spikes in nitrite concentration. These studies were performed at one location, so the areal extent of the bloom is unknown, nor has it been demonstrated conclusively that it develops in inshore waters. We collected data on rates of ammonia oxidation and the distribution of Thaumarchaeota, ammonia-oxidizing Betaproteobacteria (AOB), nitrite-oxidizing Nitrospina and environmental variables during 6 cruises aboard the UNOLS vessel R/V Savannah from April to November 2014 on transects of the South Atlantic Bight to evaluate the areal extent and timing of the bloom. This data set includes measurements of Chlorophyll-a concentration, PAR attenuation coefficient, oxygen concenrations, temperature, salinity and nitogenous nutrient concentrations (nitrite, nitrite + nitrate, ammonium, urea), and estimates of Archaea, bacteria and diatom gene concentration based on quantitative PCR.

openCC (other)Apr 2025View details →
edi60/100

Soil Bacteria and Archaea in Macrosystems Biodiversity Project at Harvard Forest 2012

Patterns of biodiversity, such as the increase toward the tropics and the peaked curve during ecological succession, are fundamental phenomena for ecology. Such patterns have multiple, interacting causes, but temperature emerges as a dominant factor across organisms from microbes to trees and mammals, and across terrestrial, marine, and freshwater environments. However, there is little consensus on the underlying mechanisms, even as global temperatures increase and the need to predict their effects becomes more pressing. The purpose of this project is to generate and test theory for how temperature impacts biodiversity through its effect on biochemical processes and metabolic rate. A combination of standardized surveys in the field and controlled experiments in the field and laboratory measure diversity of three taxa -- trees, invertebrates, and microbes -- and key biogeochemical processes of decomposition in seven forests distributed along a geographic gradient of increasing temperature from cold temperate to warm tropical. This field experiment focused on soil microbes. DNA was extracted and purified from soil cores from an array of 21 1m2 subplots. The V4 region of the 16S rRNA genes for bacteria and archaea were amplified and sequenced using Illumina MiSeq by the University of Oklahoma Institute for Environmental Genomics as part of a macrosystems biodiversity and latitude project supported by the National Science Foundation under Cooperative Agreement DEB#1065836.

openCC0Dec 2023View details →
zenodo52/100

Dominant contribution of Asgard archaea to eukaryogenesis (2024) Tobiasson, V., Koonin, E. PROCESSED DATA AND METADATA

<h1>Main data deposit for "Dominant contribution of Asgard archaea to eukaryogenesis".&nbsp;</h1> <p>Victor Tobiasson, Jacob Luo, Yuri I Wolf, Eugene V Koonin</p> <p>Computational Biology Branch, Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA</p> <p><strong>The Origin of eukaryotes is one of the key problems in evolutionary biology. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes inform and constrain evolutionary scenarios of eukaryogenesis. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, without consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis.&nbsp;</strong></p> <div> <div> <div>Version 0.3, updated 180325</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>Main data repository for:</div> <div>Dominant contribution of Asgard archaea to eukaryogenesis (2024)&nbsp;</div> <div>Tobiasson, V., Koonin, E.</div> <div>&nbsp;</div> <div>Contains all final parsed data from the main Eukaryogenesis project&nbsp;</div> <div>investigating the evolutionary ancetries of eukaryotic protein families.&nbsp;</div> <div>&nbsp;</div> <div>Currently (non-static) available at:&nbsp;</div> <div>https://www.biorxiv.org/content/10.1101/2024.10.14.618318v2</div> <div>https://assets-eu.researchsquare.com/files/rs-5352492/v1/2f9c68ae-cf3e-420a-8d29-867b6fb1a878.pdf</div> <div>&nbsp;</div> <div>All code used to generate the data present within this repository available at:&nbsp;</div> <div>https://github.com/VictorTobiasson/eukgen&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### General information</div> <div>&nbsp;</div> <div>To identify associations between prokaryotic and eukaryotic protein families, separate</div> <div>hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed&nbsp;</div> <div>using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using&nbsp;</div> <div>mmseqs2, followed by a multistep data-reduction and multiple sequence alignment (MSA)&nbsp;</div> <div>procedure to generate HMM profiles using hhsuite.&nbsp;</div> <div>&nbsp;</div> <div>A prokaryotic database of 37 million protein sequences was curated from prokaryotic&nbsp;</div> <div>genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins&nbsp;</div> <div>extracted from 146 Asgard genome assemblies. To avoid inclusion of genes present only&nbsp;</div> <div>within a narrow subset of species, possibly resulting from horizontal transfer from&nbsp;</div> <div>eukaryotes post LECA, we reconstructed the &ldquo;soft-core&rdquo; pangenome for each of the 26&nbsp;</div> <div>curated prokaryotic taxonomic classes. These pangenomes include only those genes that&nbsp;</div> <div>are present in at least 67% of the families within each class of Bacteria and Archaea.&nbsp;</div> <div>The initial eukaryotic database consisted of 30 million protein sequences from 993&nbsp;</div> <div>species taken from EukprotV3 and cleaned using mmseqs2 to remove likely prokaryotic&nbsp;</div> <div>contaminants.&nbsp;</div> <div>&nbsp;</div> <div>Both databases were clustered and MSAs constructed for all non, singleton clusters&nbsp;</div> <div>and HMM profiles created. The resulting eukaryotic HMM dataset was queried against&nbsp;</div> <div>the prokaryotic dataset using hhblits to identify sets of homologous protein sequences.&nbsp;</div> <div>Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual</div> <div>&nbsp;sequence set, hereinafter referred to as an Eukaryotic/Prokaryotic Orthologous Cluster&nbsp;</div> <div>(EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes&nbsp;</div> <div>(each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic&nbsp;</div> <div>proteins can be present in multiple EPOCs) that were used for phylogenetic tree&nbsp;</div> <div>construction, annotation, and evolutionary hypothesis testing.&nbsp;</div> <div>&nbsp;</div> <div>To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC,&nbsp;</div> <div>rather than relying on the tree topology directly, we employed a probabilistic approach&nbsp;</div> <div>for evolutionary hypothesis testing using constraint trees. We exhaustively sampled all&nbsp;</div> <div>arrangements of likely sister clades and obtained Expected Likelihood Weights (ELW) for&nbsp;</div> <div>the set of possible sister clade models. As the ELW metric is analogous to model selection&nbsp;</div> <div>confidence, here we take it to be proportional to the probability of a sampled prokaryotic&nbsp;</div> <div>clade to be the true sister group of the given eukaryotic clade among a set of competing&nbsp;</div> <div>sister clades. For each EPOC, our analysis dynamically accounts for long branch outliers&nbsp;</div> <div>and is robust to phylogenetically non-homogenous clades. This analysis is further capable&nbsp;</div> <div>of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a&nbsp;</div> <div>single datapoint for downstream analysis. Our resulting data contains EPOCs annotated&nbsp;</div> <div>using profiles generated from KEGG Orthology Groups (KOGs), each with an MSA generated&nbsp;</div> <div>using muscle5, a maximum likelihood tree inferred using IQtree2 and associated ELW values&nbsp;</div> <div>for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was&nbsp;</div> <div>performed only for those eukaryotic clades that included more than 5 distinct taxonomic&nbsp;</div> <div>labels, with at least one coming from Amorphea and one from Diaphoretickes, the two&nbsp;</div> <div>expansive eukaryotic clades considered to represent either the first or the second&nbsp;</div> <div>bifurcation in the evolution of eukaryotes. Thus, these clades likely represent genes&nbsp;</div> <div>mapping back to the LECA.</div> <div>&nbsp;</div> <div>For further details please see main publication or contact</div> <div>victor.tobiasson@nih.gov</div> <div>eugene.koonin@nih.gov</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### Included files</div> <div>&nbsp;</div> <div>Unless otherwise stated all files contained are tab separated and utf-8 encoded&nbsp;</div> <div>with the first row containing header information.&nbsp;</div> <div>All data entries encoding lists are &ldquo;|&rdquo; (pipe) separated.&nbsp;</div> <div>Fields without data values are filled with string entries of &ldquo;none&rdquo;.</div> <div>&nbsp;</div> <div>--- Databases ---</div> <div>euk72_ep.tar.gz</div> <div>prok2311_as.tar.gz</div> <div>Prok2311As_final_clusters.tsv</div> <div>Euk72Ep_final_clusters.tsv</div> <div>prok2311_as.hmmDB.tar.gz</div> <div>euk72_ep.hmmDB.tar.gz</div> <div>&nbsp;</div> <div>--- Annotation and Curation ---</div> <div>NCBI_taxonomy_species_addendum.tsv</div> <div>NCBI_taxonomy_class_addendum.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>KEGG_category_mapping.tsv</div> <div>KEGG_metadata.tsv</div> <div>&nbsp;</div> <div>--- EPOC data ---</div> <div>EPOC_data.tar.gz</div> <div>EPOC_annotation_KEGG.tsv</div> <div>EPOC_data.tsv</div> <div>EPOC_data.pangenomes_s10.tsv</div> <div>EPOC_data.pangenomes_s25.tsv</div> <div>EPOC_data.pangenomes_s67.tsv</div> <div>EPOC_data.GTDB.tsv</div> <div>&nbsp;</div> <div># euk72_ep.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files&nbsp;</div> <div>constituting the initial eukaryotic mmseqs2 database with taxonomy annotation.&nbsp;</div> <div>Constructed from a pre-selected list of 72 eukaryotic proteomes downloaded from&nbsp;</div> <div>NCBI as well as a &ldquo;clean&rdquo; version of Eukprot, lacking highly prokaryotic-like&nbsp;</div> <div>contaminant sequences.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files constituting the&nbsp;</div> <div>initial prokaryotic mmseqs2 database with taxonomy annotation. Constructed from&nbsp;</div> <div>47545 complete genomes retrieved from NCBI in November 2023.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from prok2311_as non--singleton clusters, contains 26286 profiles.</div> <div>&nbsp;</div> <div># euk72_ep.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from euk72_ep non-singleton clusters, contains 1631704 profiles.</div> <div>&nbsp;</div> <div># NCBI_taxonomy_species_addendum.tsv</div> <div>Taxonomy mapping file with manually curated &lsquo;class&rsquo; level annotation for poorly&nbsp;</div> <div>annotated species.&nbsp;</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>proposed_class_id: Manually assigned NCBI taxid</div> <div>proposed_class_label: NCBI class name</div> <div>org_name: NCBI organism name</div> <div>&nbsp;</div> <div># NCBI_taxonomy_class_addendum.tsv</div> <div>Class revision file mapping poorly populated class level entries to higher order&nbsp;</div> <div>manually curated labels. Also includes information for small classes with shallow&nbsp;</div> <div>taxonomy which are deleted from the EPOC analysis at the level of tree construction.</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>ncbi_class: NCBI taxid of rank corresponding to &lsquo;class&rsquo; following manual&nbsp;</div> <div>amendment as per NCBI_taxonomy_species_addendum.tsv</div> <div>revised_class_id: Manually assigned NCBI taxid of rank corresponding to &lsquo;class&rsquo;</div> <div>revised_class_label: Proposed cleartext name of manually revised revised_class_id&nbsp;</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Final taxonomy at NCBI rank &lsquo;class&rsquo; following revisions for all sequences in Euk72Ep or&nbsp;</div> <div>Prok2311As. These taxonomic labels are used for EPOC tree annotation.&nbsp;</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of manually revised NCBI rank &lsquo;class&rsquo; identifier for annotation</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Final taxonomy at GTDB rank &lsquo;phylum&rsquo; transferred using marker genes from GTDB release 220</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of assigne GTDB phylum</div> <div>&nbsp;</div> <div># Prok2311As_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation&nbsp;&nbsp;</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># Euk72Ep_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># EPOC_data.tar.gz</div> <div>Gunzip-ed directory containing 16035 EPOC folders. Each folder is named corresponding&nbsp;</div> <div>to the eukaryotic cluster representative which generated its profile as an ID&nbsp;</div> <div>Matches the tree_name field in EPOC_data_prok2311As.tsv</div> <div>contains the following files:</div> <div>&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta: sequences for all members of the EPOC</div> <div>&lt;EPOC_ID&gt;.merged.fasta.leaf_mapping: tsv separated file containing taxonomy and tree reduction data</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle: main cropped MSA for tree generation&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.iqtree: IQtree2 output from tree generation</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.treefile.annot: annotated newick tree file with final tree</div> <div>&lt;EPOC_ID&gt;.merged.tree_data.tsv: final parsed tree data with columns matching&nbsp; EPOC_data_prok2311As.tsv</div> <div>&nbsp;</div> <div>EPOCs with more than one possible eukaryotic sister phyla also contains&nbsp;</div> <div>a folder "constraint_analysis" with constraint tree information used for&nbsp;</div> <div>ELW value calculation.&nbsp;</div> <div>&nbsp;</div> <div># EPOC_data.tsv</div> <div>Main resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs)&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class. This is the main</div> <div>data to be used for genereting the core dataset and for data visualistation</div> <div>Contains information regarding tree breakdown, LCA membership and phylogenetic&nbsp;</div> <div>distances between all detected LCAs. Equivalent to the stacked dataframes from all&nbsp;</div> <div>EPOC directories in EPOC_data&nbsp;</div> <div>&nbsp;</div> <div>tree_name: unique index for each EPOC&nbsp;</div> <div>euk_clade_rep: unique index for each annotated eukaryotic clade within each tree_name</div> <div>euk_clade_size: number of original sequences represented by euk_clade_rep</div> <div>euk_clade_weight: metric for taxonomic purity for each euk_clade_rep</div> <div>euk_leaf_clade: boolean indicating whether euk_clade_rep contains a single leaf</div> <div>euk_LCA: lowest taxa spanning all members in euk_clade_rep</div> <div>euk_scope: list of all taxonomic classes in euk_clade_rep</div> <div>euk_scope_len: length of euk_scope list</div> <div>prok_clade_rep: unique index for each annotated prokaryotic clade for each euk_clade_rep</div> <div>prok_clade_size: number of original sequences represented by prok_clade_rep</div> <div>prok_clade_weight: metric for taxonomic purity for each prok_clade_rep</div> <div>prok_leaf_clade: boolean indicating whether prok_clade_rep contains a single leaf</div> <div>prok_taxa: lowest taxa spanning all members in prok_clade_rep</div> <div>dist: tree-distance from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>top_dist: graph-distance (node-distance) from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>raw_stem_length: tree-distance from lowest tree node containing the union of all members of prok_clade_rep and euk_clade_rep to the tree node containing all members of euk_clade_rep</div> <div>median_euk_leaf_dist: median value for all tree distances from the tree node containing all members of euk_clade_rep to the individual leaves</div> <div>stem_length: raw_stem_length/median_euk_leaf_dist</div> <div>logL: log likelihood of best constraint tree constructed</div> <div>deltaL: log likelihood difference between constraint tree for prok_clade_rep and best constraint tree constructed</div> <div>bp-RELL: validation metric from IQtree -trees, see iqtree.org</div> <div>bp-RELL_accept: as above</div> <div>p-KH: as above</div> <div>p-KH_accept: as above</div> <div>p-SH: as above</div> <div>p-SH_accept: as above</div> <div>c-ELW: as above</div> <div>c-ELW_accept: as above</div> <div>p-AU: as above</div> <div>p-AU_accept: as above</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s10.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s25.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 25% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s67.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 67% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.GTDB.tsv</div> <div>Resulting data&nbsp; from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>under revised taxonomy from GTDB based on data from Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.alpha_replicates.tsv</div> <div>Resulting data from 20 repetitions of Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>from a subset of Alphaproteobacterial-derived EPOCs.&nbsp;</div> <div>Identical file structure to EPOC_data.tsv with the addition of:</div> <div>&nbsp;</div> <div>rep: indicating technical replicate number, 0-19</div> <div>&nbsp;</div> <div># EPOC_annotation_KEGG.tsv</div> <div>Parsed HHblits output of HMM profiles generated from KEGG KOGs (KEGG Orthologous Groups)&nbsp;</div> <div>against eukaryotic profiles constituting each EPOC</div> <div>&nbsp;</div> <div>Query: query name equal to tree_name from EPOC_data</div> <div>Target: target name equal to kogid in KEGG_category_mapping and KEGG_metadata</div> <div>Prob: data from HHblits, see https://github.com/soedinglab/hh-suite/wiki</div> <div>E-value : as above</div> <div>P-value : as above</div> <div>Score: as above</div> <div>SS: as above</div> <div>Cols: as above</div> <div>Identities: as above</div> <div>Similarity: as above</div> <div>Sum_probs: as above</div> <div>Query-HMM-start: as above</div> <div>Query-HMM-end: as above</div> <div>Template-HMM-start: as above</div> <div>Template-HMM-end: as above</div> <div>Template_columns: as above</div> <div>Template_Neff : as above</div> <div>Pairwise_cov: calculated pairwise coverage from Query and Target start and end</div> <div>Description: category_name from KEGG_category_mapping</div> <div>&nbsp;</div> <div># KEGG_category_mapping.tsv</div> <div>Mapping of relevant KOG identifiers to their higher order categories as&nbsp;</div> <div>"Maps" "Modules" or "Reactions" as per KEGG see https://www.kegg.jp/kegg/pathway.html</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>category_id: KEGG map, module, or reaction number</div> <div>category_name: cleartext name for KOG identifier</div> <div>&nbsp;</div> <div># KEGG_metadata.tsv</div> <div>File mapping KOGs to BRITE classification and to additional databases of chemical properties.</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>name: cleartext name for KOG identifier</div> <div>brite_A: list of BRITE-A sets including KOG</div> <div>brite_B: list of BRITE-A sets including KOG</div> <div>brite_C: list of BRITE-A sets including KOG</div> <div>EC: list of Enzyme commission numbers associated with KOG, see https://enzyme.expasy.org/</div> <div>TC: list of transporter classification numbers associated with KOG, see https://www.tcdb.org/</div> <div>RN: list of KEGG reaction numbers associated with KOG</div> <div>CA: list of CAZY numbers associated with KOG, see http://www.cazy.org/</div> <div>GO: list of GO terms associated with KOG, see https://geneontology.org/</div> </div> <div>&nbsp;</div> </div>

opencc-by-4.0Oct 2024View details →
zenodo48/100

DADA2 formatted 16S rRNA gene sequences for both bacteria & archaea

<p><strong><em>This version is to stay up to date with the improvements and increase in 16S rRNA gene sequences (SSU) added to the GTDB release 220.&nbsp; Please read this post for the stats on the updates. </em></strong><strong><em>https://gtdb.ecogenomic.org/stats/r220 </em></strong><strong><em>.</em></strong><strong><em> </em></strong></p> <p><strong><em>There has been no change to the RDP-RefSeq reference database please use previous versions.</em></strong></p> <p><strong><em>If anyone has concerns&nbsp;with MAG extracted 16S rRNA gene contamination concerns, then I suggest that they contact the curators of GTDB themselves because it is outside of my role with these resources designed for DADA2 usage only. </em></strong></p> <p><strong><em>Another concern that was raised was the orientation of the DB sequences, to get past this problem please use the tryRC = TRUE argument in the assignTaxonomy command within DADA2, this will search your ASVs in the reverse complement as well.&nbsp;&nbsp;</em></strong></p> <p>The bacterial and archaeal 16S rRNA gene sequence databases were collated from various sources and formatted to use the "assignTaxonomy" command within the DADA2 pipeline. The data was converted to suite DADA2 format by Alishum Ali.</p> <ol> <li>Genome Taxonomy Database (GTDB): The new version of our dada2 formatted GTDB reference sequences now contains 58102 bacteria and 3672 archaea full 16S rRNA gene sequences. If you wonder why there are fewer species with 16S rRNA, that is because some metagenomics-assembled genomes (MAGs) lack the 16S gene and thus cannot be extracted.&nbsp; The database was downloaded from <a href="https://data.ace.uq.edu.au/public/gtdb/data/releases/release95/">https://data.ace.uq.edu.au/public/gtdb/data/releases/</a> on 24/10/2024. Please read the release notes and file descriptions.&nbsp;</li> </ol> <p>The formatting to DADA2 was done using simple awk bash scripts. The script takes as input a fasta file and a tab-delimited taxonomy file (slightly edited to remove special characters) and then it outputs a fasta file with all 7 taxonomy ranks separated by ";" as required for DADA2 compatibility. Additionally, we have concatenated the unique sequence GTDB ID to the species entry (but replaced the "." with an " _". We see this as an important QC step to highlight the issues/confidence associated with short-read taxonomy assignment at the finer rank levels.</p> <p>Also, this update includes two other files that you can use with the assignTaxonomy and addSpecies commands in DADA2.</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

gapseq reference sequence databases for Bacteria and Archaea

<p>The repository contains the protein sequences used by <a href="https://github.com/jotech/gapseq">gapseq</a> to predict the presence of metabolic reactions and to construct metabolic models.</p> <p>The workflow using gapseq to generate this set of reference protein sequences:</p> <p>&nbsp;</p> <p>```sh</p> <p># delete all "old" data<br>rm dat/seq/Bacteria/rev/*.fasta<br>rm dat/seq/Bacteria/unrev/*.fasta<br>rm dat/seq/Bacteria/rxn/*.fasta<br>rm dat/seq/Archaea/rev/*.fasta<br>rm dat/seq/Archaea/unrev/*.fasta<br>rm dat/seq/Archaea/rxn/*.fasta</p> <p># run gapseq find to re-download everything#<br># the genome is irrelevant as no blasting is performed ('-x')<br>gapseq find -p all -t Bacteria -n -x -U toy/ecoli.faa.gz &gt; bac_update.log 2&gt;&amp;1<br>gapseq find -p all -t Archaea -n -x -U toy/ecoli.faa.gz &gt; ar_update.log 2&gt;&amp;1</p> <p># create all sequence .tar.gz archives (rev/unrev/rxn)<br>cd dat/seq/Bacteria/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../</p> <p># create md5sum table for all tar.gz archives<br>cd dat/seq/<br>find -mindepth 2 -type f -name "*.tar.gz" -exec md5sum {} \; &gt; md5sums.txt</p> <p># create taxon-specific final archive for Zenodo upload<br>tar -czvf Bacteria.tar.gz Bacteria/*/*.tar.gz<br>tar -czvf Archaea.tar.gz Archaea/*/*.tar.gz</p> <p># Upload Bacteria.tar.gz, Archaea.tar.gz, and md5sums.txt &nbsp;to Zenodo via the web-interface</p> <p>```</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Bacteria and archaea of the Columbia and Willamette Rivers, 16S rRNA gene amplicon library metadata

<p>Bacterial and archaeal communities in the Columbia and Willamette Rivers in the Portland, OR, USA, region were characterized by 16S rRNA gene amplicon sequencing as part of the Lewis &amp; Clark College spring 2022 Microbial Ecology course. Whole-water (&gt;0.2 &micro;m) samples were collected from: the Willamette River; the Columbia River above the confluence with the Willamette; and the Columbia River just downstream of the confluence with the Willamette.</p> <p>This dataset provides additional metadata to supplement the DNA sequences archived with the NCBI SRA at&nbsp;<a href="https://www.ncbi.nlm.nih.gov/sra/PRJNA865380">https://www.ncbi.nlm.nih.gov/sra/PRJNA865380</a></p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

dudesdb_201709 - Archaea and Bacteria - RefSeq - Complete Genomes

<p>bowtie2 index and dudes database for the set of Archaeal and Bacterial complete genomes from NCBI RefSeq, dating from 2017-09. The dudes database was made based on accession version numbers (DUDesDB.py option -m "av").</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

dudesdb_201503 - Archaea and Bacteria - RefSeq - Complete Genomes

<p>bowtie2 index and dudes database (.ddb for version 0.06 and .npz for version 0.07) for the set of Archaeal and Bacterial complete genomes from NCBI RefSeq, dating from 2015-03. The dudes database was made based on accession version numbers (DUDesDB.py option -m "av").</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

Genomic insights into the Archaea inhabiting an Australian radioactive legacy site

<p><strong>Abstract</strong></p> <p>During the 1960s, small quantities of radioactive materials were co-disposed with chemical waste at the Little Forest Legacy Site (LFLS, Sydney, Australia). The microbial function and population dynamics during a rainfall event using shotgun metagenomics has been previously investigated. This revealed a broad abundance of candidate and potentially undescribed taxa in this iron-rich, radionuclide-contaminated environment.</p> <p>Here, applying genome-based metagenomic methods, we recovered 37 refined archaeal bins (&ge;50% completeness, &le;10% redundancy) from 10 different major lineages. They were mostly included in 4 proposed lineages within the DPANN supergroup (LFWA-I to IV) and <em>Methanoperedenaceae</em>.</p> <p>The new <em>Methanoperedens</em> spp. bins, together with previously published data, suggests a potentially widespread ability to use nitrate (or nitrite) and metal ions as electron acceptors during the anaerobic oxidation of methane by <em>Methanoperedens</em> spp.</p> <p>While most of the new DPANN lineages show reduced genomes with limited central metabolism typical of other DPANN, the candidate species from the proposed LFWA-III lineage show some unusual features not often present in DPANN genomes, i.e. a more comprehensive central metabolism and anabolic capabilities.While there is still some uncertainty about the capabilities of LFW-121_3 and closely related archaea for the biosynthesis of nucleotides <em>de novo</em>, and amino acids, it is to date the most promising candidate to be the first <em>bona fide</em> free-living DPANN archaeon.</p> <p><strong>Repository Contents</strong></p> <p><em>genomes.tar:</em> includes each of the reference and novel assembled MAG/bins used for analysis in the main manuscript as .tar.gz compressed folders. Each genome folder contains the output from:</p> <ul> <li>Anvi&#39;o,</li> <li>EggNOG analysis with arNOG library,</li> <li>InterProScan analysis,</li> <li>rRNA search with Barrnap,</li> <li>output from searching high heme cytochromes (&ge;10 heme binding sites in a single protein),</li> <li>CAZy search output,</li> <li>MEROPS output (BLASTp),</li> <li>PSORTb, and</li> <li>TCDB.</li> </ul> <p>In the case of the new MAGs (i.e. LFW_Bin_00*, referred in the paper as LFW-*), some additional contents are included:</p> <ul> <li>tRNAs from tRNAscan-SE, and</li> <li>Prokka annotation files.</li> </ul> <p><em>pangenomics.tar</em>: includes the Anvi&#39;o files for the pangenomic analysis of:</p> <ul> <li><em>anme-2d.tar.gz</em>: pangenome of <em>Methanoperedens</em> spp. based on 4 reference and 6 novel MAGs clustered with an MCL inflation value of 6.0. It also includes the output of ANI analysis (via pyANI) and AAI (via CompareM).</li> <li><em>LFWA-III.tar.gz</em>: pangenome analysis of the LFWA-III lineage (&#39;Gugararchaeaceae) based on 2 reference and 8 novel MAGs clustered at inflation values of 1.0, 1.5 and 2.0. It also includes the output of ANI analysis (via pyANI) and AAI (via CompareM).</li> </ul> <p><em>phylogeny.tar</em>: includes the files required for the phylogenomic/phylogenetic analyses shown in the paper and the Supplementary Information:</p> <ul> <li><em>rp44.tar.gz</em>: phylogenomic analysis based on 44 universal and archaeal-specific ribosomal proteins (Figure 1 in paper). It also includes annotation files for iTOL.</li> <li><em>NarG.tar.gz</em>: phylogeny of NarG and relate molybdopterin oxidoreductase proteins (Figure 5 in paper).</li> <li><em>LysJ-ArgD.tar.gz</em>: phylogeny of LysJ/ArgD proteins and their orthologous (Figure S3 in paper).</li> </ul>

opencc-by-4.0Aug 2019View details →
zenodo40/100

Fig. 12 in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 12. Basidiospores under scanning electron microscope of Russula japonica Hongo. A–C. GDGM79699. D. GDGM79704. E–F. GDGM79706. Scale bars = 1 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 7. Russula reticulofolia Y in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 7. Russula reticulofolia Y.Song sp. nov., holotype (GDGM79559) (A–B, E–F), GDGM79560 (C– D). A–D. Fruiting bodies. E–F. Basidiospores under scanning electron microscope. Scale bars: A–D = 1 cm; E–F = 2 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 6. Russula lacteocarpa Y in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 6. Russula lacteocarpa Y.Song sp. nov., holotype (GDGM79555). A. Basidia. B. Pleurocystidia. C. Cheilocystidia. D. Terminal cells and incrustations. E. Caulocystidia. F. Pileocystidia. Scale bars = 10 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 4. Russula cylindrica Y in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 4. Russula cylindrica Y.Song sp. nov., holotype (GDGM79551). A. Basidia. B. Pleurocystidia. C. Cheilocystidia. D. Pileocystidia. E. Caulocystidia. Scale bars = 10 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 13 in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 13. Russula japonica Hongo (GDGM79697). A. Basidia. B. Cheilocystidia. C. Pleurocystidia. D. Pileocystidia. E. Caulocystidia. Scale bars = 10 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 2 in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 2. Maximum Likelihood tree of Russula Pers. based on concatenated 5-locus (LSU, mtSSU, rpb1, rpb2 and tef1) sequences. Bootstrap values higher than 50% were shown around the nodes. Sequences generated in this study are shown in bold blue.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 1 in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 1. Maximum Likelihood tree of Russula Pers. based on ITS sequences, bootstrap values higher than 50% were shown around the nodes. Sequences generated in this study are shown in bold blue.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 9. Russula callainomarginis J.F.Liang & J in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 9. Russula callainomarginis J.F.Liang &amp; J.Song (GDGM79715). A –B. Fruiting bodies. C. Terminal elements in pileipellis. D–G. Basidia. H–I. Basidiospores under scanning electron microscope. Scale bars: A–B = 1 cm; C = 25 µm; D–G = 10 µm; H–I = 1 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 11 in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 11. Fruiting bodies of Russula japonica Hongo. A–B. GDGM79699. C–D. GDGM79710. E. GDGM79706. F. GDGM79704. Scale bars = 1 cm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 3. Russula cylindrica Y in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 3. Russula cylindrica Y.Song sp. nov., holotype GDGM79551(A–B, G–H), specimen GDGM79553 (C–D), specimen GDGM79552 (E–F). A–F. Fruiting bodies. G–H. Basidiospores under scanning electron microscope. Scale bars: A–F = 1 cm; G–H = 1 µm.

opencc-by-4.0Mar 2023View details →
zenodo40/100

Fig. 5. Russula lacteocarpa Y in Species of Russula subgenera Archaeae, Compactae and Brevipedum (Russulaceae, Basidiomycota) from Dinghushan Biosphere Reserve

Fig. 5. Russula lacteocarpa Y.Song sp. nov., holotype (GDGM79555). A –B. Fruiting bodies. C. Basidia and pleurocystidia in hymenium. D–E. Basidiospores under scanning electron microscope. Scale bars: A–B = 1 cm; C = 10 µm; D–E = 1 µm.

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record