Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.9.0
Dataset results
35 results for “Microbial genomics”
Geochemical, physicochemical, and genomic data from a continental-scale survey of microbial diversity in Antarctic soils (2003-2023)
This data package offers comprehensive insights into Antarctic soil microbial diversity and composition. From 2003 to 2023, a total of 186 samples were collected from diverse locations spanning the Antarctic Peninsula to East Antarctica, representing a wide range of environmental gradients and climatic conditions. Soils were stored at -20°C to preserve their integrity for downstream analyses. This data package integrates cultivation-independent sequencing of prokaryotic and fungal communities alongside a robust cultivation-dependent culture collection to enable direct comparisons across microbial diversity assessment methods. Accompanying geochemical, physicochemical, and environmental parameters provide critical context for biogeographical analyses, offering a valuable resource for studying microbial adaptations and community dynamics in extreme Antarctic environments.
Data for manuscript: A framework for integrating genomics, microbial traits, and ecosystem biogeochemistry
<p>Support manuscript: A framework for integrating genomics, microbial traits, and ecosystem biogeochemistry. </p> <p>Dataset includes 1) model and analysis, and 2) supplemental data. </p> <p>In the "model_analysis" file, we include the ecosys model source code, the modeling runs, and the modeling results. The detailed introduction is in README.md file. </p> <p><strong>Acknowledgments</strong></p> <p>We thank the EMERGE Biology Integration Institute Coordinators (members listed in Supplementary Information) for project guidance and management. This research is a contribution of the EMERGE Biology Integration Institute, funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070 (V.I.R., R.K.V., S.R.S., M.B.S., E.L.B., and the EMERGE Coordinators). Additional support for individual contributors included the following. Z.L. was additionally supported by Lawrence Livermore National Laboratory under the auspices of the U.S. Department of Energy under contract DE-AC52-07NA27344. W.J.R. was supported by the Belowground Biogeochemistry Scientific Focus Area and U.K. was supported by the Watershed Function Science Area, both funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research under contract no. DE-AC02-05CH11231. G.L.M. was supported by the LLNL "Microbes Persist" Soil Microbiome Scientific Focus Area SCW1632 and an associated KBase award SCW1746. N.J.B. was supported by the US Department of Energy, Office of Science (BER), Early Career Research Program (#FP00005182). B.J.W. was supported by an Australian Research Council Future Fellowship (#FT210100521). J.T. was supported by the Laboratory Directed Research and Development Program of Lawrence Berkeley National Laboratory. </p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council’s grant 4.3-2021-00164. This research used resources of the National Energy Research Scientific Computing Center (NERSC) which is a U.S. Department of Energy Office of Science user facility. This research used the Lawrencium computational cluster resource provided by the IT Division at the Lawrence Berkeley National Laboratory (Supported by the Director, Office of Science, Office of Basic Energy Sciences, of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231). </p> <p><strong>Full list of the EMERGE Biology Integration Institute Coordinators and Affiliations</strong></p> <p>Eoin L. Brodie1,2, Sarah C. Bagby3, Jeffrey P. Chanton4, Jessica G. Ernakovich5, Regis Ferriere6,7, Suzanne B. Hodgkins8, William J. Riley1, Virginia I. Rich8,9, Scott R. Saleska6, Matthew B. Sullivan8,9,10, Ruth K. Varner11, Gene W. Tyson12, Malak M. Tfaily13, Ahmed A. Zayed8,9<br> 1Climate and Ecosystem Sciences Division, Lawrence Berkeley National Laboratory; Berkeley, CA 94720, USA.<br>2Department of Environmental Science, Policy and Management, University of California; Berkeley, CA 94720, USA.<br>3Department of Biology, Case Western Reserve University; Cleveland, OH, USA, 44106<br>4Earth Ocean and Atmospheric Sciences, Florida State University; Tallahassee, FL, USA<br>5Department of Natural Resources and the Environment, University of New Hampshire;<br>Durham, NH, USA 03824<br>6Department of Ecology and Evolutionary Biology, University of Arizona; Tucson, AZ,<br>85721, USA<br>7Institut de Biologie de l’ENS, Université Paris Sciences & Lettres; Paris, 75005, France<br>8Department of Microbiology, The Ohio State University; Columbus, OH, USA, 43210<br>9Center of Microbiome Science, The Ohio State University; Columbus, Ohio 43210, USA.<br>10Department of Civil, Environmental and Geodetic engineering, The Ohio State University; Columbus, Ohio 43210, USA.<br>11Department of Earth Sciences and Institute for the Study of Earth, Oceans and Space, University of New Hampshire; Durham, NH 03824, USA.<br>12Centre for Microbiome Research, School of Biomedical Sciences, Queensland University<br>of Technology (QUT), Translational Research Institute; Woolloongabba, QLD, Australia<br>13Department of Environmental Science, University of Arizona; Tucson, AZ, 85721, USA</p>
Rumen_Microbial_Genomes_from_Cow_Fed_with_Red_Seaweed_Additives
<p>This dataset represents the 3180 non-redundant species-level rumen microbial genomes established through the integration of MAGs recovered from the rumen of cows that were fed with red seaweed or not, publicly available rumen MAGs and isolate genomes from Hungate collection. In order to match those old genome names with the new genome names used in the manuscript, please refer to supplementary table 3.</p>
Soil properties in agricultural systems affect microbial genomic traits
<p>Code and supplementary table </p>
Datasets for 'Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation'
<p>Datasets for the paper '<strong>Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation</strong>'</p> <p>Files:</p> <ul> <li>616k* - Files for the analysis of 661k bacterial genomes from the SRA (note typo 616-661k). Includes mandrake output and input files (.npz)</li> <li>gps_acc - Files for the analysis of 20k S. pneumoniae accessory genomes from the GPS project. Original accessory matrix is gps_gene_presence_absence.Rtab</li> <li>sc2million_v1* - Files for the analysis of ~1M SARS-CoV-2 genomes. sc2million_v3.npz are the input distances.</li> <li>sce<commit hash>.qdrep - Nvidia systems profile of code at that commit hash</li> <li>sce<commit hash>.ncu-rep - Nvidia kernel profile of code at that commit hash</li> </ul>
Data Pack fro VirSorter: mining viral signal from microbial genomic data
<p>This is the data pack for VirSorter, the publication of which by Roux et al. titled "<strong>VirSorter: mining viral signal from microbial genomic data</strong>" appeared in PeerJ on 2015-05-28 (<a href="https://doi.org/10.7717/peerj.985">doi:10.7717/peerj.985</a>).</p> <p>Most up-to-date tutorials and the code for VirSorter can be found at the GitHub repository <a href="https://github.com/simroux/VirSorter">https://github.com/simroux/VirSorter</a>.</p> <p>The original source of this data pack was here (last accessed 2018-02-03): <a href="http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz">http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz</a></p>
The mOTUs online database provides web-accessible genomic context to taxonomic profiling of microbial communities - Supplementary Tables
<p><strong>Supplementary Table 1:</strong></p> <p>A map between each of the genomes in mOTUs-db (3’747’151), the associated study and its metagenomic sample (in case of MAGs).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> STUDY → Unique mOTUs-db name of the study</code><br><code> IS_MAG → True if genome is a MAG, otherwise False </code><br><code> METAGENOMIC_SAMPLE → Unique name of the metagenomic sample or NA in case of non-MAG genome</code></p> <p>Example:</p> <p><code> GENOME STUDY IS_MAG METAGENOMIC_SAMPLE</code><br><code> ---------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_MAG_00000001 ACIN21-1 True ACIN21-1_SAMN05421555_METAG</code><br><code> RSGB23-1_GCA-006096615-V1_GENO_10000001 RSGB23-1 False NA</code></p> <p><strong>Supplementary Table 2:</strong></p> <p>A map between all non-MAG genomes (919’090) and their source (e.g. Refseq or JGI).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> SOURCE_SAMPLE_LINK → Link to the original location of this genome</code></p> <p>Example:</p> <p><code> #GENOME SOURCE_SAMPLE_LINK</code><br><code> --------------------------------------------------------------------------------------------------------</code><br><code> JGIG23-1_GA0055041_GENO_10000001 https://gold.jgi.doe.gov/analysis_project?id=Ga0055041</code><br><code> RSGB23-1_GCA-006717865-V1_GENO_10000001 https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_006717865.1</code></p> <p><strong>Supplementary Table 3:</strong></p> <p>A list of all metagenomic studies processed for the mOTUs-db, their number of samples, the number of reconstructed MAGs and the associated publication.</p> <p>Columns:</p> <p><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> BIOPROJECT --> Public identifier (NCBI/JGI) of metagenomic sequencing project</code><br><code> SAMPLES --> Number of metagenomic samples</code><br><code> MAGs --> Number of reconstructed MAGs</code><br><code> PUBLICATION --> Link to publication</code></p> <p>Example:</p> <p><code> STUDY BIOPROJECT SAMPLES MAGs PUBLICATION</code><br><code> -------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1 PRJEB44456 58 1,110 https://www.nature.com/articles/s42003-021-02112-2</code></p> <p><strong>Supplementary Table 4:</strong></p> <p>Mapping between mOTUs-db sample identifier, the associated biosample and the environment.</p> <p>Columns:</p> <p><code> SAMPLE --> Unique mOTUS-db sample identifier</code><br><code> BIOSAMPLE --> Public identifier (NCBI/JGI) of metagenomic sample</code><br><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> ENVIRONMENT --> Environment of metagenomic sample</code><br><code> SOURCE_SAMPLE_LINK --> Link to the original location of this sample</code></p> <p>Example:</p> <p><code> #SAMPLE BIOSAMPLE STUDY ENVIRONMENT SOURCE_SAMPLE_LINK</code><br><code> ---------------------------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_METAG SAMN05421555 ACIN21-1 marine https://www.ncbi.nlm.nih.gov/biosample/SAMN05421555/</code></p> <p><strong>Supplementary Table 5:</strong></p> <p>A list of environments covered in the mOTUs-db mapped to the respective NCBI taxonomy (if possible)</p> <p>Columns:</p> <p><code> TERM --> Unique environment name</code><br><code> NCBI TAXONOMY ID --> Link to the NCBI taxonomy</code></p> <p>Example:</p> <p><code> TERM NCBI TAXONOMY ID</code><br><code> ----------------------------------------------</code><br><code> activated sludge metagenome NCBI:txid942017</code><br><code> air metagenome NCBI:txid655179</code></p>
Supplement: Genomic and phenotypic imprints of microbial domestication on cheese starter cultures
<p>This data repository contains the latest version of the supplemental data, code and figures for the following manuscript:</p> <p>Vincent Somerville, Nadine Thierer, Remo S. Schmidt, Alexandra Roetschi, Laurianne Braillard, Monika Haueter, Hélène Berthoud, Noam Shani , Ueli von Ah, Florent Mazel & Philipp Engel 2024. “Genomic and phenotypic imprints of Neolithic domestication on Cheese Starter cultures” </p> <p>All additional genomic data is stored under the following Bioprojects on NCBI:</p> <p>PRJNA717134<br>PRJNA1048529<br>PRJNA1083966<br>PRJNA1157897</p> <p> </p> <p> </p> <p> </p>
High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes
<p>This project contains anvi'o profiles and contigs databases that is used and/or referenced from the Lee STM and Khan SA, <em>et al.</em> study titled "<strong>High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes</strong>". The pre-print of this study is available via http://dx.doi.org/10.1101/090993.</p> <p>To be able to work with the data files you will need anvi'o <strong>v2.1.0</strong> to be installed on your system. For installation instructions, or to have access to a Docker image for anvi'o, please visit this URL: http://merenlab.org/software/anvio</p> <p>Public data:</p> <ul> <li><strong>ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION.tar.gz</strong>: Data files for a quick visualization of the 97 MAGs and their distribution across the two FMT recipients. A run script in the archive explains how to use this data.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MERGED-PROFILE.tar.gz</strong>: The merged anvi'o profile for the entire data, which also contains a collection of 97 MAGs identified in the donor. The profile database contains no hierarchical clustering of contigs, however, individual MAGs can be displayed via the following notation since the collection 'MAGs' describe the organization of contigs in each MAG referenced from the dataset `ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION`, as well as from the paper: "anvi-refine -c CONTIGS.db -p PROFILE.db -C MAGs -b <em>FMT-Donor_MAG_00054</em>". All MAG names are in the supplementary tables in our paper.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MAGs-SUMMARY.tar.gz</strong>: A static HTML website that contains FASTA files for each MAG, and TAB-delimited matrices for coverage and detection values, and others. After unpacking, you can double-click the index.html file. </li> </ul>
Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments
<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2) A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p> </p> <p>Manuscript title Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>
Lemonade Creek, Yellowstone National Park, USA - Microbial Community Analysis - Cyanidiophyceae genome data for HGT analysis
<p>This dataset consists of 12 metagenome samples that were collected from one of three environments in Yellowstone National Park:</p> <ul> <li>4 samples (numbered 1, 2, 3, 4) are from the "CreekBiofilm" environment.</li> <li>4 samples (1, 2, 3, 4) are from the "Endolithic" environment.</li> <li>4 samples (1, 2, 3, 4) are from the "Soil" environment.</li> </ul> <p>We have found that there are two species of cyanidiophyceae present in these samples: one *Galdieria sulphuraria* (the `*Gsulp*` files) and one *Cyanidioschyzon merolae* (the `*Cmer*` files). For each of these species I extracted their contigs from the metagenome assembly if they had >=10% of their lengths covered by hits with >90% ID to the respective reference genome (i.e., contigs with >10% coverage of hits with >90% ID to a given reference genome). The majority of contigs have >90% hit coverage however, to prevent removal of contigs with novel sequences (arising via HGT or other processes), I used a lenient threshold of 10%. The naming of the files indicate which sample the contigs are from and which of the two cyanidiophyceae species they are putatively from. NOTE: that there are very few predicted proteins in the `YNP_CreekBiofilm_*_Gsulp*` files. This is because this environment is completely dominated by the other algal species and so we recovered very few contigs from this species from these environments.</p>
Designing a Synthetic Microbial Community through Genome Metabolic Modeling to enhance Plant-Microbe Interaction
<p>Supplementary data 1 - <strong>Reconstructed genome-scale metabolic networks from MAGs and Hosts</strong></p> <p>Supplementary data 2 - P<strong>lant growth-promoting traits among members of the minimal community</strong></p> <p> </p> <p>Manipulating the rhizosphere microbial community through beneficial microorganism inoculation has gained interest in improving crop productivity and stress resistance. Synthetic microbial communities, known as SynCom, mimic natural microbial compositions while reducing the number of components. However, achieving this goal requires a comprehensive understanding of natural microbial communities and a careful selection of compatible microorganisms with colonization traits, which still pose challenges. In this study, we employed an <em>in-silico</em> approach using genome metabolic modeling to design a synthetic microbial community aimed at improving the yield of important crop plants. We used a targeted approach to select a minimal community (MinCom) encompassing essential compounds for microbial metabolism and compounds relevant to plant interactions. This resulted in a reduction of the initial community size by approximately 4.5-fold. Notably, the MinCom retained crucial genes associated with essential plant growth-promoting traits, such as iron acquisition, EPS production, potassium solubilization, nitrogen fixation, GABA production, and IAA-related tryptophan metabolism. Furthermore, our selection process for the SymCom, based on a comprehensive understanding of microbe-microbe-plant interactions, yielded a set of six hub species that displayed notable taxonomic novelty, including members of the Eremiobacterota and Verrucomicrobiota phyla. Our study contributes to the growing body of research on synthetic microbial communities and their potential to enhance agricultural practices. The insights gained from our in-silico approach and the selection of hub species pave the way for further investigations into the development of tailored microbial communities that can optimize crop productivity and improve stress resilience in agricultural systems.</p>
Data from: Assembling microbial communities: a genomic analysis of a natural experiment in neotropical bamboo internodes
Open the record for dataset details and reuse information.
Data from: Whole-genome duplication and host genotype affect rhizosphere microbial communities
<p>The composition of microbial communities found in association with plants is influenced by host phenotype and genotype. Yet, the ways in which specific genetic architectures of host plants shape microbiomes is unknown. Genome duplication events are common in the evolutionary history of plants, influence many important plant traits, and, thus, they may affect associated microbial communities. Using experimentally induced whole genome duplication (WGD), we tested the effect of WGD on rhizosphere bacterial communities in<i> Arabidopsis thaliana</i>. We performed 16S rRNA amplicon sequencing to characterize differences between microbiomes associated with specific host genetic backgrounds (Columbia <i>vs</i>. Landsberg) and ploidy levels (diploid <i>vs</i>. tetraploid). We modeled relative abundances of bacterial taxa using the Dirichlet and multinomial distributions via a hierarchical Bayesian approach. We found that host genetic background and ploidy level affected rhizosphere community composition. We then tested to what extent microbiomes derived from a specific genetic background or ploidy level affected plant performance by inoculating sterile seedlings with microbial communities harvested from a prior generation. We found a negative effect of the tetraploid Columbia microbiome on growth of all four plant genetic backgrounds. These findings suggest an interplay between host genetic architecture and bacterial community assembly with potential ramifications for host fitness. Moreover, we uncovered an intriguing role of ploidy-level for shaping plant microbiomes. Given the prevalence of ploidy-level variation in both wild and managed plant populations, the effects on microbiomes of this aspect of host genetic architecture could be a widespread driver of differences in plant microbiomes.</p>
Distributions of alternative start codons for all microbial Refseq genomes
<p>Distributions of alternative start codons for all microbial Refseq genomes</p>
Distribution of alternative start codons for each microbial Refseq genome
<p>Distributions of alternative start codons for each microbial Refseq genome. </p>
Lemonade Creek, Yellowstone National Park, USA - Microbial Community Analysis - Genome and Transcriptome Data
<p>Genome and Transcriptome data used for analysis of microbial community function over a diurnal cycle in Lemonade Creek, Yellowstone National Park, USA.</p> <p> </p> <p><code>mags.tar</code> Non-redundant metagenome data (genome assemblies, predicted genes, and gene functional annotations).</p> <p> </p> <p>In each directory are the the following files:</p> <p>- <code>*.mRNA.faa</code> protein sequences of protein-coding genes</p> <p>- <code>*.mRNA.fna</code> nucleotide sequences of protein-coding genes</p> <p>- <code>*.mRNA.gff3</code> genomic location of protein-coding genes</p> <p>- <code>*.mRNA.emapper.tsv</code> eggNOG-mapper annotations for the protein-coding genes</p> <p>- <code>*.mRNA.interproscan.gff3</code> InterProScan annotations for the protein-coding genes</p> <p> </p> <p>In the <code>prokaryote</code> directory there are the following files:</p> <p>- <code>*.rRNA.fna</code> nucleotide sequences of rRNA genes</p> <p>- <code>*.rRNA.gff3</code> genomic location of rRNA genes</p> <p>- <code>*.tRNA.fna</code> nucleotide sequences of tRNA genes</p> <p>- <code>*.tRNA.gff3</code> genomic location of tRNA genes</p> <p>- <code>*.other.fna</code> nucleotide sequences of other genes (i.e., CRISPR, ncRNA, oriC, regulatory_region, repeat_region, tmRNA - if any were predicted)</p> <p>- <code>*.other.gff3</code> genomic location of other genes</p> <p> </p> <p><strong>Eukaryotes</strong></p> <p>Five MAGs from other eukaryotes that were assembled from a coassembly of the Soil samples.</p> <p> </p> <p><strong>Prokaryotes</strong></p> <p>The final dereplicated prokaryote MAGs (at 95% ID). The two <code>*stats*</code> files list the taxonomic information (from <code>GTDB-Tk</code>), completeness (from <code>CheckM</code>), and assembly stats (from the <code>stats.sh</code> script from the <code>bbmap</code> package) for each of the prokaryotic MAGs + the number of predicted protein-coding and non-protein-coding genes predicted in each MAG.</p> <p> </p> <p><strong>Viruses</strong></p> <p>The final dereplicated viral MAGs and vOTUs.</p> <p> </p> <p> </p> <p> </p> <p><code>read_mapping.tar</code> Abundance results from metagenome and metatranscriptome read mapping analysis against the non-redundant metagenome data and predicted genes (respectively). This analysis includes the cyanidiophyceae reference nuclear and organelle genomes.</p> <p> </p> <p><strong>mags</strong></p> <p>Results from <code>bbmaps</code> alignment of metagenome reads against a database of non-redudant metagenome MAGs + cyanidiophyceae reference nuclear and organelle genomes. <code>CoverM</code> was used to calculate MAG abundances.</p> <p> </p> <p><strong>genes</strong></p> <p><code>Salmon</code> abundance quantification of PolyA and RiboMinus metatranscriptome reads mapped against the predicted genes in the non-redudant metagenome MAGs + cyanidiophyceae reference nuclear and organelle genomes.</p>
Microbial Genomes and Metagenomes Workshop
Open the record for dataset details and reuse information.
Data from: Distinctive microbial community and genome structure in coastal seawater from a human-made port and nearby offshore island in northern Taiwan facing the Northwestern Pacific Ocean
<p><span>Pollution in human-made fishing ports caused by petroleum </span><span>from</span><span> boats, dead fish, toxic </span><span>chemicals</span><span>, and effluent </span><span>poses</span><span> a challenge to the organisms in seawater. To decipher the impact of pollution on the microbiome, we collected surface water </span><span>from</span><span> a fishing port and a nearby offshore island in northern Taiwan facing the </span><span>Northwestern Pacific Ocean. By employing 16S </span><span>rRNA gene</span><span> amplicon sequencing and whole-genome shotgun sequencing, we discovered that </span><span>Rhodobacteraceae, Vibrionaceae, and Oceanospirillaceae emerged as the dominant species in the fishing port</span><span>,</span><span> where we found many genes harboring the functions of </span><span>antibiotic</span><span> resistance (</span><span>ansamycin, nitroimidazole, and aminocoumarin), metal tolerance (copper, chromium, iron and multimetal), virulence factors (chemotaxis, flagella, T3SS1), carbohydrate metabolism (biofilm formation and remodeling of bacterial cell </span><span>walls</span><span>), nitrogen metabolism (denitrification, N<sub>2</sub> fixation, and ammonium assimilation), and ABC transporters (phosphate, lipopolysaccharide, and branched-chain amino </span><span>acids</span><span>). The dominant bacteria at the nearby offshore island (</span><span>Alteromonadaceae, Cryomorphaceae, Flavobacteriaceae, Litoricolaceae, and Rhodobacteraceae) were partly similar to those in the South China Sea and the East China Sea. Furthermore, we inferred</span><span> that</span><span> the microbial community network of </span><span>the cooccurrence</span><span> of dominant bacteria </span><span>on the</span><span> offshore island was connected to dominant bacteria in </span><span>the </span><span>fishing port by mutual</span> <span>exclusion. By examining the assembled microbial genomes collected from the coastal seawater of the fishing port, we revealed four genomic islands containing large gene-containing sequences</span><span>,</span><span> including phage integrase, DNA</span> <span>invertase, restriction enzyme, DNA gyrase inhibitor, and antitoxin HigA-1.</span><span> In this study, </span><span>we provided </span><span>clues </span><span>for the possibility of genomic islands as the units of horizontal transfer and as the tools of microbes for facilitating adaptation in a human-made port environment.</span></p>
Viral genomes, functional genes, and microbial genomes of 144 activated sludge samples taken from 54 WWTPs across 13 countries on a global scale
<p>The study dataset contains 85,114 viral genomes, 1,115,185 viral functional genes, and 3,823 microbial genomes obtained from 54 WWTPs across 13 countries. </p><p>If this study dataset is useful, please cite: Fan, X., Ji, M., Mu, D. <i>et al.</i> Global diversity and biogeography of DNA viral communities in activated sludge systems. <i>Microbiome</i> <strong>11</strong>, 234 (2023). https://doi.org/10.1186/s40168-023-01672-1</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.