Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

103

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

103 results for “metagenomes assembly”

Learn how ShareScore rates datasets ↗
zenodo52/100

Orbicella faveolata and O. franksi coral metagenome assemblies from the Lower Florida Keys region of Florida, USA

<div> <p>The enclosed files include mostly <em>Orbicella faveolata</em> and three <em>Orbicella franksi</em> coral metagenome assemblies collected from the Lower Keys in Florida&rsquo;s Coral Reef, USA. Metadata for the files is included in this repository. Apparently healthy coral tissue cores were collected between May 28 and June 21, 2021. The DNA was extracted from the host and associated microorganisms and sequenced in a paired-end 150 bp format on an Illumina NovaSeq. Trimming and quality filtering of DNA sequence reads proceeded, followed by host and photoendosymbiotic dinoflagellate DNA removal. The host-cleaned reads were assembled individually by coral sample into longer contigs using MegaHit v1.1.4. The &ldquo;Assembly_Fastas&rdquo; zipped file contains 41 metagenome assemblies from the individual <em>Orbicella faveolata</em> corals and 3 assemblies from the individual <em>Orbicella franksi&nbsp;</em>colonies for a total of 44 assemblies. In addition, these assemblies were annotated with eggnog-mapper v2.1.6 to generate both predicted gene regions and annotation output files. The &ldquo;Predicted_Gene_Fastas&rdquo; zipped file contains nucleotide fasta files of the predicted gene regions for all 44&nbsp;coral metagenome assemblies. The fasta header of each gene includes the contig ID it originated from in the associated &ldquo;Assembly_Fasta&rdquo;. The &ldquo;Predicted_Gene_Annotations&rdquo; zipped file contains either .csv or .xlsx files with the eggnog-mapper-based annotations. These files contain a &ldquo;query contig&rdquo; that corresponds to the contig ID in the fasta header of the &ldquo;Predicted_Gene_Fasta&rdquo;.&nbsp;</p> <p>In addition to individual assemblies, a co-assembly was generated that included all 41 <em>Orbicella faveolata</em> coral samples. Prior to co-assembly, further removal of eukaryotic DNA proceeded by splitting the indiviudual assemblies into eukaryotic and prokaryotic content with the program EukRep v0.6.7, followed by mapping of the host-clean reads to the eukaryotic DNA to remove them. The eukaryote-clean reads from all 41 corals were input into MegaHit to generate a co-assembly. The co-assembly is included (FLK_OFAV_MG_coassembly_final.contigs.fa). Predicted genes from the co-assembly were generated with Prodigal v2.6.3 and the nucleotide fasta of the output is included in this repository (FLK_OFAV_MG_pred.fna). Like with the indiviudal assemblies, eggnog-mapper was used to generate annotations of the predicted genes from Prodigal (FLK_OFAV_MG.emapper.annotations.xlsx).&nbsp; Additionally, the abundance of each predicted gene was generated using Salmon to map the eukaryote-clean reads to the predicted genes. The number of reads (counts) for each gene across each coral sample were aggregated as integers into one table and included in this repository (FLK_OFAV_MG_pred_NumReads.tsv).&nbsp;&nbsp;</p> </div> <div> <p>These data were processed and generated by Julie Meyer&rsquo;s Lab at the University of Florida, using funding from the Florida Department of Environmental Protection.&nbsp;&nbsp;</p> </div>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly

<p>Data and conda software environment file for the chapter &#39;<em>de novo</em> Genome Assembly&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (MAGs v2)

<p><strong>This release (MAGs v2) is a major new version of this metagenome-assembled genome (MAG) set.</strong> All previous releases on this page (which only differ in the metadata) are designated "MAGs v1." The current release (MAGs v2) uses<strong>&nbsp;</strong>CheckM2 v1.0.2 filtering (&ge;70% completeness, &le;10% contamination) to expand this dataset to include <strong>36,419 MAGs</strong>, with the following subcategories:</p> <ul> <li>Cronin_v1:&nbsp; Manually-curated subset of the "Field" category from MAGs v1.</li> <li>Cronin_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate Cronin_v1.</li> <li>Woodcroft_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate the MAGs reported in <a href="https://doi.org/10.1038/s41586-018-0338-1">Woodcroft &amp; Singleton et al. (2018)</a>.</li> <li>SIPS:&nbsp; Updated genomes from samples originating from a stable isotope probing (SIP) incubation experiment by Moira Hough et al. ("SIP" in MAGs v1), re-analyzed due to read truncation and sample linkage issues in MAGs v1.</li> <li>JGI:&nbsp; Expanded set of genomes from the Joint Genome Institute's metagenome annotation pipeline.</li> </ul> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>Emerge_MAGs_v2.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_v2_EMERGE.tsv</strong>&nbsp;- Table containing source sample names and accessions, GTDB taxonomy information, CheckM2 quality reports, NCBI GenomeBatch- and MIMAG(6.0)-formatted sample attributes and other metadata for the MAGs.&nbsp;</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data collected at the Joint Genome Institute was generated under the following awards:</p> <ul> <li>The majority of sequencing at JGI was supported by BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</li> <li>Sequencing of SIP samples was performed under the Facilities Integrating Collaborations for User Science (FICUS) initiative (proposal 503547; award DOI:&nbsp;<a href="https://doi.org/10.46936/fics.proj.2017.49950/60006215">10.46936/fics.proj.2017.49950/60006215</a>) and used resources at the DOE Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://ror.org/04rc0xn13">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities. Both facilities are sponsored by the Office of Biological and Environmental Research and operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Metagenome-Assembled Genomes Abundance & Activity Tables. Environmental Parameters Associated with the dataset.

<p>Lake Mendota, WI, USA, is a temperate lake subject to annual temperature and oxygen fluctuations. Each summer, the water column becomes anoxic (no-oxygen). In 2020, we sampled the lake at weekly intervals, at different depths (5, 10, 15, 20 and 23.5m). For each time+depth sample, we collected metagenomes, viromes and metatranscriptomes. Environmental data profiles were collected on-site for each sampling day.&nbsp;</p> <p>Following standard metagenomic binning best practices, we obtained 431 metagenomes-assembled-genomes (MAGs).</p> <p>This record comprises the microbial abundance and expression table for these MAGs, and the environmental profiles collected each day.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Gut Metagenome Assemblies for Veseli et al. 2023

<p>A collection of anvi&#39;o contigs databases for 408 human fecal metagenome assemblies from the study by Veseli et al. titled &quot;High metabolic independence is a determinant of microbial resilience in the face of gut stress&quot;. These are publicly-available gut metagenomes originally obtained from several studies of the gut microbiome. See `METAGENOMES_INFO.txt` file for references and sample SRA accessions.</p> <p>The metagenomes were assembled individually using IDBA-UD as part of the anvi&#39;o metagenomics&nbsp;workflow in anvi&#39;o v7.1-dev. As part of this workflow, they were annotated with KEGG KOfams using `anvi-run-kegg-kofams` and a KEGG snapshot from December 12, 2020&nbsp;(modules database hash value `45b7cc2e4fdc`). See manuscript and its reproducible workflow for details.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Metagenomics of a pustular microbial mat from Shark Bay, Australia: Raw sequences and assembled MAGs

<p>This data accompanies the paper, &quot;<a href="https://www.nature.com/articles/s43705-022-00128-1">Metagenomic,&nbsp;(bio)chemical, and microscopic analyses reveal the potential for the cycling of sulfated EPS in Shark Bay pustular mats</a>&quot;&nbsp;which looks at the cycling of sulfated polysaccharides in peritidal pustular mats from Shark Bay, Australia. The microbial community was sequenced, assembled, and binned. The&nbsp;raw sequencing reads used in this analysis are the following:</p> <ul> <li>SB_forward_paired_copy.fastq.gz&nbsp;</li> <li>SB_reverse_paired_copy.fastq.gz&nbsp;</li> </ul> <p>The resulting metagenome-assembled genomes (MAGs) are presented in the following folder:</p> <ul> <li>MAGs.zip</li> </ul> <p>&nbsp;</p> <ul> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)

<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 &amp; 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by&nbsp;<a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads.&nbsp;</p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a>&nbsp;(v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with&nbsp;<a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a>&nbsp;(v2.9-b1768) using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1)&nbsp;using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18)&nbsp;and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0)&nbsp;and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2)&nbsp;and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Catalog of stool metagenome-assembled genomes from patients with different cancer types

<p><strong>A non-redundant catalog of 3,816 genomes with at least 75% completeness and no more than 15% contamination assembled from metagenomes. Samples of 976 metagenomes were obtained from patients receiving immunotherapy for the treatment of different types of cancers.</strong></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes

<p>&nbsp;</p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; RELEASE MAG-2018/01<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --------------------------------------</p> <p>&nbsp;</p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p>&nbsp;</p> <p>2. LOCATION</p> <p>&nbsp;</p> <p>&nbsp;&nbsp; MAGs sequences are available under ENA project PRJEB25176.</p> <p>&nbsp;&nbsp;</p> <p>&nbsp;&nbsp;&nbsp; ENA_accession&nbsp;&nbsp; Isolate<br> &nbsp;&nbsp; &nbsp;--------------------&nbsp;&nbsp; &nbsp;--------------<br> &nbsp;&nbsp; &nbsp;ERZ494218&nbsp;&nbsp; &nbsp;AM_0118<br> &nbsp;&nbsp; &nbsp;ERZ494219&nbsp;&nbsp; &nbsp;AM_0219<br> &nbsp;&nbsp; &nbsp;ERZ494220&nbsp;&nbsp; &nbsp;AM_0226<br> &nbsp;&nbsp; &nbsp;ERZ494221&nbsp;&nbsp; &nbsp;AM_0228<br> &nbsp;&nbsp; &nbsp;ERZ494222&nbsp;&nbsp; &nbsp;AM_0233<br> &nbsp;&nbsp; &nbsp;ERZ494223&nbsp;&nbsp; &nbsp;AM_0240<br> &nbsp;&nbsp; &nbsp;ERZ494224&nbsp;&nbsp; &nbsp;AM_0244<br> &nbsp;&nbsp; &nbsp;ERZ494225&nbsp;&nbsp; &nbsp;AM_0256<br> &nbsp;&nbsp; &nbsp;ERZ494226&nbsp;&nbsp; &nbsp;AM_0268<br> &nbsp;&nbsp; &nbsp;ERZ494227&nbsp;&nbsp; &nbsp;AM_0275<br> &nbsp;&nbsp; &nbsp;ERZ494228&nbsp;&nbsp; &nbsp;AM_0466<br> &nbsp;&nbsp; &nbsp;ERZ494229&nbsp;&nbsp; &nbsp;AM_0507<br> &nbsp;&nbsp; &nbsp;ERZ494230&nbsp;&nbsp; &nbsp;AM_0510<br> &nbsp;&nbsp; &nbsp;ERZ494231&nbsp;&nbsp; &nbsp;AM_0519<br> &nbsp;&nbsp; &nbsp;ERZ494232&nbsp;&nbsp; &nbsp;AM_0528<br> &nbsp;&nbsp; &nbsp;ERZ494233&nbsp;&nbsp; &nbsp;AM_0546<br> &nbsp;&nbsp; &nbsp;ERZ494234&nbsp;&nbsp; &nbsp;AM_0608<br> &nbsp;&nbsp; &nbsp;ERZ494235&nbsp;&nbsp; &nbsp;AM_0615<br> &nbsp;&nbsp; &nbsp;ERZ494236&nbsp;&nbsp; &nbsp;AM_0616<br> &nbsp;&nbsp; &nbsp;ERZ494237&nbsp;&nbsp; &nbsp;AM_0619<br> &nbsp;&nbsp; &nbsp;ERZ494238&nbsp;&nbsp; &nbsp;AM_0621<br> &nbsp;&nbsp; &nbsp;ERZ494239&nbsp;&nbsp; &nbsp;AM_0630<br> &nbsp;&nbsp; &nbsp;ERZ494240&nbsp;&nbsp; &nbsp;AM_0643<br> &nbsp;&nbsp; &nbsp;ERZ494241&nbsp;&nbsp; &nbsp;AM_0729<br> &nbsp;&nbsp; &nbsp;ERZ494242&nbsp;&nbsp; &nbsp;AM_0764<br> &nbsp;&nbsp; &nbsp;ERZ494243&nbsp;&nbsp; &nbsp;AM_0832<br> &nbsp;&nbsp; &nbsp;ERZ494244&nbsp;&nbsp; &nbsp;AM_0849<br> &nbsp;&nbsp; &nbsp;ERZ494245&nbsp;&nbsp; &nbsp;AM_0854<br> &nbsp;&nbsp; &nbsp;ERZ494246&nbsp;&nbsp; &nbsp;AM_0876<br> &nbsp;&nbsp; &nbsp;ERZ494247&nbsp;&nbsp; &nbsp;AM_0902<br> &nbsp;&nbsp; &nbsp;ERZ494248&nbsp;&nbsp; &nbsp;AM_0936<br> &nbsp;&nbsp; &nbsp;ERZ494249&nbsp;&nbsp; &nbsp;AM_1003<br> &nbsp;&nbsp; &nbsp;ERZ494250&nbsp;&nbsp; &nbsp;AM_1104<br> &nbsp;&nbsp; &nbsp;ERZ494251&nbsp;&nbsp; &nbsp;AM_1111<br> &nbsp;&nbsp; &nbsp;ERZ494252&nbsp;&nbsp; &nbsp;AM_1205<br> &nbsp;&nbsp; &nbsp;ERZ494253&nbsp;&nbsp; &nbsp;AM_1312<br> &nbsp;&nbsp; &nbsp;ERZ494254&nbsp;&nbsp; &nbsp;AM_1409<br> &nbsp;&nbsp; &nbsp;ERZ494255&nbsp;&nbsp; &nbsp;AM_1503<br> &nbsp;&nbsp; &nbsp;ERZ494256&nbsp;&nbsp; &nbsp;AM_1603<br> &nbsp;&nbsp; &nbsp;ERZ494257&nbsp;&nbsp; &nbsp;AM_1606<br> &nbsp;&nbsp; &nbsp;ERZ494258&nbsp;&nbsp; &nbsp;AM_1801<br> &nbsp;&nbsp; &nbsp;ERZ494259&nbsp;&nbsp; &nbsp;AM_1811<br> &nbsp;&nbsp; &nbsp;ERZ494260&nbsp;&nbsp; &nbsp;AM_2104<br> &nbsp;&nbsp; &nbsp;ERZ494261&nbsp;&nbsp; &nbsp;AM_2116<br> &nbsp;&nbsp; &nbsp;ERZ494262&nbsp;&nbsp; &nbsp;AM_2124<br> &nbsp;&nbsp; &nbsp;ERZ494263&nbsp;&nbsp; &nbsp;AM_2202<br> &nbsp;&nbsp; &nbsp;ERZ494264&nbsp;&nbsp; &nbsp;AM_2207<br> &nbsp;&nbsp; &nbsp;ERZ494265&nbsp;&nbsp; &nbsp;AM_2208<br> &nbsp;&nbsp; &nbsp;ERZ494266&nbsp;&nbsp; &nbsp;AM_2324<br> &nbsp;&nbsp; &nbsp;ERZ494267&nbsp;&nbsp; &nbsp;AM_2502<br> &nbsp;&nbsp; &nbsp;ERZ494268&nbsp;&nbsp; &nbsp;AM_2804<br> &nbsp; &nbsp;</p> <p>3. ACKNOWLEDGEMENTS<br> &nbsp; &nbsp;</p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of S&atilde;o Carlos, S&atilde;o Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Cient&iacute;fico e Tecnol&oacute;gico (CNPq), as well as, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; the spanish funding organ Consejo Superior de Investigaciones Cient&iacute;ficas (CSIC).</p> <p>This study was financed in part by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento de Pessoal de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p> <p>&nbsp;</p> <p>4. CONTACT INFORMATION</p> <p>&nbsp;&nbsp; Current curators:</p> <p>&nbsp;&nbsp; - C&eacute;lio Dias Santos J&uacute;nior (celio.diasjunior@gmail.com)<br> &nbsp;&nbsp; - Flavio Henrique-Silva (dfhs@ufscar.br)<br> &nbsp;&nbsp; - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> &nbsp;</p> <p>5. COPYRIGHT NOTICE</p> <p>&nbsp;&nbsp; Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> &nbsp;&nbsp; Copyright (C) 2018 The AMnrGC consortium.</p> <p>&nbsp;&nbsp; This database is provided &ldquo;as is&rdquo; and without any warranty of any kind,<br> &nbsp;&nbsp; of openly available. You can redistribute and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p>&nbsp;&nbsp; &nbsp;https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

The Metagenome-Assembled Genome Inventory for Children (MAGIC)

<div> <div>Existing microbiota databases are biased towards adult samples, hampering accurate profiling of the infant gut microbiome. Here, we generated a **M**etagenome-**A**ssembled **G**enome **I**nventory for **C**hildren (**MAGIC**) from a large collection of bulk and viral-like particle-enriched metagenomes from 0-7 years of age, encompassing `3,299` prokaryotic and `139,624` viral species-level genomes, `8.5%` and `63.9%` of which are unique to MAGIC. MAGIC improves early-life microbiome profiling, with the greatest improvement in read mapping observed in Africans. We then identified `54` candidate keystone species, including several *Bifidobacterium spp.* and four phages, forming guilds that fluctuated in abundance with time. Their abundances were reduced in preterm infants and were associated with childhood allergies. By analyzing the *B. longum* pangenome, we found evidence of phage-mediated evolution and quorum sensing-related ecological adaptation. Together, the MAGIC database recovers genomes that enable characterization of dynamics of early-life microbiomes, identification of candidate keystone species, and strain-level study of target species.</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Assemblies of 269 Metagenomic Tara Pacific Sequencing Samples - part 2

<p>This data is the result of the metagenomic assembly of 269 sequencing samples reflecting a first subset of the Tara Pacific metagenomes. Assemblies are used in&nbsp;</p> <p>- Preprint:&nbsp;<a href="https://doi.org/10.1101/2022.04.11.487905">Endogenous viral elements reveal associations between a non-retroviral RNA virus and symbiotic dinoflagellate genomes</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.7839794">Part 1</a></p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Assemblies of 269 Metagenomic Tara Pacific Sequencing Samples - part 1

<p>This data is the result of the metagenomic assembly of 269 sequencing samples reflecting a first subset of the Tara Pacific metagenomes. Assemblies are used in&nbsp;</p> <p>- Preprint:&nbsp;<a href="https://doi.org/10.1101/2022.04.11.487905">Endogenous viral elements reveal associations between a non-retroviral RNA virus and symbiotic dinoflagellate genomes</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.7840044">Part 2</a></p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Palleja et al. 2018 Metagenome Assemblies for Veseli et al. 2023

<p>A collection of anvi&#39;o contigs databases for 57 human fecal metagenome assemblies generated for the study by Veseli et al. titled &quot;High metabolic independence is a determinant of microbial resilience in the face of gut stress&quot;. These are publicly-available gut metagenomes originally obtained from the study by Palleja et al titled &quot;Recovery of gut microbiota of healthy adults following antibiotic exposure&quot; (https://doi.org/10.1038/s41564-018-0257-9). See `PALLEJA_ET_AL_SAMPLES_INFO.txt` file for&nbsp;sample SRA accessions.</p> <p>The metagenomes were assembled individually using IDBA-UD as part of the anvi&#39;o metagenomics&nbsp;workflow in anvi&#39;o v7.1-dev. As part of this workflow, they were annotated with KEGG KOfams using `anvi-run-kegg-kofams` and a KEGG snapshot from December 12, 2020&nbsp;(modules database hash value `45b7cc2e4fdc`). See manuscript and its reproducible workflow for details.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

New Soil Metagenome-Assembled Genomes Catalogue Boosts Genetic Resources

<p><strong>Soil harbors a vast expanse of unidentified microbes, termed as microbial dark matter, presenting an untapped reservoir of microbial biodiversity and genetic resources, but has yet to be fully explored. In this study, we conducted the first large-scale excavation of soil microbial dark matter by reconstructing 40,039 metagenome-assembled genome bins (the SMAG catalog) from 3,304 soil metagenomes. We identified 16,530 of 21,077 species-level genome bins (SGBs) as unknown SGBs (uSGBs), which greatly expand archaeal and bacterial diversity across the tree of life. We also illustrate the pivotal role of uSGBs in augmenting soil microbiome&#39;s functional landscape and intra-species genome diversity, providing large proportions of the 43,169 biosynthetic gene clusters and 8,545 CRISPR-Cas genes. Additionally, we determined that uSGBs contributed 84.6% of novel viral-host associations identified from the SMAG catalog. Our results propose the SMAG catalog, a novel and expansive genomic resource that brings the soil microbial biodiversity and novel genetic resources to light.</strong></p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Metagenomic assembly and bin3C clustering result for a healthy human faecal microbiome transplant donor

<p>Metagenomic WGS assembly and Hi-C deconvolution&nbsp;of a healthy human faecal microbiome transplant donor.</p> <p>Metagenomic assembly was produced using Spades (v3.13.1).</p> <p>Extracted MAGs were produced using bin3C&nbsp;(v0.3.3) and QC&#39;d using CheckM&nbsp;(v1.0.18).</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

Metagenomics assemblies and high-quality MAGs for "Long-read metagenomics to retrieve high-quality metagenome-assembled genomes from canine feces"

<p>This dataset includes the different metagenomics assemblies analyzed and its summary (_info.txt file):</p> <p>-&nbsp;<a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/100_assembly.fasta">100_assembly.fasta</a>&nbsp;is the Flye 2.7 metagenomics assembly merging HMW and non-HMW datasets</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/75_assembly.fasta">75_assembly.fasta</a>&nbsp;is the Flye 2.7 metagenomics assembly including 75% of random data of the merged dataset.</p> <p>-&nbsp;<a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/50_assembly.fasta">50_assembly.fasta</a>&nbsp;is the Flye 2.7 metagenomics assembly including 50% of random data of the merged dataset.</p> <p>-&nbsp;<a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/HMW_assembly.fasta?versionId=749ff6fd-2642-4ad1-971a-7f3404baa595">HMW_assembly.fasta</a>&nbsp;is the Flye 2.7 metagenomics assembly for HMW dataset.</p> <p>Moreover, it also includes the eight frameshift-corrected high-quality MAGs analyzed in the manuscript.&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Metagenome assemblies and metagenome-assembled genomes from the Daphnia magna microbiota

<p>Metagenome assemblies generated from raw reads not mapping to the Daphnia magna genome for six samples assembled individually (G4, G14, S1-S4) and a coassembly of all six samples (a_assembly)&nbsp;using metaSPAdes in SPAdes v3.14. Assemblies can be found in metagenome_assemblies.zip.</p> <p>Metagenome-assembled genomes generated using VAMB v3.0.2 (vamb_bins.zip) and ProxiMeta (proximeta_bins.zip). These MAGs were taxonomically identified using GTDB-Tk v1.3 and quality checked using CheckM v1.1. Outputs from GTDB-Tk and CheckM can be found in the .tsv and .tab files, respectively.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Companion data deposit of manuscript: Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies

<p>This upload contains the metagenome assemblies and their binning results generated &amp; described in the manuscript "Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies" (preprint version: arxiv2210.00098, "Towards complete representation of bacterial contents in metagenomic samples").&nbsp;</p> <p>Mapping of sample names in the file names and the descriptors used as in the manuscript can be found in table S1, which is available along with the manuscript and also included in the supplementary_tables_and_figures tar archive here.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Metagenome-assembled Genomes of Scandinavium goeteborgense MCPNR19-05 and Erwinia aphidicola MCPNR19-06

<p>Here we provide two fasta files (MCPNR19-05.fa and MCPNR19-06.fa) which represent low-quality metagenome assembled genomes (MAGs) obtained from genomic DNA from Massospora cicadina isolate MCPNR19 azygospores collected from multiple seventeen-year cicada (Magicicada septendecim) June 2019 at Powdermill Nature Reserve, Rector, Pennsylvania.</p> <p>MCPNR19-05.fa = Scandinavium goeteborgense MCPNR19-05, a 1.82 Mb 45.61% complete MAG<br> MCPNR19-06.fa = Erwinia aphidicola MCPNR19-06, a 1.79 Mb 29.31% complete MAG<br> <br> <strong>Raw data availability</strong><br> Sequence reads are deposited under SRA project accessions <a href="https://ncbi.nlm.nih.gov/sra/SRR17553520">SRR17553520</a>-<a href="https://ncbi.nlm.nih.gov/sra/SRR17553526">SRR17553526</a> and BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA795459">PRJNA795459</a>. These MAGs are metagenomic assemblies obtained from the host Massospora cicadina (BioSample: SAMN24722893). &nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 4c Introduction to Genome Assembly

<p>Teaching data for&nbsp;practical session: &quot;4c&nbsp;Introduction to Genome Assembly&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record