Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.7.1
Dataset results
13 results for “metagenome-assembled genome (MAG)”
Metagenome-assembled genomes from Stordalen Mire, Sweden (MAGs v2)
<p><strong>This release (MAGs v2) is a major new version of this metagenome-assembled genome (MAG) set.</strong> All previous releases on this page (which only differ in the metadata) are designated "MAGs v1." The current release (MAGs v2) uses<strong> </strong>CheckM2 v1.0.2 filtering (≥70% completeness, ≤10% contamination) to expand this dataset to include <strong>36,419 MAGs</strong>, with the following subcategories:</p> <ul> <li>Cronin_v1: Manually-curated subset of the "Field" category from MAGs v1.</li> <li>Cronin_v2: MAGs from raw bin filtering on the same assemblies used to generate Cronin_v1.</li> <li>Woodcroft_v2: MAGs from raw bin filtering on the same assemblies used to generate the MAGs reported in <a href="https://doi.org/10.1038/s41586-018-0338-1">Woodcroft & Singleton et al. (2018)</a>.</li> <li>SIPS: Updated genomes from samples originating from a stable isotope probing (SIP) incubation experiment by Moira Hough et al. ("SIP" in MAGs v1), re-analyzed due to read truncation and sample linkage issues in MAGs v1.</li> <li>JGI: Expanded set of genomes from the Joint Genome Institute's metagenome annotation pipeline.</li> </ul> <p> </p> <p>FILES:</p> <ul> <li><strong>Emerge_MAGs_v2.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_v2_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB taxonomy information, CheckM2 quality reports, NCBI GenomeBatch- and MIMAG(6.0)-formatted sample attributes and other metadata for the MAGs. </li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data collected at the Joint Genome Institute was generated under the following awards:</p> <ul> <li>The majority of sequencing at JGI was supported by BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</li> <li>Sequencing of SIP samples was performed under the Facilities Integrating Collaborations for User Science (FICUS) initiative (proposal 503547; award DOI: <a href="https://doi.org/10.46936/fics.proj.2017.49950/60006215">10.46936/fics.proj.2017.49950/60006215</a>) and used resources at the DOE Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://ror.org/04rc0xn13">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities. Both facilities are sponsored by the Office of Biological and Environmental Research and operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</li> </ul>
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
Metagenomics assemblies and high-quality MAGs for "Long-read metagenomics to retrieve high-quality metagenome-assembled genomes from canine feces"
<p>This dataset includes the different metagenomics assemblies analyzed and its summary (_info.txt file):</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/100_assembly.fasta">100_assembly.fasta</a> is the Flye 2.7 metagenomics assembly merging HMW and non-HMW datasets</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/75_assembly.fasta">75_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 75% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/50_assembly.fasta">50_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 50% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/HMW_assembly.fasta?versionId=749ff6fd-2642-4ad1-971a-7f3404baa595">HMW_assembly.fasta</a> is the Flye 2.7 metagenomics assembly for HMW dataset.</p> <p>Moreover, it also includes the eight frameshift-corrected high-quality MAGs analyzed in the manuscript. </p>
Metagenome-Assembled Genome DRAM Annotations (EMERGE 97% dereplicated MAGs)
<p>This is the combined DRAM annotation outputs for the 1,864 97% dereplicated metagenome-assembled genomes from Stordalen Mire, Sweden. </p> <ul> <li>1864_97percentmags_annotations_combined.tsv.gz</li> <li>1864_97percentmags_metabolism_summary.xlsx</li> <li>product_0.html</li> <li>product_1.html</li> </ul> <p>METHODS:</p> <p>MAGs were annotated and distilled using DRAM (v1.4.0).</p> <p>FUNDING:<br> This research is a contribution of the EMERGE Biology Integration Institute ((https://emerge-bii.github.io/), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.<br> We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.<br> This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.<br> A portion of this research was performed under the Facilities Integrating Collaborations for User Science (FICUS) program (proposal: 10.46936/fics.proj.2017.49950/60006215 and 10.46936/10.25585/60001148) and used resources at the DOE Joint Genome Institute (<a href="https://www.google.com/url?q=https://ror.org/04xm1d337&sa=D&source=docs&ust=1674859614742521&usg=AOvVaw2XgXYw9eI4JIXRMKn3S9Se">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://www.google.com/url?q=https://ror.org/04rc0xn13&sa=D&source=docs&ust=1674859614742655&usg=AOvVaw3UXdoHIFmVjc-mXUhDXYQt">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</p>
Metagenome-assembled genomes(MAGs) generated by MetaCC binning
<p>MAGs generated by MetaCC binning from the human gut short-read, the wastewater (WW) short-read, the cow rumen long-read, and the sheep gut long-read metaHi-C datasets</p>
Metagenome-assembled genomes(MAGs) generated from soil dataset.
<p>MAGs generated from soil dataset with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes (MAGs), colorectal cancer (CRC)
<p>This archive contains (i) Metagenome assemblies of short-term enrichment cultures of CRC mucosal tissue microbiota, and (ii) Reconstructed metagenome-assembled genomes (MAGs) generated through binning of metagenome contigs.</p>
Fathi Camel Microbiome Project (FCMP) Fecal Metagenome-assembled Genomes (MAGs)
<p>The Fathi Camel Microbiome Project (FCMP) aims to characterize the diversity and phenotypic associations of the dromedary camel microbiome. The gut microbiome of N = 55 camels was deeply sequenced via dropped stool. The raw reads, after QC, were assembled and binned into metagenome-assembled genomes (MAGs). We include here a collection of 3165 medium-quality or higher prokaryotic MAGs by MiMAG-like criteria (completeness >= 50%, contamination <= 5%). </p>
Metagenome-assembled genomes(MAGs) generated from CRC human gut (PRJEB27928).
<p>MAGs generated from CRC human gut (PRJEB27928) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from dog gut (PRJEB20308).
<p>MAGs generated from dog gut (PRJEB20308) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from ocean (PRJEB1787).
<p>MAGs generated from ocean (PRJEB1787) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Prokaryotic gene catalog, prokaryotic Metagenome-Assembled Genomes (MAGs) and taxonomic profiling of metagenomic data of NEREA Augmented Observatory
<p>The NEREA_metaG directory is dedicated to the in-depth analysis of NEREA microbial communities using metagenomic sequencing data. </p> <p><strong>Gene catalog:</strong> This directory contains the gene catalog compiled from metagenomic data, which includes: Protein and nucleotide sequence files for genes; Cluster files grouping similar genes; Annotation files mapping genes to KEGG pathways; Normalized gene abundance profiles.</p> <div><strong>MAGs:</strong> Directory for Metagenome-Assembled Genomes (MAGs). It contains comprehensive annotation files for the MAGs, providing insights into gene functions, metabolic pathways, and other genomic features. It also contains the individual MAGs categorized by sample origin. Each MAG is stored in a compressed FASTA format.</div> <p><strong>mOTUs</strong>: Contains files related to microbial taxonomic units identified and quantified using the mOTUs profiler. </p>
Metagenome-assembled genomes (MAGs) from Acropora pharaonis holobiont
<p>These metagenome-assembled genomes were assembled from the Acropora pharonis metagenomes. Reads were simultaneously mapped to <em>Acropora</em>, <em>Cladocopium</em> genomes before the reads were individually assembled and binned. The resulting genomes were dereplicated and filtered for contamination and completeness. A high-quality MAG for Endozoicomonas was obtained and used along with paired metabolomic data to investigate the roles Endozoicomonas play in the coral holobiont during thermal stress. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.