Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
71
datasets available to search
ShareScore release 0.7.1
Dataset results
71 results for “Metagenome assembled genome”
Antarctic endolithic bacterial metagenome-assembled genomes
<p>Bacterial assembled genomes and annotation data from the Antarctic cryptoendolithic communities collected during the XXXI (2015-16) Italian Antarctic Expedition.</p> <p>The dataset consists of 4 zip archives and 3 files (comma-separated values). Here is a brief summary of their contents:</p> <ul> <li><strong>MAGs: </strong>high quality (HQ) and medium quality (MQ) bacterial metagenome assembled genomes.</li> <li><strong>MAGs_metadata: </strong>completeness, contamination, length, N50, GTDB classification for each MAG.</li> <li><strong>MAGs_HQ_CDS:</strong> translated coding sequences for each high quality MAG.</li> <li><strong>MAGs_HQ_Annotation: </strong>EggNOG annotation files. For each high quality MAG, the following files are included: <ul> <li>eggnog.emapper.annotations: the final EggNOG annotation;</li> <li>eggnog.emapper.hmm_hits: list of significant hits to eggNOG Orthologous Groups</li> <li>eggnog.emapper.seed_orthologs: best match of each query within the best Orthologous Group (OG) reported in the eggnog.emapper.hmm_hits file<strong>.</strong></li> </ul> </li> <li><strong>Jiangella_Antarctica: </strong><em>Candidatus Jiangella antarctica</em> representative genome (UniValnordMG_2_bin.36.fa) and the extracted ribosomal RNA genes (rRNA.fasta).</li> <li><strong>Order_MSA: </strong>protein multiple sequence alignments using the 120 GTDB bacterial marker genes. These alignments were used to estimate divergence times on orders containing at least 4 CBS, for a total of 19 orders.</li> <li><strong>Samples_accession</strong>: table that relates to the NCBI deposition of the shotgun metagenomes, the following info are included: <ul> <li>NCBI Sequence Read Archive (SRA)</li> <li>BioProject accession numbers</li> <li>JGI Integrated Microbial Genomes & Microbiomes site IDs</li> <li>N50 values</li> <li>Metadata</li> </ul> </li> <li><strong>Samples_metadata: </strong>geographic coordinates, temperature, relative humidity and sampling date are reported.</li> </ul>
High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes
<p>This project contains anvi'o profiles and contigs databases that is used and/or referenced from the Lee STM and Khan SA, <em>et al.</em> study titled "<strong>High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes</strong>". The pre-print of this study is available via http://dx.doi.org/10.1101/090993.</p> <p>To be able to work with the data files you will need anvi'o <strong>v2.1.0</strong> to be installed on your system. For installation instructions, or to have access to a Docker image for anvi'o, please visit this URL: http://merenlab.org/software/anvio</p> <p>Public data:</p> <ul> <li><strong>ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION.tar.gz</strong>: Data files for a quick visualization of the 97 MAGs and their distribution across the two FMT recipients. A run script in the archive explains how to use this data.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MERGED-PROFILE.tar.gz</strong>: The merged anvi'o profile for the entire data, which also contains a collection of 97 MAGs identified in the donor. The profile database contains no hierarchical clustering of contigs, however, individual MAGs can be displayed via the following notation since the collection 'MAGs' describe the organization of contigs in each MAG referenced from the dataset `ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION`, as well as from the paper: "anvi-refine -c CONTIGS.db -p PROFILE.db -C MAGs -b <em>FMT-Donor_MAG_00054</em>". All MAG names are in the supplementary tables in our paper.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MAGs-SUMMARY.tar.gz</strong>: A static HTML website that contains FASTA files for each MAG, and TAB-delimited matrices for coverage and detection values, and others. After unpacking, you can double-click the index.html file. </li> </ul>
Non-redundant metagenome-assembled genomes of activated sludge reactors at different disturbances and scales
<p>Metagenome-assembled genomes (MAGs) are microbial genomes reconstructed from metagenomic data and can be assigned to known taxa or lead to uncovering novel ones. MAGs can provide insights into how microbes interact with the environment. Here, we performed genome-resolved metagenomics on sequencing data from four studies using sequencing batch reactors at microcosm (~25 mL) and mesocosm (~4 L) scales inoculated with sludge from full-scale wastewater treatment plants. These studies investigated how microbial communities in such plants respond to two environmental disturbances: the presence of toxic 3-chloroaniline and changes in organic loading rate. We report 839 non-redundant MAGs with at least 50% completeness and 10% contamination (MIMAG medium-quality criteria). From these, 399 are of putative high-quality, while sixty-seven meet the MIMAG high-quality criteria. MAGs in this catalogue represent the microbial communities in sixty-eight laboratory-scale reactors used for the disturbance experiments, and in the full-scale wastewater treatment plant which provided the source sludge. This dataset can aid meta-studies aimed at understanding the responses of microbial communities to disturbances, particularly as ecosystems confront rapid environmental changes.</p>
Metagenome-Assembled Genomes of 2_2_Ac_Mat
<p>The dataset is featured in the data report titled "MAGnificent Microbes: Metagenome-Assembled Genomes of Marine Microorganisms in Mats from a Submarine Groundwater Discharge Site in Mabini, Batangas, Philippines." The study utilized shotgun metagenomics to examine the diversity and functional profiles of marine microorganisms in microbial mats from an SGD-influenced site in Mabini. The dataset includes extracted metagenome-assembled genomes (MAGs) along with their annotations using RAST.</p>
Zostera marina leaf associated bacterial metagenome assembled genomes
<p>Metagenome assembled genomes (MAGs) associated with:</p> <p>A genomic resource for exploring bacterial-viral dynamics in seagrass ecosystems</p> <p>Analysis, code, intermediate and supporting files are archived here: <a href="https://doi.org/10.5281/zenodo.14226514">10.5281/zenodo.14226514</a></p> <p>Viral sequences from this work are archived here: <a href="https://doi.org/10.5281/zenodo.14226038">10.5281/zenodo.14226038</a><br><br>This archive contains:<br>(i) Fifty-six fasta files representing the MAGs described in the above titled work with > 80% completion and < 10% contamination based on CheckM2 metrics<br>(ii) Metadata file describing the MAGs (i.e., subset of Table S3 from the above work)</p>
Metagenome-assembled genomes(MAGs) generated from soil dataset.
<p>MAGs generated from soil dataset with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Twenty-five metagenome assembled genomes recovered from the gut microbiome of the domestic ferret, Mustela putorius
<p>This dataset is composed of 25 unique metagenome assembled genomes (MAGs) recovered from the gut microbiome of three domestic ferrets (<em>Mustela putorius</em>). Details on both MAG and host ferret metadata, as well as information on sample collection, DNA sequencing, and bioinformatic processing can be found in the American Society for Microbiology Resource Announcement by Amundson et al. (in prep). </p>
HiFi Metagenomic Sequencing Enables Assembly of Accurate and Complete Genomes from Human Gut Microbiota.
<p>We reported 102 complete metagenome assembled genomes (cMAGs) from five human fecal HiFi sequencing samples.</p> <p>102_cMAGs_fna.tar.gz: Fasta sequence files of 102 cMAGs.</p> <p>gc_skew_figures.tar.gz: GC-skew pattern figures of 102 cMAGs. (SVG format)</p> <p>coverage_plots.tar.gz: Genome coverage plot of 102 cMAGs.</p>
Dereplicated Metagenome assembled genomes (MAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 55 metagenome assembled genomes (MAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments
<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2) A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p> </p> <p>Manuscript title Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>
Metagenome-assembled genomes obtained from fecal and salivary microbiomes of pancreatic cancer patients and controls
<p>7,546 MAGs obtained from fecal and salivary metagenomes of pancreatic cancer patients and controls</p>
Early-life human gut metagenome-assembled genomes and proteins catalogs
<p>The description of the files:</p> <p>(1) The 32,277 genomes include six parts: ELGG_part_1.zip, ELGG_part_2.zip, ELGG_part_3.zip, ELGG_part_4.zip, ELGG_part_5.zip, ELGG_part_6.zip.</p> <p>(2) The 2,172 representative species: ELGG_representatives_2172.zip.</p> <p>(3) The ELGP catalog clustered at 95% amino acid identity: ELGP_95.faa.gz.</p> <p> </p>
Metagenome-assembled genomes (MAGs), colorectal cancer (CRC)
<p>This archive contains (i) Metagenome assemblies of short-term enrichment cultures of CRC mucosal tissue microbiota, and (ii) Reconstructed metagenome-assembled genomes (MAGs) generated through binning of metagenome contigs.</p>
Fathi Camel Microbiome Project (FCMP) Fecal Metagenome-assembled Genomes (MAGs)
<p>The Fathi Camel Microbiome Project (FCMP) aims to characterize the diversity and phenotypic associations of the dromedary camel microbiome. The gut microbiome of N = 55 camels was deeply sequenced via dropped stool. The raw reads, after QC, were assembled and binned into metagenome-assembled genomes (MAGs). We include here a collection of 3165 medium-quality or higher prokaryotic MAGs by MiMAG-like criteria (completeness >= 50%, contamination <= 5%). </p>
Metagenome Assembled Genomes (MAGs) from faecal microbiomes of great tits and blue tits
<h2><span>Overview:</span></h2> <p><span>The vertebrate gut microbiome plays crucial roles in host health and disease. However, there is limited data on the microbiomes of wild birds, most of which is restricted to barcode sequences. We therefore explored the use of shotgun metagenomics on the faecal microbiomes of two wild bird species widely used as model organisms in ecological studies: the great tit (<em>Parus major</em>) and the Eurasian blue tit (<em>Cyanistes caeruleus</em>). High and Medium quality Metagenome Assembled Genomes (MAGs) were assembled from these metagenomes and are made available as a catalogue in this archive.</span></p> <h2><span>Methods:</span></h2> <p><span><span>Metagenomic reads were trimmed, and quality controlled using FastP configured to a minimum phred score of 20 and minimum length of 50 bp</span><span>. In order to avoid contamination of the bins by eukaryotic sequences, Tiara v1.0.3 was used to classify contigs longer than 3.000 kb into their high-level kingdoms, allowing to exclude sequences of a eukaryotic or of an organelle origin, and only retaining all unclassified contigs and prokaryotic contigs for the binning step. <br>Contigs were binned using MaxBin2 v2.2.7 , SemiBin2 v2.1.0 and Metabat2 v 2.15 independently. The bins were refined using DasTool v 1.1.7 using a min score threshold of 0.3. The quality of the refined bins was obtained using CheckM2 v 1.0.2, and any bin with a contamination above 10% were excluded. The final MAGs were classified as Low-quality (<50% completeness, <10% contamination), medium-quality (>50% completeness, <10% contamination) and high-quality (>90% completeness, <5% contamination), as recommended by the MIMAG specification . Finally, the MAGs were dereplicated using an dRep v 3.4.3 with an ANI of 95% and classified using gtdb-tk v2.4.0 using the gtdb database release220.<br></span></span></p> <p><span><span>Files:</span></span></p> <ul> <li><span><span>The <strong>MAGs_sequences_v1.0.0 </strong>contains the fasta sequence for the individual MAGs assembled in this project</span></span></li> <li><span><span>The <strong>MAGs_catalogue_v1.0.0.xlsx</strong> contains a description of the quality, taxonomic annotation and characteristics of each MAGs in the dataset</span></span></li> </ul>
Wheat phyllosphere metagenome assembled genomes collected in Ringsted, Denmark
<p><span>We present a completely novel </span><span>and comprehensive wheat phyllosphere metagenomic dataset of 211 samples and </span><span>1261 MAGs. This dataset represents a significant contribution to the field, as it </span><span>provides insights into the poorly studied microbial communities associated with </span><span>wheat leaf surfaces, an ecosystem of considerable agricultural importance.</span></p>
Metagenome-Assembled Genomes and Annotations for McGivern et al
<h3>Files:</h3> <ul> <li><code>reactorEMERGE_annotations.txt</code>: DRAM annotations for MAGs</li> <li><code>gene_lengths.txt</code>: gene length file used to calculate geTMM</li> <li><code>genes.gff.tar.gz</code>: gff file needed for metaT processing</li> <li><code>genes.faa.tar.gz</code>: amino acid sequences for MAG genes</li> <li><code>genes.fna.tar.gz</code>: nucleotide sequences for MAG genes, used as database for metaT mapping</li> </ul>
Viral metagenome assembled genomes (vMAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 111 viral metagenome assembled genomes (vMAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Metagenome assembled genome (MAG) annotations for Columbia River sediment bacteria and archaea
<p>Excel spreadsheet containing all annotations for metagenome assembled genomes (MAGs) that form part of a publication to be submitted titled: "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Novel canine high-quality metagenome-assembled genomes by long-read metagenomics together with Hi-C proximity ligation
<p>We characterized a canine fecal sample of a healthy dog by combining a long-read metagenomics assembly (Nanopore sequencing) with Hi-C cross-linking data, and further correction of the frameshift errors. We retrieved and characterized 27 HQ MAGs and seven MQ MAGs considering MIMAG criteria.</p> <p>Find in this repository the final Hi-C genomics bins (CanMAG_XX-HiCbin.fa), including both the genome and the extra-chromosomal elements within the bin. </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.