Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.9.0
Dataset results
103 results for “metagenomes assembly”
Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments
<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2) A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p> </p> <p>Manuscript title Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>
Metagenome-assembled genomes obtained from fecal and salivary microbiomes of pancreatic cancer patients and controls
<p>7,546 MAGs obtained from fecal and salivary metagenomes of pancreatic cancer patients and controls</p>
Early-life human gut metagenome-assembled genomes and proteins catalogs
<p>The description of the files:</p> <p>(1) The 32,277 genomes include six parts: ELGG_part_1.zip, ELGG_part_2.zip, ELGG_part_3.zip, ELGG_part_4.zip, ELGG_part_5.zip, ELGG_part_6.zip.</p> <p>(2) The 2,172 representative species: ELGG_representatives_2172.zip.</p> <p>(3) The ELGP catalog clustered at 95% amino acid identity: ELGP_95.faa.gz.</p> <p> </p>
Metagenome-assembled genomes (MAGs), colorectal cancer (CRC)
<p>This archive contains (i) Metagenome assemblies of short-term enrichment cultures of CRC mucosal tissue microbiota, and (ii) Reconstructed metagenome-assembled genomes (MAGs) generated through binning of metagenome contigs.</p>
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Real and mock datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>real and mock</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Simulated datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>simulated</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Fathi Camel Microbiome Project (FCMP) Fecal Metagenome-assembled Genomes (MAGs)
<p>The Fathi Camel Microbiome Project (FCMP) aims to characterize the diversity and phenotypic associations of the dromedary camel microbiome. The gut microbiome of N = 55 camels was deeply sequenced via dropped stool. The raw reads, after QC, were assembled and binned into metagenome-assembled genomes (MAGs). We include here a collection of 3165 medium-quality or higher prokaryotic MAGs by MiMAG-like criteria (completeness >= 50%, contamination <= 5%). </p>
Metagenome Assembled Genomes (MAGs) from faecal microbiomes of great tits and blue tits
<h2><span>Overview:</span></h2> <p><span>The vertebrate gut microbiome plays crucial roles in host health and disease. However, there is limited data on the microbiomes of wild birds, most of which is restricted to barcode sequences. We therefore explored the use of shotgun metagenomics on the faecal microbiomes of two wild bird species widely used as model organisms in ecological studies: the great tit (<em>Parus major</em>) and the Eurasian blue tit (<em>Cyanistes caeruleus</em>). High and Medium quality Metagenome Assembled Genomes (MAGs) were assembled from these metagenomes and are made available as a catalogue in this archive.</span></p> <h2><span>Methods:</span></h2> <p><span><span>Metagenomic reads were trimmed, and quality controlled using FastP configured to a minimum phred score of 20 and minimum length of 50 bp</span><span>. In order to avoid contamination of the bins by eukaryotic sequences, Tiara v1.0.3 was used to classify contigs longer than 3.000 kb into their high-level kingdoms, allowing to exclude sequences of a eukaryotic or of an organelle origin, and only retaining all unclassified contigs and prokaryotic contigs for the binning step. <br>Contigs were binned using MaxBin2 v2.2.7 , SemiBin2 v2.1.0 and Metabat2 v 2.15 independently. The bins were refined using DasTool v 1.1.7 using a min score threshold of 0.3. The quality of the refined bins was obtained using CheckM2 v 1.0.2, and any bin with a contamination above 10% were excluded. The final MAGs were classified as Low-quality (<50% completeness, <10% contamination), medium-quality (>50% completeness, <10% contamination) and high-quality (>90% completeness, <5% contamination), as recommended by the MIMAG specification . Finally, the MAGs were dereplicated using an dRep v 3.4.3 with an ANI of 95% and classified using gtdb-tk v2.4.0 using the gtdb database release220.<br></span></span></p> <p><span><span>Files:</span></span></p> <ul> <li><span><span>The <strong>MAGs_sequences_v1.0.0 </strong>contains the fasta sequence for the individual MAGs assembled in this project</span></span></li> <li><span><span>The <strong>MAGs_catalogue_v1.0.0.xlsx</strong> contains a description of the quality, taxonomic annotation and characteristics of each MAGs in the dataset</span></span></li> </ul>
Wheat phyllosphere metagenome assembled genomes collected in Ringsted, Denmark
<p><span>We present a completely novel </span><span>and comprehensive wheat phyllosphere metagenomic dataset of 211 samples and </span><span>1261 MAGs. This dataset represents a significant contribution to the field, as it </span><span>provides insights into the poorly studied microbial communities associated with </span><span>wheat leaf surfaces, an ecosystem of considerable agricultural importance.</span></p>
Metagenome-Assembled Genomes and Annotations for McGivern et al
<h3>Files:</h3> <ul> <li><code>reactorEMERGE_annotations.txt</code>: DRAM annotations for MAGs</li> <li><code>gene_lengths.txt</code>: gene length file used to calculate geTMM</li> <li><code>genes.gff.tar.gz</code>: gff file needed for metaT processing</li> <li><code>genes.faa.tar.gz</code>: amino acid sequences for MAG genes</li> <li><code>genes.fna.tar.gz</code>: nucleotide sequences for MAG genes, used as database for metaT mapping</li> </ul>
Viral metagenome assembled genomes (vMAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 111 viral metagenome assembled genomes (vMAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Metagenome assembled genome (MAG) annotations for Columbia River sediment bacteria and archaea
<p>Excel spreadsheet containing all annotations for metagenome assembled genomes (MAGs) that form part of a publication to be submitted titled: "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Novel canine high-quality metagenome-assembled genomes by long-read metagenomics together with Hi-C proximity ligation
<p>We characterized a canine fecal sample of a healthy dog by combining a long-read metagenomics assembly (Nanopore sequencing) with Hi-C cross-linking data, and further correction of the frameshift errors. We retrieved and characterized 27 HQ MAGs and seven MQ MAGs considering MIMAG criteria.</p> <p>Find in this repository the final Hi-C genomics bins (CanMAG_XX-HiCbin.fa), including both the genome and the extra-chromosomal elements within the bin. </p> <p> </p>
Metagenome-assembled genomes(MAGs) generated from CRC human gut (PRJEB27928).
<p>MAGs generated from CRC human gut (PRJEB27928) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from dog gut (PRJEB20308).
<p>MAGs generated from dog gut (PRJEB20308) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Metagenome-assembled genomes(MAGs) generated from ocean (PRJEB1787).
<p>MAGs generated from ocean (PRJEB1787) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Fasta format protein sequences from assembled kyphosid fish gut metagenomes
<p>Predicted proteins sequences from kyposid fish gut metagenomic samples F5, F6, F7, and F8, obtained as described in the following study:</p> <p>Podell S, Oliver A, Kelly LW, Sparagon W, Plominsky, A, Nelson RS, Laurens LML, Augyte, S, Sims NA, Nelson CE, Allen EE. Herbivorous fish microbiome adaptations to sulfated dietary polysaccharides (2023)<br> manuscript submitted.</p>
HairSplitter: separating strains in metagenome assemblies with long reads
<p>Datasets, command lines and assemblies analyzed in the manuscript "HairSplitter: separating strains in metagenome assemblies with long reads".</p>
Oceanic Prokaryotes Metagenome-Assembled Genomes reconstructed using metagenomic distances
<p>Sets of reconstructed Metagenome-Assembled Genomes (MAGs) from Tara Oceans dataset. The reconstructed MAGs belong to Magneto paper: https://doi.org/10.1128/msystems.00432-22</p> <p>The dataset is composed of 93 oceanic metagenomes sampled from non-polar oceanic regions.</p> <p>The Metagenomic Distance MAGs were reconstructed following a co-assembly protocol driven by nucleotidic composition similarity, as detailed in the publication.</p> <p>The Oceanic Region MAGs were reconstructed by co-assembly of samples belonging to the same Oceanic Regions.</p> <p>The file clusters.tsv sum up the metagenomic distance cluster and the oceanic region each sample belongs to.</p>
Traing Data for "Assembly of metagenomic sequencing data" tutorial
<p>Metagenomics involves the extraction, sequencing and analysis of combined genomic DNA from <strong>entire microbiome</strong> samples. It includes then DNA from <strong>many different organisms</strong>, with different taxonomic background.</p> <p>Reconstructing the genomes of microorganisms in the sampled communities is critical step in analyzing metagenomic data. To do that, we can use <strong>assembly</strong> and assemblers, <em>i.e.</em> computational programs that stich together the small fragments of sequenced DNA produced by sequencing instruments.</p> <p>Assembling seems intuitively similar to putting together a jigsaw puzzle. Essentially, it looks for reads “that work together” or more precisely, reads that overlap. Tasks like this are <strong>not straightforward</strong>, but rather complex because of the complexity of the genomics (specially the repeats), the missing pieces and the errors introduced during sequencing.</p> <p>In this tutorial, we will learn how to run metagenomic assembly tool and evaluate the quality of the generated assemblies. To do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a>. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.