Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

915 results for “metagenomics”

Learn how ShareScore rates datasets ↗
zenodo36/100

HiFi Metagenomic Sequencing Enables Assembly of Accurate and Complete Genomes from Human Gut Microbiota.

<p>We reported 102 complete metagenome assembled genomes (cMAGs) from five human fecal HiFi sequencing samples.</p> <p>102_cMAGs_fna.tar.gz: Fasta sequence files of 102 cMAGs.</p> <p>gc_skew_figures.tar.gz: GC-skew pattern figures of 102 cMAGs. (SVG format)</p> <p>coverage_plots.tar.gz: Genome coverage plot of 102 cMAGs.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Dereplicated Metagenome assembled genomes (MAGs) from Columbia River hyporheic sediments

<p>Fasta file containing 55&nbsp;metagenome assembled genomes (MAGs) from&nbsp;publication to be submitted titled&nbsp;&quot;<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments&quot;.&nbsp;</strong></p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments

<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2)&nbsp;A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p>&nbsp;</p> <p>Manuscript title&nbsp;Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

MIntO: a Modular and Scalable Pipeline for Microbiome Metagenomic and Metatranscriptomic Meta-omics Data Integration

<p>To illustrate the use of MIntO, a set of 91 human fecal metagenomes from the Inflammatory Bowel Disease Multi&rsquo;omics Database was selected (IBDMDB).&nbsp;We selected six participants diagnosed as non-IBD (P6018 (nIBD1), M2072 (nIBD2)); Crohn&rsquo;s disease (H4006 (CD1) and H4020 (CD2)); and ulcerative colitis (H4019 (UC1) and H4035 (UC2)) that were followed for one year each.&nbsp;</p> <p>Here, we present the results from the <em>genome-based assembly-dependent</em>&nbsp;mode, where we used 91 metagenomic high-quality reads<strong> </strong>to recover 163&nbsp;high-quality MAGs,&nbsp;which constituted a set of non-redundant genomes.</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

A dynamic ancestral graph model and GPU-based simulation of a community based on metagenomic sampling

<p>In this paper we present an ancestral graph model of the evolution of a guild in an ecological community. The model is based on a metagenomic sampling design in that a random sample is taken at the community, as opposed the taxon, level and species are discovered by genetic sequencing. The specific implementation of the model envisions an ecological guild that was founded by colonization at some point in the past that then potentially undergoes diversification by natural selection. Within the graph, species emerge and evolve through the diversification process and their densities in the graph are dynamic and governed by both ecological drift and random genetic drift, as well as differential viability. We employ the 3% sequence divergence rule at a marker locus to identify Operational Taxonomic Units. We then explore approaches to see if there are indirect signals of the diversification process, including population genetic and ecological approaches. In terms of population genetics, we study the joint site frequency spectrum of OTUs, as well its associated statistics. In terms of ecology, we study the species (or OTU) abundance distribution. For both we observe deviations from neutrality, which indicates that there may be signals of diversifying selection in metagenomic studies under certain conditions. The model is available as a GPU-based computer program in C/C++ and using OpenCL, with the long-term goal of adding functionality iteratively to model large-scale eco-evolutionary processes for metagenomic data.</p>

opencc-zeroMar 2022View details →
zenodo36/100

A metagenomic-based study of two sites from the Barbadian reef system

<p><strong>A metagenomic-based study of two sites from the Barbadian reef system&nbsp;</strong></p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Metagenome-assembled genomes obtained from fecal and salivary microbiomes of pancreatic cancer patients and controls

<p>7,546 MAGs obtained from fecal and salivary metagenomes of pancreatic cancer patients and controls</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Pacbio of Broiler chicken: cecum metagenome

<p>The goal is to sequence and assemble broiler chicken&#39;s intestine microbial meta-genome. The data we show is sequenced from cecum microbial by PacBio.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AMR Benchmarking dataset - Metagenomics

<p>Metagenomic benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. These genomes were then randomly assigned a relative abundance, had additional AMR genes randomly inserted (to cover all AMR genes in CARD v3.1.4) and metagenomic Illumina reads simulated from them.<br> <br> Metagenomic simulation was performed using: https://github.com/fmaguire/AMR_Metagenome_Simulator and the entire process can be repeated using the <a href="https://zenodo.org/api/files/ba12e743-ae3d-43c0-98bd-2075d6f50340/metagenome_benchmark.sh">metagenome_benchmark.sh </a>script included above.<br> &nbsp;</p> <p><strong>Files</strong><br> `amr_benchmarking_metagenome.csv` contains the metadata the input genome accessions, paths, and simulated copy number used for creation of the AMR metagenome.</p> <p>`AMR_metagenome_labels.tsv` a two column csv containing names of all reads that are derived from an AMR gene and an identifier for the corresponding AMR gene. AMR genes are identified using CARD Antibiotic Resistance Ontology (ARO), with a suffix listing any SNVs for nmutation related resistance genes.<br> &nbsp;</p> <p>`simulated_metagenome.fna.gz` contains the full &quot;assembled&quot; true metagenomic contigs (derived directly from the input genome assemblies amplified to the correct copy number).</p> <p>`metagenome_unsorted.bed` contains the location of AMR genes in the full &quot;assembled&quot; true metagenomic contigs<br> <br> `simulated_metagenome_{1,2}.fq.gz` contain the simulated paired end metagenomics reads</p> <p>`simulated_metagenome_error_free.bam` contains the error-free mapping location from which simulated reads were derived</p> <table> <tbody> <tr> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> </tr> </tbody> </table>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Supp. Info. for Further analysis of metagenomic datasets containing GD and GX pangolin CoVs indicates widespread contamination, undermining pangolin host attribution

<p>Supplemanty Information for&nbsp;<strong>Further analysis of metagenomic datasets containing GD and GX pangolin CoVs indicates widespread contamination, undermining pangolin host attribution</strong></p> <p>Files:</p> <p>Supp_Info_1_PRJNA641544_DG14_DG18.xlsx</p> <p>Supp_Info_3_PRJNA606875_SRR11093270_reads_blast_nt_seq5_hsps1_PCT80_E0.05_hsps.txt</p> <p>Supp_Info_4_PRJNA573298_Analysis.xlsx</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Early-life human gut metagenome-assembled genomes and proteins catalogs

<p>The description of the files:</p> <p>(1) The 32,277 genomes include&nbsp;six parts:&nbsp;ELGG_part_1.zip,&nbsp;ELGG_part_2.zip,&nbsp;ELGG_part_3.zip,&nbsp;ELGG_part_4.zip,&nbsp;ELGG_part_5.zip,&nbsp;ELGG_part_6.zip.</p> <p>(2) The 2,172 representative&nbsp;species: ELGG_representatives_2172.zip.</p> <p>(3) The&nbsp;ELGP&nbsp;catalog&nbsp;clustered at 95% amino acid identity:&nbsp;ELGP_95.faa.gz.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Metagenome-assembled genomes (MAGs), colorectal cancer (CRC)

<p>This archive contains (i) Metagenome assemblies of short-term enrichment cultures of CRC mucosal tissue microbiota, and (ii) Reconstructed metagenome-assembled genomes (MAGs) generated through binning of metagenome contigs.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Genome-resolved metagenomics using short-, long-read and metaHiC sequencing

<p>Reference-quality metagenome-assembled genomes (MAGs) are the key to exploring microbial compositions and microbe-phenotype associations. They can be recovered by different sequencing technologies and computational tools, which need an unbiased and comprehensive assessment to identify best practices. This work systematically evaluates 40 distinct strategies to recover high-quality MAGs generated by eight assemblers, eight metagenomics binners, and four sequencing technologies, including short-, long-read and metaHiC sequencing. We notice that the hybrid assemblies of short- and long-reads outperform either short- or long-read assemblies and generate more contigs with high contiguity. When the hybrid assemblies are combined with metaHiC-based binning (Hybrid-HiC), more high-quality MAGs with higher taxonomic diversity are recovered, and more tRNA and rRNA genes, phages, plasmids and antibiotic resistance genes are identified in the mock, simulated and real datasets.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Supplementary informations of endogenous starter culture and cheese metagenomes

<p>The first description of the virome composition in Brazilian artisanal Canastra cheese and the phage-bacterial interactions in this food system. Here you can access supplemental methods and results (docx) and supplemental data&nbsp;(MAGs contigs, novel 987 group phage contigs, and MAGs spacers).</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Metagenomic data for Bathymodiolus symbionts deposited in IMG (2017)

<p>Metagenomic data for the sulfur- and methane-oxidizing symbionts of <em>Bathymodiolus</em> mussels and different sponge species deposited in the Integrated Microbial Genomes (IMG) database of the DOE Joint Genome Institute (http://img.jgi.doe.gov/)  until October 2017.</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Metagenomic data for gutless oligochaete symbionts (2017)

<p>This list contains accession numbers to:</p> <p>-Metatranscriptomic data of gutless oligochaete Olavius algarvensisworms depositied in European Nucleotide Archive ENA</p> <p>-Metaproteomic data of gutless oligochaete Olavius algarvensisworms deposited in proteomics data repositories</p> <p>-Draft genomes of endosymbionts from various gutless oligochaete hosts, in submission process.</p> <p>If you are interested in these datasets, please let us know.</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification

<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a>:</p> <ul> <li>CAMI2_reference_database.tar.gz_1 to CAMI2_reference_database.tar.gz_6:&nbsp;Merge them into a single file using the&nbsp;<em>cat</em> command. The folder produced by decompressing this file is the reference database for CAMI2 (RefSeq sequence filenames are in the file CAMI2_reference_database_16864_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification

<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a>&nbsp;(see Figure 1 of the article):</p> <ul> <li>37345_filtered_training_and_test_genomes_sequence_files.tar.gz_1 to 37345_filtered_training_and_test_genomes_sequence_files.tar.gz_6:&nbsp;Merge them into a single file using the&nbsp;<em>cat</em> command. The folder produced by decompressing this file contains the filtered RefSeq sequence files of all 37345 training and test genomes, which were used to build the uniform reference database and generate the test reads (The filenames are in the file 37345_filtered_training_and_test_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Culex pipiens merged anvi'o profiles from midgut and ovary metagenomes

<p>Anvi&rsquo;o merged profile&nbsp;databases for <em>Culex pipiens</em> midgut and ovary samples.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Wolbachia MAGs from Culex pipiens midgut and ovary metagenomes

<p><em>Wolbachia</em> MAGs (fasta files) from <em>Culex pipiens</em> midgut and ovary samples.&nbsp;</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record