Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

915 results for “metagenomics”

Learn how ShareScore rates datasets ↗
zenodo52/100

Orbicella faveolata and O. franksi coral metagenome assemblies from the Lower Florida Keys region of Florida, USA

<div> <p>The enclosed files include mostly <em>Orbicella faveolata</em> and three <em>Orbicella franksi</em> coral metagenome assemblies collected from the Lower Keys in Florida&rsquo;s Coral Reef, USA. Metadata for the files is included in this repository. Apparently healthy coral tissue cores were collected between May 28 and June 21, 2021. The DNA was extracted from the host and associated microorganisms and sequenced in a paired-end 150 bp format on an Illumina NovaSeq. Trimming and quality filtering of DNA sequence reads proceeded, followed by host and photoendosymbiotic dinoflagellate DNA removal. The host-cleaned reads were assembled individually by coral sample into longer contigs using MegaHit v1.1.4. The &ldquo;Assembly_Fastas&rdquo; zipped file contains 41 metagenome assemblies from the individual <em>Orbicella faveolata</em> corals and 3 assemblies from the individual <em>Orbicella franksi&nbsp;</em>colonies for a total of 44 assemblies. In addition, these assemblies were annotated with eggnog-mapper v2.1.6 to generate both predicted gene regions and annotation output files. The &ldquo;Predicted_Gene_Fastas&rdquo; zipped file contains nucleotide fasta files of the predicted gene regions for all 44&nbsp;coral metagenome assemblies. The fasta header of each gene includes the contig ID it originated from in the associated &ldquo;Assembly_Fasta&rdquo;. The &ldquo;Predicted_Gene_Annotations&rdquo; zipped file contains either .csv or .xlsx files with the eggnog-mapper-based annotations. These files contain a &ldquo;query contig&rdquo; that corresponds to the contig ID in the fasta header of the &ldquo;Predicted_Gene_Fasta&rdquo;.&nbsp;</p> <p>In addition to individual assemblies, a co-assembly was generated that included all 41 <em>Orbicella faveolata</em> coral samples. Prior to co-assembly, further removal of eukaryotic DNA proceeded by splitting the indiviudual assemblies into eukaryotic and prokaryotic content with the program EukRep v0.6.7, followed by mapping of the host-clean reads to the eukaryotic DNA to remove them. The eukaryote-clean reads from all 41 corals were input into MegaHit to generate a co-assembly. The co-assembly is included (FLK_OFAV_MG_coassembly_final.contigs.fa). Predicted genes from the co-assembly were generated with Prodigal v2.6.3 and the nucleotide fasta of the output is included in this repository (FLK_OFAV_MG_pred.fna). Like with the indiviudal assemblies, eggnog-mapper was used to generate annotations of the predicted genes from Prodigal (FLK_OFAV_MG.emapper.annotations.xlsx).&nbsp; Additionally, the abundance of each predicted gene was generated using Salmon to map the eukaryote-clean reads to the predicted genes. The number of reads (counts) for each gene across each coral sample were aggregated as integers into one table and included in this repository (FLK_OFAV_MG_pred_NumReads.tsv).&nbsp;&nbsp;</p> </div> <div> <p>These data were processed and generated by Julie Meyer&rsquo;s Lab at the University of Florida, using funding from the Florida Department of Environmental Protection.&nbsp;&nbsp;</p> </div>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Simulated metagenomes with quality and abundance distributions derived from real samples

<p>Species abundances and quality values were derived from the following list of samples:</p> <pre><code>SAMEA2466896 SAMEA2466916 SAMEA2466952 SAMEA2466953 SAMEA2466965 SAMEA2466996 SAMEA2467015 SAMEA2467039 SAMEA2621010 SAMEA2621033 SAMEA2621107 SAMEA2621155 SAMEA2621229 SAMEA2621247 SAMEA2621300 SAMEA2622357 </code></pre> <p>Reference abundances (.abund files) were generated using <a href="https://github.com/motu-tool/mOTUs_v2">mOTUs profiler</a>.<br> Metagenomes were simulated with <a href="https://sourceforge.net/projects/cmessi/">cMESSi</a> using <a href="http://progenomes.embl.de/data/repGenomes/representatives.contigs.fasta.gz">proGenomes&#39; representative contigs</a> for species and the aforementioned abundances. In cases where a <em>ref_mOTU_v2</em> corresponded to more than one genome, the abundance of said <em>ref_mOTU</em> was distributed equally over all genomes.<br> GFF location files were produced using location information generated by cMESSi.<br> Two variants of truth values were obtained by intersecting coordinates of simulated reads with coordinates of <a href="http://eggnogdb.embl.de">eggNOG</a> orthologous groups (OG at NOG level) as predicted by <a href="https://github.com/jhcepas/eggnog-mapper">eggNOG-mapper</a>.</p> <ol> <li>.cog-simulated files contain the NOG distribution that was effectively simulated, <em>i.e.</em> a count of the number of reads overlapping with genes annotated with each NOG. A read overlapping multiple genes is considered for each gene. If a gene possesses multiple NOG annotations, each annotation gets assigned the total number of overlapping reads. Longer genes will (in expectation) generate more reads, all else being equal.</li> <li>.cog-distribution file contains the expected distribution for every NOG on all samples. The number of genes annotated with each NOG is multiplied by the abundance of the corresponding species. Length of the gene is not taken into account.</li> </ol> <p>If you use this dataset, please cite: <a href="https://www.biorxiv.org/node/111718.full">NG-meta-profiler: fast processing of metagenomes using NGLess, a domain-specific language</a></p>

opencc-by-4.0Jan 2019View details →
zenodo48/100

Public metagenome datasets annotated using SingleM

<p>These data underlie the community profiles shown at <a href="https://sandpiper.qut.edu.au">https://sandpiper.qut.edu.au</a></p> <p>&nbsp;</p> <h2>Changelog</h2> <p>version 1.0.0</p> <ul> <li>Public metagenomes published before Feb 20, 2025 were analysed using SingleM pipe v0.18.3 (the default R220 metapackage), and then renewed using an R226 metapackage (S5.4.0.GTDB_r226.metapackage_20250331).</li> </ul> <p>version 0.3.0</p> <ul> <li>Update profiles to use GTDB R220, generated using SingleM renew v0.17.0.</li> </ul> <p>version 0.2.0</p> <ul> <li>Initial version. Created using a GTDB R214-based reference SingleM metapackage S3.2.1.GTDB_r214.metapackage_20231006 based on public datasets available Dec 15, 2021.</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly

<p>Data and conda software environment file for the chapter &#39;<em>de novo</em> Genome Assembly&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Authentication

<div> <p>Data and conda software environment file for the chapter 'Authentication' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to Git(Hub)

<p>Data and conda software environment file for the chapter &#39;Introduction to Git(Hub)&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to R and the Tidyverse

<p>Data and conda software environment file for the chapter &#39;Introduction to R and the Tidyverse&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to the Command Line

<p>Data and conda software environment file for the chapter 'Introduction to the Command Line' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to Python and Pandas

<p>Data and conda software environment file for the chapter &#39;Introduction to Python and Pandas&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Phylogenomics

<p>Data and conda software environment file for the chapter &#39;Phylogenomics&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Genome Mapping

<p>Data and conda software environment file for the chapter &#39;Genome Mapping&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Taxonomic Profiling, OTU Tables, and Visualisation

<p>Data and conda software environment file for the chapter &#39;Taxonomic Profiling, OTU Tables, and Visualisation&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Ancient Metagenomic Pipelines

<p>Data and conda software environment file for the chapter &#39;Ancient Metagenomic Pipelines&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Contamination

<p>Data and conda software environment file for the chapter 'Contamination' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Metagenome quality metrics and taxonomical annotation visualization through the integration of MAGFlow and BIgMAG (Sup. Material)

<p>Dataset encompassing:</p> <ul> <li>The recovered MAGs by 6 different metagenomics pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) using a mock community as input (SRR8359173 and SRR9328980), complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.&nbsp;</li> <li>The MAGs produced by nf-core/mag using rice/rhizosphere sequenced libraries (PRJNA663614, PRJNA448773 and PRJNA645385) in either single assembly/single binning or co-assembly/co-binning mode, complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.</li> <li>Scripts, commands and configuration files to run the different pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) and reproduce the experimental conditions.</li> <li>Outputs, commands and scripts to run Metabinner and Semibin in their default configuration using the rice soil samples co-assembly, along with the MAGFlow (v1.1.0) output to compare these binners against MetaBAT2.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (MAGs v2)

<p><strong>This release (MAGs v2) is a major new version of this metagenome-assembled genome (MAG) set.</strong> All previous releases on this page (which only differ in the metadata) are designated "MAGs v1." The current release (MAGs v2) uses<strong>&nbsp;</strong>CheckM2 v1.0.2 filtering (&ge;70% completeness, &le;10% contamination) to expand this dataset to include <strong>36,419 MAGs</strong>, with the following subcategories:</p> <ul> <li>Cronin_v1:&nbsp; Manually-curated subset of the "Field" category from MAGs v1.</li> <li>Cronin_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate Cronin_v1.</li> <li>Woodcroft_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate the MAGs reported in <a href="https://doi.org/10.1038/s41586-018-0338-1">Woodcroft &amp; Singleton et al. (2018)</a>.</li> <li>SIPS:&nbsp; Updated genomes from samples originating from a stable isotope probing (SIP) incubation experiment by Moira Hough et al. ("SIP" in MAGs v1), re-analyzed due to read truncation and sample linkage issues in MAGs v1.</li> <li>JGI:&nbsp; Expanded set of genomes from the Joint Genome Institute's metagenome annotation pipeline.</li> </ul> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>Emerge_MAGs_v2.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_v2_EMERGE.tsv</strong>&nbsp;- Table containing source sample names and accessions, GTDB taxonomy information, CheckM2 quality reports, NCBI GenomeBatch- and MIMAG(6.0)-formatted sample attributes and other metadata for the MAGs.&nbsp;</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data collected at the Joint Genome Institute was generated under the following awards:</p> <ul> <li>The majority of sequencing at JGI was supported by BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</li> <li>Sequencing of SIP samples was performed under the Facilities Integrating Collaborations for User Science (FICUS) initiative (proposal 503547; award DOI:&nbsp;<a href="https://doi.org/10.46936/fics.proj.2017.49950/60006215">10.46936/fics.proj.2017.49950/60006215</a>) and used resources at the DOE Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://ror.org/04rc0xn13">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities. Both facilities are sponsored by the Office of Biological and Environmental Research and operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Metagenome-Assembled Genomes Abundance & Activity Tables. Environmental Parameters Associated with the dataset.

<p>Lake Mendota, WI, USA, is a temperate lake subject to annual temperature and oxygen fluctuations. Each summer, the water column becomes anoxic (no-oxygen). In 2020, we sampled the lake at weekly intervals, at different depths (5, 10, 15, 20 and 23.5m). For each time+depth sample, we collected metagenomes, viromes and metatranscriptomes. Environmental data profiles were collected on-site for each sampling day.&nbsp;</p> <p>Following standard metagenomic binning best practices, we obtained 431 metagenomes-assembled-genomes (MAGs).</p> <p>This record comprises the microbial abundance and expression table for these MAGs, and the environmental profiles collected each day.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Gut Metagenome Assemblies for Veseli et al. 2023

<p>A collection of anvi&#39;o contigs databases for 408 human fecal metagenome assemblies from the study by Veseli et al. titled &quot;High metabolic independence is a determinant of microbial resilience in the face of gut stress&quot;. These are publicly-available gut metagenomes originally obtained from several studies of the gut microbiome. See `METAGENOMES_INFO.txt` file for references and sample SRA accessions.</p> <p>The metagenomes were assembled individually using IDBA-UD as part of the anvi&#39;o metagenomics&nbsp;workflow in anvi&#39;o v7.1-dev. As part of this workflow, they were annotated with KEGG KOfams using `anvi-run-kegg-kofams` and a KEGG snapshot from December 12, 2020&nbsp;(modules database hash value `45b7cc2e4fdc`). See manuscript and its reproducible workflow for details.</p>

opencc-by-4.0Apr 2023View details →
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of metagenomic DNA sequence analyses of bacterial and archaeal water column communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2012

In contrast to temperate systems, Arctic lagoons that span the Alaska Beaufort Sea coast face extreme seasonality. Nine months of ice cover up to ∼1.7 m thick is followed by a spring thaw that introduces an enormous pulse of freshwater, nutrients, and organic matter into these lagoons over a relatively brief 2–3 week period. Prokaryotic communities link these subsidies to lagoon food webs through nutrient uptake, heterotrophic production, and other biogeochemical processes, but little is known about how the genomic capabilities of these communities respond to seasonal variability. This study characterizes the metabolic capabilities of microbial communities across three seasons in two lagoons and one open coastal site along the eastern Alaska Beaufort Sea coast. We used metagenomic DNA sequence data of bacterial and archaeal water column communities to identify genes of relevant biogeochemical pathways. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA642637 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA642637. This data package is associated with the following publication: Baker, Kristina D., Colleen T. E. Kellogg, James W. McClelland, Kenneth H. Dunton, and Byron C. Crump. “The Genomic Capabilities of Microbial Communities Track Seasonal Variation in Environmental Conditions of Arctic Lagoons.” Frontiers in Microbiology 12 (2021). https://doi.org/10.3389/fmicb.2021.601901. Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provi

openCC0Apr 2021View details →
zenodo44/100

MACREL software benchmark data set: Simulated metagenomes with sequencing quality, errors profile and abundance distributions derived from real samples

<p>These metagenomes were used in the benchmarking of FACS pipeline, and were designed after NGLess benchmark dataset (doi.org/10.5281/zenodo.2560288).&nbsp; Metagenomes were simulated with <a href="https://www.niehs.nih.gov/research/resources/software/biostatistics/art/index.cfm">ART-bin-MountRainier-2016.06.05</a> using real abundance profiles (.abund files) available <a href="https://doi.org/10.5281/zenodo.2560288">elsewhere</a>, and <a href="http://progenomes1.embl.de/data/repGenomes/representatives.contigs.fasta.gz">proGenomes&#39; representative contigs</a> as reference genomes. There are available metagenomes with 40, 60 and 80 M (million of reads) based in the reference genomes and abundances of the following samples:</p> <pre><code>SAMEA2466916 SAMEA2466953 SAMEA2466965 SAMEA2621107 SAMEA2621229 SAMEA2621247</code></pre> <p>To convert them from the CRAM format back to fastq files:</p> <pre><code> ## 1. converting from cram to bam format: samtools view -b -T refgenome.fa -o file.bam file.cram ## 2. sorting the bam file: samtools sort -n file.bam -o input_sorted.bam # sort reads by identifier-name (-n) ## 3. converting from bam to fastq format: bedtools bamtofastq -i input_sorted.bam -fq output_r1.fastq -fq2 output_r2.fastq </code></pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record