Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

39 results for “Shotgun metagenomics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Evaluation of an adapted semi-automated DNA extraction for human salivary shotgun metagenomics

<p>This deposit contains :</p> <p>- a&nbsp;RMarkdown filte containing the&nbsp;codes for the mcirobial analysis of saliva samples</p> <p>- the html report with codes,&nbsp;results and figures</p> <p>- a RData containing microbial datasets (MSp species abundance table, genus, family and phylum abundance tables, matrix of genes correlations, taxonomy)</p> <p>- a RData containing associated metadata&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

High resolution shotgun metagenomics: the more data, the better?

<p>This data archive contains results generated using a high resolution shotgun metagenomics (HRSM) bioinformatic pipeline (ShotgunMG - https://jtremblay.github.io/shotgunmg.html) for the following projects:</p> <p>Human gut microbiome dataset: PRJNA588513</p> <p>Antarctic soil dataset: PRJNA513362</p> <p>Agricultural soil dataset: PRJNA513362</p> <p>Mock communities: PRJNA873699</p> <p>These analyses were performed in the context of evaluating if shallow shotgun metagenomic sequencing is an adequate approach to analyze SM sequencing data using a HRSM pipeline.</p> <p>Briefly, these PRJNA projects were analyzed using identical bioinformatic procedures, but using various initial raw sequencing data loads.</p> <p>Notable end results in each archive includes: de novo co-assembly (fasta files), contigs and genes abubance matrices. Beta-diversity (Bray-Curtis dissimilarity matrices), alpha diversity (richness, chao1, Simpson and Shannon indexes matrices), taxonomic summaries from the kingdom to up to the species level and Metagenome Assembled Genomes (MAGs).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Shotgun metagenomic sequencing dataset of a synthetic mock community containing 20 genomes spiked-in at even and staggered concentrations.

<p>Shotgun metagenomics (SM) sequencing is a popular method used in microbial ecology to obtain insights on microbial community structure and function potential in a given biological system without the need to cultivate microorganisms. The dataset described in this article describes technical triplicates of shotgun metagenomic sequence libraries generated from two purified and titrated mixes of 20 distinct reference bacterial genomes for which key characteristics such as genome size, sequence and spiked-in concentrations are known. In one of the genomic DNA mix, each genome is spiked-in at similar concentrations (representing an even microbial community) and in the other, genomes are spiked-in at different concentrations with some genomes highly abundant and other in low quantity, mimicking an uneven microbial community DNA extract. In order to be interpretable, SM sequencing data needs to be properly analyzed by complex analytical bioinformatic pipelines. Environments investigated with this method can range from simple to very complex. Typically, microbial communities contain microbes that are ubiquitous and some others much rarer. Analysis of rare microbes in a complex microbial community are challenging to perform as their sequencing signals get submerged by the microbial genomes that are more abundant. In this context, it is critical to have access to sequencing data of simple mock communities of mixes of well characterized genomes in order to develop and validate bioinformatic methods that aim to accurately analyze microbial communities.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Databases for MyCodentifier: A tool for routine identification of nontuberculous mycobacteria using MGIT enriched shotgun metagenomics.

<p>Databases used for MyCodentifier a Nextflow pipeline to identify Mycobacterium tuberculosis complex (MTBC)&nbsp;and Nontuberculous mycobacteria (NTM)&nbsp;species from Next-generation sequencing (NGS)&nbsp;data.<br> <br> <strong>Short description:</strong><br> The pipeline is constructed using nextflow as workflow manager running in a docker container. It is able to identify species of MTBC/NTM from positive Mycobacterial Growth Indicator Tube (MGIT) cultures. To do so it uses an hsp65 database for fast identification coupled with a Metagenomic method using centrifuge to identify on genome level. For TB it also is able to identify subspecies. Results are presented in automated pdf and html reports.</p> <table> <caption><strong>Databases</strong></caption> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Short Description</strong></td> </tr> <tr> <td>20220726_ref.tar.gz</td> <td>7 major mycobacterial genomes as centrifuge classification database, used for reference-based mapping and genotype resistance prediction</td> </tr> <tr> <td>20220726_wgs_centrifuge_db_Radboudumc_MB.tar.gz</td> <td>centrifuge classification database using Tortoli <em>et al</em> 2017 Mycobacterium strains + additional strains</td> </tr> <tr> <td>genomes.tar.gz</td> <td>7 major mycobacterial genomes, annotation and Genbank files. Files are paired with 20220726_ref.tar.gz</td> </tr> <tr> <td>snpEff.tar.gz</td> <td>7 major mycobacterial genomes annotation models for snpEff.</td> </tr> <tr> <td>Tortoli_etal_hsp65.tar.gz</td> <td>KMA database of hsp65 gene extractions of the&nbsp;Tortoli <em>et al</em> 2017 Mycobacterium strains.</td> </tr> <tr> <td> <p>Used in the study:<br> p_compressed+h+v.tar.gz (12/06/2016)</p> </td> <td> <p>Databases available via&nbsp;<a>ftp://ftp.ccb.jhu.edu/pub/infphilo/centrifuge/data</a>&nbsp;or&nbsp;<a href="https://ccb.jhu.edu/software/centrifuge/manual.shtml#custom-database">https://ccb.jhu.edu/software/centrifuge/manual.shtml#custom-database</a></p> </td> </tr> </tbody> </table> <p><strong>MyCodentifier Github:</strong></p> <p><a href="https://jordycoolen.github.io/MyCodentifier/">https://jordycoolen.github.io/MyCodentifier/</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data - use case

<p>Data products returned by&nbsp;<a href="https://github.com/emo-bon/MetaGOflow">metaGOflow</a> (<a href="https://github.com/emo-bon/MetaGOflow/releases/tag/v1.0.0">v1.0.0</a>) and packed as a Research Object&nbsp;(RO) Crate, when performed with:</p> <ul> <li>a <strong>seawater metagenomic sample </strong>(TARA OCEAN,&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/ERR599171">ERR599171</a>)</li> <li>a <strong>fish gut&nbsp;</strong>sample (<a href="https://www.ebi.ac.uk/ena/browser/view/ERR4765907">ERR4765907</a>)</li> <li>a<strong> human gut </strong>sample (<a href="https://www.ebi.ac.uk/ena/browser/view/SRR9654976">SRR9654976</a>)</li> </ul> <p>This Zenodo repo accompanies the metaGOflow paper and more about the analysis of this sample can be found there.</p> <p>You can also have a look at some visual components of the workflow at this <a href="https://data.emobon.embrc.eu/MetaGOflow/">GitHub page</a>.&nbsp;</p> <p>The source code of metaGOflow is available through <a href="http://github.com/emo-bon/MetaGOflow">GitHub</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Supplementary Material for publication "Bifidobacteria Define Gut Microbiome Profiles of Golden Lion Tamarin (Leontopithecus rosalia} and Marmoset Callithrix sp. Metagenomic Shotgun Pools

<p>Supplementary Tables and Figure for the publication&nbsp;&quot;Bifidobacteria Define Gut Microbiome Profiles of Golden Lion Tamarin <em>Leontopithecus rosalia</em>&nbsp;and Marmoset <em>Callithrix</em> sp. Metagenomic Shotgun Pools&quot;</p>

opencc-by-4.0Aug 2023View details →
dryad32/100

Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics

<p>Amplicon metabarcoding is an established technique to analyse the taxonomic composition of communities of organisms using high-throughput DNA sequencing, but there are doubts about its ability to quantify the relative proportions of the species, as opposed to the species list. Here, we bypass the enrichment step and avoid the PCR-bias, by directly sequencing the extracted DNA using shotgun metagenomics. This approach is common practice in prokaryotes, but not in eukaryotes, because of the low number of sequenced genomes of eukaryotic species. We tested the metagenomics approach using insect species whose genome is already sequenced and assembled to an advanced degree. We shotgun-sequenced, at low-coverage DNA, 18 species of insects in 22 single-species and 6 mixed-species libraries and mapped the reads against 110 reference genomes of insects. We used the single-species libraries to calibrate the process of assignation of reads to species and the libraries created from species mixtures to evaluate the ability of the method to quantify the relative species abundance. Our results showed that the shotgun metagenomic method is easily able to set apart closely-related insect species, like four species of <i>Drosophila</i> included in the artificial libraries. However, to avoid the counting of rare misclassified reads in samples, it was necessary to use a rather stringent detection limit of 0.001, so species with a lower relative abundance are ignored. We also identified that approximately half the raw reads were informative for taxonomic purposes. Finally, using the mixed-species libraries, we showed that it was feasible to quantify with confidence the relative abundance of individual species in the mixtures.</p>

opencc-zeroJan 2021View details →
zenodo32/100

MetaChick: characterization of the chicken caecal metagenome by deep shotgun sequencing

<p></p><h1>Data sources</h1><br>This dataset was constructed using the samples of the MetaChick project (phase 1) corresponding to the cecal content of 340 animals. Sequencing data and associated metadata have been submitted to INSDC (bioproject: PRJEB38174).<br><h1>Sequencing data QC and metagenomic assembly</h1><br>First, sequencing adapters removal and read trimming was performed with fastxtend. Reads mapped on the host genome (GRCg7b GCA_016699485.1) with bowtie2 were removed with samtools. Finally, metagenomic assembly was performed with metaSPAdes v3.14.1. Contigs of less than 1500 bp were removed.<br><h1>MAGs recovery</h1><br>MAGs were generated with MetaBAT 2 (multi-coverage mode) and MAGs quality was assessed with CheckM. MAGs with completeness &lt; 70% or contamination &gt; 5% or N50 &lt; 8Kb were discarded. Pairwise Average Nucleotide Identity (ANI) was computed for all recovered MAGs with fastANI and dereplication at species level (ANI cutoff = 95%).<br><h1>Non-redundant gene catalog</h1><br>Genes were predicted on all contigs from metagenomic assemblies with Prodigal (parameters : -m -p meta). Genes were pooled and clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0) by choosing those from the longest contigs as representatives.<br><h1>MSPs recovery</h1><br>A raw gene abundance table (13,6M genes quantified in 340 samples) was generated with meteorMeteor. Then, co-abundant genes were binned in Metagenomic Species Pan-genomes (MSPs, i.e. gene clusters that likely belong to the same microbial species) using MSPminer.<br><h1>MAGs and MSPs taxonomic annotation</h1><br>Dereplicated MAGs were annotated with GTDB-Tk based on GTDB r214. Then, MAGs taxonomic annotation was propagated to the corresponding MSPs.<br><h1>Construction of the phylogenetic tree</h1><br>39 universal phylogenetic markers genes were extracted from the dereplicated MAGs with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB38174 (cohort used in catalogue assembly) and PRJEB29033, PRJEB33338, PRJEB53667 (independent cohort not used in assembly).<p></p>

opencc-zeroDec 2021View details →
zenodo32/100

MicroReset: characterization of the rabbit (Oryctolagus cuniculus) fecal metagenome and resistome by deep shotgun sequencing

<p></p><h1>Data source</h1><br>The dataset was generated from 30 rabbit fecal samples subjected to deep shotgun metagenomic sequencing. The sequencing data is available under the BioProject PRJEB50625.<br>Metagenomic Assembly<br>Raw sequencing reads were first pre-processed using fastp for adapter removal and quality trimming. Host-derived reads were filtered out by mapping to the rabbit reference genome (GCF_000001635.27) using Bowtie2 and removing mapped reads with Samtools. Each sample was individually assembled using metaSPAdes. Contigs shorter than 1,500 bp were excluded from downstream analysis.<br><h1>MAG Recovery</h1><br>Reads from each sample were mapped to all 30 assemblies (30×30 mappings) using Bowtie2. The resulting alignments were sorted and indexed with Samtools. Contig coverage across all samples was computed using `jgi_summarize_bam_contig_depths`. Binning was performed with MetaBAT 2 and SemiBin v1.3. MAG quality was assessed with CheckM. Only high-quality MAGs (≥70% completeness, ≤5% contamination, N50 ≥ 8 kb) were retained.<br>Non-Redundant Gene Catalog<br>Gene prediction was carried out using Prodigal on all contigs from the current study (with `-m -p meta`). Genes shorter than 90 bp or lacking start/stop codons were discarded. The remaining genes from both sources were pooled and clustered using CD-HIT-EST (parameters: `-c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0`). The longest contigs were used to select representative genes.<br><h1>MSP Recovery</h1><br>Shotgun reads from the 30 samples were aligned to the non-redundant gene catalog using the Meteor suite, generating a gene abundance matrix (5.7 million genes × 30 samples). Co-abundant genes were grouped into 1,053 Metagenomic Species Pan-genomes (MSPs) using MSPminer.<br><h1>Taxonomic Annotation of MSPs</h1><br>MAGs representing each species were taxonomically annotated using GTDB-Tk with GTDB release r214. The resulting taxonomy was propagated to the corresponding MSPs.<br><h1>Phylogenetic Tree Construction</h1><br>A set of 39 universal phylogenetic marker genes was extracted from the 1,053 MSPs (or their corresponding MAGs, when available) using fetchMGs. Each marker was independently aligned using MUSCLE, and the alignments were concatenated and trimmed using trimAl (parameter: `-automated1`). A maximum-likelihood phylogenetic tree was constructed with FastTreeMP (parameters: `-gamma -pseudo -spr -mlacc 3 -slownni`).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters) for PRJEB50625 (cohort used in catalogue assembly).<p></p>

opencc-zeroDec 2021View details →
zenodo32/100

Supplementary material 7 from: Tedersoo L, Anslan S, Bahram M, Põlme S, Riit T, Liiv I, Kõljalg U, Kisand V, Nilsson RH, Hildebrand F, Bork P, Abarenkov K (2015) Shotgun metagenomes and multiple primer pair-barcode combinations of amplicons reveal biases in metabarcoding analyses of fungi. MycoKeys 10: 1-43. https://doi.org/10.3897/mycokeys.10.4852

Table S7. Taxonomic classification of the rDNA of fungal.: Explanation note: Taxonomic classification of the rDNA of fungal shotgun metagenome.

opencc-by-4.0May 2015View details →
zenodo32/100

Supplementary material 3 from: Tedersoo L, Anslan S, Bahram M, Põlme S, Riit T, Liiv I, Kõljalg U, Kisand V, Nilsson RH, Hildebrand F, Bork P, Abarenkov K (2015) Shotgun metagenomes and multiple primer pair-barcode combinations of amplicons reveal biases in metabarcoding analyses of fungi. MycoKeys 10: 1-43. https://doi.org/10.3897/mycokeys.10.4852

Table S3. Data set of the SSU V4 and V5 barcodes.: Explanation note: Data set of the SSU V4 and V5 barcodes.

opencc-by-4.0May 2015View details →
zenodo32/100

Supplementary material 1 from: Tedersoo L, Anslan S, Bahram M, Põlme S, Riit T, Liiv I, Kõljalg U, Kisand V, Nilsson RH, Hildebrand F, Bork P, Abarenkov K (2015) Shotgun metagenomes and multiple primer pair-barcode combinations of amplicons reveal biases in metabarcoding analyses of fungi. MycoKeys 10: 1-43. https://doi.org/10.3897/mycokeys.10.4852

Table S1. Characteristics of soil samples.: Explanation note: Characteristics of soil samples used in this study.

opencc-by-4.0May 2015View details →
zenodo32/100

Supplementary material 6 from: Tedersoo L, Anslan S, Bahram M, Põlme S, Riit T, Liiv I, Kõljalg U, Kisand V, Nilsson RH, Hildebrand F, Bork P, Abarenkov K (2015) Shotgun metagenomes and multiple primer pair-barcode combinations of amplicons reveal biases in metabarcoding analyses of fungi. MycoKeys 10: 1-43. https://doi.org/10.3897/mycokeys.10.4852

Table S6. Data set of the LSU D1, D2, and D3 barcodes.: Explanation note: Data set of the LSU D1, D2, and D3 barcodes.

opencc-by-4.0May 2015View details →
zenodo32/100

Supplementary material 2 from: Tedersoo L, Anslan S, Bahram M, Põlme S, Riit T, Liiv I, Kõljalg U, Kisand V, Nilsson RH, Hildebrand F, Bork P, Abarenkov K (2015) Shotgun metagenomes and multiple primer pair-barcode combinations of amplicons reveal biases in metabarcoding analyses of fungi. MycoKeys 10: 1-43. https://doi.org/10.3897/mycokeys.10.4852

Table S2. Taxonomic composition and clustering of the mock community sample.: Explanation note: Taxonomic composition and clustering of the mock community sample.

opencc-by-4.0May 2015View details →
dryad32/100

Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics

Open the record for dataset details and reuse information.

publicJan 2021View details →
zenodo28/100

Supplementary material 8 from: Garrido-Sanz L, Senar MÀ, Piñol J (2020) Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics. Metabarcoding and Metagenomics 4: e48281. https://doi.org/10.3897/mbmg.4.48281

: Data type: Excel table

opencc-zeroFeb 2020View details →
zenodo28/100

Supplementary material 5 from: Garrido-Sanz L, Senar MÀ, Piñol J (2020) Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics. Metabarcoding and Metagenomics 4: e48281. https://doi.org/10.3897/mbmg.4.48281

: Data type: Excel table

opencc-zeroFeb 2020View details →
zenodo28/100

Supplementary material 4 from: Garrido-Sanz L, Senar MÀ, Piñol J (2020) Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics. Metabarcoding and Metagenomics 4: e48281. https://doi.org/10.3897/mbmg.4.48281

: Data type: Excel table

opencc-zeroFeb 2020View details →
zenodo28/100

Supplementary material 3 from: Garrido-Sanz L, Senar MÀ, Piñol J (2020) Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics. Metabarcoding and Metagenomics 4: e48281. https://doi.org/10.3897/mbmg.4.48281

: Data type: Excel table

opencc-zeroFeb 2020View details →
zenodo28/100

Supplementary material 6 from: Garrido-Sanz L, Senar MÀ, Piñol J (2020) Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics. Metabarcoding and Metagenomics 4: e48281. https://doi.org/10.3897/mbmg.4.48281

: Data type: Excel table

opencc-zeroFeb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record