Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.7.1
Dataset results
31 results for “shotgun sequencing”
Shotgun metagenomic sequencing dataset of a synthetic mock community containing 20 genomes spiked-in at even and staggered concentrations.
<p>Shotgun metagenomics (SM) sequencing is a popular method used in microbial ecology to obtain insights on microbial community structure and function potential in a given biological system without the need to cultivate microorganisms. The dataset described in this article describes technical triplicates of shotgun metagenomic sequence libraries generated from two purified and titrated mixes of 20 distinct reference bacterial genomes for which key characteristics such as genome size, sequence and spiked-in concentrations are known. In one of the genomic DNA mix, each genome is spiked-in at similar concentrations (representing an even microbial community) and in the other, genomes are spiked-in at different concentrations with some genomes highly abundant and other in low quantity, mimicking an uneven microbial community DNA extract. In order to be interpretable, SM sequencing data needs to be properly analyzed by complex analytical bioinformatic pipelines. Environments investigated with this method can range from simple to very complex. Typically, microbial communities contain microbes that are ubiquitous and some others much rarer. Analysis of rare microbes in a complex microbial community are challenging to perform as their sequencing signals get submerged by the microbial genomes that are more abundant. In this context, it is critical to have access to sequencing data of simple mock communities of mixes of well characterized genomes in order to develop and validate bioinformatic methods that aim to accurately analyze microbial communities.</p>
Shallow shotgun sequencing of the microbiome recapitulates 16S amplicon results and provides functional insights
<p>Prevailing 16S rRNA gene-amplicon methods for characterizing the bacterial microbiome of wildlife are economical, but result in coarse taxonomic classifications, are subject to primer and 16S copy number biases, and do not allow for direct estimation of microbiome functional potential. While deep shotgun metagenomic sequencing can overcome many of these limitations, it is prohibitively expensive for large sample sets. We evaluated the ability of shallow shotgun metagenomic sequencing to characterize taxonomic and functional patterns in the fecal microbiome of a model population of feral horses (Sable Island, Canada). Since 2007, this unmanaged population has been the subject of an individual-based, long-term ecological study. Using deep shotgun metagenomic sequencing, we determined the sequencing depth required to accurately characterize the horse microbiome. In comparing conventional versus high-throughput shotgun metagenomic library preparation techniques, we validate the use of more cost-effective lab methods. Finally, we characterize similarities between 16S amplicon and shallow shotgun characterization of the microbiome and demonstrate that the latter recapitulates biological patterns first described in a published amplicon dataset. Unlike amplicon data, we further demonstrate how shallow shotgun metagenomic data provide useful insights about microbiome functional potential which support previously hypothesized diet effects in this study system.</p>
Shallow shotgun sequencing of the microbiome recapitulates 16S amplicon results and provides functional insights
Open the record for dataset details and reuse information.
Kraken analysis from shotgun sequenced museum specimens via krona plot visualization
<p>This dataset contains html files with krona plots from Kraken analysis on shotgun sequenced museum specimens.</p>
Data from: Microsatellite markers from the Ion Torrent: a multi-species contrast to 454 shotgun sequencing
The development and screening of microsatellite markers have been accelerated by next-generation sequencing (NGS) technology and in particular GS-FLX pyro-sequencing (454). More recent platforms such as the PGM semiconductor sequencer (Ion Torrent) offer potential benefits such as dramatic reductions in cost, but to date have not been well utilized. Here, we critically compare the advantages and disadvantages of microsatellite development using PGM semiconductor sequencing and GS-FLX pyro-sequencing for two gymnosperm (a conifer and a cycad) and one angiosperm species. We show that these NGS platforms differ in the quantity of returned sequence data, unique microsatellite data and primer design opportunities, mostly consistent with the differences in read length. The strength of the PGM lies in the large amount of data generated at a comparatively lower cost and time. The strength of GS-FLX lies in the return of longer average length sequences and therefore greater flexibility in producing markers with variable product length, due to longer flanking regions, which is ideal for capillary multiplexing. These differences need to be considered when choosing a NGS method for microsatellite discovery. However, the ongoing improvement in read lengths of the NGS platforms will reduce the disadvantage of the current short read lengths, particularly for the PGM platform, allowing greater flexibility in primer design coupled with the power of a larger number of sequences.
MetaChick: characterization of the chicken caecal metagenome by deep shotgun sequencing
<p></p><h1>Data sources</h1><br>This dataset was constructed using the samples of the MetaChick project (phase 1) corresponding to the cecal content of 340 animals. Sequencing data and associated metadata have been submitted to INSDC (bioproject: PRJEB38174).<br><h1>Sequencing data QC and metagenomic assembly</h1><br>First, sequencing adapters removal and read trimming was performed with fastxtend. Reads mapped on the host genome (GRCg7b GCA_016699485.1) with bowtie2 were removed with samtools. Finally, metagenomic assembly was performed with metaSPAdes v3.14.1. Contigs of less than 1500 bp were removed.<br><h1>MAGs recovery</h1><br>MAGs were generated with MetaBAT 2 (multi-coverage mode) and MAGs quality was assessed with CheckM. MAGs with completeness < 70% or contamination > 5% or N50 < 8Kb were discarded. Pairwise Average Nucleotide Identity (ANI) was computed for all recovered MAGs with fastANI and dereplication at species level (ANI cutoff = 95%).<br><h1>Non-redundant gene catalog</h1><br>Genes were predicted on all contigs from metagenomic assemblies with Prodigal (parameters : -m -p meta). Genes were pooled and clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0) by choosing those from the longest contigs as representatives.<br><h1>MSPs recovery</h1><br>A raw gene abundance table (13,6M genes quantified in 340 samples) was generated with meteorMeteor. Then, co-abundant genes were binned in Metagenomic Species Pan-genomes (MSPs, i.e. gene clusters that likely belong to the same microbial species) using MSPminer.<br><h1>MAGs and MSPs taxonomic annotation</h1><br>Dereplicated MAGs were annotated with GTDB-Tk based on GTDB r214. Then, MAGs taxonomic annotation was propagated to the corresponding MSPs.<br><h1>Construction of the phylogenetic tree</h1><br>39 universal phylogenetic markers genes were extracted from the dereplicated MAGs with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB38174 (cohort used in catalogue assembly) and PRJEB29033, PRJEB33338, PRJEB53667 (independent cohort not used in assembly).<p></p>
MicroReset: characterization of the rabbit (Oryctolagus cuniculus) fecal metagenome and resistome by deep shotgun sequencing
<p></p><h1>Data source</h1><br>The dataset was generated from 30 rabbit fecal samples subjected to deep shotgun metagenomic sequencing. The sequencing data is available under the BioProject PRJEB50625.<br>Metagenomic Assembly<br>Raw sequencing reads were first pre-processed using fastp for adapter removal and quality trimming. Host-derived reads were filtered out by mapping to the rabbit reference genome (GCF_000001635.27) using Bowtie2 and removing mapped reads with Samtools. Each sample was individually assembled using metaSPAdes. Contigs shorter than 1,500 bp were excluded from downstream analysis.<br><h1>MAG Recovery</h1><br>Reads from each sample were mapped to all 30 assemblies (30×30 mappings) using Bowtie2. The resulting alignments were sorted and indexed with Samtools. Contig coverage across all samples was computed using `jgi_summarize_bam_contig_depths`. Binning was performed with MetaBAT 2 and SemiBin v1.3. MAG quality was assessed with CheckM. Only high-quality MAGs (≥70% completeness, ≤5% contamination, N50 ≥ 8 kb) were retained.<br>Non-Redundant Gene Catalog<br>Gene prediction was carried out using Prodigal on all contigs from the current study (with `-m -p meta`). Genes shorter than 90 bp or lacking start/stop codons were discarded. The remaining genes from both sources were pooled and clustered using CD-HIT-EST (parameters: `-c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0`). The longest contigs were used to select representative genes.<br><h1>MSP Recovery</h1><br>Shotgun reads from the 30 samples were aligned to the non-redundant gene catalog using the Meteor suite, generating a gene abundance matrix (5.7 million genes × 30 samples). Co-abundant genes were grouped into 1,053 Metagenomic Species Pan-genomes (MSPs) using MSPminer.<br><h1>Taxonomic Annotation of MSPs</h1><br>MAGs representing each species were taxonomically annotated using GTDB-Tk with GTDB release r214. The resulting taxonomy was propagated to the corresponding MSPs.<br><h1>Phylogenetic Tree Construction</h1><br>A set of 39 universal phylogenetic marker genes was extracted from the 1,053 MSPs (or their corresponding MAGs, when available) using fetchMGs. Each marker was independently aligned using MUSCLE, and the alignments were concatenated and trimmed using trimAl (parameter: `-automated1`). A maximum-likelihood phylogenetic tree was constructed with FastTreeMP (parameters: `-gamma -pseudo -spr -mlacc 3 -slownni`).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters) for PRJEB50625 (cohort used in catalogue assembly).<p></p>
Comparison between 16S rRNA and shotgun sequencing in colorectal cancer, advanced colorectal lesions, and healthy human gut microbiota
<div> <p><span><span>Background</span></span><span><span>: Gut dysbiosis has been associated with colorectal cancer (CRC), the third most prevalent cancer in the world. </span><span>This study compares microbiota taxonomic and abundance results obtained by 16S rRNA gene sequencing (16S) and whole shotgun metagenomic sequencing to investigate their reliability for bacteria profiling. The experimental design included 156 human stool samples from healthy controls, advanced (high-risk) colorectal lesion patients (HRL), and CRC cases</span><span>, with each sample sequenced using both 16S and shotgun methods</span><span>. We thoroughly compared both sequencing technologies at the species, genus, and family annotation levels, the abundance differences in these taxa, sparsity, alpha and beta diversities, ability to train prediction models, and the similarity of the microbial signature derived from these models.</span></span><span> </span></p> </div> <div> <p><span><span>Results</span></span><span><span>: </span><span>As expected, the results showed that </span></span><span><span>16S detects only part of the gut microbiota community revealed by shotgun, although some genera were only profiled by 16S. The </span></span><span><span>16S </span><span>abundance data was sparser and </span><span>exhibited</span><span> lower alpha diversity. In lower taxonomic ranks, shotgun and 16S highly differed, </span><span>partially</span><span> due to a disagreement in reference databases. When considering only shared taxa, the abundance was positively correlated between the two strategies. We also found a moderate correlation between the shotgun and 16S alpha-diversity measures, as well as their </span><span>PCoAs</span><span>. </span><span>Regarding</span><span> the machine learning models, only some of the shotgun models showed some degree of predictive power in an independent test set, but we could not </span><span>demonstrate</span><span> a clear superiority of one technology over the other. Microbial signatures from both sequencing techniques reveal</span><span>ed</span><span> taxa previously associated with CRC development, e.g., </span></span><span><span>Parvimonas</span><span> micra</span></span><span><span>.</span></span><span> </span></p> </div> <div> <p><span><span>Conclusions</span></span><span><span>: </span></span><span><span>Shotgun and 16S sequencing provide two different lenses to examine microbial communities.</span><span> While we have </span><span>demonstrated</span><span> that they can unravel common patterns (including microbial signatures), </span><span>shotgun often gives</span><span> a more detailed snapshot than 16S, both in depth and breadth. </span><span>Instead</span><span>,</span> <span>16S will tend to show only part of the picture, giving greater weight to dominant bacteria in a sample.</span> <span>Therefore, w</span><span>e recommend choosing one or another sequencing technique before launching a study.</span><span> Specifically, </span><span>s</span><span>hotgun sequencing is preferred for stool microbiome samples and in-depth analyses, while 16S is </span><span>more </span><span>suitable for tissue samples</span><span> and</span><span> studies with </span><span>targeted</span> <span>aims</span><span>.</span></span><span> </span></p> </div>
Supplementary material 5 from: Sullivan JP, Hopkins CD, Pirro S, Peterson R, Chakona A, Mutizwa TI, Mukweze Mulelenu C, Alqahtani FH, Vreven E, Dillman CB (2022) Mitogenome recovered from a 19 th Century holotype by shotgun sequencing supplies a generic name for an orphaned clade of African weakly electric fishes (Osteoglossomorpha, Mormyridae). ZooKeys 1129: 163-196. https://doi.org/10.3897/zookeys.1129.90287
Measurements and counts of Heteromormyrus specimens and relevant mormyrid types from existing literature and taken from photographs & radiographs of newly sequenced individuals
Data from: Genome skimming by shotgun sequencing helps resolve the phylogeny of a pantropical tree family
Open the record for dataset details and reuse information.
Data from: Development and characterization of thirty-three microsatellite markers for the Patagonian sprat, Sprattus fuegensis (Jenyns, 1842), using paired-end Illumina shotgun sequencing
Open the record for dataset details and reuse information.
Data from: Microsatellite markers from the Ion Torrent: a multi-species contrast to 454 shotgun sequencing
Open the record for dataset details and reuse information.
Data from: Breakdown of phylogenetic signal: a survey of microsatellite densities in 454 shotgun sequences from 154 non model eukaryote species
Microsatellites are ubiquitous in Eukaryotic genomes. A more complete understanding of their origin and spread can be gained from a comparison of their distribution within a phylogenetic context. Although information for model species is accumulating rapidly, it is insufficient due to a lack of species depth, thus intragroup variation is necessarily ignored. As such, apparent differences between groups may be overinflated and generalizations cannot be inferred until an analysis of the variation that exists within groups has been conducted. In this study, we examined microsatellite coverage and motif patterns from 454 shotgun sequences of 154 Eukaryote species from eight distantly related phyla (Cnidaria, Arthropoda, Onychophora, Bryozoa, Mollusca, Echinodermata, Chordata and Streptophyta) to test if a consistent phylogenetic pattern emerges from the microsatellite composition of these species. It is clear from our results that data from model species provide incomplete information regarding the existing microsatellite variability within the Eukaryotes. A very strong heterogeneity of microsatellite composition was found within most phyla, classes and even orders. Autocorrelation analyses indicated that while microsatellite contents of species within clades more recent than 200 Mya tend to be similar, the autocorrelation breaks down and becomes negative or non-significant with increasing divergence time. Therefore, the age of the taxon seems to be a primary factor in degrading the phylogenetic pattern present among related groups. The most recent classes or orders of Chordates still retain the pattern of their common ancestor. However, within older groups, such as classes of Arthropods, the phylogenetic pattern has been scrambled by the long independent evolution of the lineages.
Data from: Characterization of microsatellite loci for the Gulf Coast waterdog (Necturus beyeri) using paired-end Illumina shotgun sequencing and cross-amplification in other Necturus
[No abstract filled]
Supplementary material 3 from: Sullivan JP, Hopkins CD, Pirro S, Peterson R, Chakona A, Mutizwa TI, Mukweze Mulelenu C, Alqahtani FH, Vreven E, Dillman CB (2022) Mitogenome recovered from a 19 th Century holotype by shotgun sequencing supplies a generic name for an orphaned clade of African weakly electric fishes (Osteoglossomorpha, Mormyridae). ZooKeys 1129: 163-196. https://doi.org/10.3897/zookeys.1129.90287
Cyt b plus nuclear markers phylogenetic analysis
Supplementary material 2 from: Sullivan JP, Hopkins CD, Pirro S, Peterson R, Chakona A, Mutizwa TI, Mukweze Mulelenu C, Alqahtani FH, Vreven E, Dillman CB (2022) Mitogenome recovered from a 19 th Century holotype by shotgun sequencing supplies a generic name for an orphaned clade of African weakly electric fishes (Osteoglossomorpha, Mormyridae). ZooKeys 1129: 163-196. https://doi.org/10.3897/zookeys.1129.90287
Heteromormyrus pauciradiatus holotype NMW 22417 mitogenome reconstruction
Supplementary material 1 from: Sullivan JP, Hopkins CD, Pirro S, Peterson R, Chakona A, Mutizwa TI, Mukweze Mulelenu C, Alqahtani FH, Vreven E, Dillman CB (2022) Mitogenome recovered from a 19 th Century holotype by shotgun sequencing supplies a generic name for an orphaned clade of African weakly electric fishes (Osteoglossomorpha, Mormyridae). ZooKeys 1129: 163-196. https://doi.org/10.3897/zookeys.1129.90287
List of Heteromormyrus specimens examined and/or included in molecular analyses
Supplementary material 4 from: Sullivan JP, Hopkins CD, Pirro S, Peterson R, Chakona A, Mutizwa TI, Mukweze Mulelenu C, Alqahtani FH, Vreven E, Dillman CB (2022) Mitogenome recovered from a 19 th Century holotype by shotgun sequencing supplies a generic name for an orphaned clade of African weakly electric fishes (Osteoglossomorpha, Mormyridae). ZooKeys 1129: 163-196. https://doi.org/10.3897/zookeys.1129.90287
COI phylogenetic analysis
Data from: Characterization of microsatellite loci for the Gulf Coast waterdog (Necturus beyeri) using paired-end Illumina shotgun sequencing and cross-amplification in other Necturus
Open the record for dataset details and reuse information.
Data from: Sequencing degraded DNA from non-destructively sampled museum specimens for RAD-tagging and low-coverage shotgun phylogenetics
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.