Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

40

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

40 results for “16S sequencing data”

Learn how ShareScore rates datasets ↗
zenodo44/100

16S rRNA sequencing gene datasets for CRC data

<p>Used datasets:&nbsp;</p> <table> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>Dataset</th> <th>16S rRNA Region</th> <th>Control (n)</th> <th>Adenoma (n)</th> <th>CRC (n)</th> <th>Available metadata</th> </tr> </thead> <tbody> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4823848/">Baxter</a></td> <td>V4</td> <td>171</td> <td>198</td> <td>120</td> <td>Gender, age, weight, height, BMI, country, race</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4221363/">Zackular</a></td> <td>V4</td> <td>30</td> <td>30</td> <td>30</td> <td>Gender, age, weight, height, BMI, country, race, FOBT, medication</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4299606/">Zeller</a></td> <td>V4</td> <td>50</td> <td>38</td> <td>41</td> <td>Gender, age, BMI, country, FOBT</td> </tr> <tr> <td><strong>TOTAL</strong></td> <td>V4</td> <td>251</td> <td>266</td> <td>191</td> <td><em>All of the above</em></td> </tr> </tbody> </table> </th> </tr> </thead> <tbody> <tr> <td>&nbsp;</td> </tr> </tbody> </table> <p>Data processing &amp; sharing</p> <p>All datasets were processed using&nbsp;<a href="https://docs.qiime2.org/2021.11/">qiime2</a>&nbsp;pipeline with&nbsp;<a href="https://benjjneb.github.io/dada2/">DADA2</a>&nbsp;for Sequence quality control and feature table construction and&nbsp;<a href="https://www.arb-silva.de/">SILVA</a>&nbsp;database for taxonomic assignment, and then a <em>phyloseq </em>object was constructed.</p> <ul> <li>Abundance table at genus level is in file <em>genus.csv</em> (Sample counts with NO filtering).</li> <li>Clean metadata is in <em>metadata.csv</em> file (Countries: CA - Canada. USA - United States of America. FRA - France.)</li> <li>Phyloseq object is in file <em>physeq.RDS</em> (Saved as an RDS object in R)</li> </ul> <p>More information is&nbsp;<a href="https://hackmd.io/nbsLqCLlSNSRFc5RBX9c5Q?view">here</a>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

16S rRNA sequences data set of Dicronocephalus species in this study.

Explanation note: This 16S rRNA data includes 46 individual sequences of the examined Dicronocephalus species in this study.

opencc-by-4.0Feb 2017View details →
zenodo40/100

16S rRNA Sequencing Data of Fecal Microbiota in an Italian Cohort of Patients with CDKL5 Deficiency Disorder

<h3>Summary of the study&nbsp;</h3> <p>CDKL5 deficiency disorder (CDD) is a neurodevelopmental condition characterized by global developmental delay, early-onset seizures, intellectual disability, visual and motor impairments, distinct from Rett Syndrome (RTT) due to the absence of a clear regression period. Gastrointestinal (GI) disturbances and signs of subclinical immune dysregulation are common in CDD patients, yet the underlying causes are unknown. Recent studies hint at a possible link between neurological disorders and gut microbiota, an unexplored area in CDD. In this groundbreaking study, we examined fecal microbiota in CDD patients and their healthy relatives, revealing differences in bacterial diversity and composition. We further investigated microbiota changes based on various factors, including the severity of GI issues, seizure frequency, sleep disorders, food intake type, neuro-behavioral features (assessed through the RTT Behaviour Questionnaire &ndash; RSBQ), and ambulation capacity.&nbsp;</p> <p>Our findings suggest a potential connection between CDD, microbiota, and symptom severity. This study represents the first exploration of the gut-microbiota-brain axis in CDD patients, contributing to the growing body of research on the role of gut microbiota in neurodevelopmental disorders. It opens doors to potential interventions targeting intestinal microbes to enhance the well-being of individuals with CDD.</p> <h3>Mehods</h3> <p>The Dataset represent the raw data (.fastq) obtained from the sequencing of the fecal samples from 17 Italian Patients with CDD, and 17 Healthy Relatives (i.e. siblings or mother), collected at a single time-point.</p> <p>Samples from Patients affected by CDD are called CDD, samples from Healthy Relatives are called HC-CDD (i.e. healthy controls of patients affected by CDD). For details about the sample names see the &ldquo;Explanation Table&rdquo;.</p> <p>Bacterial DNA was extracted from fecal samples using the QIAmp Powerfexal DNA Kit (Qiagen, Germany) following the manufacturer's protocol. The 16S rRNA sequencing and analysis was performed by a service offered by Zymo Research (Germany).</p> <p><em>Targeted Library Preparation</em>: The DNA samples were prepared for targeted sequencing with the Quick-16S&trade; NGS Library Prep Kit (Zymo Research). The primer sets used were Quick-16S&trade; Primer Set V3-V4 (Zymo Research). The sequencing library was prepared using an innovative library preparation process in which PCR reactions were performed in real-time PCR machines to control cycles and therefore limit PCR chimera formation. The final PCR products were quantified with qPCR fluorescence readings and pooled together based on equal molarity. The final pooled library was cleaned up with the Select-a-Size DNA Clean &amp; Concentrator&trade;, then quantified with TapeStation&reg; (Agilent Technologies, Santa Clara, CA) and Qubit&reg; (Thermo Fisher Scientific, Waltham, WA).&nbsp;&nbsp;</p> <p><em>Sequencing:</em> The final library was sequenced on Illumina&reg; MiSeq&trade; with a v3 reagent kit (600 cycles).&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

OTU level 16s sequence data for "Algae drive convergent bacterial community assembly when nutrients are scarce"

<p>16s sequence data at the OTU level for the experiments conducted in&nbsp;&quot;Algae drive convergent bacterial community assembly when nutrients are scarce&quot;</p> <p>The file is in fasta format, which can be read by many software packages including biopython, R, and SILVA&#39;s alignment, classification and tree service.</p> <p>The sequence ids can be used to find the phylogeny and OTU ID of the sequences from Supplementary dataset 4.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

A Comprehensive Assessment of Demographic, Environmental and Host Genetic Associations with Gut Microbiome Diversity in Healthy Individuals (16S rRNA gene sequencing data)

<p>Microbiome data accompanying manuscript &quot;A Comprehensive Assessment of Demographic, Environmental and Host Genetic Associations with Gut Microbiome Diversity in Healthy Individuals&quot;. Data is available for alpha- and beta- diversity, as well as&nbsp;for individual taxa both in binary and quantitative&nbsp;phenotypic representation.&nbsp;Data is available for 827 individuals that gave consent for their data to be shared outside of the Milieu int&eacute;rieur consortium.&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S in A herpetological survey of western Zambia

Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S rRNA of Phrynobatrachus natalensis. Numbers above branches are non-parametric bootstrap support values. Specimen vouchers or GenBank accession numbers are shown in parentheses. Colored polygons highlight the clades comprising specimens from this study. (*) Nearest sample from type locality of Phrynobatrachus natalensis; (**) Haplotype groups A and B in Zimkus and Schick (2010).

opencc-by-4.0Aug 2019View details →
zenodo36/100

Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon" - PART I

<p>Raw Fast5 data for &quot;Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon&quot;. See Supplementary Table 2 for associating each sample to its barcode.</p> <p>- FC1_1 includes data for the HM mock community from BEI resources and skin microbiome of the chin in dogs.</p> <p>- FC1_2 includes data for the dorsal skin samples</p> <p>- FC2 includes data for the Zymobiomics mock community&nbsp;and Staphylococcus pseudintermedius isolate</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

16S rRNA Sequence Data, Brazilian Coffee Soils

<p>16S sequencing data for DNA extracted from soils from Brazillian coffee farms. Sequenced on Illumina MiSeq with primers from Caporaso (2011, 2012).</p>

opencc-zeroAug 2014View details →
dryad36/100

16S rDNA sequencing data for characterizing the endosymbionts of leaf curl plum aphid ( Brachycaudus helichrysi) clones

<p>Asexual lineages often exhibit broad distributions and can thrive in extreme habitats compared to their sexual counterparts. Several hypotheses can be proposed to explain this pattern. Asexual lineages could be versatile genotypes with wide environmental tolerance, enabling their dispersal and persistence across large geographic areas. Alternatively, asexual genotypes could be ecological specialists that thrive in specific environments and outcompete relatives colonizing distantly related areas with similar conditions in the process. Several aphid species feature widespread obligate asexual lineages, commonly known as "superclones". Yet it is often unknown whether these clones are widespead ecological generalists or  successful specialists. To explore these hypotheses, we examined climatic niche differentiation among six globally distributed obligate asexual lineages of the cosmopolitan aphid pest, <em>Brachycaudus helichrysi</em>. To insure that we were investigating the aphid genotype niche and not a by-product of their association with endosymbionts mediating thermal tolerance, we first verified that clones hosted similar endosymbiont communities. Subsequently, we conducted multivariate analyses on clone occurrence data on a worldwide scale. Our results revealed that despite their global distribution, <em>B. helichrysi</em> superclones occupy different climatic niches. This study represents the first evidence that aphid superclones distribution can be mediated by distinctive ranges of climatic tolerance.</p>

opencc-zeroApr 2024View details →
dryad36/100

16S rRNA sequencing data: Altered microbiota, impaired quality of life, malabsorption, infection, and inflammation in CVID patients with diarrhoea

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

16S rRNA gene sequencing data from: Breastmilk IgG engages the neonatal immune system to instruct immune responses to gut antigens

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad36/100

16S rDNA sequencing data for characterizing the endosymbionts of leaf curl plum aphid ( Brachycaudus helichrysi) clones

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad32/100

Data from: 16S rRNA amplicon sequencing for epidemiological surveys of bacteria in wildlife

The human impact on natural habitats is increasing the complexity of human-wildlife interactions and leading to the emergence of infectious diseases worldwide. Highly successful synanthropic wildlife species, such as rodents, will undoubtedly play an increasingly important role in transmitting zoonotic diseases. We investigated the potential for recent developments in 16S rRNA amplicon sequencing to facilitate the multiplexing of the large numbers of samples needed to improve our understanding of the risk of zoonotic disease transmission posed by urban rodents in West Africa. In addition to listing pathogenic bacteria in wild populations, as in other high-throughput sequencing (HTS) studies, our approach can estimate essential parameters for studies of zoonotic risk, such as prevalence and patterns of coinfection within individual hosts. However, the estimation of these parameters requires cleaning of the raw data to mitigate the biases generated by HTS methods. We present here an extensive review of these biases and of their consequences, and we propose a comprehensive trimming strategy for managing these biases. We demonstrated the application of this strategy using 711 commensal rodents, including 208 Mus musculus domesticus, 189 Rattus rattus, 93 Mastomys natalensis, and 221 Mastomys erythroleucus, collected from 24 villages in Senegal. Seven major genera of pathogenic bacteria were detected in their spleens: Borrelia, Bartonella, Mycoplasma, Ehrlichia, Rickettsia, Streptobacillus, and Orientia. Mycoplasma, Ehrlichia, Rickettsia, Streptobacillus, and Orientia have never before been detected in West African rodents. Bacterial prevalence ranged from 0% to 90% of individuals per site, depending on the bacterial taxon, rodent species, and site considered, and 26% of rodents displayed coinfection. The 16S rRNA amplicon sequencing strategy presented here has the advantage over other molecular surveillance tools of dealing with a large spectrum of bacterial pathogens without requiring assumptions about their presence in the samples. This approach is therefore particularly suitable to continuous pathogen surveillance in the context of disease-monitoring programs.

opencc-zeroDec 2015View details →
zenodo32/100

16S rDNA sequence data of research--Different periods of high-fat diet during peri-pregnancy cause variant offspring gut microbiota

<p>&nbsp; These are the 16S rRNA gene sequencing data based on&nbsp;the Illumina HiSeq2500 platform. Group a, b, c, d stand for group CD-CD, CD-HFD, HFD-CD, HFD-HFD separately. Each mouse of one group&nbsp;had two samples such as a1_1 and a1_2.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

16S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory

<p><strong>Metabarcoding - 16S ASV generation and taxonomic assignment</strong>. Raw 16S paired-end sequences were subjected to a data quality control step and subsequently imported into the QIIME2 pipeline v.2022.2.0. Leftover primers and adapter sequences were removed through cutadapt. The amplicon sequence variants (ASV) table, which represent true biological sequences within each sample, was generated using the denoised-paired method including truncation, denoising, dereplication, merging, and chimera filtering of the DADA2 (Divisive Amplicon Denoising Algorithm 2) plugin inside QIIME2. Default parameters were used with the exception of the forward and reverse sequence length (--p-trunc-len-f and --p-trunc-len-r), that were set to 220 and 180, respectively. Processed reads that passed all these filters were used for taxonomy classification. The V4-V5 regions were extracted from the pre-formatted reference sequences and taxonomy file built on the SILVA 138 99% OTUs database and the vsearch v.2.6.2 global alignment implemented in QIIME2 was used.</p>

opencc-by-4.0Jul 2024View details →
dryad32/100

Data from: Rapid species-level identification of vaginal and oral lactobacilli using MALDI-TOF MS analysis and 16S rDNA sequencing

Background: Lactobacillus represents a large genus with different implications for the human host. Specific lactobacilli are considered to maintain vaginal health and to protect from urogenital infection. The presence of Lactobacillus species in carious lesions on the other hand is associated with progressive caries. Despite their clinical significance, species-level identification of lactobacilli still poses difficulties and mostly involves a combination of different phenotypic and genotypic methods. This study evaluated rapid MALDI-TOF MS analysis of vaginal and oral Lactobacillus isolates in comparison to 16S rDNA analysis. Results: Both methods were used to analyze 77 vaginal and 21 oral Lactobacillus isolates. The concordance of both methods was at 96% with five samples discordantly identified. Fifteen different Lactobacillus species were found in the vaginal samples, primarily L. iners, L. crispatus, L. jensenii and L. gasseri. In the oral samples 11 different species were identified, mostly L. salivarius, L. gasseri, L. rhamnosus and L. paracasei. Overall, the species found belonged to six different phylogenetic groups. For several samples, MALDI-TOF MS analysis only yielded scores indicating genus-level identification. However, in most cases the species found agreed with the 16S rDNA analysis result. Conclusion: MALDI-TOF MS analysis proved to be a reliable and fast tool to identify lactobacilli to the species level. Even though some results were ambiguous while 16S rDNA sequencing yielded confident species identification, accuracy can be improved by extending reference databases. Thus, mass spectra analysis provides a suitable method to facilitate monitoring clinically relevant Lactobacillus species.

opencc-zeroDec 2014View details →
dryad32/100

16S V4 raw read count data; 16S reads metadata; new MHC class II allele sequences

<p>Pathogen-mediated selection at the major histocompatibility complex (MHC) is thought to promote MHC-based mate choice in vertebrates. Mounting evidence implicates odour in conveying MHC genotype, but the underlying mechanisms remain uncertain. MHC effects on odour may be mediated by odour-producing symbiotic microbes whose community structure is shaped by MHC genotype. In birds, preen oil is the primary source of body odour and similarity at MHC predicts similarity in preen oil composition. Hypothesizing that this relationship is mediated by symbiotic microbes, we characterized MHC genotype, preen gland microbial communities, and preen oil chemistry of song sparrows (<i>Melospiza melodia</i>). Consistent with the microbial mediation hypothesis, pairwise similarity at MHC predicted similarity in preen gland microbiota. Overall microbial similarity did not predict chemical similarity of preen oil, counter to this hypothesis. However, permutation testing identified a maximally predictive set of microbial taxa that best reflect MHC genotype, and another set of taxa that best predict preen oil chemical composition. The relative strengths of relationships between MHC and microbes, microbes and preen oil, and MHC and preen oil suggest that MHC may affect host odour both directly and indirectly. Thus, birds may assess MHC genotypes based on both host-associated and microbially-mediated odours.</p>

opencc-zeroOct 2021View details →
dryad32/100

16S V4 raw read count data; 16S reads metadata; new MHC class II allele sequences

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad32/100

Data from: A comparison between transcriptome sequencing and 16S metagenomics for detection of bacterial pathogens in wildlife

Open the record for dataset details and reuse information.

publicAug 2015View details →
dryad32/100

Data from: Rapid species-level identification of vaginal and oral lactobacilli using MALDI-TOF MS analysis and 16S rDNA sequencing

Open the record for dataset details and reuse information.

publicNov 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record