Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

130

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

130 results for “Human microbiome”

Learn how ShareScore rates datasets ↗
zenodo52/100

Full-length and split homologs of human proteins in the gut microbiome

<p>These files were generated as part of the manuscript "Human xenobiotic metabolism proteins have full-length and split homologs in the gut microbiome" (submitted).</p> <p>The .tar file contains .ipc files that are tables of full-length (full_humcover3.ipc) and split homologs (part_humcover3.ipc) of human proteins in the gut microbiome, organized by alignment coverage threshold. For example, the directory `HumanUPR_0.67_src_20000_70` contains results obtained at a 67% alignment coverage threshold for the bacterial protein, and 70% for the human protein. Note that our pipeline collapses full-length alignments to the same UHGP-90 protein family into a single entry per species, with the number of genomes reported in the column nGenomes. Split homologs are not collapsed because genomic context is used to define them, and this context may differ across individual genomes.</p> <p>These files are in Arrow <a href="https://arrow.apache.org/docs/python/ipc.html#ipc">IPC</a> format, which provides compression and fast I/O for large tables. We recommend reading them using <a href="https://pola.rs/">pola.rs</a> or the <a href="https://arrow.apache.org/docs/r/">R Arrow</a> package. In particular, because the full-length homolog table is large, you may wish to work with it without loading it into memory, which can be accomplished using&nbsp;<a href="https://docs.pola.rs/api/python/dev/reference/api/polars.scan_ipc.html">scan_ipc</a> in pola.rs or <a href="https://arrow.apache.org/docs/r/reference/open_dataset.html">open_dataset</a> in R Arrow.</p> <p>We also provide gzipped .csv format datasets of full-length (pgkb_FH_drugs.csv.gz) and split (pgkb_SH_drugs.csv.gz) homologs, at the default 67% alignment coverage threshold for bacterial and 70% for human proteins, organized by their&nbsp;<a href="https://www.pharmgkb.org/">PharmGKB</a> annotations. For each drug annotated in PharmGKB as being metabolized by a human protein with full-length or split homologs, we provide the human protein(s) responsible, its xenobiotic enzyme class, the bacterial protein homolog(s), length and percent identity of the alignment, and either the specific genome (g, split homologs only) or the number of genomes (nGenomes, full homologs only). Xenobiotic enzyme classes are defined as in Figure 4 of the manuscript, with the additional classes "nucl" (nucleobase-containing metabolic proteins not annotated to any other class), "redox" (oxidoreductases not annotated to any other class), and "other" (all remaining proteins).</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Human Microbiome Compendium dataset

<p>The Human Microbiome Compendium is an ongoing project to build a large collection of human microbiome sequencing data processed with a uniform pipeline. Currently, the compendium contains 16S rRNA amplicon sequencing data for human gut microbiome samples retrieved from the Sequence Read Archive. Our&nbsp;website at <strong><a href="https://microbiomap.org">microbiomap.org</a></strong> has more information about the project and links to related resources.</p> <p>This data is freely available under the&nbsp;<a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license</a> (<strong>CC BY 4.0</strong>). If you use it in your work, please cite our publication:</p> <p>Abdill, Richard J., Samantha P. Graham, Vincent Rubinetti, et al. &ldquo;Integration of 168,000 Samples Reveals Global Patterns of the Human Gut Microbiome.&rdquo; Cell 188, no. 4 (2025): 1100&ndash;18. <a href="https://doi.org/10.1016/j.cell.2024.12.017" target="_blank" rel="noopener">https://doi.org/10.1016/j.cell.2024.12.017</a></p> <p>If you are using this dataset in combination with your own results, it's important to note that the taxonomic classifications may differ between releases, as documented in CHANGELOG.md. The most recent release (1.1.1) includes assignments made using&nbsp;<strong><a href="https://www.arb-silva.de/news/view/2024/07/11/silva-release-1382/">SILVA 138.2</a></strong> (SSU Ref NR 99) and <strong>Greengenes2</strong> (2022.10 backbone).</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Upcycling Human Excrement: The Gut Microbiome to Soil Microbiome Axis (supporting data)

<div> <div>This archive contains the supporting data and code for <a href="https://doi.org/10.1093/ismeco/ycaf089" target="_blank" rel="noopener">Meilander et al., 2024:&nbsp;<em>Upcycling Human Excrement: The Gut Microbiome to Soil Microbiome Axis</em></a>.</div> <div>&nbsp;</div> <div><strong>Clicking the links below will open the corresponding files using QIIME 2 View (<a href="https://view.qiime2.org" target="_blank" rel="noopener">https://view.qiime2.org</a>).&nbsp;</strong></div> <div>&nbsp;</div> <div> <div> <div>Summaries of master data files:</div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/asv-table.qzv/content" target="_blank" rel="noopener">Summary of master feature table (<code>asv-table.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/sample-metadata.qzv/content">Tabulated view of sample metadata (<code>sample-metadata.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/records/15390940/files/asv-seqs-ms10.qzv?download=1" target="_blank" rel="noopener">Summary of ASV sequences observed in at least 10 samples: (<code>asv-seqs-ms10.qzv</code>)</a></div> <div>&nbsp;</div> <div>PCoA plots:</div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/braycurtis.qzv/content" target="_blank" rel="noopener">Bray-Curtis Emperor plot (<code>braycurtis.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/jaccard.qzv/content" target="_blank" rel="noopener">Jaccard Emperor plot (<code>jaccard.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/unweighted_unifrac.qzv/content" target="_blank" rel="noopener">Unweighted UniFrac Emperor plot (<code>unweighted_unifrac.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/weighted_unifrac.qzv/content" target="_blank" rel="noopener">Weighted UniFrac Emperor plot (<code>weighted_unifrac.qzv</code>)</a></div> <div>&nbsp;</div> <div>Taxonomy barplots:</div> <div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/records/15390940/files/taxa-bar-plots-bucket2-gtdb-r214.1-weighted-stool-taxonomy.qzv?download=1">Taxonomy bar plot for Bucket 2 only (<code>taxa-bar-plots-bucket2-gtdb-r214.1-weighted-stool-taxonomy.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/records/15390940/files/taxa-bar-plots-bucket3-gtdb-r214.1-weighted-stool-taxonomy.qzv?download=1" target="_blank" rel="noopener">Taxonomy bar plot for Bucket 3 only (<code>taxa-bar-plots-bucket3-gtdb-r214.1-weighted-stool-taxonomy.qzv</code>)</a></div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/taxa-bar-plots-gtdb-r214.1-weighted-stool-taxonomy.qzv/content" target="_blank" rel="noopener">Taxonomy bar plot for all samples (<code>taxa-bar-plots-gtdb-r214.1-weighted-stool-taxonomy.qzv</code></a>)</div> </div> <div>&nbsp;</div> <div>q2-fmt "raincloud plots":</div> <div><a href="https://view.qiime2.org/visualization/?src=https://zenodo.org/api/records/13887457/files/hec-raincloud.qzv/content">Raincloud plot (<code>hec-raincloud.qzv</code>)</a></div> <div>&nbsp;</div> </div> </div> <div>&nbsp;</div> <div>The linked <code>.qzv</code> files are also contained in the <code>gut-to-soil-qiime2.zip</code> zip file, along with all relevant data artifacts (<code>.qza</code> files).</div> <div>The <code>.qzv</code> files are also maintained outside of the <code>.zip</code> file to facilitate their viewing with QIIME 2 View.</div> </div> <div>&nbsp;</div> <div> <div>Code for generating figures 1 and 2 (and corresponding supplemental figures):</div> <div><code>gut-to-soil-manuscript-figures-main.zip</code> (also see: <a href="https://github.com/caporaso-lab/gut-to-soil-manuscript-figures" target="_blank" rel="noopener">https://github.com/caporaso-lab/gut-to-soil-manuscript-figures</a>)</div> <div>&nbsp;</div> <div>Code for generating ridgeline plots (Figure S6):</div> <div><code>gut-to-soil-ridgeline-plots-main.zip</code> (also see: <a href="https://github.com/caporaso-lab/gut-to-soil-ridgeline-plots" target="_blank" rel="noopener">https://github.com/caporaso-lab/gut-to-soil-ridgeline-plots</a>)</div> </div> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Data from: Public human microbiome data are dominated by highly developed countries

<p>Supplementary tables and datasets associated with the <em>PLOS Biology </em>publication &quot;Public human microbiome data are dominated by highly developed countries.&quot; See README.txt for a description of the files and the data fields they contain.</p> <p>Update, 8 Apr 2022: The &quot;figures.md&quot; file described in the readme was inadvertently left out of the initial upload. It has been added here.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Metagenomic assembly and bin3C clustering result for a healthy human faecal microbiome transplant donor

<p>Metagenomic WGS assembly and Hi-C deconvolution&nbsp;of a healthy human faecal microbiome transplant donor.</p> <p>Metagenomic assembly was produced using Spades (v3.13.1).</p> <p>Extracted MAGs were produced using bin3C&nbsp;(v0.3.3) and QC&#39;d using CheckM&nbsp;(v1.0.18).</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

Putative mobilized colistin resistance (mcr) genes co-occurring with other antibiotic resistance genes are widespread in the human gut microbiome

<p><strong>The dataset from the article&nbsp;</strong><strong>Putative mobilized colistin resistance (mcr) genes co-occurring with other antibiotic resistance genes are widespread in the human gut microbiome</strong></p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Decoding host-microbiome interactions through co-expression network analysis within the non-human primate intestine

<p>Supplementary Table&nbsp;Captions:</p> <p>Supplementary Table S9. Evaluation and parameter determination of host and microbiome RNA read classification using simulation datasets</p> <p>Supplementary Table S10. 40 pathways significantly upregulated in the cecum as compared to the transverse colon</p> <p>Supplementary Table S11. Host-microbiome gene co-expression network edges</p> <p>Supplementary Table S12. Host-host gene co-expression network edges</p> <p>Supplementary Table S13. Microbiome-microbiome gene co-expression network edges</p> <p>Supplementary Table S14. List of genes included in each gene module identified from the gene co-expression network</p> <p>Supplementary Table S15. Results of enrichment analysis for each gene module identified from the gene co-expression network</p> <p>Supplementary Table S16. The top 32 bacterial species in terms of expression abundance based on metatranscriptome profiles</p> <p>Supplementary Table S17. Number of microbiome RNA reads annotated by the KEGG database</p> <p>Supplementary Table S18. Results of enrichment analysis of gene modules for each parameter</p> <p>Supplementary Table S19. Evaluation of modules in each parameter of Newman algorithm</p> <p>Supplementary Table S20. Evaluation of modules in each parameter of Louvain algorithm</p> <p>Supplementary Table S21. Evaluation of modules in each parameter of Leiden algorithm</p> <p>Supplementary Table S22. Evaluation of modules in each parameter of WGCNA</p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Data for: Community assembly of the human piercing microbiome

<p>Predicting how biological communities respond to disturbance requires understanding the forces that govern their assembly. We propose using human skin piercings as a model system for studying community assembly after rapid environmental change. Local skin sterilization provides a 'clean slate' within the novel ecological niche created by the piercing. Stochastic assembly processes can dominate skin microbiomes due to the influence of environmental exposure on local dispersal, but deterministic processes might play a greater role within occluded skin piercings if piercing habitats impose strong selection pressures on colonizing species. Here we explore the human ear-piercing microbiome and demonstrate that community assembly is predominantly stochastic but becomes significantly more deterministic with time, producing increasingly diverse and ecologically complex communities. We also observed changes in two dominant and medically relevant antagonists (<em>Cutibacterium acnes </em>and <em>Staphylococcus epidermidis</em>), consistent with competitive exclusion induced by a transition from sebaceous to moist environments. By exploiting this common yet uniquely human practice, we show that skin piercings are not just culturally significant but also represent ecosystem engineering on the human body. The novel habitats and communities that skin piercing produce may provide general insights into biological responses to environmental disturbance with implications for both ecosystem and human health.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Dog gut gene catalog. Supplemental data for "Similarity of the dog and human gut microbiomes in gene content and response to diet"

<p>Gene catalogue for the dog gut microbiome including</p> <ol> <li>FASTA file of nucleotide sequences (including padding, see coords file for exact coordinates)</li> <li>FASTA file of amino-acid sequences</li> <li>coords file (gene coordinates)</li> <li>Taxonomic predictions</li> <li>Functional predictions</li> </ol> <p>See the paper &quot;<em>Similarity of the dog and human gut microbiomes in gene content and response to diet</em>&quot; by Coelho et al. in Microbiome for details. We ask that you cite that publication when using this dataset in published literature</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

bin3C - Cluster report and CheckM result for bin3C solution of the real human gut microbiome

<p>Supplementary data table&nbsp;S3 from the manuscript</p> <p>bin3C : Exploiting Hi-C sequencing data to accurately resolve metagenome-assembled genomes (MAGs)</p> <p>A real human gut microbiome was deconvoluted using bin3C. Subsequently, bin3C produced a report detailing per-cluster statistics for the entire solution. The largest 296 clusters were then analyzed with CheckM and joined to the report.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics"

<p>NGS Data Accompanying &quot;Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics&quot;, currently in review.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

MAGs and gapseq models for auxotrophy predictions in the human gut microbiome

<p>This dataset contains MAGs, their DNA sequence, genome statistics, quantification per sample, and their metabolic model reconstructions from two human population cohorts from northern Germany.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Enterosignatures define common bacterial guilds in the human gut microbiome (data)

<p>Additional data related to the manuscript <em>Enterosignatures define common bacterial guilds in the human gut microbiome.</em></p> <p>The deposit contains the following files:</p> <ul> <li>GMR_dataset.zip - data and results related to enterosignature computation in the GMR dataset: raw data, enterosignature composition, metabolic potentials of the associated metagenomic species</li> </ul> <p><strong>GMR_dataset</strong><br> ├── NMF_results <em># results of NMF decomposition, from k = 2 signatures to 10, 5 being the optimal number</em><br> │&nbsp;&nbsp; ├── 10_H.tsv<br> │&nbsp;&nbsp; ├── ....<br> │&nbsp;&nbsp; ├── 9_W.tsv<br> │&nbsp;&nbsp; └── README.txt<br> ├── age_metadata.tsv <em># age of individuals in the GMR dataset</em><br> ├── drama_MGS_gtdbtk_taxonomy.tsv <em># taxonomy of the MGS </em><br> ├── drama_genus_level_abundance.tsv <em># genus-level abundance table used for NMF</em><br> └── metabo_data <em># Metabolic potential of the MGS associated to the enterosignatures</em><br> &nbsp;&nbsp;&nbsp; ├── cazymes_by_genome.tsv <em># cazymes</em><br> &nbsp;&nbsp;&nbsp; ├── gene_numbers.tsv <em># number of genes</em><br> &nbsp;&nbsp;&nbsp; ├── kegg_metabolic_processes_by_genome.tsv <em># Kegg annotations</em><br> &nbsp;&nbsp;&nbsp; ├── level1_onto_metabolites.tsv <em># Metacyc classes of metabolites, highest level</em><br> &nbsp;&nbsp;&nbsp; ├── level2_onto_metabolites.tsv <em># Metacyc classes of metabolites, second level</em><br> &nbsp;&nbsp;&nbsp; ├── metabolic_producers_full_community.tsv <em># predicted producers of metabolites, all genomes considered</em><br> &nbsp;&nbsp;&nbsp; ├── metabolic_producers_withinES.tsv <em># predicted producers of metabolites, within an ES</em><br> &nbsp;&nbsp;&nbsp; └── westerndiet.sbml <em># nutrients considered for metabolic modelling</em></p> <ul> <li>BMIS_dataset.zip - data and results related to enterosignature computation in the BMIS dataset</li> </ul> <p><strong>BMIS_dataset</strong><br> ├── bmis_es_composition.tsv <em># ES assignments of BMIS samples (ES computed on the GMR dataset)</em><br> ├── bmis_es_et_assignments.tsv<em> # ES and enterotype assignments of BMIS samples</em><br> └── bmis_genus-level_abundance.tsv&nbsp; <em># genus-level abundance matrix used for ES abundance computation</em></p> <ul> <li>NonWestern_dataset.zip - data and results related to enterosignature computation in the non-western dataset</li> </ul> <p><strong>NonWestern_dataset</strong><br> ├── nonwestern_es_composition.tsv <em># ES assignments of Non-western samples (ES computed on the GMR dataset)</em><br> └── nonwestern_genus_abundance_normalised.tsv <em># genus-level abundance matrix used for ES abundance computation</em></p> <p>The data is also available in the following repository:<em> </em><a href="https://gitlab.inria.fr/cfrioux/enterosignature-paper/">https://gitlab.inria.fr/cfrioux/enterosignature-paper/</a>.</p> <p>See also <a href="https://enterosignatures.quadram.ac.uk/">https://enterosignatures.quadram.ac.uk/</a>.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Human Genetics Influences Microbiome Composition Involved in Asthma Exacerbations despite Inhaled Corticosteroid Treatment

<p>Additional Supplement Data of the manuscript&nbsp;Perez-Garcia J, Espuela-Ortiz A, Hern&aacute;ndez-P&eacute;rez JM, et al. Human Genetics Influences Microbiome Composition Involved in Asthma Exacerbations despite Inhaled Corticosteroid Treatment (in press).&nbsp;<em>J Allergy Clin Immunol</em>. 2023;S0091-6749(23)00748-0. doi:10.1016/j.jaci.2023.05.021.</p> <p>This Supplement Data contains the following items:</p> <table> <tbody> <tr> <td>Additional Table 1</td> <td>List of independent mbQTLs associated with p&lt;1x10<sup>-5</sup> in the mbGWAS from saliva samples in Europeans (discovery).</td> </tr> <tr> <td>Additional Table 2</td> <td>List of independent mbQTLs associated with p&lt;1x10<sup>-5</sup> in the mbGWAS from pharyngeal samples in Europeans (discovery).</td> </tr> <tr> <td>Additional Table 3</td> <td>List of independent mbQTLs associated with p&lt;1x10<sup>-5</sup> in the mbGWAS from nasal samples in Europeans (discovery).</td> </tr> <tr> <td>Additional Table 4</td> <td>Summary results of enrichment analyses using the Enrichr platform.</td> </tr> <tr> <td>Additional Table 5</td> <td>Summary results of stratified enrichment analyses by biological sample type.</td> </tr> <tr> <td>Additional Table 6</td> <td>List of independent mbQTLs associated with p&lt;1x10<sup>-5</sup> in the mbGWAS from saliva samples in African Americans (replication phase).</td> </tr> <tr> <td>Additional Table 7</td> <td>List of independent mbQTLs associated with p&lt;1x10<sup>-5</sup> in the mbGWAS from saliva samples in Latinos (replication phase).</td> </tr> <tr> <td>Additional Table 8</td> <td>Summary results of replication enrichment analyses in saliva samples.</td> </tr> <tr> <td>Additional Table 9</td> <td>Summary statistics of the mbQTL analysis of <em>Streptococcus</em> in nasal samples.</td> </tr> <tr> <td>Additional Table 10</td> <td>Summary statistics of the mbQTL analysis of <em>Campylobacter</em> in pharyngeal samples.</td> </tr> <tr> <td>Additional Table 11</td> <td>Summary statistics of the mbQTL analysis of <em>Tannerella</em> in pharyngeal samples.</td> </tr> <tr> <td>Additional Table 12</td> <td>Summary table of the different genetic models tested for each mbQTL pair.</td> </tr> </tbody> </table>

opencc-by-4.0Jun 2023View details →
dryad40/100

Gene-specific selective sweeps are pervasive across human gut microbiomes

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

Data for: Community assembly of the human piercing microbiome

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad40/100

Supplementary information from: An extensive archaeological dental calculus dataset spanning 5000 years for ancient human oral microbiome research

Open the record for dataset details and reuse information.

publicJul 2025View details →
zenodo36/100

VMGC - Human Vaginal Microbiome Genome Collection

<p>The VMGC is a large-scale reference genome resource of the human vaginal microbiome, including over 33,000 genomes derived from 786 prokaryotes, 11 fungi, and 4,263 viruses associated with the human vagina. In terms of representation, the VMGC demonstrates high efficiency in capturing microbial sequences, with a median mapping rate of 91.7% across 4,472 vaginal metagenomic samples obtained from 14 countries.</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Microbiome of domestic water from rural communities in the southern caribbean: water quality and human health implications

<p><strong>LIST OF SUPPLEMENTARY DATA FILES</strong></p> <p>Chapter 4 - The microbiome of domestic water in the rainwater harvesting-dependent island community of Carriacou, Grenada in the Southern Caribbean</p> <p>Chapter 5 - Microbial composition of drinking water in Speightstown, Barbados</p> <p>Chapter 6 - Bacterial diversity of drinking water in Nariva, Trinidad</p> <p>Chapter 7 - A preliminary assessment of protozoans in various water sources in Nariva, Trinidad</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

MicrobiomeHD: the human gut microbiome in health and disease

<p><strong>Overview</strong></p> <p>MicrobiomeHD is a standardized database of human gut microbiome studies in health and disease. This database includes publicly available 16S data from published case-control studies and their associated patient metadata. Raw sequencing data for each study was downloaded and processed through a standardized pipeline.</p> <p>To be included in MicrobiomeHD, datasets have:</p> <ul> <li>publicly available raw sequencing data (fastq or fasta)</li> <li>publicly available metadata with at least case and control labels for each patient</li> </ul> <p>Currently, MicrobiomeHD is focused on stool samples. Additional samples may be included in certain datasets, as indicated in the metadata.</p> <p><strong>Files</strong></p> <p>Additional information about the datasets included in this MicrobiomeHD release are in the MicrobiomeHD github repo <a href="https://github.com/cduvallet/microbiomeHD">https://github.com/cduvallet/microbiomeHD</a>, in the file <em>db/dataset_info.yaml</em>. Top-level identifiers correspond to dataset IDs labeled by disease_first-author. For the most part, sample sizes in the yaml file are those that were described in the papers, and may not exactly reflect the actual data (due to missing/extra data, samples which didn&#39;t pass quality control, etc).</p> <p>Each dataset was downloaded and processed through a standardized pipeline. The raw processing results are available in the *.tar.gz files here. Each file has the same directory structure and files, as described in the pipeline documentation: <a href="http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html">http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html</a>.</p> <p>Specific files of interest in each *.tar.gz folder include:</p> <ul> <li><strong>summary_file.txt</strong>: this file contains a summary of all parameters used to process the data</li> <li><strong>datasetID.metadata.txt</strong>: the metadata associated with the samples. Note that some samples in the metadata may not have sequencing data, and vice versa.</li> <li><strong>RDP/datasetID.otu_table.100.denovo.rdp_assigned</strong>: the 100% OTU tables with Latin taxonomic names assigned using the RDP classifier (c = 0.5).</li> <li><strong>datasetID.otu_seqs.100.fasta</strong>: representative sequences for each OTU in the 100% OTU table. OTU labels in the OTU table end with <code>d__denovoID</code> - these denovoIDs correspond to the sequences in this file.</li> <li><strong>README.txt</strong>: additional information about steps taken to download and process each dataset, as needed.</li> </ul> <p>The raw data was acquired as described in the supplementary materials of Duvallet et al.&#39;s &quot;Meta analysis of microbiome studies identifies shared and disease-specific patterns&quot; and, when available, the respective dataset README files.</p> <p>Raw sequencing data was processed with the Alm lab&#39;s in-house 16S processing pipeline: <a href="https://github.com/thomasgurry/amplicon_sequencing_pipeline">https://github.com/thomasgurry/amplicon_sequencing_pipeline</a></p> <p>Pipeline documentation is available at: <a href="http://amplicon-sequencing-pipeline.readthedocs.io/">http://amplicon-sequencing-pipeline.readthedocs.io/</a></p> <p>Metadata was extracted from the original papers and/or data sources, and formatted manually. When possible, these steps are documented in each dataset&#39;s associated README.txt file.</p> <p><strong>Contributing</strong></p> <p>MicrobiomeHD is a resource that can be used to extract disease-specific microbiome signals in individual case-control studies. Many microbes respond non-specifically to health and disease, and the majority of bacterial associations within individual studies overlap with this non-specific response. Researchers should cross-check their results with the data presented here to ensure that their identified microbial associations are specific to their disease under study.</p> <p>We provide an updated list of non-specific microbes here, as well as the raw OTU tables for anyone who wishes to reproduce and adapt this analysis to their study question.</p> <p>If you would like to include your case-control dataset in MicrobiomeHD, please email ejalm[at]mit.edu and duvallet[at]mit.edu.</p> <p>For us to process your data through our standard pipeline, you will need to provide the following files and information about your data:</p> <ul> <li>raw sequencing data in fastq or fasta format (preferably fastq)</li> <li>information about which processing steps will be required (e.g. removing primers or barcodes, merging paired-end reads, etc)</li> <li>sample IDs associated with the sequencing data (either mapped to barcodes still in the sequences, or to each de-multiplexed sequencing file)</li> <li>case/control metadata of each sample</li> <li>other relevant metadata (e.g. sampling site, if not all samples are stool; sampling time point, if multiple samples per patient were taken; etc)</li> </ul> <p>By using MicrobiomeHD in your own analyses, you agree to contribute your dataset to this database and to make your raw sequencing data (i.e. fastq files) publicly available.</p> <p><strong>Citing MicrobiomeHD</strong></p> <p>The MicrobiomeHD database and original publications for each of these datasets are described in Duvallet et al. (2017): <a href="http://dx.doi.org/10.1038/s41467-017-01973-8">http://dx.doi.org/10.1038/s41467-017-01973-8</a></p> <p>Duvallet, C., Gibbons, S. M., Gurry, T., Irizarry, R. A., &amp; Alm, E. J. (2017). Meta-analysis of gut microbiome studies identifies disease-specific and shared responses. <em>Nature communications</em>, 8(1), 1784.</p> <p>If you use any of these datasets in your analysis, please cite both MicrobiomeHD (Duvallet et al. (2017)) and the original publication for each dataset that you use.</p> <p>The code used to process and analyze this data in the paper is available on github: <a href="https://github.com/cduvallet/microbiomeHD">https://github.com/cduvallet/microbiomeHD</a></p> <p><strong>Files</strong></p> <p><em>Data files</em></p> <p><strong>file-S3.nonspecific_genera.txt</strong>: Supplemental Table 3 from Duvallet et al. (2017), listing the non-specific health- and disease-associated microbes.<br> <strong>dataset_info.yaml</strong>: yaml file with additional dataset metadata.</p> <p><em>Datasets</em></p> <p>Note that MicrobiomeHD contains all 28 datasets from Duvallet et al. (2017), as well as additional datasets which did not meet the inclusion criteria for the meta-analysis presented in the paper. Additional information about the datasets included in this MicrobiomeHD release are in the original publications and the MicrobiomeHD github repo https://github.com/cduvallet/microbiomeHD, and in the file <em>dataset_info.yaml</em>.</p> <p>The sample sizes listed here reflect what was reported in the original publications. Some may have discrepancies between what is reported and what is in the actual data due to missing data, quality issues, barcode mismatches, etc.</p> <ul> <li><strong>asd_son_results.tar.gz</strong> (<em>asd_son</em>): NT: 44, ASD: 59 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0137725</li> </ul> </li> <li><strong>autism_kb_results.tar.gz</strong> (<em>asd_kang</em>): H: 20, ASD: 20 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0068322</li> </ul> </li> <li><strong>cdi_schubert_results.tar.gz</strong> (<em>cdi_schubert</em>): H: 155, nonCDI: 89, CDI: 94 <ul> <li>http://dx.doi.org/10.1128/mBio.01021-14</li> </ul> </li> <li><strong>cdi_vincent_v3v5_results.tar.gz</strong> (<em>cdi_vincent</em>): H: 25, CDI: 25 <ul> <li>http://dx.doi.org/10.1186/2049-2618-1-18</li> </ul> </li> <li><strong>cdi_youngster_results.tar.gz</strong> (<em>cdi_youngster</em>): H: 4, CDI: 19 <ul> <li>http://dx.doi.org/10.1093/cid/ciu135</li> </ul> </li> <li><strong>crc_baxter_results.tar.gz</strong> (<em>crc_baxter</em>): adenoma: 198, H: 172, CRC: 120 <ul> <li>http://dx.doi.org/10.1186/s13073-016-0290-3</li> </ul> </li> <li><strong>crc_xiang_results.tar.gz</strong> (<em>crc_chen</em>): H: 22, CRC: 21 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0039743</li> </ul> </li> <li><strong>crc_zackular_results.tar.gz</strong> (<em>crc_zackular</em>): adenoma: 30, H: 30, CRC: 30 <ul> <li>http://dx.doi.org/10.1158/1940-6207.CAPR-14-0129</li> </ul> </li> <li><strong>crc_zeller_results.tar.gz</strong> (<em>crc_zeller</em>): H: 75, CRC: 41 <ul> <li>http://dx.doi.org/10.15252/msb.20145645</li> </ul> </li> <li><strong>crc_zhao_results.tar.gz</strong> (<em>crc_wang</em>): H: 56, CRC: 46 <ul> <li>http://dx.doi.org/10.1038/ismej.2011.109}</li> </ul> </li> <li><strong>edd_singh_results.tar.gz</strong> (<em>edd_singh</em>): STEC: 28, CAMP: 71, SALM: 66, SHIG: 34, H: 75 <ul> <li>http://dx.doi.org/10.1186/s40168-015-0109-2</li> </ul> </li> <li><strong>hiv_dinh_results.tar.gz</strong> (<em>hiv_dinh</em>): H: 16, HIV: 21 <ul> <li>http://dx.doi.org/10.1093/infdis/jiu409</li> </ul> </li> <li><strong>hiv_lozupone_results.tar.gz</strong> (<em>hiv_lozupone</em>): H: 13, HIV: 25 <ul> <li>http://dx.doi.org/10.1016/j.chom.2013.08.006</li> </ul> </li> <li><strong>hiv_noguerajulian_results.tar.gz</strong> (<em>hiv_noguerajulian</em>): H: 34, HIV: 206 <ul> <li>https://doi.org/10.1016%2Fj.ebiom.2016.01.032</li> </ul> </li> <li><strong>ibd_alm_results.tar.gz</strong> (<em>ibd_papa</em>): IBDundef: 1, nonIBD: 24, UC: 43, CD: 23 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0039242</li> </ul> </li> <li><strong>ibd_engstrand_maxee_results.tar.gz</strong> (<em>ibd_willing</em>): CCD: 12, H: 35, ICD: 15, UC: 16, ICCD: 2 <ul> <li>http://dx.doi.org/10.1053/j.gastro.2010.08.049</li> </ul> </li> <li><strong>ibd_gevers_2014_results.tar.gz</strong> (<em>ibd_gevers</em>): H: 31, CD: 224 <ul> <li>http://dx.doi.org/10.1016/j.chom.2014.02.005</li> </ul> </li> <li><strong>ibd_huttenhower_results.tar.gz</strong> (<em>ibd_morgan</em>): H: 18, UC: 48, CD: 62 <ul> <li>http://dx.doi.org/10.1186/gb-2012-13-9-r79</li> </ul> </li> <li><strong>mhe_zhang_results.tar.gz</strong> (<em>liv_zhang</em>): CIRR: 25, H: 26, MHE: 26 <ul> <li>http://dx.doi.org/10.1038/ajg.2013.221</li> </ul> </li> <li><strong>nash_chan_results.tar.gz</strong> (<em>nash_wong</em>): H: 22, NASH: 16 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0062885</li> </ul> </li> <li><strong>nash_ob_baker_results.tar.gz</strong> (<em>nash_ob_zhu</em>): H: 16, NASH: 22, OB: 25 <ul> <li>http://dx.doi.org/10.1002/hep.26093</li> </ul> </li> <li><strong>ob_escobar_results.tar.gz</strong> (<em>ob_escobar</em>): OW: 10, H: 10, OB: 10 <ul> <li>https://doi.org/10.1186/s12866-014-0311-6</li> </ul> </li> <li><strong>ob_goodrich_results.tar.gz</strong> (<em>ob_goodrich</em>): OW: 322, H: 433, OB: 183 <ul> <li>http://dx.doi.org/10.1016/j.cell.2014.09.053</li> </ul> </li> <li><strong>ob_gordon_2008_v2_results.tar.gz</strong> (<em>ob_turnbaugh</em>): H: 61, OB: 219 <ul> <li>http://dx.doi.org/10.1038/nature07540</li> </ul> </li> <li><strong>ob_jumpertz_results.tar.gz</strong> (<em>ob_jumpertz</em>): H: 12, OB: 9 <ul> <li>http://ajcn.nutrition.org/content/early/2011/05/03/ajcn.110.010132</li> </ul> </li> <li><strong>ob_ross_results.tar.gz</strong> (<em>ob_ross</em>): H: 26, OB: 37 <ul> <li>http://dx.doi.org/10.1186/s40168-015-0072-y</li> </ul> </li> <li><strong>ob_wu_results.tar.gz</strong> (<em>ob_wu</em>): bmi_data: 101 <ul> <li>http://dx.doi.org/10.1126/science.1208344</li> </ul> </li> <li><strong>ob_zeevi_results.tar.gz</strong> (<em>ob_zeevi</em>): bmi_data: 870 <ul> <li>http://dx.doi.org/10.1016/j.cell.2015.11.001</li> </ul> </li> <li><strong>ob_zupancic_results.tar.gz</strong> (<em>ob_zupancic</em>): H: 167, OB: 117 <ul> <li>http://dx.doi.org/10.1371/journal.pone.0043052</li> </ul> </li> <li><strong>par_scheperjans_results.tar.gz</strong> (<em>par_scheperjans</em>): H: 72, PAR: 72 <ul> <li>http://dx.doi.org/10.1002/mds.26069</li> </ul> </li> <li><strong>ra_littman_results.tar.gz</strong> (<em>art_scher</em>): H: 28, NORA: 44, CRA: 26, PSA: 16 <ul> <li>http://dx.doi.org/10.7554/eLife.01202</li> </ul> </li> <li><strong>t1d_alkanani_results.tar.gz</strong> (<em>t1d_alkanani</em>): T1D: 21, H: 55, T1D_new-onset: 35 <ul> <li>http://dx.doi.org/10.2337/db14-1847</li> </ul> </li> <li><strong>t1d_mejialeon_results.tar.gz</strong> (<em>t1d_mejialeon</em>): T1D: 21, H: 8 <ul> <li>http://dx.doi.org/10.1038/srep03814</li> </ul> </li> </ul> <p><strong>Version changes</strong></p> <p>Version 3</p> <ul> <li>added missing ob_escobar metadata</li> <li>added ob_jumpertz, ob_zeevi, and ob_wu</li> <li>added README.txt files to all folders, with info about data downloading and processing steps</li> <li>removed deprecated quality_control folders from all dataset results</li> <li>changed Supplemental File S3 to the most updated version of non-specific genera (as published in Duvallet et al 2017)</li> </ul> <p>Version 2</p> <ul> <li>added crc_zhu and ob_escobar datasets</li> <li>added list of core genera and dataset_info.yaml</li> </ul>

opencc-by-nc-4.0May 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record