Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.7.1
Dataset results
915 results for “metagenomics”
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Real and mock datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>real and mock</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Simulated datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>simulated</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Annual dynamics and metagenomics of marine vesicles: a layer of complexity in the dissolved organic fraction
<p>The data consists of Illumina raw reads obtained from prokaryotes from coastal Mediterranean seawaters in Alicante, Spain, over the course of one year. Illumina DNA prep kit (IDT, Ref. 20026930) was used following the manufacturer's protocol to generate and sequence the libraries for the obtained DNA samples from prokaryotic metagenomes. Subsequently, these libraries were sequenced using HiSeq X technology (150 PE, 1 lane) by two different facilities: Macrogen in Seoul, Republic of Korea, and the Genomics Unit of the Center for Genomic Regulation (CRG) in Barcelona, Spain. The dataset was used in order to ascertain the main microbial producers involved in the release of DNA packaged in EVs in our annual sampling.</p> <p>The samples are named according to the convention 'prokmeta[Month], where 'prokmeta' signifies prokaryotic metagenome, '[Month]' indicates the month of sampling, and "1" represents the forward , while (2) represent the reverse read. For instance, 'prokmetaDEC.1' denotes the prokaryotic metagenome for sample of December, with '1' indicating the forward read, while 'prokmetaDEC.2' represents the corresponding reverse read. This naming scheme is consistent across all samples in the dataset from sample of November 2021 to sample of June 2022.</p> <p> </p>
Annual dynamics and metagenomics of marine vesicles: a layer of complexity in the dissolved organic fraction.
<p>The dataset consists of third part of data for the same paper " Annual dynamics and metagenomics of marine vesicles: a layer of complexity in the dissolved organic fraction." Data were uploaded separately due to the limited space available for uploading which is f 50.00 GB and since i have more I could not upload them together.</p> <p>The data are Illumina raw reads obtained from vesicle fractions of 20%, isolated from Coastal Mediterranean seawaters in Alicante, Spain, over the course of one year. Illumina DNA prep kit (IDT, Ref. 20026930) was used following the manufacturer's protocol to generate and sequence the libraries for the obtained DNA samples from EVs 20% fraction metagenomes. Subsequently, these libraries were sequenced using HiSeq X technology (150 PE, 1 lane) by two different facilities: Macrogen in Seoul, Republic of Korea, and the Genomics Unit of the Center for Genomic Regulation (CRG) in Barcelona, Spain. The dataset was used in order to ascertain the main microbial organisms involved in the production of EVs and packaging of DNA into those EVs in our annual sampling and the types of genes that were packaged in vessicles isolated from the Coastal Mediterranean seawaters (Functional annotation of genes packaged in vesicles).</p> <p> The data are labeled as "EVs_20_DEC_1.fastq" which denotes the vesicle 20% Optiprep fraction for sample collected in December, with "1" indicating the forward read, while "EVs_20_DEC_2.fastq" represents the corresponding reverse read. This naming scheme is consistent across all vesicle fraction samples in the dataset.</p>
Data supporting publication "Metagenomic Immunoglobulin Sequencing (MIG-Seq) Exposes Patterns of IgA Antibody Binding in the Healthy Human Gut Microbiome"
<p>Data supporting publication "Metagenomic Immunoglobulin Sequencing (MIG-Seq) Exposes Patterns of IgA Antibody Binding in the Healthy Human Gut Microbiome"</p>
Annual dynamics and metagenomics of marine vesicles: a layer of complexity in the dissolved organic fraction
<p>The data consists second part of data for the same paper " Annual dynamics and metagenomics of marine vesicles: a layer of complexity in the dissolved organic fraction." Data were uploaded separately due to the limited space available for uploading.</p> <p>The data are Illumina raw reads obtained from vesicle fractions of 25%, isolated from coastal Mediterranean seawaters in Alicante, Spain, over the course of one year. Illumina DNA prep kit (IDT, Ref. 20026930) was used following the manufacturer's protocol to generate and sequence the libraries for the obtained DNA samples from EVs fractions metagenomes. Subsequently, these libraries were sequenced using HiSeq X technology (150 PE, 1 lane) by two different facilities: Macrogen in Seoul, Republic of Korea, and the Genomics Unit of the Center for Genomic Regulation (CRG) in Barcelona, Spain. The dataset was used in order to ascertain the main microbial organisms involved in the production of EVs and packaging of DNA into those EVs in our annual sampling and the types of genes that were packaged in vessicles isolated from the Coastal Mediterranean seawaters (Functional annotation of genes packaged in vesicles).</p> <p> The data are labeled as "EVs_25_DEC_1.fastq" which denotes the vesicle 25% Optiprep fraction for sample collected in December, with "1" indicating the forward read, while "EVs_25_DEC_2.fastq" represents the corresponding reverse read. This naming scheme is consistent across all vesicle fraction samples in the dataset. </p>
Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing - RASE Database for EuSCAPE
<p>RASE databases used for the prediction of antibiotic phenotype in the paper titled "Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing". A database for <em>Klebsiella pneuomiae </em>constructed from EuSCAPE isolates<em>.</em></p>
Fathi Camel Microbiome Project (FCMP) Fecal Metagenome-assembled Genomes (MAGs)
<p>The Fathi Camel Microbiome Project (FCMP) aims to characterize the diversity and phenotypic associations of the dromedary camel microbiome. The gut microbiome of N = 55 camels was deeply sequenced via dropped stool. The raw reads, after QC, were assembled and binned into metagenome-assembled genomes (MAGs). We include here a collection of 3165 medium-quality or higher prokaryotic MAGs by MiMAG-like criteria (completeness >= 50%, contamination <= 5%). </p>
Metagenomics education in a modular CURE format positively affects students' scientific discovery perception and data analytical skills
<p>The targeted metagenomics study performed by Pollock et al. <em>(Pollock 2018)</em> was part of the Global Coral Microbiome Project <em>(Vega Thurber Lab 2014)</em>. Since the original publication of Pollock et al. <em>(Pollock 2018) </em>was written for a highly specialized academic public, we provided Bsc. students with a short introduction of the paper and guided them through the metadata table stored in the <em>gcmp16S_map_r25.txt </em>file under the <em>GCMP Australia sequence data, OTU tables, and metadata</em> folder (<a href="https://doi.org/10.6084/m9.figshare.c.3855466.v2">https://doi.org/10.6084/m9.figshare.c.3855466.v2</a>).</p> <p>In order to process the fastq-files and obtain OTU tables, a custom-made pipeline on a Galaxy instance <em>(Galaxy Community)</em> was used. This pipeline contained data analytical tools obtained from the Naturalis Biodiversity Center GitHub environment (<a href="https://github.com/naturalis">https://github.com/naturalis</a>). In order to align with the newest Galaxy best practices, we updated the tools and links to the original as well as the updated GitHub and toolshed versions are included in the references list.</p> <p>The names of the original raw fastq-files contain many dots which can hamper file-type recognition during downstream analyses in Galaxy. Therefore, prior to distributing the fastq-files to the students, these names were renamed using the ManageZIP tool <em>(The BLFS Development Team 1999-2023)</em>. Dots were replaced with underscores, with an exception made for the last 2 file type extension dots. For example, the file name <em>E1.2.Tur.pelt.1.20140814.M_S71_L001_R1_001.fastq.gz</em> was changed into <em>E1_2_Tur_pelt_1_20140814_M_S71_L001_R1_001_1.fastq.gz</em>. Similarly, the filenames in the metadata table were adjusted accordingly. The total collection of renamed file names is available through this zenodo repository as wel as the accompanying metadata file.</p>
Simulated Human gut metagenomic samples to benchmark mOTUs v2
<p>We simulated ten human gut metagenomic samples to assess the taxonomic quantification accuracy of the mOTUs tool (<a href="http://motu-tool.org/">link</a>). In this directory you can find the metagenomic samples, the gold standard (used to produce them) and the profiles obtained with four metagenomic profiler tools.</p> <p>Check README.txt for more information.</p>
Simulation data for "MicroPro: using metagenomic unmapped reads to provide insights into human microbiota and disease associations"
<p>This is the simulation data used in the analysis of microbiome-disease association using MicroPro pipeline. Samples 0-24 and 25-49 are cases and controls respectively.</p>
Dataset S2 - Viral metagenomics in the clinical realm: lessons learned from a Swiss-wide ring trial
<p>Dataset S2. FASTQ datasets for increment 2.</p> <p>Supplemental material of article "Viral metagenomics in the clinical realm: lessons learned from a Swiss-wide ring trial".</p>
Dataset S1 - Viral metagenomics in the clinical realm: lessons learned from a Swiss-wide ring trial
<p>Dataset S1. SIB common database.</p> <p>Supplementary material from article "Viral metagenomics in the clinical realm: lessons learned from a Swiss-wide ring trial".</p>
Addiitonal Files for The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling
<p>Additional Files 1-3 for "The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling"</p> <p> </p>
Metagenome Assembled Genomes (MAGs) from faecal microbiomes of great tits and blue tits
<h2><span>Overview:</span></h2> <p><span>The vertebrate gut microbiome plays crucial roles in host health and disease. However, there is limited data on the microbiomes of wild birds, most of which is restricted to barcode sequences. We therefore explored the use of shotgun metagenomics on the faecal microbiomes of two wild bird species widely used as model organisms in ecological studies: the great tit (<em>Parus major</em>) and the Eurasian blue tit (<em>Cyanistes caeruleus</em>). High and Medium quality Metagenome Assembled Genomes (MAGs) were assembled from these metagenomes and are made available as a catalogue in this archive.</span></p> <h2><span>Methods:</span></h2> <p><span><span>Metagenomic reads were trimmed, and quality controlled using FastP configured to a minimum phred score of 20 and minimum length of 50 bp</span><span>. In order to avoid contamination of the bins by eukaryotic sequences, Tiara v1.0.3 was used to classify contigs longer than 3.000 kb into their high-level kingdoms, allowing to exclude sequences of a eukaryotic or of an organelle origin, and only retaining all unclassified contigs and prokaryotic contigs for the binning step. <br>Contigs were binned using MaxBin2 v2.2.7 , SemiBin2 v2.1.0 and Metabat2 v 2.15 independently. The bins were refined using DasTool v 1.1.7 using a min score threshold of 0.3. The quality of the refined bins was obtained using CheckM2 v 1.0.2, and any bin with a contamination above 10% were excluded. The final MAGs were classified as Low-quality (<50% completeness, <10% contamination), medium-quality (>50% completeness, <10% contamination) and high-quality (>90% completeness, <5% contamination), as recommended by the MIMAG specification . Finally, the MAGs were dereplicated using an dRep v 3.4.3 with an ANI of 95% and classified using gtdb-tk v2.4.0 using the gtdb database release220.<br></span></span></p> <p><span><span>Files:</span></span></p> <ul> <li><span><span>The <strong>MAGs_sequences_v1.0.0 </strong>contains the fasta sequence for the individual MAGs assembled in this project</span></span></li> <li><span><span>The <strong>MAGs_catalogue_v1.0.0.xlsx</strong> contains a description of the quality, taxonomic annotation and characteristics of each MAGs in the dataset</span></span></li> </ul>
OMA - Metagenomic analysis datasets
<p><strong>Datasets to test and run LOMA and SOMA metagenomic analysis pipelines. </strong></p> <p><strong>Pipelines:</strong></p> <ul> <li><em><a href="https://github.com/ukhsa-collaboration/LOMA/" target="_blank" rel="noopener">LOMA</a></em> - Long-read (Oxford Nanopore only) metagenomic analysis pipeline. </li> <li><a href="https://github.com/ukhsa-collaboration/SOMA/" target="_blank" rel="noopener"><em>SOMA</em></a> - Short-read metagenomic analysis pipeline.</li> </ul> <p><strong>Files:</strong></p> <div> <ul> <li><em>r220.msh</em> <ul> <li>Mash database created from the Genome Taxonomy Database (GTDB) (Release 220).</li> </ul> </li> </ul> <p> </p> <ul> <li><em>ERR7287988.subset.fastq.gz</em> <ul> <li>Non-randomly subset (~10% original dataset size) of Oxford Nanopore sequencing reads of the ZymoBIOMICS HMW DNA Standard. Derived from a publicly available dataset. The original reads are available <a href="https://www.ebi.ac.uk/ena/browser/view/ERR7287988">here</a> and the relevant publication is available <a href="https://www.nature.com/articles/s41592-022-01539-7">here</a>.</li> </ul> </li> </ul> </div>
Wheat phyllosphere metagenome assembled genomes collected in Ringsted, Denmark
<p><span>We present a completely novel </span><span>and comprehensive wheat phyllosphere metagenomic dataset of 211 samples and </span><span>1261 MAGs. This dataset represents a significant contribution to the field, as it </span><span>provides insights into the poorly studied microbial communities associated with </span><span>wheat leaf surfaces, an ecosystem of considerable agricultural importance.</span></p>
Metagenome-Assembled Genomes and Annotations for McGivern et al
<h3>Files:</h3> <ul> <li><code>reactorEMERGE_annotations.txt</code>: DRAM annotations for MAGs</li> <li><code>gene_lengths.txt</code>: gene length file used to calculate geTMM</li> <li><code>genes.gff.tar.gz</code>: gff file needed for metaT processing</li> <li><code>genes.faa.tar.gz</code>: amino acid sequences for MAG genes</li> <li><code>genes.fna.tar.gz</code>: nucleotide sequences for MAG genes, used as database for metaT mapping</li> </ul>
Detection of Neoplasms by Metagenomic Sequencing of Cerebrospinal Fluid
<p>Detection of Neoplasms by Metagenomic Sequencing of Cerebrospinal Fluid</p> <p>Wei Gu, MD, PhD<sup>1,2</sup>*, Andreas M. Rauschecker, MD, PhD<sup>3</sup>, Elaine Hsu, BS<sup>1</sup>, Kelsey C. Zorn, BS<sup>4</sup>, Yasemin Sucu, BS<sup>1</sup>, Scot Federman, BS<sup>1</sup>, Allan Gopez, BS<sup>1</sup>, Shaun Arevalo, BS<sup>1</sup>, Hannah A. Sample, BS<sup>4</sup>, Eric Talevich, PhD<sup>5</sup>, Eric D. Nguyen, MD, PhD<sup> 1</sup>, Marc Gottschall, BS<sup>1</sup>, Bardia Nourbakhsh, MD, MAS<sup>6</sup>, Carl A. Gold, MD, MS<sup>7</sup>, Bruce A.C. Cree, MD, PhD, MAS<sup>8</sup>, Vanja Douglas, MD<sup>8</sup>, Megan B. Richie, MD<sup>8</sup>, Maulik P. Shah, MD, MHS<sup>8</sup>, S. Andrew Josephson, MD<sup>8</sup>, Jeffrey M. Gelfand, MD, MAS<sup>8</sup>, Steve Miller, MD, PhD<sup> 1</sup>, Linlin Wang, MD<sup>1</sup>, Tarik Tihan, MD, PhD<sup>9</sup>, Joseph L. DeRisi, PhD<sup>4,10</sup>, Charles Y. Chiu, MD, PhD<sup>11,12</sup>, Michael R. Wilson, MD, MAS<sup>8</sup>*</p> <p> </p> <p><sup>1</sup>Department of Laboratory Medicine, University of California San Francisco, CA 94107, USA</p> <p><sup>2</sup>Department of Pathology, Stanford University, CA 94305, USA</p> <p><sup>3</sup>Department of Radiology and Biomedical Imaging, University of California San Francisco, CA 94107, USA</p> <p><sup>4</sup>Department of Biochemistry and Biophysics, University of California San Francisco, CA 94107, USA</p> <p><sup>5</sup>DNANexus, Mountain View, CA 94040, USA</p> <p><sup>6</sup>Department of Neurology, Johns Hopkins University, MD 21287, USA</p> <p><sup>7</sup>Department of Neurology, Stanford University, CA 94305, USA</p> <p><sup>8</sup>Weill Institute for Neurosciences, Department of Neurology, University of California San Francisco, CA 94107, USA</p> <p><sup>9</sup>Department of Pathology, University of California San Francisco, CA 94107, USA</p> <p><sup>10</sup>Chan Zuckerberg Biohub, San Francisco, CA 94107, USA</p> <p><sup>11</sup>UCSF-Abbott Viral Diagnostics and Discovery Center, San Francisco, CA 91407, USA</p> <p><sup>12</sup>Department of Medicine, Division of Infectious Diseases, University of California San Francisco, CA 94107, USA</p> <p> </p> <p>* Corresponding authors:</p> <p>Michael Wilson, MD, MAS</p> <p>UCSF Department of Neurology</p> <p>675 Nelson Rising Lane, NS212</p> <p>San Francisco, CA 94158</p> <p><a href="mailto:Michael.Wilson@ucsf.edu">Michael.Wilson@ucsf.edu</a></p> <p> </p> <p>Wei Gu, MD, PhD</p> <p>Stanford University, Department of Pathology</p> <p>3373 Hillview Ave, 218</p> <p>Palo Alto, CA 94304</p> <p>650-723-1914</p> <p><a href="mailto:weigu@stanford.edu">weigu@stanford.edu</a></p>
Viral metagenome assembled genomes (vMAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 111 viral metagenome assembled genomes (vMAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.