Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.7.1
Dataset results
915 results for “metagenomics”
27 MAGs from the Family of Endozoicomonadaceae derived from Tara Pacific Metagenomes
<p>This dataset contains 27 MAGs from the family of Endozoicomonadaceae generated from a subset of Tara Pacific metagenomes.</p> <p>Contextual information of the MAGs can be found in the associated publication: <strong>Ecology of Endozoicomonadaceae in three coral species across the Pacific Ocean</strong>, Hochart et al, submitted</p> <p> </p>
TMI cheese metagenomics sourmash output files
<p>Sourmash output files from signature sketches, compare, gather, and taxonomy of time-series cheese metagenomes generated using both short- and long-read sequencing technologies. Originally generated as part of the pub <a href="https://research.arcadiascience.com/pub/data-set-metagenomics-timecourse-cheese/release/1">"Paired long- and short-read metagenomics of cheese rind microbial communities at multiple time points"</a>. Analysis code of this data provided on Github at <a href="https://github.com/Arcadia-Science/metagenomics-of-cheese-rind-microbiomes">https://github.com/Arcadia-Science/metagenomics-of-cheese-rind-microbiomes</a>. </p> <p>To directly analyze this data, download and unzip the tarball. Clone the Github repository and point to the 2023-05-23-TMI-processed-data directory within a project using the `scripts/TMI-sourmash-explore.R` script. Install the packages used in the script if necessary. </p>
179 high quality metagenome-assembled genomes sequences and annotations
<p>We analyzed seven sediment samples collected adjacent to ferromanganese nodules from the Clarion–Clipperton Fracture Zone (CCFZ) in the eastern Pacific Ocean. Through deep metagenomic sequencing, assembly, and binning, we reconstructed 179 high quality metagenome-assembled genomes (MAGs). This archive contains these genomes sequences and annotations. </p>
Dalton reservoir metagenome assemblies
<p>Assembly data for metagenomes from the Dalton reservoir, using metaSPADES [v3.14.1; (Nurk <em>et al</em>., 2017)], as described in the paper "Isolation and characterization of a novel Lambda-like phage infecting the bloom-forming cyanobacteria <em>Cylindrospermopsis raciborskii". </em>https://doi.org/10.1111/1462-2920.15908</p> <p>Code for assembly can be found at: https://github.com/danschw/phage_e</p> <p> </p>
Simulated Illumina metagenomic reads
<p>We simulated metagenomic Illumina sequencing reads to a mixture ratio that approximates that found in patient sputa, albeit with a slightly higher mycobacterial component. In total, 0.9 gigabases were generated, at proportions: 46% each for bacteria and human, 6\% <em>Mycobacterium tuberculosis </em>complex (MTBC), and 1% each for virus and non-tuberculous mycobacteria (NTM).</p> <p>The reference genomes that reads were simulated from for these groups were gathered as follows. The references for the virus group were obtained using kraken's (v2.1.2) --download-library functionality. The viral library was downloaded on June 15 2023. The human genome from which the reads were simulated was KOREF_S1v2.1 (RefSeq accession GCA_020497085.1), with contigs shorter than 10kbp removed. The bacterial references were obtained by first downloading the bacteria library through kraken, followed by a subsampling due to the size (166Gb) of the resulting FASTA file. We subsampled the file by first removing sequences with a length <50kbp. We then extracted each sequence into its own FASTA file under a directory for the genus of the sequence - excluding the <em>Mycobacterium</em> genus. Genera were randomly subsampled to contain a maximum of 1000 assemblies. Each genus was then reduced to a representative subset using Assembly Dereplicator (commit 2dfcb14; https://github.com/rrwick/Assembly-Dereplicator) by keeping only 10% of the assemblies for each genus (-f 0.1). The NTM references selected were <em>M. abscessus</em> (accession GCF_017190695.1), <em>M. avium</em> (GCF_020735285.1), <em>M. kansasii</em> (GCA_014701265.1), <em>M. ulcerans</em> (GCF_000013925.1), <em>M. intracellulare</em> (GCF_016756075.1), <em>M. terrae</em> (GCF_010727125.1), and <em>M. fortuitum</em> (GCF_001307545.1). The MTBC reference is a lineage 1 assembly (GCF_932530395.1).</p> <p>Illumina reads were simulated with ART (v2016.06.05). We simulated paired reads from a MiSeq v3 system (-ss MSv3) with a read length of 150, a mean fragment length of 250 and fragment length standard deviation 10 (-l 150 -m 250 -s 10).</p> <p>We removed simulated Illumina reads with any ambiguous base.</p>
Simulated Nanopore metagenomic reads
<p>We simulated metagenomic Nanopore sequencing reads to a mixture ratio that approximates that found in patient sputa, albeit with a slightly higher mycobacterial component. In total, 4.5 gigabases were generated, at proportions: 46% each for bacteria and human, 6\% <em>Mycobacterium tuberculosis </em>complex (MTBC), and 1% each for virus and non-tuberculous mycobacteria (NTM).</p> <p>The reference genomes that reads were simulated from for these groups were gathered as follows. The references for the virus group were obtained using kraken's (v2.1.2) --download-library functionality. The viral library was downloaded on June 15 2023. The human genome from which the reads were simulated was KOREF_S1v2.1 (RefSeq accession GCA_020497085.1), with contigs shorter than 10kbp removed. The bacterial references were obtained by first downloading the bacteria library through kraken, followed by a subsampling due to the size (166Gb) of the resulting FASTA file. We subsampled the file by first removing sequences with a length <50kbp. We then extracted each sequence into its own FASTA file under a directory for the genus of the sequence - excluding the <em>Mycobacterium</em> genus. Genera were randomly subsampled to contain a maximum of 1000 assemblies. Each genus was then reduced to a representative subset using Assembly Dereplicator (commit 2dfcb14; https://github.com/rrwick/Assembly-Dereplicator) by keeping only 10% of the assemblies for each genus (-f 0.1). The NTM references selected were <em>M. abscessus</em> (accession GCF_017190695.1), <em>M. avium</em> (GCF_020735285.1), <em>M. kansasii</em> (GCA_014701265.1), <em>M. ulcerans</em> (GCF_000013925.1), <em>M. intracellulare</em> (GCF_016756075.1), <em>M. terrae</em> (GCF_010727125.1), and <em>M. fortuitum</em> (GCF_001307545.1). The MTBC reference is a lineage 1 assembly (GCF_932530395.1).</p> <p>We used Badreads (v0.4.0) to produce the simulated Nanopore reads for each group, specifying the number of bases in the appropriate proportions mentioned above. For all groups we specified no junk or random reads and 0.5% chimeric reads. In addition, for the MTBC, virus, and NTM groups we used a non-default length option --length 4000,3000 to produce reads with mean length 4000bp and a standard deviation of 3000. Defaults were used for all other options (the default error model is trained on real R10.4.1 Nanopore reads from 2023).</p> <p>We filtered the simulated Nanopore reads to remove any read with a length <500bp or an ambiguous nucleotide (non-ACGT).</p>
metaFlye: scalable long-read metagenome assembly using repeat graphs
<p>Supplementary files for the manuscript titled "metaFlye: scalable long-read metagenome assembly using repeat graphs".</p> <p>The archive includes generated assemblies and the corresponding metaQUAST evaluations.</p>
Metagenomic & Metabolomic Study: Bifidobacterium Probiotic Effects on Gut Microbiota & SCFA in Preterm NICU Infants
ClinicalTrials.gov study NCT07213414. IPD Sharing: NO. Countries: 1. Publications: 1.
Metabolic and Metagenomic Effects of Intestinal Microbiome Repopulation in Unexplained Atherosclerosis
ClinicalTrials.gov study NCT04410003. IPD Sharing: NO. Countries: 1. Publications: 1.
Metagenomic and genomic data associated with the tooth-cavity hair and endogenous DNA of the Tsavo lions
Open the record for dataset details and reuse information.
Poison frog microbiome metagenomics data
Open the record for dataset details and reuse information.
Metagenomics tools employed in microbiome research, 2018-2023, per peer-reviewed publications
Open the record for dataset details and reuse information.
Metagenomes and metagenome-assembled genomes from Onthophagus taurus
Open the record for dataset details and reuse information.
A dynamic ancestral graph model and GPU-based simulation of a community based on metagenomic sampling
Open the record for dataset details and reuse information.
Supplemental data for: Longitudinal, multi-platform metagenomics yields a high-quality genomic catalog and guides an in vitro model for cheese communities
Open the record for dataset details and reuse information.
Data from: Mitochondrial metagenomics reveals the ancient origin and phylodiversity of soil mites and provides a phylogeny of the Acari
Open the record for dataset details and reuse information.
Data from: An accessible metagenomic strategy allows for better characterization of invertebrate bulk samples
Open the record for dataset details and reuse information.
Metagenomics: A viable tool for reconstructing herbivore diet
Open the record for dataset details and reuse information.
Metagenomic and genomic data of symbiotic bacteria from wood-boring shipworms <em>Lyrodus pedicellatus</em> and <em>Teredo bartschi</em>
Open the record for dataset details and reuse information.
Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.