Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.7.1
Dataset results
915 results for “metagenomics”
Data from: Mitochondrial metagenomics reveals the ancient origin and phylodiversity of soil mites and provides a phylogeny of the Acari
<p>High-throughput DNA methods hold great promise for phylogenetic analysis of lineages that are difficult to study with conventional molecular and morphological approaches. The mites (Acari), and in particular the highly diverse soil-dwelling lineages, are among the least known branches of the metazoan Tree-of-Life. We extracted numerous minute mites from soils in an area of mixed forest and grassland in southern Iberia. Selected specimens representing the full morphological diversity were shotgun sequenced in bulk, followed by genome assembly of short reads from the mixture, which produced >100 mitochondrial genomes representing diverse acarine lineages. Phylogenetic analyses in combination with taxonomically limited mitogenomes available publicly resulted in plausible trees defining basal relationships of the Acari. Several critical nodes were supported by ancestral-state reconstructions of mitochondrial gene rearrangements. Molecular calibration placed the minimum age for the common ancestor of the superorder Acariformes, which includes most soil-dwelling mites, to the Cambrian-Ordovician (likely within 455–552 Mya), while the origin of the superorder Parasitiformes was placed later in the Carboniferous-Permian. Most family-level taxa within the Acariformes were dated to the Jurassic and Triassic. The ancient origin of Acariformes and the early diversification of major extant lineages linked to the soil are consistent with a pioneering role for mites in building the earliest terrestrial ecosystems.</p>
Supplementary data (simulated metagenome set 2) to accompany "phyloFlash – Rapid SSU rRNA profiling and targeted assembly from metagenomes"
<p>Comparison of SSU rRNA read extraction and targeted assembly from simulated shotgun metagenome of closely related Bacteorides strains.</p> <p>The phyloFlash software is available from https://github.com/HRGV/phyloFlash. Examples were generated with phyloFlash v3.3b.</p>
Supplementary data (simulated metagenome set 3) to accompany "phyloFlash – Rapid SSU rRNA profiling and targeted assembly from metagenomes"
<p>Comparison of SSU rRNA read extraction and targeted assembly from simulated shotgun metagenome of closely related Bacteorides strains.</p> <p>The phyloFlash software is available from https://github.com/HRGV/phyloFlash. Examples were generated with phyloFlash v3.3b.</p>
Antarctic endolithic bacterial metagenome-assembled genomes
<p>Bacterial assembled genomes and annotation data from the Antarctic cryptoendolithic communities collected during the XXXI (2015-16) Italian Antarctic Expedition.</p> <p>The dataset consists of 4 zip archives and 3 files (comma-separated values). Here is a brief summary of their contents:</p> <ul> <li><strong>MAGs: </strong>high quality (HQ) and medium quality (MQ) bacterial metagenome assembled genomes.</li> <li><strong>MAGs_metadata: </strong>completeness, contamination, length, N50, GTDB classification for each MAG.</li> <li><strong>MAGs_HQ_CDS:</strong> translated coding sequences for each high quality MAG.</li> <li><strong>MAGs_HQ_Annotation: </strong>EggNOG annotation files. For each high quality MAG, the following files are included: <ul> <li>eggnog.emapper.annotations: the final EggNOG annotation;</li> <li>eggnog.emapper.hmm_hits: list of significant hits to eggNOG Orthologous Groups</li> <li>eggnog.emapper.seed_orthologs: best match of each query within the best Orthologous Group (OG) reported in the eggnog.emapper.hmm_hits file<strong>.</strong></li> </ul> </li> <li><strong>Jiangella_Antarctica: </strong><em>Candidatus Jiangella antarctica</em> representative genome (UniValnordMG_2_bin.36.fa) and the extracted ribosomal RNA genes (rRNA.fasta).</li> <li><strong>Order_MSA: </strong>protein multiple sequence alignments using the 120 GTDB bacterial marker genes. These alignments were used to estimate divergence times on orders containing at least 4 CBS, for a total of 19 orders.</li> <li><strong>Samples_accession</strong>: table that relates to the NCBI deposition of the shotgun metagenomes, the following info are included: <ul> <li>NCBI Sequence Read Archive (SRA)</li> <li>BioProject accession numbers</li> <li>JGI Integrated Microbial Genomes & Microbiomes site IDs</li> <li>N50 values</li> <li>Metadata</li> </ul> </li> <li><strong>Samples_metadata: </strong>geographic coordinates, temperature, relative humidity and sampling date are reported.</li> </ul>
Flexible metagenome analysis using the MGX framework -- Benchmark data
<p>Synthetic benchmark metagenomes and annotations used to benchmark taxonomic classification approaches</p> <p>in https://doi.org/10.1186/s40168-018-0460-1</p>
UHGG v1 database for inStrain genome resolved metagenomic analysis
<p>A series of files that are useful for profiling metagenomic communities with the program inStrain.</p>
Resurrection of a global, metagenomically defined gokushovirus
<p>Gokushoviruses are single-stranded, circular DNA bacteriophages found in metagenomic datasets from diverse ecosystems wordwide, including human gut microbiomes. Despite their ubiquity and abundance, little is known about their biology or host range: isolates are exceedingly rare, known only from three obligate intracellular bacterial genera. By synthesizing circularized phage genomes from prophages embedded in diverse enteric bacteria, we produced gokushoviruses in an experimentally tractable model system, allowing us to investigate their features and biology. We demonstrate <a>that virions can reliably</a> infect and lysogenize hosts by hijacking a conserved chromosome-dimer resolution system. Sequence motifs required for lysogeny are detectable in other metagenomically defined gokushoviruses; however, we show that even partial motifs enable phages to persist cytoplasmically without leading to collapse of their host culture. This ability to employ multiple, disparate survival strategies is likely key to the long-term persistence and global distribution of <em>Gokushovirinae</em>.</p>
Searching for anthrax in the New York City subway metagenome.
<p>You can view the write up at the following link: http://read-lab-confederation.github.io/nyc-subway-anthrax-study/</p> <p>This data set includes the scripts and write up of the following GitHub repository: https://github.com/Read-Lab-Confederation/nyc-subway-anthrax-study</p> <p> </p> <p>In January 2015 Chris Mason and his team published<sup>1</sup> an in-depth analysis of metagenomic<sup>2</sup> data(environmental shotgun DNA sequence) from samples isolated from public surfaces in the New York City (NYC) subway system. Along with a ton of really interesting findings, the authors claimed to have detected DNA from the bacterial biothreat pathogens <em>Bacillus anthracis</em> (which causes anthrax) and <em>Yersinia pestis</em>(causes plague) in some of the samples. This predictably led to a huge interest from the press and scientists on social media. The authors followed up with an re-analysis of the data on microbe.net<sup>3</sup>, where they showed some results that suggested the tools that they were using for species identification overcalled anthrax and plague.</p> <p><em>B. anthracis</em> is a Gram-positive bacterium that forms tough spores as part of its lifecycle. The 5.2 M basepair (Mb) main chromosome is very similar to those of other bacteria in species informally called the ‘<em>Bacillus cereus</em> group’<sup>4</sup> (including <em>B. cereus</em>, <em>B. thuringiensis</em> and <em>B. mycoides</em>). <em>Bacillus cereus</em> group strains in general are commonly found in soil but <em>B. anthracis</em> itself is very rare and generally associated with livestock grazing sites with a past history of anthrax.</p> <p>What sets <em>B. anthracis</em> apart from close relatives is the presence of two plasmids: pXO1 (181kb), which carries the lethal toxin genes and pXO2 (94kb), which includes genes for a protective capsule. Without one of these plasmids, <em>B. anthracis</em> is considered attenuated in virulence and unable to cause classic anthrax. Other <em>B. cereus</em> group bacteria can have plasmids very similar to pXO1 and pXO2 but missing the important virulence genes. Rarely, other <em>B. cereus</em> group carry pXO1 and appear to cause anthrax-like disease. Its a confusing situation, not helped by the current overly-narrow species definitions. This recent review<sup>5</sup> gives more information.</p> <p>The NYC subway metagenome study raised very timely questions about using unbiased DNA sequencing for pathogen detection. We were interested in this dataset as soon as the publication appeared and started looking deeper into why the analysis software gave false positive results and indeed what exactly was found in the subway samples. We decided to wrap up the results of our preliminary analysis and put it on this site. This report focuses on the results for <em>B. anthracis</em> but we also did some preliminary work on <em>Y.pestis</em> and may follow up on this later.</p> <ol> <li>http://www.sciencedirect.com/science/article/pii/S2405471215000022</li> <li>http://en.wikipedia.org/wiki/Metagenomics</li> <li>http://microbe.net/2015/02/17/the-long-road-from-data-to-wisdom-and-from-dna-to-pathogen/</li> <li>http://genome.cshlp.org/content/22/8/1512</li> <li>http://www.annualreviews.org/doi/abs/10.1146/annurev.micro.091208.073255</li> </ol> <p> </p> <p> </p>
High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes
<p>This project contains anvi'o profiles and contigs databases that is used and/or referenced from the Lee STM and Khan SA, <em>et al.</em> study titled "<strong>High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes</strong>". The pre-print of this study is available via http://dx.doi.org/10.1101/090993.</p> <p>To be able to work with the data files you will need anvi'o <strong>v2.1.0</strong> to be installed on your system. For installation instructions, or to have access to a Docker image for anvi'o, please visit this URL: http://merenlab.org/software/anvio</p> <p>Public data:</p> <ul> <li><strong>ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION.tar.gz</strong>: Data files for a quick visualization of the 97 MAGs and their distribution across the two FMT recipients. A run script in the archive explains how to use this data.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MERGED-PROFILE.tar.gz</strong>: The merged anvi'o profile for the entire data, which also contains a collection of 97 MAGs identified in the donor. The profile database contains no hierarchical clustering of contigs, however, individual MAGs can be displayed via the following notation since the collection 'MAGs' describe the organization of contigs in each MAG referenced from the dataset `ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION`, as well as from the paper: "anvi-refine -c CONTIGS.db -p PROFILE.db -C MAGs -b <em>FMT-Donor_MAG_00054</em>". All MAG names are in the supplementary tables in our paper.<br> </li> <li><strong>ANVIO-FMT-D-R01-R02-MAGs-SUMMARY.tar.gz</strong>: A static HTML website that contains FASTA files for each MAG, and TAB-delimited matrices for coverage and detection values, and others. After unpacking, you can double-click the index.html file. </li> </ul>
Galaxy Training Date for "Analyses of metagenomic data - The global picture"
<p>These training datasets are part of a Galaxy Training Network tutorial that analyzes metagenomic (amplicon and WGS) data. These datasets are extracted of a project studying the Argentinean agricultural pampean soils (https://www.ebi.ac.uk/metagenomics/projects/SRP016633). </p>
Non-redundant metagenome-assembled genomes of activated sludge reactors at different disturbances and scales
<p>Metagenome-assembled genomes (MAGs) are microbial genomes reconstructed from metagenomic data and can be assigned to known taxa or lead to uncovering novel ones. MAGs can provide insights into how microbes interact with the environment. Here, we performed genome-resolved metagenomics on sequencing data from four studies using sequencing batch reactors at microcosm (~25 mL) and mesocosm (~4 L) scales inoculated with sludge from full-scale wastewater treatment plants. These studies investigated how microbial communities in such plants respond to two environmental disturbances: the presence of toxic 3-chloroaniline and changes in organic loading rate. We report 839 non-redundant MAGs with at least 50% completeness and 10% contamination (MIMAG medium-quality criteria). From these, 399 are of putative high-quality, while sixty-seven meet the MIMAG high-quality criteria. MAGs in this catalogue represent the microbial communities in sixty-eight laboratory-scale reactors used for the disturbance experiments, and in the full-scale wastewater treatment plant which provided the source sludge. This dataset can aid meta-studies aimed at understanding the responses of microbial communities to disturbances, particularly as ecosystems confront rapid environmental changes.</p>
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity
<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>
Metagenomic analysis of gut microbiome illuminates the mechanisms and evolution of lignocellulose degradation in mangrove herbivorous crabs
<p><strong>Background:</strong></p> <p>Sesarmid crabs dominate mangrove habitat as the major primary consumers, which facilitates the trophic link and nutrient recycling in the ecosystem. Therefore, the adaptations and mechanisms of sesarmid crabs to herbivory is not only crucial to terrestrialization and its evolutionary success, but also to the healthy functioning of mangrove forest ecosystems. Although endogenous cellulases expressions were reported in crab species, it remains unknown if the endogenous enzymes alone can complete the whole lignocellulolytic pathway, or they also depend on the contribution from their intestinal microbiome. We attempt to investigate the role of gut symbiotic microbes of mangrove-feeding sesarmid crabs in plant digestion using a comparative metagenomic approach.</p> <p><strong>Results:</strong></p> <p>Metagenomics analyses on 43 crab gut samples from 23 species of mangrove crabs revealed a wide coverage of 127 CAZy families and nine KOs targeting lignocellulose and their derivatives in all species analyzed, including predominantly carnivorous species, suggesting the crab species gut microbiome have lignocellulolytic capacity regardless of dietary preference. Microbial cellulase, hemicellulase and pectinase genes in herbivorous and detritivorous crabs were differentially more abundant when compared to omnivorous and carnivorous crabs, indicating the importance of gut symbionts in lignocellulose degradation in mangrove crabs and the enrichment of lignocellulolytic microbes in response to diet with higher lignocellulose content. The herbivorous and detritivorous crabs showed highly similar CAZyme composition compared to dissimilarities observed in taxonomic profiles observed in both groups, suggesting a stronger selection force to gut microbiota by its functional capacity than by taxonomy. The gut microbiota in herbivorous sesarmid crabs were also enriched with nitrogen reduction and fixation genes, implying possible roles of the gut microbiota in supplementing nitrogen that is deficient in plant diet.</p> <p><strong>Conclusions:</strong></p> <p>Endosymbiotic cellulolytic microbes play an important role in lignocellulose degradation in most crab species but their abundance is strongly correlated with dietary preference, and they are highly enriched in herbivorous sesarmids, thus enhancing their capacity for digestion of mangrove leaves. Dietary preference is a stronger driver in determining the microbial CAZyme composition and taxonomic profile in mangrove crab microbiome, resulting in functional redundancy of endosymbiotic microbes. Our results showed that crabs implement a mixed mode of digestion utilizing both endogenous and microbial enzymes in lignocellulose degradation, as observed in most of the more advanced herbivorous invertebrate species.</p>
Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches
<p><strong>Background</strong></p> <p>The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.</p> <p><strong>Results</strong></p> <p>Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.</p> <p><strong>Conclusions</strong></p> <p>We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.</p>
Metagenome-Assembled Genomes of 2_2_Ac_Mat
<p>The dataset is featured in the data report titled "MAGnificent Microbes: Metagenome-Assembled Genomes of Marine Microorganisms in Mats from a Submarine Groundwater Discharge Site in Mabini, Batangas, Philippines." The study utilized shotgun metagenomics to examine the diversity and functional profiles of marine microorganisms in microbial mats from an SGD-influenced site in Mabini. The dataset includes extracted metagenome-assembled genomes (MAGs) along with their annotations using RAST.</p>
Zostera marina leaf associated bacterial metagenome assembled genomes
<p>Metagenome assembled genomes (MAGs) associated with:</p> <p>A genomic resource for exploring bacterial-viral dynamics in seagrass ecosystems</p> <p>Analysis, code, intermediate and supporting files are archived here: <a href="https://doi.org/10.5281/zenodo.14226514">10.5281/zenodo.14226514</a></p> <p>Viral sequences from this work are archived here: <a href="https://doi.org/10.5281/zenodo.14226038">10.5281/zenodo.14226038</a><br><br>This archive contains:<br>(i) Fifty-six fasta files representing the MAGs described in the above titled work with > 80% completion and < 10% contamination based on CheckM2 metrics<br>(ii) Metadata file describing the MAGs (i.e., subset of Table S3 from the above work)</p>
Metagenome-assembled genomes(MAGs) generated from soil dataset.
<p>MAGs generated from soil dataset with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(pretrain).tar.gz. </p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>
Insight into the ecology of vaginal bacteria through integrative analyses of metagenomic and metatranscriptomic data
<p>Datasets and code used to analyze and and to prepare figures for: "Insight into the ecology of vaginal bacteria through integrative analyses of metagenomic and metatranscriptomic data", France et al 2022.</p>
Twenty-five metagenome assembled genomes recovered from the gut microbiome of the domestic ferret, Mustela putorius
<p>This dataset is composed of 25 unique metagenome assembled genomes (MAGs) recovered from the gut microbiome of three domestic ferrets (<em>Mustela putorius</em>). Details on both MAG and host ferret metadata, as well as information on sample collection, DNA sequencing, and bioinformatic processing can be found in the American Society for Microbiology Resource Announcement by Amundson et al. (in prep). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.