Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.7.1
Dataset results
915 results for “metagenomics”
Supplementary data of endogenous starter culture and cheese metagenomes
<p>The first description of the virome composition in Brazilian artisanal Canastra cheese and the phage-bacterial interactions in this food system. Here you can access the contigs that compose the 16 metagenome-assembled genomes (MAG) explored in paper and their spacers sequences.</p>
Saw Kill river (NY, USA) metagenomics and environmental variables
<p>Microbial community structure and diversity in waterways are altered by wastewater treatment plant (WWTP) discharge which can introduce human-associated microbes, antimicrobial resistance (AMR), and significantly change environmental variables. To better understand these interactions, we investigated the impact of the Bard College WWTP on microbial communities collected from the surface water and sediment from four sites, two sites above and two sites below the Bard College outflow, as well as the outflow itself, this was performed over a period of five months. We measured physico-chemical parameters such as temperature, turbidity, conductivity, dissolved oxygen, and salinity as well as the bioindicators <em>Escherichia</em> <em>coli</em>, total coliforms, and <em>Enterococcus</em> sp. concentration, endotoxins, the <em>intI1</em> gene marker, and the total bacterial abundance through <em>16S</em> <em>rRNA</em> gene.</p>
High resolution shotgun metagenomics: the more data, the better?
<p>This data archive contains results generated using a high resolution shotgun metagenomics (HRSM) bioinformatic pipeline (ShotgunMG - https://jtremblay.github.io/shotgunmg.html) for the following projects:</p> <p>Human gut microbiome dataset: PRJNA588513</p> <p>Antarctic soil dataset: PRJNA513362</p> <p>Agricultural soil dataset: PRJNA513362</p> <p>Mock communities: PRJNA873699</p> <p>These analyses were performed in the context of evaluating if shallow shotgun metagenomic sequencing is an adequate approach to analyze SM sequencing data using a HRSM pipeline.</p> <p>Briefly, these PRJNA projects were analyzed using identical bioinformatic procedures, but using various initial raw sequencing data loads.</p> <p>Notable end results in each archive includes: de novo co-assembly (fasta files), contigs and genes abubance matrices. Beta-diversity (Bray-Curtis dissimilarity matrices), alpha diversity (richness, chao1, Simpson and Shannon indexes matrices), taxonomic summaries from the kingdom to up to the species level and Metagenome Assembled Genomes (MAGs).</p>
Orbicella faveolata coral metagenome assemblies from the ECA region of Florida, USA
<p>The enclosed files include <em>Orbicella faveolata</em> coral metagenome assemblies collected from the Coral Ecosystem Conservation Area (ECA) in southeast Florida, USA. Metadata for the files including region of collection and associated NCBI accession numbers is included in this repository as the metadata file. Apparently healthy coral tissue cores were collected between May 28 and June 21, 2021. The DNA was extracted from the host tissue and mucus and sequenced in a paired-end 150 bp format on an Illumina NovaSeq. Trimming and quality filtering of DNA sequences proceeded, followed by host and endosymbiotic dinoflagellate DNA removal. The host-cleaned reads were assembled individually by coral sample into longer contigs using MegaHit v1.1.4. The “Assembly_Fastas” zipped file contains 45 metagenome assemblies from the individual <em>Orbicella faveolata</em> corals. In addition, these metagenome assemblies were annotated with eggnog-mapper v2.1.6 to generate both predicted gene regions and annotation output files. The “Predicted_Gene_Fastas” zipped file contains nucleotide fasta files of the predicted gene regions for the 45 coral metagenome assemblies. The fasta header of each gene includes the contig ID it originated from in the associated “Assembly_Fastas”. The “Predicted_Gene_Annotations” zipped file contains either .csv or .xlsx files with the eggnog-mapper-based annotations. These files contain a “query contig” that corresponds to the contig ID in the fasta header of the “Predicted_Gene_Fastas”. </p> <p>These data were processed and generated by Julie Meyer’s Lab at the University of Florida, using funding from the Florida Department of Environmental Protection.</p>
Figure 3 in Metagenomic study of the communities of bacterial endophytes in the desert plant Senna Italica and their role in abiotic stress resistance in the plant
Figure 3. Alfa rarefaction curve observed based on observed species (OTUs) value. The curve has shown flatter to the right, which indicates the comparatively high species richness of the senna italica samples. Roots samples: Roots.1, Roots.2, and Roots.3. Leaves samples: Leaves.1, Leaves.2, and Leaves.3 are associated with Senna italica.
Figure 1 in Metagenomic study of the communities of bacterial endophytes in the desert plant Senna Italica and their role in abiotic stress resistance in the plant
Figure 1. (A) Results of clustering: Assembling a group of organisms (The organisms in the same group are similar). (B) The number of OTUs generated for each sample. The Root.1 sample had the most OTUs of 24, while the Leave.1 sample had the fewest of 13. Roots samples: Roots.1, Roots.2, and Roots.3. Leaves samples: Leaves.1, Leaves.2, and Leaves.3 are associated with Senna italica.
Figure 6. A in Metagenomic study of the communities of bacterial endophytes in the desert plant Senna Italica and their role in abiotic stress resistance in the plant
Figure 6. A. The phylum level in Bacteria (bar chart), the bacterial composition of the different samples was similar, while the distribution of each phylum varied in all samples. Based on the V3-V4 region of the 16S rRNA region. Bacterial communities at the phylum classification among the samples (pie chart), as a percentage of the total bacteria isolated from roots and leaves endophyte region. Based on the full-length 16S rRNA sequences. (B) The number of Actinobacteria among the samples. (C) The number of Proteobacteria among the samples. (D) The number of unclassified phyla among the samples. (E) The number of Firmicutes phyla among the samples. (F) The number of Cyanobacteria/Chloroplast among the samples. Roots samples: Roots.1, Roots.2, and Roots.3. Leaves samples: Leaves.1, Leaves.2, and Leaves.3 are associated with Senna italica.
Figure 5 in Metagenomic study of the communities of bacterial endophytes in the desert plant Senna Italica and their role in abiotic stress resistance in the plant
Figure 5. Phylogenetic tree based on 16S rRNA gene sequences representing the diversity of endophytic bacterial communities associated with the leaves and roots from the desert medicinal plant Senna italica "at the Phylum level". The tree was constructed using the "one-click" mode in Phylogeny.fr.(Dereeper et al., 2008).
Figure 6 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 6. Analysis of principal coordinates of sampling sectors according to the distribution of bacterial families reported at 40% contribution according to SIMPER analysis.
Figure 2 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 2. Good quality DNA samples from bacterial populations divided by sampling time factor in soils with pre-sowing levels (A), Lepidium meyenii hypocotyl development (B) and post-harvest (C).
Figure 7 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 7. Dendogram of bacterial families at 40% contribution according to SIMPER analysis in fields disturbed by Lepidium meyenii culture under the effect of the two factors under study (use pressure and sampling period).
Figure 8 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 8. Analysis of the behavior of the clusters of bacterial families significantly differentiated (p<0.05) according to the SIMPROF analysis at a 40% contribution of the total of families registered in soils under the factors use pressure and sampling period.
Figure 3 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 3. Band migration in bacterial populations according to the V3 - V4 region of the 16S bacterial rRNA genes, for the 12 samples divided by the sampling time factor in soils with pre-sowing levels (A),Lepidium meyenii hypocotyl development (B) and post-harvest (C).
Figure 5 in Bacterial diversity in high Andean grassland soils disturbed with Lepidium meyenii crops evaluated by metagenomics
Figure 5. Non-metric MSD and cluster analysis of sampling sectors divided by use pressure and sampling period factors.
Results of a Galaxy metagenomic analysis of bee gut microbiome data from PRJNA977416
<p>This dataset contains the outputs of a metagenomic Galaxy workflow run on the raw data of the project PRJNA977416, including the CSV file of associated metadata and the workflow.ga used for the analysis.</p> <p>Firstly, it has information on taxonomic assignment with :</p> <ul> <li>the reports of all samples for Kraken2, Bracken, and MetaPhlan taxonomic profilers. </li> <li>two tabular files obtained with Taxpasta, which merge samples and standardize taxonomic abundances.</li> <li>for the Bracken standardised abundance, a file with the measures of alpha diversity calculated </li> <li>two HTML files giving access to the Krona diagram for this taxonomic composition.</li> </ul> <p>Secondly, it contains functional informations with :</p> <ul> <li>a tabular file with the relative abundance of all GO terms for all samples</li> <li>a directory detailing pathways and genes families detected.</li> </ul>
Example 16S Metagenomics Dataset
<p>This is a 16S Metagenomics example dataset obtained by transforming data originally from Batista et al. (2015). It consists of data corresponding to 2 conditions (WT untreated and WT after Streptomycin treatment) with 5 replicates each, where exactly 10000 reads were obtained from the original forward raw reads of each sample. This dataset includes a metadata file, sequencing reads as well as the greengenes reference dataset, which is given here for convenience and reproducibility.</p>
Training datasets for 16S Metagenomics analysis with FROGS
<p>This training dataset is from 2 imaginary microbiome samples. Each one is from a paired end 16S amplicon sequencing run and contains 2 fastq files (forwards and reverse.)</p> <p>It is a useful dataset for demonstrating:</p> <ul> <li>16S metagenomics analysis techniques</li> <li>Differences between microbiome samples</li> </ul>
Example data for "Sunbeam: an extensible pipeline for analyzing metagenomic sequencing experiments" [Version 2]
<p>This repository contains the example datasets analyzed in the Sunbeam paper, version 2. Please see the current <a href="http://sunbeam.readthedocs.io/en/latest/quickstart.html">Sunbeam Quickstart Guide</a> for up-to-date instructions on installing and running Sunbeam.</p>
Common photosinthetic enzymes from 174 metagenomes from the Malaspina Expedition 2010 (Ortega et al. 2019)
<p>Predicted genes corresponding to the four most common enzymes present in photosynthetic organisms: NADH:ubiquinone reductase (H+-translocating), N-acetyl-gamma-glutamyl-phosphate reductase, DNA-directed RNA polymerase and non-specific serine/threonine protein kinase of 174 metagenomes sequenced during the Malaspina 2010 global expedition.</p> <p>From: Alejandra Ortega, Nathan R Geraldi, Intikhab Alam, Allan A Kamau, Silvia G Acinas, Ramiro Logares, Josep M Gasol, Ramon Massana, Dorte Krause-Jensen and Carlos M Duarte. Important contribution of macroalgae to oceanic carbon sequestration. Nature Geoscience.</p>
TOPC_bin_586 metagenome assembled genome (MAG)
<p><strong>Contig, gene sequences and functional annotation of the <em>TOPC_bin_586</em> metagenome assembled genome (MAG)</strong></p> <p>Data available:</p> <ol> <li>Nucleotide sequences of the contigs composing the MAG [<em>topc.bin.586.fna</em>]</li> <li>Amino acid sequences of the genes (open reading frames, ORFs) [<em>topc.bin.586_ORFs.faa</em>]</li> <li>Functional annotation table (tab-delimited) for the ORFs [<em>topc.bin.586_ORFs_annotation.tsv</em>]</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.