Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
24
datasets available to search
ShareScore release 0.7.1
Dataset results
24 results for “mock community”
Shotgun metagenomic sequencing dataset of a synthetic mock community containing 20 genomes spiked-in at even and staggered concentrations.
<p>Shotgun metagenomics (SM) sequencing is a popular method used in microbial ecology to obtain insights on microbial community structure and function potential in a given biological system without the need to cultivate microorganisms. The dataset described in this article describes technical triplicates of shotgun metagenomic sequence libraries generated from two purified and titrated mixes of 20 distinct reference bacterial genomes for which key characteristics such as genome size, sequence and spiked-in concentrations are known. In one of the genomic DNA mix, each genome is spiked-in at similar concentrations (representing an even microbial community) and in the other, genomes are spiked-in at different concentrations with some genomes highly abundant and other in low quantity, mimicking an uneven microbial community DNA extract. In order to be interpretable, SM sequencing data needs to be properly analyzed by complex analytical bioinformatic pipelines. Environments investigated with this method can range from simple to very complex. Typically, microbial communities contain microbes that are ubiquitous and some others much rarer. Analysis of rare microbes in a complex microbial community are challenging to perform as their sequencing signals get submerged by the microbial genomes that are more abundant. In this context, it is critical to have access to sequencing data of simple mock communities of mixes of well characterized genomes in order to develop and validate bioinformatic methods that aim to accurately analyze microbial communities.</p>
Genotyping by sequencing for estimating relative abundances of diatom taxa in mock communities
Open the record for dataset details and reuse information.
IMP simulated mock community data set
<p>This file contains the simulated mock (SM) metagenomic and metatranscriptomic dataset along with the original genomes used for simulation used within the article:</p> <p><strong>IMP: a reproducible pipeline for reference-independent integrated metagenomic and metatranscriptomic analyses</strong></p> <p>Shaman Narayanasamy<sup>†</sup>, Yohan Jarosz<sup>†</sup>, Emilie E.L. Muller, Cédric C. Laczny, Malte Herold, Anne Kaysen, Anna Heintz-Buschart, Nicolás Pinel, Patrick May, and Paul Wilmes<sup>*</sup></p> <p>Preprint: http://biorxiv.org/content/early/2016/02/10/039263</p> <p>The folder contains two subfolders (MG and MT), each containing the simulated metagenomic (MG) data and simulated metatranscriptomic (MT) data respectively. The files within these folders are in FASTQ format. The methods for generating these simulated data sets are described in the article.<br> </p> <p>The genomes and the resulting simulated metatranscriptomic data was generated and analysed within the article:</p> <p><strong>Comparison of assembly algorithms for improving rate of metatranscriptomic functional annotation</strong></p> <p>Albi Celaj, Janet Markle, Jayne Danska and John Parkinson; 2014; doi:10.1186/2049-2618-2-39</p> <p> </p> <p>It was provided upon request by the first author Albi Celaj, with permission to share the data. Please cite the aforementioned publication if this simulated metatranscriptomic data is used data is used.</p>
In silico mock communities for evaluation of taxonomic profilers across prokaryotes and viruses
<p><em>In silico </em>mock communities generated with CAMISIM for benchmarking the performance of taxonomic profilers across prokaryotic (50 communities), eukaryotic (30 communities), and viral communities (10 communities) of the human microbiome. Metagenomes were generated using CAMISIM (Fritz et al., 2019), which simulates 2.1 Gb of Illumina 2 ×150 bp paired end reads with the default HiSeq 2500 error profile and a mean insert size of 200 bp. To assess profiling performance for a range of sequencing depths, the 50 <em>in silico</em> metagenomes were also rarefied with seqtk (-s100) to sequencing depths of 20, 5, 2, 1, 0.5, 0.25 and 0.1 million read pairs. Counts are provided for rarefied metagenomes.</p> <p><strong>Prokaryotic communities<br></strong>For prokaryotic benchmarking, 10 body site-representative prokaryotic metagenomes were simulated for each of the following five body sites: adult gut, infant gut, oral, skin, and vagina. Genome accession ids for prokaryotic species found in each human body site were identified from published literature (Bäckhed et al., 2015; Proctor et al., 2019; Saheb Kashaf et al., 2021). </p> <p>Adult Gut: pro_gut_adult.zip<br>Infant Gut: pro_gut_infant.zip<br>Oral: pro_oral.zip<br>Skin: pro_skin_1.zip, pro_skin_2.zip, pro_skin_3.zip<br>Vaginal: pro_vaginal.zip</p> <p>Downsized counts: </p> <p><strong>Eukaryotic communities<br></strong>30 eukaryotic <em>in silico</em> metagenomes comprising up to 200 randomly sampled genomes from a set of 113 eukaryotic species (See Supplementary Table 2 from the paper) corresponding to the eukaryotic species within both CHAMP and MetaPhlAn 4 (Blanco-Míguez et al., 2023) databases.</p> <p>Eukaryotic data is deposited here: <a href="https://doi.org/10.5281/zenodo.12090449" target="_blank" rel="noopener">doi: 10.5281/zenodo.12090449</a></p> <p><strong>Viral communities</strong><br>10 viral communities were simulated with 95% of the reads from bacteria and 5% of the reads originating from phages. Each community consisted of 200 randomly selected bacterial genomes from GTDB with species-level annotation and 200 viral genomes from the Gut Phage Database (GPD, Camarillo-Guerrero et al., 2021). </p> <p>Counts: phage_communities_counts.zip<br>FastQ, forward reads: camisimu_[1-10].fq.1.gz<br>FastQ, reverse reads: camisimu_[1-10].fq.2.gz</p> <p><strong>References</strong></p> <p>Bäckhed, F., Roswall, J., Peng, Y., Feng, Q., Jia, H., Kovatcheva-Datchary, P., et al. (2015). Dynamics and Stabilization of the Human Gut Microbiome during the First Year of Life. <em>Cell Host Microbe</em> 17, 690–703. doi: 10.1016/J.CHOM.2015.04.004</p> <p>Blanco-Míguez, A., Beghini, F., Cumbo, F., McIver, L. J., Thompson, K. N., Zolfo, M., et al. (2023). Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. <em>Nature Biotechnology 2023 41:11</em> 41, 1633–1644. doi: 10.1038/s41587-023-01688-w</p> <p>Camarillo-Guerrero, L. F., Almeida, A., Rangel-Pineros, G., Finn, R. D., and Lawley, T. D. (2021). Massive expansion of human gut bacteriophage diversity. <em>Cell</em> 184, 1098. doi: 10.1016/J.CELL.2021.01.029</p> <p>Fritz, A., Hofmann, P., Majda, S., Dahms, E., Dröge, J., Fiedler, J., et al. (2019). CAMISIM: Simulating metagenomes and microbial communities. <em>Microbiome</em> 7, 1–12. doi: 10.1186/S40168-019-0633-6/FIGURES/5</p> <p>Proctor, L. (2019). Priorities for the next 10 years of human microbiome research. <em>Nature 2021 569:7758</em> 569, 623–625. doi: 10.1038/d41586-019-01654-0</p> <p>Saheb Kashaf, S., Proctor, D. M., Deming, C., Saary, P., Hölzer, M., Mullikin, J., et al. (2021). Integrating cultivation and metagenomics for a multi-kingdom view of skin microbiome diversity and functions. <em>Nature Microbiology 2021 7:1</em> 7, 169–179. doi: 10.1038/s41564-021-01011-w</p>
Fish mock community with 41 species from 13 orders
<p>Tissue extracts of 41 North American fish species were obtained from the Ministère des Forêts, de la Faune et des Parcs (Québec). The selected species were chosen to represent a diversity of families within the Actinopterygii, and to include some species known to be present at the Experimental Lakes Area. Muscle or fin tissue was extracted using Qiagen Blood and Tissue kits and equimolarised to 15ng/µl. Library preparation and next-generation sequencing (NGS) was performed by equimolarising the DNA and combining two replicate mock community libraries which were then PCR amplified and sequenced. Sequencing was conducted using 2x300bp Illumina MiSeq at the McGill University and Génome Québec Innovation Centre, Montréal.<br> </p>
Efficacy of metabarcoding for identification of fish eggs evaluated with mock communities
<p>There is urgent need for effective and efficient monitoring of marine fish populations. Monitoring eggs and larval fish may be more informative that traditional fish surveys since ichthyoplankton surveys reveal the reproductive activities of fish populations, which directly impact their population trajectories. Ichthyoplankton surveys have turned to molecular methods (DNA barcoding & metabarcoding) for identification of eggs and larval fish due to challenges of morphological identification. In this study we examine the effectiveness of using metabarcoding methods on mock communities of known fish egg DNA. We constructed six mock communities with known ratios of species. In addition we analyzed two samples from a large field collection of fish eggs and compared metabarcoding results with traditional DNA barcoding results. We examine the ability of our metabarcoding methods to detect species and relative proportion of species identified in each mock community. We found that our metabarcoding methods were able to detect species at very low input proportions, however levels of successful detection depended on the markers used in amplification, suggesting that the use of multiple markers is desirable. Variability in our quantitative results may result from amplification bias as well as interspecific variation in mitochondrial DNA copy number. Our results demonstrate that there remain significant challenges to using metabarcoding for estimating proportional species composition; however the results provide important insights into understanding how to interpret metabarcoding data. This study will aid in the continuing development of efficient molecular methods of biological monitoring for fisheries management.</p>
Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities
The formation of chimeric sequences can create significant methodological bias in PCR-based DNA metabarcoding analyses. During mixed-template amplification of barcoding regions, chimera formation is frequent and well documented. However, profiling of fungal communities typically uses the more variable rDNA region ITS. Due to a larger research community, tools for chimera detection have been developed mainly for the 16S/18S markers. However, these tools are widely applied to the ITS region without verification of their performance. We examined the rate of chimera formation during amplification and 454 sequencing of the ITS2 region from fungal mock communities of different complexities. We evaluated the chimera detecting ability of two common chimera-checking algorithms: Perseus and UCHIME. Large proportions of the chimeras reported were false positives. No false negatives were found in the dataset. Verified chimeras accounted for only 0.2% of the total ITS2 reads, which is considerably less than what is typically reported in 16S and 18S metabarcoding analyses. Verified chimeric "parent sequences" had significantly higher percent identity to one another than to random members of the mock communities. Community complexity increased the rate of chimera formation. GC content was higher around the verified chimeric break points, potentially facilitating chimera formation through base pair mismatching in the neighboring regions of high similarity in the chimeric region. We conclude that the hypervariable nature of the ITS region seem to buffer the rate of chimera formation in comparison to other, less variable barcoding regions, due to shorter regions of high sequence similarity.
In silico mock communities for evaluation of taxonomic profilers across eukaryotes in the human microbiome
<p><em>In silico </em>mock communities generated with CAMISIM for benchmarking the performance of taxonomic profilers across prokaryotic (50 communities), eukaryotic (30 communities), and viral communities (10 communities) of the human microbiome. Metagenomes were generated using CAMISIM (Fritz et al., 2019), which simulates 2.1 Gb of Illumina 2 ×150 bp paired end reads with the default HiSeq 2500 error profile and a mean insert size of 200 bp.</p> <p><strong>Eukaryotic communities<br></strong>30 eukaryotic <em>in silico</em> metagenomes comprising up to 200 randomly sampled genomes from a set of 113 eukaryotic species (See Supplementary Table 2 from the paper) corresponding to the eukaryotic species within both CHAMP and MetaPhlAn 4 (Blanco-Míguez et al., 2023) databasess.</p> <p><strong>Prokaryotic and viral communities</strong></p> <p>In silico data for prokaryotic and viral communities from the human microbiome can be found here: <a href="https://doi.org/10.5281/zenodo.10777404">doi: 10.5281/zenodo.10777404</a></p> <p><strong>References</strong></p> <p>Blanco-Míguez, A., Beghini, F., Cumbo, F., McIver, L. J., Thompson, K. N., Zolfo, M., et al. (2023). Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. <em>Nature Biotechnology 2023 41:11</em> 41, 1633–1644. doi: 10.1038/s41587-023-01688-w</p> <p>Fritz, A., Hofmann, P., Majda, S., Dahms, E., Dröge, J., Fiedler, J., et al. (2019). CAMISIM: Simulating metagenomes and microbial communities. <em>Microbiome</em> 7, 1–12. doi: 10.1186/S40168-019-0633-6/FIGURES/5</p>
Mock community REMEI
<p>Mock community sequenced with Illumina for quality control purposes. </p>
Supplementary material 9 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Processed TaXon tables of each primer pair (subtracted negative controls and filtered for fish and lamprey taxa OTUs)
Supplementary material 1 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Pairwise comparison of the log-transformed reads of the non-normalized mock community (MC1) compared to the DNA concentration (ng/ul) of each species
Supplementary material 3 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Sampled specimens and their respective species assignment collected for the fish mock community, extraction date, collection site, and concentration after DNA extraction
Supplementary material 4 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
List of all species reported from Germany, their occurrence status, and their presence in the mock community (data from fishbase.org)
Supplementary material 2 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Pairwise comparison of the log-transformed reads of the non-normalized mock community (MC1) compared to log-transformed reads of the normalized mock community (MC2) of each species
Fish mock community with 41 species from 13 orders
Open the record for dataset details and reuse information.
Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities
Open the record for dataset details and reuse information.
Efficacy of metabarcoding for identification of fish eggs evaluated with mock communities
Open the record for dataset details and reuse information.
Fish DNA mock communities of São Francisco and Jequitinhonha rivers
<p>This files correspont to mock communities composed of genomic DNA obtained from fish species from the Jequitinhonha and São Francisco rivers. The mock communities were amplified with three 12S markers - MiFish, NeoFish and Teleo. Theses samples were designed to enable the comparison of marker amplification and taxonomic classification performance.</p>
Supplementary material 8 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Unmodified TaXon tables of each primer pair
Supplementary material 7 from: Macher T-H, Schütz R, Yildiz A, Beermann AJ, Leese F (2023) Evaluating five primer pairs for environmental DNA metabarcoding of Central European fish species based on mock communities. Metabarcoding and Metagenomics 7: e103856. https://doi.org/10.3897/mbmg.7.103856
Protocol for the adapted NucleoMag Tissue Kit
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.