Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

915 results for “metagenomics”

Learn how ShareScore rates datasets ↗
zenodo40/100

Datasets for Paper "MetagenomicKG: a knowledge graph for metagenomic applications"

<p>This repository contains some required data that is used for building MetagenomicKG. Please see more details in <a href="https://github.com/KoslickiLab/MetagenomicKG">https://github.com/KoslickiLab/MetagenomicKG</a>.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Companion data deposit of manuscript: Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies

<p>This upload contains the metagenome assemblies and their binning results generated &amp; described in the manuscript "Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies" (preprint version: arxiv2210.00098, "Towards complete representation of bacterial contents in metagenomic samples").&nbsp;</p> <p>Mapping of sample names in the file names and the descriptors used as in the manuscript can be found in table S1, which is available along with the manuscript and also included in the supplementary_tables_and_figures tar archive here.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

MARMICRODB database for taxonomic classification of (marine) metagenomes

<p><strong>UPDATE (April 2024):<br>Please note that the combined fasta file (MARMICRODB.faa.bz2) associated with version 1.0.0 was mistakenly uploaded from an early prototype of the database.&nbsp;</strong>As a result, some taxonomy identifiers appended to the sequence records are mismatched with the names.dmp, nodes.dmp, and the MARMICRODB_catalog.tsv files. <br><br><strong>Importantly, the taxon IDs in the FM-compressed Kaiju database are correct and valid for use with the names.dmp and nodes.dmp files. </strong>Any analysis using the compressed Kaiju database (MARMICRODB.fmi) from version 1.0.0 will remain valid. However, if you have used the combined fasta file (MARMICRODB.faa.bz2) to generate your own Kaiju database and you used the taxdump files from version 1.0.0 I cannot guarantee the validity of those results. It is <em><strong>strongly recommended</strong></em> that you rebuild the combined fasta file from scratch using the FTP links to the NCBI assemblies provided in the MARMICRODB_catalog.tsv file.</p> <p>The combined fasta file has been removed from this version of the MARMICRODB.</p> <p><strong>Disclaimer:<br></strong>MARMICRODB is optimized for metagenomic samples from the marine environment, in particular, planktonic microbes from the pelagic euphotic zone. We expect this database may also be useful for classifying other types of marine metagenomic samples (for example, mesopelagic, bathypelagic, or even benthic or marine host-associated), but it has not been tested for this purpose. The original purpose of this database was to quantify clades/ecotypes of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 in metagenomes from Tara Oceans Expedition and the GEOTRACES project. We carefully annotated and quality controlled genomes from these five groups, but the processing of the other marine taxa was automated and unsupervised. Taxonomy for other groups was copied over from the Genome Taxonomy Database version RS83 (GTDB) [19,39] and NCBI Taxonomy [23] so any inconsistencies in those databases at the time of access will be propagated to MARMICRODB. If the user&rsquo;s goal is to focus on a particular organism/clade that we did not curate in the database then the user should spend some time curating those genomes (ie checking for contamination, dereplicating, building a genome phylogeny for custom taxonomy node assignment). Currently the custom taxonomy is hardcoded in the MARMICRODB.fmi index, but if users wish to modify MARMICRODB by adding or removing genomes, or reconfiguring taxonomic ranks the names.dmp and nodes.dmp files can easily be modified. However, the Kaiju index will need to be rebuilt, which will require a high-performance compute cluster.</p> <p><strong>Introduction:</strong><br>This sequence database (MARMICRODB) was introduced in the publication JW Becker, SL Hogle, K Rosendo, and SW Chisholm. 2019. Co-culture and biogeography of <em>Prochlorococcus</em> and SAR11. ISME J. doi:10.1038/s41396-019-0365-4. Please see the original publication and its associated supplementary material for the original description of this resource.&nbsp;</p> <p><strong>Motivation:</strong><br>We needed a reference database to annotate shotgun metagenomes from the Tara Oceans project [1] the <a href="https://www.geotraces.org/">GEOTRACES</a>&nbsp;cruises GA02, GA03, GA10, and GP13 and the HOT and BATS time series [2]. Our interests are primarily in quantifying and annotating the free-living, oligotrophic bacterial groups <em>Prochlorococcus</em>, <em>Pelagibacterales</em>/SAR11, SAR116, and SAR86 from these samples using the protein classifier tool Kaiju [3]. Kaiju&rsquo;s sensitivity and classification accuracy depends on the composition of the reference database, and the highest sensitivity is achieved when the reference database contains a comprehensive representation of expected taxa from an environment/sample of interest. However, the speed of the algorithm decreases as the database size increases. Therefore, we aimed to create a reference database that maximized the representation of sequences from marine bacteria, archaea, and microbial eukaryotes while minimizing (but not excluding) the sequences from clinical, industrial, and terrestrial host-associated samples.</p> <p><strong>Results/Description:</strong><br>MARMICRODB consists of 56 million sequence non-redundant protein sequences from 18769 bacterial/archaeal/eukaryote genome and transcriptome bins and 7492 viral genomes optimized for use with the protein homology classifier Kaiju [3]. To ensure maximum representation of marine bacteria, archaea, and microbial eukaryotes, we included translated genes/transcripts from 5397 representative &ldquo;specI&rdquo; species clusters from the proGenomes database [4]; 113 transcriptomes from the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) [5]; 10509 metagenome assembled genomes from the Tara Oceans expedition [6,7], the Red Sea [8], the Baltic Sea [9], and other aquatic and terrestrial sources [10]; 994 isolate genomes from the Genomic Encyclopedia of Bacteria and Archaea [11]; 7492 viral genomes from NCBI RefSeq [12]; 786 bacterial and archaeal genomes from MarRef [13]; and 677 marine single cell genomes [14]. In order to annotate metagenomic reads at the clade/ecotype level (subspecies) for the focal taxa&nbsp;<em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116, we generated custom MARMICRODB taxonomies based on curated genome phylogenies for each group. The curated phylogenies, Kaiju formatted Burrows-Wheeler index, translated genes, the custom taxonomy hierarchy, an <a href="https://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">interactive kronaplot of the taxonomic composition</a>, and scripts and instructions for how to use or rebuild the resource is available from <a href="https://doi.org/10.5281/zenodo.3520509">10.5281/zenodo.3520509</a>.&nbsp;</p> <p><strong>Methods:</strong><br>The curation and quality control of MARMICRODB single-cell, metagenome-assembled, and isolate genomes was performed as described in [15]. Briefly, we downloaded all MARMICRODB genomes as raw nucleotide assemblies from NCBI. We determined an initial genome taxonomy for these assemblies using checkM with the default lineage workflow [16]. All genome bins met the completion/contamination thresholds outlined in prior studies [7,17]. For single cell and metagenome-assembled genomes, especially those from Tara Oceans Mediterranean sea samples [18], we use the GTDB-Tk classification workflow [19] to verify the taxonomic fidelity of each genome bin. We then selected genomes with a checkM taxonomic assignment of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 for further analysis and confirmed taxonomic assignment using blast matches to known Prochlorococcus/Synechococcus ITS sequences and by matching 16S sequences to the SILVA database [20]. To refine our estimates of completeness/contamination of <em>Prochlorococcus</em> genome bins we created a custom set of 730 single copy protein families (available from <a href="https://doi.org/10.5281/zenodo.3719132">10.5281/zenodo.3719132</a>) from closed, isolate <em>Prochlorococcus</em> genomes [21] for quality assessments with checkM. For <em>Synechococcus</em> we used the CheckM taxonomic-specific workflow with the genus <em>Synechococcus</em>. After the custom CheckM quality control, we excluded any genome bins from downstream analysis that had an estimated quality &lt; 30, defined as %completeness &ndash; 5x %contamination resulting in 18769 genome/transcriptome bins. We predicted genes in the resulting genome bins using prodigal [22] and excluded protein sequences with lengths less than 20 and greater than 20000 amino acids, removed non-standard amino acid residues, and condensed redundant protein sequences to a single representative sequence to which we assigned a lowest common ancestor (LCA) taxonomy identifier from the NCBI taxonomy database [23]. The resulting protein sequences were compiled and used to build a Kaiju [3] search database.&nbsp;</p> <p>The above filtering criteria resulted in 605 <em>Prochlorococcus</em>, 96 <em>Synechococcus</em>, 186 SAR11/<em>Pelagibacterales</em>, 60 SAR86, and 59 SAR116 high-quality genome bins. We constructed a high-quality fixed reference phylogenetic tree for each taxonomic group based on genomes manually selected for completeness and the phylogenetic diversity. For example the <em>Prochlorococcus</em> and <em>Synechococcus</em> genomes for the fixed reference phylogeny are estimated &gt; 90% complete, and SAR11 genomes are estimated &gt; 70% complete. We created multiple sequence alignments of phylogenetically conserved genes from these genomes using the GTDB-Tk pipeline [19] with default settings. The pipeline identifies conserved proteins (120 bacterial proteins) and generates concatenated multi-protein alignments [17] from the genome assemblies using hmmalign from the hmmer software suite. We further filtered the resulting alignment columns using the bacterial and archaeal alignment masks from [17] (http://gtdb.ecogenomic.org/downloads). We removed columns represented by fewer than 50% of all taxa and/or columns with no single amino acid residue occurring at a frequency greater than 25%. We trimmed the alignments using trimal [24] with the automated -gappyout option to trim columns based on their gap distribution. We inferred reference phylogenies using multithreaded RAxML [25] with the GAMMA model of rate heterogeneity, empirically determined base frequencies, and the LG substitution model [26](PROTGAMMALGF). Branch support is based on 250 resampled bootstrap trees. This tree was then pruned to only allow a maximum average distance to the closest leaf (ADCL) of 0.003 to reduce the phylogenetic redundancy in the tree [27]. We then &ldquo;placed&rdquo; genomes that either did not pass completeness threshold or were considered phylogenetically redundant by ADCL within the fixed reference phylogeny for each group using pplacer [28] representing each placed genome as a pendant edge in the final tree. We then examined the resulting tree and manually selected clade/ecotype cutoffs to be as consistent as possible with clade definitions previously outlined for these groups [29&ndash;32]. We then gave clades from each taxonomic group custom taxonomic identifiers and we added these identifiers to the MARMICRODB Kaiju taxonomic hierarchy.</p> <p><strong>Software/databases used:</strong><br>checkM v1.0.11[16]<br>HMMERv3.1b2 (http://hmmer.org/)<br>prodigal v2.6.3 [22]<br>trimAl v1.4.rev22 [24]<br>AliView v1.18.1 [33] [34]<br>Phyx v0.1 [35]<br>RAxML v8.2.12 [36]<br>Pplacer v1.1alpha [28]<br>GTDB-Tk v0.1.3 [19]<br>Kaiju v1.6.0 [34]<br>GTDB RS83 (https://data.ace.uq.edu.au/public/gtdb/data/releases/release83/83.0/)<br>NCBI Taxonomy (accessed 2018-07-02) [23]<br>TIGRFAM v14.0 [37]<br>PFAM v31.0 [38]</p> <p><strong>Use example:</strong><br>Because we used custom taxonomic MARMICRODB users will find many reads assigned to non-standard NCBI taxonomy identifiers. However, these reads are easily parsable using the custom names.dmp and nodes.dmp files included with the database. We include a brief description of how to do this below.</p> <p>I typically run Kaiju like:</p> <pre><code>kaiju -z 20 -a greedy -e 5 -m 11 -s 65 -E 0.05 -x \ -t nodes.dmp -f MARMICRODB.fmi \ -i inputfile_R1.fastq.gz \ -j inputfile_R2.fastq.gz \ -o MYOUTPUT.kaiju</code></pre> <p>To obtain a parseable report that lists the custom taxonomic ranks from the nodes.dmp and names.dmp files run kaiju2krona on the output.</p> <pre><code>kaiju2krona -t nodes.dmp -n names.dmp -i MYOUTPUT.kaiju -o MYOUTPUT.kaiju.krona </code></pre> <p>This report shows counts assigned to each node in the custom taxonomy and will also include the names for each rank. You can easily parse this programmatically using a scripting language like python or by using unix utilities.</p> <p><strong>File descriptions:</strong></p> <p><em>MARMICRODB_catalog.tsv</em><br>Tabular file of NCBI assembly accessions and associated taxonomic information for every genome in MARMICRODB. Also includes literature references for each genome where available. Header description:</p> <ol> <li>genome: Unique identifier for each genome</li> <li>full_name: full organism name where available</li> <li>source: literature reference where available</li> <li>taxid: NCBI taxonomy ID for the assembly accession</li> <li>MARMICRODBtaxid: taxonomy ID used in the custom Kaiju database</li> <li>lineage_assignment: taxonomic lineage assignment from NCBI</li> <li>domain: archaea, bacteria, or eukaryote</li> <li>taxgroup: short descriptive group</li> <li>taxclade: higher resolution clade assignment where available</li> <li>habitat_source: whether genome derives from marine or aquatic source</li> <li>sequence_type: isolate, single cell genome (sag), metagenome assembled genome (mag), or transcriptome in case of eukaryotes</li> <li>assembly_ftp:&nbsp;NCBI ftp for assembly</li> <li>gbk_acc: assembly genbank or refseq accession number</li> <li>gtdb_taxonomy: taxonomic lineage assignment from GTDB-Tk v0.1.3 against GTDB v83</li> </ol> <p><em>MARMICRODB_kronaplot.html</em><br><a href="http://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">Interactive Kronaplot</a> for the exploration of taxonomic composition of MARMICRODB</p> <p><em>MARMICRODB.faa.bz2</em><br>Fasta file of all protein sequences in MARMICRODB</p> <p><em>scripts.tar.gz</em><br>directory containing scripts for generating Kaiju formatted database</p> <p><em>phylogenies.tar.gz</em><br>directory containing detailed phylogenies for SAR11, Prochlorococcus, SAR86, and SAR116</p> <p><em>MARMICRODB.fmi</em><br>Kaiju index for MARMICRODB</p> <p><em>nodes.dmp</em><br>nodes file for taxonomic assignment with Kaiju</p> <p><em>names.dmp</em><br>names file for generating Kaiju reports</p> <p><strong>References:</strong><br>1. &nbsp;&nbsp; &nbsp;Karsenti E, Acinas SG, Bork P, Bowler C, De Vargas C, Raes J, et al. A holistic approach to marine eco-systems biology. PLoS Biol. 2011;9: e1001177.<br>2. &nbsp;&nbsp; &nbsp;Biller SJ, Berube PM, Dooley K, Williams M, Satinsky BM, Hackl T, et al. Marine microbial metagenomes sampled across space and time. Scientific Data. 2018;5: 180176.<br>3. &nbsp;&nbsp; &nbsp;Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>4. &nbsp;&nbsp; &nbsp;Mende DR, Letunic I, Huerta-Cepas J, Li SS, Forslund K, Sunagawa S, et al. proGenomes: a resource for consistent functional and taxonomic annotations of prokaryotic genomes. Nucleic Acids Res. 2017;45: D529&ndash;D534.<br>5. &nbsp;&nbsp; &nbsp;Keeling PJ, Burki F, Wilcox HM, Allam B, Allen EE, Amaral-Zettler LA, et al. The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): illuminating the functional diversity of eukaryotic life in the oceans through transcriptome sequencing. PLoS Biol. 2014;12: e1001889.<br>6. &nbsp;&nbsp; &nbsp;Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>7. &nbsp;&nbsp; &nbsp;Tully BJ, Graham ED, Heidelberg JF. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Sci Data. 2018;5: 170203.<br>8. &nbsp;&nbsp; &nbsp;Haroon MF, Thompson LR, Parks DH, Hugenholtz P, Stingl U. A catalogue of 136 microbial draft genomes from Red Sea metagenomes. Sci Data. 2016;3: 160050.<br>9. &nbsp;&nbsp; &nbsp;Hugerth LW, Larsson J, Alneberg J, Lindh MV, Legrand C, Pinhassi J, et al. Metagenome-assembled genomes uncover a global brackish microbiome. Genome Biol. 2015;16: 279.<br>10. &nbsp;&nbsp; &nbsp;Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017.<br>11. &nbsp;&nbsp; &nbsp;Mukherjee S, Seshadri R, Varghese NJ, Eloe-Fadrosh EA, Meier-Kolthoff JP, G&ouml;ker M, et al. 1,003 reference genomes of bacterial and archaeal isolates expand coverage of the tree of life. Nat Biotechnol. 2017;35: 676&ndash;683.<br>12. &nbsp;&nbsp; &nbsp;Haft DH, DiCuccio M, Badretdin A, Brover V, Chetvernin V, O&rsquo;Neill K, et al. RefSeq: an update on prokaryotic genome annotation and curation. Nucleic Acids Res. 2018;46: D851&ndash;D860.<br>13. &nbsp;&nbsp; &nbsp;Klemetsen T, Raknes IA, Fu J, Agafonov A, Balasundaram SV, Tartari G, et al. The MAR databases: development and implementation of databases specific for marine metagenomics. Nucleic Acids Res. 2018;46: D692&ndash;D699.<br>14. &nbsp;&nbsp; &nbsp;Berube PM, Biller SJ, Hackl T, Hogle SL, Satinsky BM, Becker JW, et al. Single cell genomes of Prochlorococcus, Synechococcus, and sympatric microbes from diverse marine environments. Scientific Data. 2018;5: 180154.<br>15. &nbsp;&nbsp; &nbsp;Becker JW, Hogle SL, Rosendo K, Chisholm SW. Co-culture and biogeography of Prochlorococcus and SAR11. ISME J. 2019. doi:10.1038/s41396-019-0365-4<br>16. &nbsp;&nbsp; &nbsp;Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25: 1043&ndash;1055.<br>17. &nbsp;&nbsp; &nbsp;Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017. doi:10.1038/s41564-017-0012-7<br>18. &nbsp;&nbsp; &nbsp;Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>19. &nbsp;&nbsp; &nbsp;Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2019. doi:10.1093/bioinformatics/btz848<br>20. &nbsp;&nbsp; &nbsp;Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41: D590&ndash;6.<br>21. &nbsp;&nbsp; &nbsp;Biller SJ, Berube PM, Berta-Thompson JW, Kelly L, Roggensack SE, Awad L, et al. Genomes of diverse isolates of the marine cyanobacterium Prochlorococcus. Sci Data. 2014;1: 140034.<br>22. &nbsp;&nbsp; &nbsp;Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics. 2010;11: 119.<br>23. &nbsp;&nbsp; &nbsp;Federhen S. The NCBI Taxonomy database. Nucleic Acids Res. 2012;40: D136&ndash;43.<br>24. &nbsp;&nbsp; &nbsp;Capella-Guti&eacute;rrez S, Silla-Mart&iacute;nez JM, Gabald&oacute;n T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25: 1972&ndash;1973.<br>25. &nbsp;&nbsp; &nbsp;Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688&ndash;2690.<br>26. &nbsp;&nbsp; &nbsp;Le SQ, Gascuel O. An improved general amino acid replacement matrix. Mol Biol Evol. 2008;25: 1307&ndash;1320.<br>27. &nbsp;&nbsp; &nbsp;Matsen FA, Gallagher A, McCoy C. Minimizing the average distance to a closest leaf in a phylogenetic tree. arXiv [q-bio.PE]. 2012. Available: http://arxiv.org/abs/1205.6867<br>28. &nbsp;&nbsp; &nbsp;Matsen F a., Kodner RB, Armbrust EV. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree. BMC Bioinformatics. 2010;11: 538.<br>29. &nbsp;&nbsp; &nbsp;Biller SJ, Berube PM, Lindell D, Chisholm SW. Prochlorococcus: the structure and function of collective diversity. Nat Rev Microbiol. 2014;13: 13&ndash;27.<br>30. &nbsp;&nbsp; &nbsp;Giovannoni SJ. SAR11 Bacteria: The Most Abundant Plankton in the Oceans. Ann Rev Mar Sci. 2016.<br>31. &nbsp;&nbsp; &nbsp;Dupont CL, Rusch DB, Yooseph S, Lombardo M-J, Alexander Richter R, Valas R, et al. Genomic insights to SAR86, an abundant and uncultivated marine bacterial lineage. ISME J. 2012;6: 1186&ndash;1199.<br>32. &nbsp;&nbsp; &nbsp;Yang S-J, Kang I, Cho J-C. Expansion of Cultured Bacterial Diversity by Large-Scale Dilution-to-Extinction Culturing from a Single Seawater Sample. Microb Ecol. 2016;71: 29&ndash;43.<br>33. &nbsp;&nbsp; &nbsp;Larsson A. AliView: a fast and lightweight alignment viewer and editor for large datasets. Bioinformatics. 2014;30: 3276&ndash;3278.<br>34. &nbsp;&nbsp; &nbsp;Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>35. &nbsp;&nbsp; &nbsp;Brown JW, Walker JF, Smith SA. Phyx: phylogenetic tools for unix. Bioinformatics. 2017;33: 1886&ndash;1888.<br>36. &nbsp;&nbsp; &nbsp;Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688&ndash;2690.<br>37. &nbsp;&nbsp; &nbsp;Haft DH, Selengut JD, Richter RA, Harkins D, Basu MK, Beck E. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 2013;41: D387&ndash;95.<br>38. &nbsp;&nbsp; &nbsp;Finn RD, Coggill P, Eberhardt RY, Eddy SR, Mistry J, Mitchell AL, et al. The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res. 2016;44: D279&ndash;85.<br>39. &nbsp;&nbsp; &nbsp;Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil P-A, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol. 2018.<br>&nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Metagenomes: gene calls

<p>Gene calls for all contigs in 1,782 metagenomes.</p> <p>Gene calls were made by Prodigal.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

"Genome binning of viral entities from bulk metagenomics data" - CAMISIM simulated datasets and genomes

<p><strong>Genome binning of viral entities from bulk metagenomics data</strong></p> <p>&nbsp;</p> <p><strong>Authors</strong></p> <p><strong>Joachim Johansen1,2, Damian R. Plichta2, Jakob Nybo Nissen1,3, Marie Louise Jespersen1,4, Shiraz A. Shah5, Ling Deng6, Jakob Stokholm5,6, Hans Bisgaard5, Dennis Sandris Nielsen6, S&oslash;ren S&oslash;rensen7, Simon Rasmussen1</strong></p> <p>&nbsp;</p> <p><strong>Affiliations</strong></p> <p>1 Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen N, Denmark</p> <p>2 Infectious Disease and Microbiome Program, Broad Institute of MIT and Harvard, Cambridge, MA, USA</p> <p>3 Statens Serum Institut, Viral &amp; Microbial Special diagnostics, Copenhagen, Denmark</p> <p>4 National Food Institute, Technical University of Denmark, Kongens Lyngby, Denmark</p> <p>5 Copenhagen Prospective Studies on Asthma in Childhood (COPSAC), Herlev and Gentofte Hospital, University of Copenhagen, Copenhagen, Denmark</p> <p>6 Section of Food Microbiology and Fermentation, Department of Food Science, Faculty of Science, University of Copenhagen, Copenhagen, Denmark</p> <p>7 Section of Microbiology, Department of Biology, University of Copenhagen, Copenhagen, Denmark</p> <p><strong>Methods description</strong></p> <p>We compared the viral binning performance of VAMB and MetaBAT2 using the official CAMI consortium method to create assemblies and metagenome profiles. To this end we generated 3 different metagenome compositions with up to 308 reference genomes; one mixed with bacteria, plasmids and viruses to test binning in complex samples i.e. high diversity (1), one with only crass-like viruses to test binning with highly similar viruses i.e. high relatedness (2) and a set of small-viruses (&lt;6,000 bp) including members of the Microviridae family to address the bias of size (3). Bacterial genomes were gathered from NCBIs refseq genome repository 2021, plasmids from the PLSDB database (v. 2021_06_23)&nbsp;and viral genomes from the recent MGV database.&nbsp;&nbsp;</p> <p>Dataset A contained a mixture of bacteria (N=8), plasmids (N=20) and viruses (N=280) to test binning in complex samples, i.e. high diversity. Dataset B contained only crass-like viruses (N=80) to test binning with highly similar viruses i.e. high relatedness. Dataset C contained small-viruses (N=50, &lt;6,000 bp) of the Microviridae family to address the bias of size. Bacterial genomes were sampled from the Refseq genome repository 2021, plasmids from the PLSDB database&nbsp; and viral genomes from the recent MGV database (Nayfach, et al. Nature Microbiology 2021).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Metagenome-assembled Genomes of Scandinavium goeteborgense MCPNR19-05 and Erwinia aphidicola MCPNR19-06

<p>Here we provide two fasta files (MCPNR19-05.fa and MCPNR19-06.fa) which represent low-quality metagenome assembled genomes (MAGs) obtained from genomic DNA from Massospora cicadina isolate MCPNR19 azygospores collected from multiple seventeen-year cicada (Magicicada septendecim) June 2019 at Powdermill Nature Reserve, Rector, Pennsylvania.</p> <p>MCPNR19-05.fa = Scandinavium goeteborgense MCPNR19-05, a 1.82 Mb 45.61% complete MAG<br> MCPNR19-06.fa = Erwinia aphidicola MCPNR19-06, a 1.79 Mb 29.31% complete MAG<br> <br> <strong>Raw data availability</strong><br> Sequence reads are deposited under SRA project accessions <a href="https://ncbi.nlm.nih.gov/sra/SRR17553520">SRR17553520</a>-<a href="https://ncbi.nlm.nih.gov/sra/SRR17553526">SRR17553526</a> and BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA795459">PRJNA795459</a>. These MAGs are metagenomic assemblies obtained from the host Massospora cicadina (BioSample: SAMN24722893). &nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Analysis of pangolin metagenomic datasets reveals significant contamination, raising concerns for pangolin CoV host attribution

<p>Supplementary Figures, Information and Data to accompany:</p> <p>Analysis of pangolin metagenomic datasets reveals significant contamination, raising concerns for pangolin CoV host attribution<br> <em>Adrian Jones, Daoyu Zhang, Yuri Deigin&nbsp;and Steven C. Quay</em></p> <p>https://arxiv.org/abs/2108.08163</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Comprehensive discovery of CRISPR-targeted terminally redundant sequences in the human gut metagenome: viruses, plasmids, and more

<p>S1 Data</p> <p>Dataset including the discovered CRISPR spacers, direct repeats, protospacers, co-occurrence-based spacer clustering results, predicted protein sequences, built HMMs, database comparison results, phylogenetic analysis results, predicted targeting hosts, and CRISPR-targeted TR sequences.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Colorectal cancer and adenoma metagenomes

<p>Normalized counts of species identified by shotgun sequencing in 156 stool samples of 51 colorectal cancer patients, 54 patients with adenoma, and 51 controls.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Code and Data associated with "Discovery of positive and purifying selection in metagenomic time series of hypermutator microbial populations"

<p>Code and data sufficient to reproduce analyses in&nbsp;&quot;Discovery of positive and purifying selection in metagenomic time series of hypermutator microbial populations&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Files for publication "Microseek: A Protein-Based Metagenomic Pipeline for Virus Diagnostic and Discovery"

<p><strong>Context</strong></p> <p>These files correspond to the article&nbsp;&ldquo;Microseek: A Protein-Based Metagenomic Pipeline for Virus Diagnostic and Discovery&rdquo; submitted to Genes.</p> <p>&nbsp;</p> <p><strong>File content</strong></p> <ul> <li>input_data-empty_matrices:&nbsp;50M-read Tissues and Plasma matrices, no spike;</li> <li>input_data-matrices_spiked_known_viruses:&nbsp;50M-read Tissues and Plasma matrices spiked with six known virus at d1, d10, d100;</li> <li>input_data-matrices_spiked_neo_viruses:&nbsp;50M-read Tissues and Plasma matrices spiked with 3 Neopneumoviruses at d1 and d10;</li> <li>input_data-neo_viruses:&nbsp;Nucleotide and protein sequences of 3 Neopeumoviruses</li> <li>input_data-tick_sample:&nbsp;raw data of a Rhipicephalus tick sample known to be infected with the Cataloi Tick Quaranjavirus (CTQV)</li> <li>input_data-negative_control:&nbsp;raw data of the negative control (water)</li> <li>output_microseek:&nbsp;Microseek outputs, raw results and results after background filtration</li> </ul> <p>&nbsp;</p> <p><strong>File listing&nbsp;</strong></p> <pre><code class="language-bash">input_data-empty_matrices.tar.xz ├── plasma.fastq └── tissue.fastq input_data-matrices_spiked_known_viruses ├── d1 │   ├── spiked_plasma.fastq │   └── spiked_tissue.fastq ├── d10 │   ├── spiked_plasma.fastq │   └── spiked_tissue.fastq └── d100 ├── spiked_plasma.fastq └── spiked_tissue.fastq input_data-matrices_spiked_neo_viruses.tar.xz ├── d1 │   ├── plasma_spiked_with_neo1.fastq │   ├── plasma_spiked_with_neo2.fastq │   ├── plasma_spiked_with_neo3.fastq │   ├── tissue_spiked_with_neo1.fastq │   ├── tissue_spiked_with_neo2.fastq │   └── tissue_spiked_with_neo3.fastq └── d10 ├── plasma_spiked_with_neo1.fastq ├── plasma_spiked_with_neo2.fastq ├── plasma_spiked_with_neo3.fastq ├── tissue_spiked_with_neo1.fastq ├── tissue_spiked_with_neo2.fastq └── tissue_spiked_with_neo3.fastq input_data-neo_viruses.tar.xz ├── genes │   ├── neo_1.fasta │   ├── neo_2.fasta │   └── neo_3.fasta └── proteins ├── neo_1.fasta ├── neo_2.fasta └── neo_3.fasta input_data-tick_sample.tar.xz └── Cataloi_S1_R1_001.fastq.xz input_data-negative_control.tar.xz └── negative_control.fastq.xz output_microseek.tar.xz ├── empty_matrices │   ├── matrix_plasma │   └── matrix_tissue ├── matrices_spiked_known_viruses │   ├── filtered │   │   ├── d100_plasma │   │   ├── d100_tissue │   │   ├── d10_plasma │   │   ├── d10_tissue │   │   ├── d1_plasma │   │   └── d1_tissue │   └── non_filtered │   ├── d100_plasma │   ├── d100_tissue │   ├── d10_plasma │   ├── d10_tissue │   ├── d1_plasma │   └── d1_tissue ├── matrices_spiked_neo_viruses │   ├── filtered │   │   ├── plasma_spiked_with_neo1_at_d1 │   │   ├── plasma_spiked_with_neo1_at_d10 │   │   ├── plasma_spiked_with_neo2_at_d1 │   │   ├── plasma_spiked_with_neo2_at_d10 │   │   ├── plasma_spiked_with_neo3_at_d1 │   │   ├── plasma_spiked_with_neo3_at_d10 │   │   ├── tissue_spiked_with_neo1_at_d1 │   │   ├── tissue_spiked_with_neo1_at_d10 │   │   ├── tissue_spiked_with_neo2_at_d1 │   │   ├── tissue_spiked_with_neo2_at_d10 │   │   ├── tissue_spiked_with_neo3_at_d1 │   │   └── tissue_spiked_with_neo3_at_d10 │   └── non_filtered │   ├── plasma_spiked_with_neo1_at_d1 │   ├── plasma_spiked_with_neo1_at_d10 │   ├── plasma_spiked_with_neo2_at_d1 │   ├── plasma_spiked_with_neo2_at_d10 │   ├── plasma_spiked_with_neo3_at_d1 │   ├── plasma_spiked_with_neo3_at_d10 │   ├── tissue_spiked_with_neo1_at_d1 │   ├── tissue_spiked_with_neo1_at_d10 │   ├── tissue_spiked_with_neo2_at_d1 │   ├── tissue_spiked_with_neo2_at_d10 │   ├── tissue_spiked_with_neo3_at_d1 │   └── tissue_spiked_with_neo3_at_d10 ├── negative_control └── tick_sample </code></pre> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 3b2 Introduction to Python and Pandas

<p>Teaching data for&nbsp;practical session: &quot;3b2&nbsp;Introduction to Python and Pandas&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 3b1 Introduction to R and the Tidyverse

<p>Teaching data for&nbsp;practical session: &quot;3b1&nbsp;Introduction to R and the Tidyverse&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 2b Introduction to Github

<p>Teaching data for&nbsp;practical session: &quot;2b Introduction to GitHub&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a> or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code class="language-bash">tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 2c Introduction to AncientMetagenomeDir

<p>Teaching data for&nbsp;practical session: &quot;2c&nbsp;Introduction to AncientMetagenomeDir&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 5b Introduction to Phylogenomics

<p>Teaching data for&nbsp;practical session: &quot;5b&nbsp;Introduction to Phylogenomics&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 2d Introduction to nf-core/eager

<p>Teaching data for&nbsp;practical session: &quot;2d&nbsp;Introduction to nf-core/eager&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 4b Introduction to Genome Mapping

<p>Teaching data for&nbsp;practical session: &quot;4b&nbsp;Introduction to Genome Mapping&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 3c Introduction to Taxonomic Profiling

<p>Teaching data for&nbsp;practical session: &quot;3c&nbsp;Introduction to Taxonomic Profiling&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 4c Introduction to Genome Assembly

<p>Teaching data for&nbsp;practical session: &quot;4c&nbsp;Introduction to Genome Assembly&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record