Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
445
datasets available to search
ShareScore release 0.9.0
Dataset results
445 results for “taxonomic classification”
FIGURE 9 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 9. SEM illustrations of pygidium and apex of elytra, dorsal aspect, of Zeanillus species. A— Z. phyllobius; B— Z. punctigerus; C— Z. pellucidus; D— Z. lescheni. Legend: ea—apex of right elytron; ed 8 — apical seta; pgd—pygidium; suelytral suture. Scale bars = 0.05 mm.
FIGURE 10 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 10. SEM illustrations of structural features of legs of Zeanillus species, various aspects. A – E right protarsus and protibia, ventral aspect: A, D – E—females; B – C—males. A— Z. phyllobius; B— Z. punctigerus; C— Z. pellucidus; D— Z. pallidus; E— Z. brouni. F – G, H, J right and I, K left mesotibia: F— Z. phyllobius, ventral aspect; G— Z. punctigerus, ventral aspect; H— Z. pellucidus, medial aspect; I— Z. pallidus, ventral aspect; J— Z. lescheni, ventral aspect; K— Z. brouni, ventral aspect. L – M, O right and N, P – Q left metatibia: L— Z. phyllobius, ventral aspect; M— Z. punctigerus, ventral aspect; N— Z. pellucidus, medial aspect; O— Z. pallidus, ventral aspect; P— Z. lescheni, ventral aspect; Q— Z. brouni, medial aspect. Legend: ac—antenna cleaner; as—adhesive seta; asp—anterior spur; asr—anterior setal row; cls—clip seta; msb—mesotibial brush; msms—mesotibial modified seta; mss—mesotibial spur; mtb—metatibial brush; mtms—metatibial modified seta; mtsmetatibial spur; psp—posterior spur; psr—posterior setal row; sb—setal band; ta 1 – ta 4 — tarsomeres 1 – 4. Scale bars = 0.05 mm.
FIGURE 8 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 8. SEM illustrations of chaetotaxy of elytra, dorso-lateral aspect, of Zeanillus species. A— Z. punctigerus; B— Z. pallidus; C— Z. pellucidus; D— Z. lescheni. Legend: ed 2 — scutellar seta; ed 3 — 1 st discal seta; ed 4 – 5 — 2 nd discal seta; ed 6 – 7 — 3 d discal seta; ed 8 — apical seta; eo 1 – 9 — setae 1 – 9 from the umbilical series; smes—subapical marginal elytral seta; sssubapical sinuation. Scale bars = 0.2 mm.
FIGURE 6 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 6. SEM illustrations of prothorax, ventral aspect, of Zeanillus species. A— Z. phyllobius; B— Z. punctigerus; C— Z. pallidus; D— Z. pellucidus; E— Z. lescheni; F— Z. brouni. Legend: pas—prosternal ambulatory seta; pep—proepipleuron; pes—proepisternum; prcx—procoxa; ps—prosternum; psp—prosternal intercoxal process. Scale bars = 0.1 mm.
FIGURE 5 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 5. SEM illustrations of pronotum, dorsal aspect, of Zeanillus species. A— Z. phyllobius; B— Z. punctigerus; C— Z. pallidus; D— Z. pellucidus; E— Z. lescheni; F— Z. nanus. Legend: aps—apicolateral pronotal seta; bdt—basilateral pronotal denticle; bs—basilateral pronotal seta; bsn—basilateral pronotal sinuation; ed 2 — scutellar seta; ls—midlateral pronotal seta; mg—marginal pronotal gutter; sct—scutellum. Scale bars = 0.1 mm.
FIGURE 2 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 2. SEM illustrations of chaetotaxy of head and pronotum, dorso-lateral aspect, of Zeanillus species. A— Z. phyllobius; B— Z. pallidus; C— Z. pellucidus; D— Z. lescheni. Legend: aps—apicolateral pronotal seta; cs—clypeal seta; fsfrontal seta; ls—midlateral pronotal seta; pos—postorbital seta; ssa—anterior supraorbital seta; ssp—posterior supraorbital seta. Scale bars = 0.1 mm.
FIGURE 1 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 1. SEM illustrations of head, dorsal aspect, of Zeanillus species. A— Z. phyllobius; B— Z. punctigerus; C— Z. pallidus; D— Z. pellucidus; E— Z. lescheni; F— Z. nanus. Legend: cl—clypeus; cs—clypeal seta; fcc—fronto-clypeal carina; fs—frontal seta; lb—labrum; mp 3 — maxillary palpomere 3; mp 4 — maxillary palpomere 4; pos—postorbital seta; ssa—anterior supraorbital seta; ssp—posterior supraorbital seta. Scale bars = 0.1 mm.
FIGURE 4 in A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution
FIGURE 4. SEM illustrations of labial complex, ventral aspect, of Zeanillus species. A— Z. phyllobius; B— Z. punctigerus; C— Z. pallidus; D— Z. pellucidus; E— Z. lescheni; F— Z. brouni. Legend: alms—anterolateral mental seta; gsc—glossal sclerite; lms—lateral mental seta; lss—lateral submental seta; m—mentum; mp 3 — maxillary palpomere 3; mp 4 — maxillary palpomere 4; mt—mental tooth; mss—mental-submental suture; pms—paramedial mental seta; prss—primary basal submental seta; sm—submentum. Scale bars = 0.1 mm.
MARMICRODB database for taxonomic classification of (marine) metagenomes
<p><strong>UPDATE (April 2024):<br>Please note that the combined fasta file (MARMICRODB.faa.bz2) associated with version 1.0.0 was mistakenly uploaded from an early prototype of the database. </strong>As a result, some taxonomy identifiers appended to the sequence records are mismatched with the names.dmp, nodes.dmp, and the MARMICRODB_catalog.tsv files. <br><br><strong>Importantly, the taxon IDs in the FM-compressed Kaiju database are correct and valid for use with the names.dmp and nodes.dmp files. </strong>Any analysis using the compressed Kaiju database (MARMICRODB.fmi) from version 1.0.0 will remain valid. However, if you have used the combined fasta file (MARMICRODB.faa.bz2) to generate your own Kaiju database and you used the taxdump files from version 1.0.0 I cannot guarantee the validity of those results. It is <em><strong>strongly recommended</strong></em> that you rebuild the combined fasta file from scratch using the FTP links to the NCBI assemblies provided in the MARMICRODB_catalog.tsv file.</p> <p>The combined fasta file has been removed from this version of the MARMICRODB.</p> <p><strong>Disclaimer:<br></strong>MARMICRODB is optimized for metagenomic samples from the marine environment, in particular, planktonic microbes from the pelagic euphotic zone. We expect this database may also be useful for classifying other types of marine metagenomic samples (for example, mesopelagic, bathypelagic, or even benthic or marine host-associated), but it has not been tested for this purpose. The original purpose of this database was to quantify clades/ecotypes of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 in metagenomes from Tara Oceans Expedition and the GEOTRACES project. We carefully annotated and quality controlled genomes from these five groups, but the processing of the other marine taxa was automated and unsupervised. Taxonomy for other groups was copied over from the Genome Taxonomy Database version RS83 (GTDB) [19,39] and NCBI Taxonomy [23] so any inconsistencies in those databases at the time of access will be propagated to MARMICRODB. If the user’s goal is to focus on a particular organism/clade that we did not curate in the database then the user should spend some time curating those genomes (ie checking for contamination, dereplicating, building a genome phylogeny for custom taxonomy node assignment). Currently the custom taxonomy is hardcoded in the MARMICRODB.fmi index, but if users wish to modify MARMICRODB by adding or removing genomes, or reconfiguring taxonomic ranks the names.dmp and nodes.dmp files can easily be modified. However, the Kaiju index will need to be rebuilt, which will require a high-performance compute cluster.</p> <p><strong>Introduction:</strong><br>This sequence database (MARMICRODB) was introduced in the publication JW Becker, SL Hogle, K Rosendo, and SW Chisholm. 2019. Co-culture and biogeography of <em>Prochlorococcus</em> and SAR11. ISME J. doi:10.1038/s41396-019-0365-4. Please see the original publication and its associated supplementary material for the original description of this resource. </p> <p><strong>Motivation:</strong><br>We needed a reference database to annotate shotgun metagenomes from the Tara Oceans project [1] the <a href="https://www.geotraces.org/">GEOTRACES</a> cruises GA02, GA03, GA10, and GP13 and the HOT and BATS time series [2]. Our interests are primarily in quantifying and annotating the free-living, oligotrophic bacterial groups <em>Prochlorococcus</em>, <em>Pelagibacterales</em>/SAR11, SAR116, and SAR86 from these samples using the protein classifier tool Kaiju [3]. Kaiju’s sensitivity and classification accuracy depends on the composition of the reference database, and the highest sensitivity is achieved when the reference database contains a comprehensive representation of expected taxa from an environment/sample of interest. However, the speed of the algorithm decreases as the database size increases. Therefore, we aimed to create a reference database that maximized the representation of sequences from marine bacteria, archaea, and microbial eukaryotes while minimizing (but not excluding) the sequences from clinical, industrial, and terrestrial host-associated samples.</p> <p><strong>Results/Description:</strong><br>MARMICRODB consists of 56 million sequence non-redundant protein sequences from 18769 bacterial/archaeal/eukaryote genome and transcriptome bins and 7492 viral genomes optimized for use with the protein homology classifier Kaiju [3]. To ensure maximum representation of marine bacteria, archaea, and microbial eukaryotes, we included translated genes/transcripts from 5397 representative “specI” species clusters from the proGenomes database [4]; 113 transcriptomes from the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) [5]; 10509 metagenome assembled genomes from the Tara Oceans expedition [6,7], the Red Sea [8], the Baltic Sea [9], and other aquatic and terrestrial sources [10]; 994 isolate genomes from the Genomic Encyclopedia of Bacteria and Archaea [11]; 7492 viral genomes from NCBI RefSeq [12]; 786 bacterial and archaeal genomes from MarRef [13]; and 677 marine single cell genomes [14]. In order to annotate metagenomic reads at the clade/ecotype level (subspecies) for the focal taxa <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116, we generated custom MARMICRODB taxonomies based on curated genome phylogenies for each group. The curated phylogenies, Kaiju formatted Burrows-Wheeler index, translated genes, the custom taxonomy hierarchy, an <a href="https://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">interactive kronaplot of the taxonomic composition</a>, and scripts and instructions for how to use or rebuild the resource is available from <a href="https://doi.org/10.5281/zenodo.3520509">10.5281/zenodo.3520509</a>. </p> <p><strong>Methods:</strong><br>The curation and quality control of MARMICRODB single-cell, metagenome-assembled, and isolate genomes was performed as described in [15]. Briefly, we downloaded all MARMICRODB genomes as raw nucleotide assemblies from NCBI. We determined an initial genome taxonomy for these assemblies using checkM with the default lineage workflow [16]. All genome bins met the completion/contamination thresholds outlined in prior studies [7,17]. For single cell and metagenome-assembled genomes, especially those from Tara Oceans Mediterranean sea samples [18], we use the GTDB-Tk classification workflow [19] to verify the taxonomic fidelity of each genome bin. We then selected genomes with a checkM taxonomic assignment of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 for further analysis and confirmed taxonomic assignment using blast matches to known Prochlorococcus/Synechococcus ITS sequences and by matching 16S sequences to the SILVA database [20]. To refine our estimates of completeness/contamination of <em>Prochlorococcus</em> genome bins we created a custom set of 730 single copy protein families (available from <a href="https://doi.org/10.5281/zenodo.3719132">10.5281/zenodo.3719132</a>) from closed, isolate <em>Prochlorococcus</em> genomes [21] for quality assessments with checkM. For <em>Synechococcus</em> we used the CheckM taxonomic-specific workflow with the genus <em>Synechococcus</em>. After the custom CheckM quality control, we excluded any genome bins from downstream analysis that had an estimated quality < 30, defined as %completeness – 5x %contamination resulting in 18769 genome/transcriptome bins. We predicted genes in the resulting genome bins using prodigal [22] and excluded protein sequences with lengths less than 20 and greater than 20000 amino acids, removed non-standard amino acid residues, and condensed redundant protein sequences to a single representative sequence to which we assigned a lowest common ancestor (LCA) taxonomy identifier from the NCBI taxonomy database [23]. The resulting protein sequences were compiled and used to build a Kaiju [3] search database. </p> <p>The above filtering criteria resulted in 605 <em>Prochlorococcus</em>, 96 <em>Synechococcus</em>, 186 SAR11/<em>Pelagibacterales</em>, 60 SAR86, and 59 SAR116 high-quality genome bins. We constructed a high-quality fixed reference phylogenetic tree for each taxonomic group based on genomes manually selected for completeness and the phylogenetic diversity. For example the <em>Prochlorococcus</em> and <em>Synechococcus</em> genomes for the fixed reference phylogeny are estimated > 90% complete, and SAR11 genomes are estimated > 70% complete. We created multiple sequence alignments of phylogenetically conserved genes from these genomes using the GTDB-Tk pipeline [19] with default settings. The pipeline identifies conserved proteins (120 bacterial proteins) and generates concatenated multi-protein alignments [17] from the genome assemblies using hmmalign from the hmmer software suite. We further filtered the resulting alignment columns using the bacterial and archaeal alignment masks from [17] (http://gtdb.ecogenomic.org/downloads). We removed columns represented by fewer than 50% of all taxa and/or columns with no single amino acid residue occurring at a frequency greater than 25%. We trimmed the alignments using trimal [24] with the automated -gappyout option to trim columns based on their gap distribution. We inferred reference phylogenies using multithreaded RAxML [25] with the GAMMA model of rate heterogeneity, empirically determined base frequencies, and the LG substitution model [26](PROTGAMMALGF). Branch support is based on 250 resampled bootstrap trees. This tree was then pruned to only allow a maximum average distance to the closest leaf (ADCL) of 0.003 to reduce the phylogenetic redundancy in the tree [27]. We then “placed” genomes that either did not pass completeness threshold or were considered phylogenetically redundant by ADCL within the fixed reference phylogeny for each group using pplacer [28] representing each placed genome as a pendant edge in the final tree. We then examined the resulting tree and manually selected clade/ecotype cutoffs to be as consistent as possible with clade definitions previously outlined for these groups [29–32]. We then gave clades from each taxonomic group custom taxonomic identifiers and we added these identifiers to the MARMICRODB Kaiju taxonomic hierarchy.</p> <p><strong>Software/databases used:</strong><br>checkM v1.0.11[16]<br>HMMERv3.1b2 (http://hmmer.org/)<br>prodigal v2.6.3 [22]<br>trimAl v1.4.rev22 [24]<br>AliView v1.18.1 [33] [34]<br>Phyx v0.1 [35]<br>RAxML v8.2.12 [36]<br>Pplacer v1.1alpha [28]<br>GTDB-Tk v0.1.3 [19]<br>Kaiju v1.6.0 [34]<br>GTDB RS83 (https://data.ace.uq.edu.au/public/gtdb/data/releases/release83/83.0/)<br>NCBI Taxonomy (accessed 2018-07-02) [23]<br>TIGRFAM v14.0 [37]<br>PFAM v31.0 [38]</p> <p><strong>Use example:</strong><br>Because we used custom taxonomic MARMICRODB users will find many reads assigned to non-standard NCBI taxonomy identifiers. However, these reads are easily parsable using the custom names.dmp and nodes.dmp files included with the database. We include a brief description of how to do this below.</p> <p>I typically run Kaiju like:</p> <pre><code>kaiju -z 20 -a greedy -e 5 -m 11 -s 65 -E 0.05 -x \ -t nodes.dmp -f MARMICRODB.fmi \ -i inputfile_R1.fastq.gz \ -j inputfile_R2.fastq.gz \ -o MYOUTPUT.kaiju</code></pre> <p>To obtain a parseable report that lists the custom taxonomic ranks from the nodes.dmp and names.dmp files run kaiju2krona on the output.</p> <pre><code>kaiju2krona -t nodes.dmp -n names.dmp -i MYOUTPUT.kaiju -o MYOUTPUT.kaiju.krona </code></pre> <p>This report shows counts assigned to each node in the custom taxonomy and will also include the names for each rank. You can easily parse this programmatically using a scripting language like python or by using unix utilities.</p> <p><strong>File descriptions:</strong></p> <p><em>MARMICRODB_catalog.tsv</em><br>Tabular file of NCBI assembly accessions and associated taxonomic information for every genome in MARMICRODB. Also includes literature references for each genome where available. Header description:</p> <ol> <li>genome: Unique identifier for each genome</li> <li>full_name: full organism name where available</li> <li>source: literature reference where available</li> <li>taxid: NCBI taxonomy ID for the assembly accession</li> <li>MARMICRODBtaxid: taxonomy ID used in the custom Kaiju database</li> <li>lineage_assignment: taxonomic lineage assignment from NCBI</li> <li>domain: archaea, bacteria, or eukaryote</li> <li>taxgroup: short descriptive group</li> <li>taxclade: higher resolution clade assignment where available</li> <li>habitat_source: whether genome derives from marine or aquatic source</li> <li>sequence_type: isolate, single cell genome (sag), metagenome assembled genome (mag), or transcriptome in case of eukaryotes</li> <li>assembly_ftp: NCBI ftp for assembly</li> <li>gbk_acc: assembly genbank or refseq accession number</li> <li>gtdb_taxonomy: taxonomic lineage assignment from GTDB-Tk v0.1.3 against GTDB v83</li> </ol> <p><em>MARMICRODB_kronaplot.html</em><br><a href="http://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">Interactive Kronaplot</a> for the exploration of taxonomic composition of MARMICRODB</p> <p><em>MARMICRODB.faa.bz2</em><br>Fasta file of all protein sequences in MARMICRODB</p> <p><em>scripts.tar.gz</em><br>directory containing scripts for generating Kaiju formatted database</p> <p><em>phylogenies.tar.gz</em><br>directory containing detailed phylogenies for SAR11, Prochlorococcus, SAR86, and SAR116</p> <p><em>MARMICRODB.fmi</em><br>Kaiju index for MARMICRODB</p> <p><em>nodes.dmp</em><br>nodes file for taxonomic assignment with Kaiju</p> <p><em>names.dmp</em><br>names file for generating Kaiju reports</p> <p><strong>References:</strong><br>1. Karsenti E, Acinas SG, Bork P, Bowler C, De Vargas C, Raes J, et al. A holistic approach to marine eco-systems biology. PLoS Biol. 2011;9: e1001177.<br>2. Biller SJ, Berube PM, Dooley K, Williams M, Satinsky BM, Hackl T, et al. Marine microbial metagenomes sampled across space and time. Scientific Data. 2018;5: 180176.<br>3. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>4. Mende DR, Letunic I, Huerta-Cepas J, Li SS, Forslund K, Sunagawa S, et al. proGenomes: a resource for consistent functional and taxonomic annotations of prokaryotic genomes. Nucleic Acids Res. 2017;45: D529–D534.<br>5. Keeling PJ, Burki F, Wilcox HM, Allam B, Allen EE, Amaral-Zettler LA, et al. The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): illuminating the functional diversity of eukaryotic life in the oceans through transcriptome sequencing. PLoS Biol. 2014;12: e1001889.<br>6. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>7. Tully BJ, Graham ED, Heidelberg JF. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Sci Data. 2018;5: 170203.<br>8. Haroon MF, Thompson LR, Parks DH, Hugenholtz P, Stingl U. A catalogue of 136 microbial draft genomes from Red Sea metagenomes. Sci Data. 2016;3: 160050.<br>9. Hugerth LW, Larsson J, Alneberg J, Lindh MV, Legrand C, Pinhassi J, et al. Metagenome-assembled genomes uncover a global brackish microbiome. Genome Biol. 2015;16: 279.<br>10. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017.<br>11. Mukherjee S, Seshadri R, Varghese NJ, Eloe-Fadrosh EA, Meier-Kolthoff JP, Göker M, et al. 1,003 reference genomes of bacterial and archaeal isolates expand coverage of the tree of life. Nat Biotechnol. 2017;35: 676–683.<br>12. Haft DH, DiCuccio M, Badretdin A, Brover V, Chetvernin V, O’Neill K, et al. RefSeq: an update on prokaryotic genome annotation and curation. Nucleic Acids Res. 2018;46: D851–D860.<br>13. Klemetsen T, Raknes IA, Fu J, Agafonov A, Balasundaram SV, Tartari G, et al. The MAR databases: development and implementation of databases specific for marine metagenomics. Nucleic Acids Res. 2018;46: D692–D699.<br>14. Berube PM, Biller SJ, Hackl T, Hogle SL, Satinsky BM, Becker JW, et al. Single cell genomes of Prochlorococcus, Synechococcus, and sympatric microbes from diverse marine environments. Scientific Data. 2018;5: 180154.<br>15. Becker JW, Hogle SL, Rosendo K, Chisholm SW. Co-culture and biogeography of Prochlorococcus and SAR11. ISME J. 2019. doi:10.1038/s41396-019-0365-4<br>16. Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25: 1043–1055.<br>17. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017. doi:10.1038/s41564-017-0012-7<br>18. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>19. Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2019. doi:10.1093/bioinformatics/btz848<br>20. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41: D590–6.<br>21. Biller SJ, Berube PM, Berta-Thompson JW, Kelly L, Roggensack SE, Awad L, et al. Genomes of diverse isolates of the marine cyanobacterium Prochlorococcus. Sci Data. 2014;1: 140034.<br>22. Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics. 2010;11: 119.<br>23. Federhen S. The NCBI Taxonomy database. Nucleic Acids Res. 2012;40: D136–43.<br>24. Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25: 1972–1973.<br>25. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>26. Le SQ, Gascuel O. An improved general amino acid replacement matrix. Mol Biol Evol. 2008;25: 1307–1320.<br>27. Matsen FA, Gallagher A, McCoy C. Minimizing the average distance to a closest leaf in a phylogenetic tree. arXiv [q-bio.PE]. 2012. Available: http://arxiv.org/abs/1205.6867<br>28. Matsen F a., Kodner RB, Armbrust EV. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree. BMC Bioinformatics. 2010;11: 538.<br>29. Biller SJ, Berube PM, Lindell D, Chisholm SW. Prochlorococcus: the structure and function of collective diversity. Nat Rev Microbiol. 2014;13: 13–27.<br>30. Giovannoni SJ. SAR11 Bacteria: The Most Abundant Plankton in the Oceans. Ann Rev Mar Sci. 2016.<br>31. Dupont CL, Rusch DB, Yooseph S, Lombardo M-J, Alexander Richter R, Valas R, et al. Genomic insights to SAR86, an abundant and uncultivated marine bacterial lineage. ISME J. 2012;6: 1186–1199.<br>32. Yang S-J, Kang I, Cho J-C. Expansion of Cultured Bacterial Diversity by Large-Scale Dilution-to-Extinction Culturing from a Single Seawater Sample. Microb Ecol. 2016;71: 29–43.<br>33. Larsson A. AliView: a fast and lightweight alignment viewer and editor for large datasets. Bioinformatics. 2014;30: 3276–3278.<br>34. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>35. Brown JW, Walker JF, Smith SA. Phyx: phylogenetic tools for unix. Bioinformatics. 2017;33: 1886–1888.<br>36. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>37. Haft DH, Selengut JD, Richter RA, Harkins D, Basu MK, Beck E. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 2013;41: D387–95.<br>38. Finn RD, Coggill P, Eberhardt RY, Eddy SR, Mistry J, Mitchell AL, et al. The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res. 2016;44: D279–85.<br>39. Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil P-A, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol. 2018.<br> </p>
Evaluation of tools for the taxonomical classification of viruses
<p>50G (1942_50GL), 500G (1943_500GL) and 1000G (1752_1000GL) sets are simulated reads previously reported Tangherlini et al. [1], and represent large reads like contigs.</p> <p>FISH-I (FISH_1_Q_sS_dup_HNoM_RbNoM), PB3(PB3_Q_sS_dup_HNoM_RbNoM), I5-8(ISA32_R1_Q_CDhit_HmNoM_RbNoM and ISA32_R1_Q_CDhit_HmNoM_RbNoM) and 121-1 (ISA44_R1_Q_CDhit_HmNoM_RbNoM and ISA44_R1_Q_CDhit_HmNoM_RbNoM) sets are real metagenomics reads previously reported Taboada et al. [2,3].</p> <p>Eukaryotic, Prokaryotic, Unclassified, Bacterial and Human sets are simulated reads, generated by Grinder [4] software and represent short reads like those of Illumina technology.<br> </p>
Linked collectors and determiners for: A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution.
Natural history specimen data linked to collectors and determiners held within, "A taxonomic review of the anilline genus Zeanillus Jeannel (Coleoptera: Carabidae: Bembidiini) of New Zealand, with descriptions of seven new species, re-classification of the species, and notes on their biogeography and evolution". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/b7651a7b-df24-4347-9585-0633b88dba51">https://bionomia.net/dataset/b7651a7b-df24-4347-9585-0633b88dba51</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/b7651a7b-df24-4347-9585-0633b88dba51">https://gbif.org/dataset/b7651a7b-df24-4347-9585-0633b88dba51</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: Classification, Natural History, and Evolution of the Subfamily Peloniinae O (Coleoptera: Cleroidea: Cleridae). Part IX. Taxonomic revision of the New World genus Muisca S.
Natural history specimen data linked to collectors and determiners held within, "Classification, Natural History, and Evolution of the Subfamily Peloniinae O (Coleoptera: Cleroidea: Cleridae). Part IX. Taxonomic revision of the New World genus Muisca S". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/bedafe83-1cab-4329-8c1d-0152874f3907">https://bionomia.net/dataset/bedafe83-1cab-4329-8c1d-0152874f3907</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/bedafe83-1cab-4329-8c1d-0152874f3907">https://gbif.org/dataset/bedafe83-1cab-4329-8c1d-0152874f3907</a>. Formatted as a Frictionless Data package.
Figures 5-8. Habitus and male genitalia. 5 in Classification, natural history, and evolution of the subfamily Peloniinae Opitz (Coleoptera: Cleroidea: Cleridae). Part VI. New taxonomic placement for Pelonium sexpunctatum Kirsch
Figures 5-8. Habitus and male genitalia. 5) Habitus of D. sexpunctatus. 6) Habitus of D. gallerucoides. 7) Habitus of D. pallidus. 8) Male genitalia of D. pallidus.
Figures 1-4 in Classification, natural history, and evolution of the subfamily Peloniinae Opitz (Coleoptera: Cleroidea: Cleridae). Part VI. New taxonomic placement for Pelonium sexpunctatum Kirsch
Figures 1-4. Structures and map of Diutius species. 1) Antenna of D. sexpunctatus. 2) Pronotum of D. sexpunctatus. 3) Pronotum of D. gallerucoides. 4) Geographic distribution of Diutius species as noted.
Investigation of machine learning algorithms for taxonomic classification of marine metagenomes
<p>Training, testing, and blind datasets used for machine learning algorithms for taxonomic classification of marine metagenomes:</p> <ol> <li><strong>K12.kmers.txt</strong> - 12bp k-mer vocabulary constructed by Jellyfish v1.1.11 from 47,894 genomes in GTDB release 202</li> <li><strong>MarRef_1.6.tsv</strong> - Metadata file downloaded from MarRef v1.6</li> <li><strong>MarRef.genustrain.fasta</strong> - Training set from MarRef v1.6 (seed=808) used for genus classification</li> <li><strong>MarRef.genustest.fasta</strong> - Testing set from MarRef v1.6 (seed=747) used for genus classification </li> <li><strong>MarRef.speciestrain.fasta</strong> - Training set from MarRef v1.6 (seed=808) used for species classification</li> <li><strong>MarRef.speciestest.fasta</strong> - Testing set from MarRef v1.6 (seed=747) used for species classification</li> <li><strong>MarRef.traintest.key.tsv</strong> - Table containing MarRef accession, GenBank accession, GenBank taxonomy ID, taxonomic information, and labels used for species and genus testing and training</li> <li><strong>anonymous_reads_*.fq</strong> - Blind datasets (1-10) in interleaved fastq format</li> <li><strong>reads_mapping_*.tsv</strong> - Key for blind datasets 1-10. Each sequence header is mapped to its corresponding MarRef accession and NCBI taxonomic ID.</li> </ol>
Taxonomic classification
<p>Refseq tailed phages with fasta headers containing family and genus names</p>
Plate 6 in Taxonomic Revision and Classification of Extant Holococcolithophores Previously Placed in the Genus Anthosphaera Kamptner emend. Kleijne 1991
Plate 6. Syracosphaera elevata Archontikis, Young, Cros sp. nov. HOL Scale bars = 1μm 1. Holotype; flattened coccosphere with high-arched CFCs and tiny and partially calcified BCs. 2. Paratype; coccosphere with CFCs that bear distal processes made of delicate microcrystals (arrow a), and BCs with a loose assembly of irregularly arranged microcrystals at their distal cover (arrow b). 3. Paratype; BCs' central structure may lose its crystal assembly (arrow). 4. Paratype.
Plate 5 in Taxonomic Revision and Classification of Extant Holococcolithophores Previously Placed in the Genus Anthosphaera Kamptner emend. Kleijne 1991
Plate 5. Syracosphaera rotaconica Archontikis, Young, Cros sp. nov. HOL Scale bars = 1μm 1. Holotype; flattened coccosphere with chiral-conical BCs that usually show a small apical boss on top of the central distal cover. BCs' central structure is made of numerous crystal struts, progressively meeting and forming the distal cover. 2. Detail of figure 1 showing the ultrastructure of BCs. 3. Paratype; coccosphere showing the delicate structure of CFCs (see arrow). 4. Paratype.
Plate 2 in Taxonomic Revision and Classification of Extant Holococcolithophores Previously Placed in the Genus Anthosphaera Kamptner emend. Kleijne 1991
Plate 2. Syracosphaera marginiporata Knappertsbusch HOL Scale bars = 1μm 1. Complete coccosphere with CFCs that show distal processes with straight sides (arrow a) and lateral columns of crystallites (arrow b). 2. Collapsed coccosphere with BCS that leave multiple openings at the central structure (arrow a) and a broken CFC in which the column is attached to the rim (arrow b). 3. Flattened coccosphere with CFCs that bear high processes with rounded to somewhat straight sides (arrow a), sustained by columns of large crystallites (arrow b). 4. Detailed view of specimen showing a broken CFCs with the two columns of crystals (see arrow) attached to the rim and next to the broken distal process (see arrow).
Fig. 2 in Taxonomic Revision and Classification of Extant Holococcolithophores Previously Placed in the Genus Anthosphaera Kamptner emend. Kleijne 1991
Fig. 2. Schematic representation and terminology of Anthosphaera coccosphere and holococcolith types.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.