Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
52
datasets available to search
ShareScore release 0.9.0
Dataset results
52 results for “taxonomic database”
Taxonomic and ecological database of trees of Western Ghats - TreeGhatsData
<p><em>TreeGhatsData</em> is a compilation of lists of tree taxa found in Western Ghats, South India:</p> <ul> <li>taxa for which the word "tree" appears in habit description in the book <em>Flowering plants of the Western Ghats</em> edited by the Tropical Botanic Garden Research Institute (TBGRI), including planted or cultivated taxa (Nayar, Beegam, and Sibi. 2014);</li> <li>tree taxa described after 2014 in journal articles;</li> <li>taxon names used in forest surveys published by the French Institute of Pondicherry (IFP), in journal articles from 2000, and in the Atlas of endemics of the Western Ghats (Ramesh and Pascal 1997);</li> <li>taxon names reported with "tree" habit in Indian Biodiversity Portal (http://indiabiodiversity.org/).</li> </ul> <p>For each plant name, <em>TreeGhatsData</em> includes the following taxonomic information: family, genus epithet, species epithet, infrataxon rank, infrataxon epithet, authority. Both the family name used in TBGRI book and the corresponding family name according to Angiosperm Phylogeny Group system III (APGIII; Bremer et al. 2009) are provided.</p> <p><em>TreeGhatsData</em> includes the taxonomic status, the reference name and the authority according to TBGRI flora, along with taxonomic status from The Plant List version 1.1 (http://www.theplantlist.org/). From these two sources, a taxonomic status is suggested for each taxon name, with corresponding reference names and authorities.</p> <p><em>TreeGhatsData</em> also includes ecological and biogeographic information from TBGRI and completed by the botanists of French Institute of Pondicherry (IFP).</p> <p>Because most vegetation surveys do not provide taxon names at infraspecific level, <em>TreeGhatsData</em> includes both the infraspecific taxa mentioned in Western Ghats and the corresponding specific binomial names.</p> <p><em>TreeGhatsData</em> is provided as a CSV file with comma separator.</p> <p><strong>Related references</strong></p> <p>Bremer, B., Bremer, K., Chase, M. W., Fay, M. F., Reveal, J. L., Soltis, D. E., Soltis, P. S., Stevens, P. F., Anderberg, A. A., Moore, M. J., Olmstead, R. G., Rudall, P. J., Sytsma, K. J., Tank, D. C., Wurdack, K., Xiang, J. Q. Y. & Zmarzty, S. (2009) An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG III. Botanical Journal of the Linnean Society, 161, 105-121.</p> <p>Nayar, T., Rasiya Beegam, A. & Sibi, M. (2014) Flowering plants of the Western Ghats, India, Volume 1 Dicots; Volume 2 Monocots. Jawaharlal Nehru Tropical Botanic Garden and Research Institute.</p> <p>Ramesh, B. & Pascal, J.-P. (1997) Atlas of endemics of the Western Ghats (India): distribution of tree species in the evergreen and semi-evergreen forests. French Institute of Pondicherry, Pondicherry, India.</p>
The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.
<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Université Claude Bernard Lyon 1, VetAgro Sup, Research Team “Bacterial Opportunistic Pathogens and Environment” (BPOE), 69280 Marcy L’Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bonté et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer’s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O'LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project “Iouqmer” EST 2016/1/120, l'Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rhône, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: Céline COLINON, Romain MARTI, Emilie BOURGEOIS, Sébastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257–4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161–168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bonté, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153–164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>
Let the giants be named - Taxonomic description database for the isopod genus Bathynomus for the DELTA system
<p><span>Morphological characters have been extracted from recent relevant literature on <em>Bathynomus</em>, mainly from the six previously known species from the Atlantic Ocean and gulf of Mexico </span><span>(Magalhães and Young 2003; Lowry and Dempsey 2006; Shipley et al. 2016; Huang et al. 2022)</span><span>. To simultaneously maximize compactness and taxonomic value of the species description, mainly characters that have previously been interpreted as diagnostic at the species level have been included, omitting others which played no role in species delimitation. A taxonomic database has been newly set up using DELTA </span><span>(Dallwitz 1980, 1993; Dallwitz et al. 2006)</span><span>. Character states for all presently known Atlantic congeners have been scored to generate a natural language description and a new identification key </span><span>(Dallwitz 1974)</span><span> for the region as well as a species diagnosis. The species diagnosis has been composed based on the characters used in the identification key.</span></p> <p> </p>
DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region
<p>This dataset has been prepared as part of the Interreg Alpine Space project Eco-AlpsWater (ASP569) - <em>Innovative Ecological Assessment and Water Management Strategy for the Protection of Ecosystem Services in Alpine Lakes and Rivers</em>, <a href="https://www.alpine-space.eu/projects/eco-alpswater/en/home">https://www.alpine-space.eu/projects/eco-alpswater/en/home</a></p> <p>Individual archives include 16S rRNA (cyanobacteria) and 18S rRNA (microalgae) FASTA sequences and associated blastn results obtained from the high throughput sequencing of plankton and biofilm bulk/eDNA samples collected in 2019 in 37 lakes and 22 rivers across the Alpine region. These are supporting files for the paper by Salmaso et al., 2022. DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region. Science of the Total Environment, in press.</p>
European Ivies (Hedera L., Araliaceae) Point Occurrence Database with Taxonomic Certainty
<p>We present two databases and six spatial layers recording biodiversity information of the six species of ivies (<em>Hedera </em>L., Araliaceae) native to W Europe (<em>Hedera azorica, H. canariensis, H. helix, H. hiberncia, H. iberica, H. maderensis)</em>. Each database covers the entire native distribution of each species. Therefore, the databases document the distribution and occurrence of all the European <em>Hedera </em>taxa except for <em>H. pastuchovii </em>subsp.<em> cypria</em> which is a restricted endemic of the south-west of the island of Cyprus. </p> <ul> <li>The first database (TaxRev) includes taxonomic, geographic and habitat information from the morphological revision of 2,276 individuals from 1,280 populations. 866 of the records also included point-occurrence data. This database represents the entire native distribution and the morphological variation of each species.</li> <li>The second database (MixOcc) includes the spatial-point occurrence of the six species across their entire native distribution ranges. This database was compiled with the 880 records from the TaxRev database (records with high taxonomic certainty, as they all were examined by the taxonomist of the genus) plus 2,372 records from curated online databases selected from the European regions with low expected taxonomic uncertainty (C and E Europe and the Macaronesian Islands). As a result the database have high taxonomic quality (certainty and coverage) and good geographical coverage for Europe at a large-scale except for France and Ireland.</li> <li>The uploaded files related to the TaxRev database are as follows:</li> </ul> <p>Hedera_TaxRevDatabase_Field description: a cvs file with the description of the 71 variables included in the database</p> <p>Hedera_TaxRevDatabase_Records: a cvs file with the database (71 variables, 1,280 records)</p> <ul> <li>The uploaded files related to the MixOcc database are as follows:</li> </ul> <p>Hedera_MixOccDatabase_Field description: a cvs file with the description of the 11 variables included in the database</p> <p>Hedera_MixDatabase_Records: a cvs file with the database (11 variables, 3,252 records)</p> <p>Finally, we also upload 20 layers including the point-occurrence maps obtained from the MixOcc database. Six species maps (one per species), five additional maps of <em>H. canariensis</em> (one per island), eight additional maps of <em>H. azorica</em> (one per island) and a combined map including the six species. In all of them, we distinguish the records from individuals morphologically reviewed by the taxonomist of the genus and those obtained from online repositories and not reviewed by the taxonomist:</p> <ul> <li>Hedera azorica_MixOccDatabase_Map</li> <li>Hedera_azorica_map_Corvo</li> <li>Hedera_azorica_map_Faial</li> <li>Hedera_azorica_map_Flores</li> <li>Hedera_azorica_map_Graciosa</li> <li>Hedera_azorica_map_Pico</li> <li>Hedera_azorica_map_Santa Maria</li> <li>Hedera_azorica_map_Sao Jorge</li> <li>Hedera_azorica_map_Sao Miguel</li> <li>Hedera_azorica_map_Terceira</li> <li>Hedera canariensis_MixOccDatabase_Map</li> <li>Hedera_canariensis_map_El Hierro</li> <li>Hedera_canariensis_map_Gran Canaria</li> <li>Hedera_canariensis_map_La Gomera</li> <li>Hedera_canariensis_map_La Palma</li> <li>Hedera_canariensis_map_Tenerife</li> <li>Hedera helix_MixOccDatabase_Map</li> <li>Hedera hibernica_TaxRevDatabase_Map</li> <li>Hedera iberica_TaxRevDatabase_Map</li> <li>Hedera maderensis_MixOccDatabase_Map</li> <li>Hedera_MixOccDatabase_Map</li> </ul> <p>The records which allow us to improve geographic coverage without compromising taxonomic certainty. The databases and the resulting spatial layers have high taxonomic and geographic certainty and a good geographic coverage for ivies in Europe.</p>
MARMICRODB database for taxonomic classification of (marine) metagenomes
<p><strong>UPDATE (April 2024):<br>Please note that the combined fasta file (MARMICRODB.faa.bz2) associated with version 1.0.0 was mistakenly uploaded from an early prototype of the database. </strong>As a result, some taxonomy identifiers appended to the sequence records are mismatched with the names.dmp, nodes.dmp, and the MARMICRODB_catalog.tsv files. <br><br><strong>Importantly, the taxon IDs in the FM-compressed Kaiju database are correct and valid for use with the names.dmp and nodes.dmp files. </strong>Any analysis using the compressed Kaiju database (MARMICRODB.fmi) from version 1.0.0 will remain valid. However, if you have used the combined fasta file (MARMICRODB.faa.bz2) to generate your own Kaiju database and you used the taxdump files from version 1.0.0 I cannot guarantee the validity of those results. It is <em><strong>strongly recommended</strong></em> that you rebuild the combined fasta file from scratch using the FTP links to the NCBI assemblies provided in the MARMICRODB_catalog.tsv file.</p> <p>The combined fasta file has been removed from this version of the MARMICRODB.</p> <p><strong>Disclaimer:<br></strong>MARMICRODB is optimized for metagenomic samples from the marine environment, in particular, planktonic microbes from the pelagic euphotic zone. We expect this database may also be useful for classifying other types of marine metagenomic samples (for example, mesopelagic, bathypelagic, or even benthic or marine host-associated), but it has not been tested for this purpose. The original purpose of this database was to quantify clades/ecotypes of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 in metagenomes from Tara Oceans Expedition and the GEOTRACES project. We carefully annotated and quality controlled genomes from these five groups, but the processing of the other marine taxa was automated and unsupervised. Taxonomy for other groups was copied over from the Genome Taxonomy Database version RS83 (GTDB) [19,39] and NCBI Taxonomy [23] so any inconsistencies in those databases at the time of access will be propagated to MARMICRODB. If the user’s goal is to focus on a particular organism/clade that we did not curate in the database then the user should spend some time curating those genomes (ie checking for contamination, dereplicating, building a genome phylogeny for custom taxonomy node assignment). Currently the custom taxonomy is hardcoded in the MARMICRODB.fmi index, but if users wish to modify MARMICRODB by adding or removing genomes, or reconfiguring taxonomic ranks the names.dmp and nodes.dmp files can easily be modified. However, the Kaiju index will need to be rebuilt, which will require a high-performance compute cluster.</p> <p><strong>Introduction:</strong><br>This sequence database (MARMICRODB) was introduced in the publication JW Becker, SL Hogle, K Rosendo, and SW Chisholm. 2019. Co-culture and biogeography of <em>Prochlorococcus</em> and SAR11. ISME J. doi:10.1038/s41396-019-0365-4. Please see the original publication and its associated supplementary material for the original description of this resource. </p> <p><strong>Motivation:</strong><br>We needed a reference database to annotate shotgun metagenomes from the Tara Oceans project [1] the <a href="https://www.geotraces.org/">GEOTRACES</a> cruises GA02, GA03, GA10, and GP13 and the HOT and BATS time series [2]. Our interests are primarily in quantifying and annotating the free-living, oligotrophic bacterial groups <em>Prochlorococcus</em>, <em>Pelagibacterales</em>/SAR11, SAR116, and SAR86 from these samples using the protein classifier tool Kaiju [3]. Kaiju’s sensitivity and classification accuracy depends on the composition of the reference database, and the highest sensitivity is achieved when the reference database contains a comprehensive representation of expected taxa from an environment/sample of interest. However, the speed of the algorithm decreases as the database size increases. Therefore, we aimed to create a reference database that maximized the representation of sequences from marine bacteria, archaea, and microbial eukaryotes while minimizing (but not excluding) the sequences from clinical, industrial, and terrestrial host-associated samples.</p> <p><strong>Results/Description:</strong><br>MARMICRODB consists of 56 million sequence non-redundant protein sequences from 18769 bacterial/archaeal/eukaryote genome and transcriptome bins and 7492 viral genomes optimized for use with the protein homology classifier Kaiju [3]. To ensure maximum representation of marine bacteria, archaea, and microbial eukaryotes, we included translated genes/transcripts from 5397 representative “specI” species clusters from the proGenomes database [4]; 113 transcriptomes from the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) [5]; 10509 metagenome assembled genomes from the Tara Oceans expedition [6,7], the Red Sea [8], the Baltic Sea [9], and other aquatic and terrestrial sources [10]; 994 isolate genomes from the Genomic Encyclopedia of Bacteria and Archaea [11]; 7492 viral genomes from NCBI RefSeq [12]; 786 bacterial and archaeal genomes from MarRef [13]; and 677 marine single cell genomes [14]. In order to annotate metagenomic reads at the clade/ecotype level (subspecies) for the focal taxa <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116, we generated custom MARMICRODB taxonomies based on curated genome phylogenies for each group. The curated phylogenies, Kaiju formatted Burrows-Wheeler index, translated genes, the custom taxonomy hierarchy, an <a href="https://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">interactive kronaplot of the taxonomic composition</a>, and scripts and instructions for how to use or rebuild the resource is available from <a href="https://doi.org/10.5281/zenodo.3520509">10.5281/zenodo.3520509</a>. </p> <p><strong>Methods:</strong><br>The curation and quality control of MARMICRODB single-cell, metagenome-assembled, and isolate genomes was performed as described in [15]. Briefly, we downloaded all MARMICRODB genomes as raw nucleotide assemblies from NCBI. We determined an initial genome taxonomy for these assemblies using checkM with the default lineage workflow [16]. All genome bins met the completion/contamination thresholds outlined in prior studies [7,17]. For single cell and metagenome-assembled genomes, especially those from Tara Oceans Mediterranean sea samples [18], we use the GTDB-Tk classification workflow [19] to verify the taxonomic fidelity of each genome bin. We then selected genomes with a checkM taxonomic assignment of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 for further analysis and confirmed taxonomic assignment using blast matches to known Prochlorococcus/Synechococcus ITS sequences and by matching 16S sequences to the SILVA database [20]. To refine our estimates of completeness/contamination of <em>Prochlorococcus</em> genome bins we created a custom set of 730 single copy protein families (available from <a href="https://doi.org/10.5281/zenodo.3719132">10.5281/zenodo.3719132</a>) from closed, isolate <em>Prochlorococcus</em> genomes [21] for quality assessments with checkM. For <em>Synechococcus</em> we used the CheckM taxonomic-specific workflow with the genus <em>Synechococcus</em>. After the custom CheckM quality control, we excluded any genome bins from downstream analysis that had an estimated quality < 30, defined as %completeness – 5x %contamination resulting in 18769 genome/transcriptome bins. We predicted genes in the resulting genome bins using prodigal [22] and excluded protein sequences with lengths less than 20 and greater than 20000 amino acids, removed non-standard amino acid residues, and condensed redundant protein sequences to a single representative sequence to which we assigned a lowest common ancestor (LCA) taxonomy identifier from the NCBI taxonomy database [23]. The resulting protein sequences were compiled and used to build a Kaiju [3] search database. </p> <p>The above filtering criteria resulted in 605 <em>Prochlorococcus</em>, 96 <em>Synechococcus</em>, 186 SAR11/<em>Pelagibacterales</em>, 60 SAR86, and 59 SAR116 high-quality genome bins. We constructed a high-quality fixed reference phylogenetic tree for each taxonomic group based on genomes manually selected for completeness and the phylogenetic diversity. For example the <em>Prochlorococcus</em> and <em>Synechococcus</em> genomes for the fixed reference phylogeny are estimated > 90% complete, and SAR11 genomes are estimated > 70% complete. We created multiple sequence alignments of phylogenetically conserved genes from these genomes using the GTDB-Tk pipeline [19] with default settings. The pipeline identifies conserved proteins (120 bacterial proteins) and generates concatenated multi-protein alignments [17] from the genome assemblies using hmmalign from the hmmer software suite. We further filtered the resulting alignment columns using the bacterial and archaeal alignment masks from [17] (http://gtdb.ecogenomic.org/downloads). We removed columns represented by fewer than 50% of all taxa and/or columns with no single amino acid residue occurring at a frequency greater than 25%. We trimmed the alignments using trimal [24] with the automated -gappyout option to trim columns based on their gap distribution. We inferred reference phylogenies using multithreaded RAxML [25] with the GAMMA model of rate heterogeneity, empirically determined base frequencies, and the LG substitution model [26](PROTGAMMALGF). Branch support is based on 250 resampled bootstrap trees. This tree was then pruned to only allow a maximum average distance to the closest leaf (ADCL) of 0.003 to reduce the phylogenetic redundancy in the tree [27]. We then “placed” genomes that either did not pass completeness threshold or were considered phylogenetically redundant by ADCL within the fixed reference phylogeny for each group using pplacer [28] representing each placed genome as a pendant edge in the final tree. We then examined the resulting tree and manually selected clade/ecotype cutoffs to be as consistent as possible with clade definitions previously outlined for these groups [29–32]. We then gave clades from each taxonomic group custom taxonomic identifiers and we added these identifiers to the MARMICRODB Kaiju taxonomic hierarchy.</p> <p><strong>Software/databases used:</strong><br>checkM v1.0.11[16]<br>HMMERv3.1b2 (http://hmmer.org/)<br>prodigal v2.6.3 [22]<br>trimAl v1.4.rev22 [24]<br>AliView v1.18.1 [33] [34]<br>Phyx v0.1 [35]<br>RAxML v8.2.12 [36]<br>Pplacer v1.1alpha [28]<br>GTDB-Tk v0.1.3 [19]<br>Kaiju v1.6.0 [34]<br>GTDB RS83 (https://data.ace.uq.edu.au/public/gtdb/data/releases/release83/83.0/)<br>NCBI Taxonomy (accessed 2018-07-02) [23]<br>TIGRFAM v14.0 [37]<br>PFAM v31.0 [38]</p> <p><strong>Use example:</strong><br>Because we used custom taxonomic MARMICRODB users will find many reads assigned to non-standard NCBI taxonomy identifiers. However, these reads are easily parsable using the custom names.dmp and nodes.dmp files included with the database. We include a brief description of how to do this below.</p> <p>I typically run Kaiju like:</p> <pre><code>kaiju -z 20 -a greedy -e 5 -m 11 -s 65 -E 0.05 -x \ -t nodes.dmp -f MARMICRODB.fmi \ -i inputfile_R1.fastq.gz \ -j inputfile_R2.fastq.gz \ -o MYOUTPUT.kaiju</code></pre> <p>To obtain a parseable report that lists the custom taxonomic ranks from the nodes.dmp and names.dmp files run kaiju2krona on the output.</p> <pre><code>kaiju2krona -t nodes.dmp -n names.dmp -i MYOUTPUT.kaiju -o MYOUTPUT.kaiju.krona </code></pre> <p>This report shows counts assigned to each node in the custom taxonomy and will also include the names for each rank. You can easily parse this programmatically using a scripting language like python or by using unix utilities.</p> <p><strong>File descriptions:</strong></p> <p><em>MARMICRODB_catalog.tsv</em><br>Tabular file of NCBI assembly accessions and associated taxonomic information for every genome in MARMICRODB. Also includes literature references for each genome where available. Header description:</p> <ol> <li>genome: Unique identifier for each genome</li> <li>full_name: full organism name where available</li> <li>source: literature reference where available</li> <li>taxid: NCBI taxonomy ID for the assembly accession</li> <li>MARMICRODBtaxid: taxonomy ID used in the custom Kaiju database</li> <li>lineage_assignment: taxonomic lineage assignment from NCBI</li> <li>domain: archaea, bacteria, or eukaryote</li> <li>taxgroup: short descriptive group</li> <li>taxclade: higher resolution clade assignment where available</li> <li>habitat_source: whether genome derives from marine or aquatic source</li> <li>sequence_type: isolate, single cell genome (sag), metagenome assembled genome (mag), or transcriptome in case of eukaryotes</li> <li>assembly_ftp: NCBI ftp for assembly</li> <li>gbk_acc: assembly genbank or refseq accession number</li> <li>gtdb_taxonomy: taxonomic lineage assignment from GTDB-Tk v0.1.3 against GTDB v83</li> </ol> <p><em>MARMICRODB_kronaplot.html</em><br><a href="http://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">Interactive Kronaplot</a> for the exploration of taxonomic composition of MARMICRODB</p> <p><em>MARMICRODB.faa.bz2</em><br>Fasta file of all protein sequences in MARMICRODB</p> <p><em>scripts.tar.gz</em><br>directory containing scripts for generating Kaiju formatted database</p> <p><em>phylogenies.tar.gz</em><br>directory containing detailed phylogenies for SAR11, Prochlorococcus, SAR86, and SAR116</p> <p><em>MARMICRODB.fmi</em><br>Kaiju index for MARMICRODB</p> <p><em>nodes.dmp</em><br>nodes file for taxonomic assignment with Kaiju</p> <p><em>names.dmp</em><br>names file for generating Kaiju reports</p> <p><strong>References:</strong><br>1. Karsenti E, Acinas SG, Bork P, Bowler C, De Vargas C, Raes J, et al. A holistic approach to marine eco-systems biology. PLoS Biol. 2011;9: e1001177.<br>2. Biller SJ, Berube PM, Dooley K, Williams M, Satinsky BM, Hackl T, et al. Marine microbial metagenomes sampled across space and time. Scientific Data. 2018;5: 180176.<br>3. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>4. Mende DR, Letunic I, Huerta-Cepas J, Li SS, Forslund K, Sunagawa S, et al. proGenomes: a resource for consistent functional and taxonomic annotations of prokaryotic genomes. Nucleic Acids Res. 2017;45: D529–D534.<br>5. Keeling PJ, Burki F, Wilcox HM, Allam B, Allen EE, Amaral-Zettler LA, et al. The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): illuminating the functional diversity of eukaryotic life in the oceans through transcriptome sequencing. PLoS Biol. 2014;12: e1001889.<br>6. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>7. Tully BJ, Graham ED, Heidelberg JF. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Sci Data. 2018;5: 170203.<br>8. Haroon MF, Thompson LR, Parks DH, Hugenholtz P, Stingl U. A catalogue of 136 microbial draft genomes from Red Sea metagenomes. Sci Data. 2016;3: 160050.<br>9. Hugerth LW, Larsson J, Alneberg J, Lindh MV, Legrand C, Pinhassi J, et al. Metagenome-assembled genomes uncover a global brackish microbiome. Genome Biol. 2015;16: 279.<br>10. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017.<br>11. Mukherjee S, Seshadri R, Varghese NJ, Eloe-Fadrosh EA, Meier-Kolthoff JP, Göker M, et al. 1,003 reference genomes of bacterial and archaeal isolates expand coverage of the tree of life. Nat Biotechnol. 2017;35: 676–683.<br>12. Haft DH, DiCuccio M, Badretdin A, Brover V, Chetvernin V, O’Neill K, et al. RefSeq: an update on prokaryotic genome annotation and curation. Nucleic Acids Res. 2018;46: D851–D860.<br>13. Klemetsen T, Raknes IA, Fu J, Agafonov A, Balasundaram SV, Tartari G, et al. The MAR databases: development and implementation of databases specific for marine metagenomics. Nucleic Acids Res. 2018;46: D692–D699.<br>14. Berube PM, Biller SJ, Hackl T, Hogle SL, Satinsky BM, Becker JW, et al. Single cell genomes of Prochlorococcus, Synechococcus, and sympatric microbes from diverse marine environments. Scientific Data. 2018;5: 180154.<br>15. Becker JW, Hogle SL, Rosendo K, Chisholm SW. Co-culture and biogeography of Prochlorococcus and SAR11. ISME J. 2019. doi:10.1038/s41396-019-0365-4<br>16. Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25: 1043–1055.<br>17. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017. doi:10.1038/s41564-017-0012-7<br>18. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>19. Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2019. doi:10.1093/bioinformatics/btz848<br>20. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41: D590–6.<br>21. Biller SJ, Berube PM, Berta-Thompson JW, Kelly L, Roggensack SE, Awad L, et al. Genomes of diverse isolates of the marine cyanobacterium Prochlorococcus. Sci Data. 2014;1: 140034.<br>22. Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics. 2010;11: 119.<br>23. Federhen S. The NCBI Taxonomy database. Nucleic Acids Res. 2012;40: D136–43.<br>24. Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25: 1972–1973.<br>25. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>26. Le SQ, Gascuel O. An improved general amino acid replacement matrix. Mol Biol Evol. 2008;25: 1307–1320.<br>27. Matsen FA, Gallagher A, McCoy C. Minimizing the average distance to a closest leaf in a phylogenetic tree. arXiv [q-bio.PE]. 2012. Available: http://arxiv.org/abs/1205.6867<br>28. Matsen F a., Kodner RB, Armbrust EV. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree. BMC Bioinformatics. 2010;11: 538.<br>29. Biller SJ, Berube PM, Lindell D, Chisholm SW. Prochlorococcus: the structure and function of collective diversity. Nat Rev Microbiol. 2014;13: 13–27.<br>30. Giovannoni SJ. SAR11 Bacteria: The Most Abundant Plankton in the Oceans. Ann Rev Mar Sci. 2016.<br>31. Dupont CL, Rusch DB, Yooseph S, Lombardo M-J, Alexander Richter R, Valas R, et al. Genomic insights to SAR86, an abundant and uncultivated marine bacterial lineage. ISME J. 2012;6: 1186–1199.<br>32. Yang S-J, Kang I, Cho J-C. Expansion of Cultured Bacterial Diversity by Large-Scale Dilution-to-Extinction Culturing from a Single Seawater Sample. Microb Ecol. 2016;71: 29–43.<br>33. Larsson A. AliView: a fast and lightweight alignment viewer and editor for large datasets. Bioinformatics. 2014;30: 3276–3278.<br>34. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>35. Brown JW, Walker JF, Smith SA. Phyx: phylogenetic tools for unix. Bioinformatics. 2017;33: 1886–1888.<br>36. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>37. Haft DH, Selengut JD, Richter RA, Harkins D, Basu MK, Beck E. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 2013;41: D387–95.<br>38. Finn RD, Coggill P, Eberhardt RY, Eddy SR, Mistry J, Mitchell AL, et al. The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res. 2016;44: D279–85.<br>39. Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil P-A, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol. 2018.<br> </p>
Fig. 9 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 9. Web page of management of common names in Database of Korean National Species List. This web page displays the list of common names registered in the database.
Fig. 8 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 8. Web pages of management function of special list of species defined by law. (A) Presents the web interface of searching taxon with five functions. (B) shows the list of registered species defined by law with Korean name.
Fig. 10 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 10. Web pages of management of references in Database of Korean National Species List. (A) Shows the list of references with function to assign references to taxa. (B) Web interface of function assigning reference to taxa.
Fig. 7 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 7. Web page of list of special list of species defined by law. This web page provides management function of special list of species defined by law.
Fig. 5 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 5. Introduction page of Korean National Species List. Introduction page of National Species List of Korea in the platform for biodiversity in Korea (http://www.kbr.go.kr/content/view.do?menuKey = 446&contentKey = 14).
Fig. 4 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 4. Web interface of list of taxa with search options. (A) Displays complex interface for searching taxon. (B) is the list view of taxa searched. (C) shows the example of brief information of taxon which will be appeared by clicking KTSN number in the list view.
Fig. 2 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 2. Web interface of taxon in Database of Korean National Species List. (A) Provides hierarchical structure of taxon with the tree interface. (B) shows basic information of taxon including KTSN, species properties, representative common names, and status of KTSN. (C) displays ranks, names, common names, identifiers, origins and etc. (D) displays information higher taxa based on hierarchical systems. (E) shows the list of synonyms. (F) is the list of references related to this taxon. (G) is the list of all common names except representative name.
Fig. 12 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 12. System structure of Database of Korean National Species List. Grey rectangles represent systems related to the Database of National Species List of Korea and blue rectangles indicates systems using species information via Database of National Species List of Korea outside of NIBR. Black arrows indicate utilization of taxonomic information from the Database of National Species List of Korea and Grey arrows present communication with institutes outside of NIBR. Dotted grey arrow means commination channel will be established soon.
Fig. 1 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 1. Database structure of Database of Korean National Species List. (A) Presents relationship of major entities of the Database of Korean National Species List. Black thick lines are n:m relationship and blue arrows show detailed content types of each entity. (B) displays example of hierarchical relations of higher taxa originated from two different systems, KNSL and APG IV.
Fig. 6 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 6. Download page of Korean National Species List. This web page provides the download link of National Species List of Korea in the main page of the platform for biodiversity in Korea (http://www.kbr.go.kr/).
Fig. 3 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 3. Tree interface of taxon. (A) Shows interface to search by management groups, rank, and name. Yellow arrow indicates fold- er icon which open all higher nodes of the taxon up and grey arrow shows check icon which provides detailed information of the taxon. (B) displays tree interface for hierarchical structure of taxa. Plus icon indicated by blue arrow means open the lower taxa in the specific taxon, and each node (orange arrow) represents each taxon.
Fig. 11 in Database of National Species List of Korea: the taxonomical systematics platform for managing scientific names of Korean native species
Fig. 11. Web interface of reporting tools in Database of Korean National Species List. (A) Presents the basket of report. Dotted box presents option of generating reports. (B) shows the list of tasks registered for generating report. (C) displays three different formats of report generated by the Database of National Species List of Korea.
The mOTUs online database provides web-accessible genomic context to taxonomic profiling of microbial communities - Supplementary Tables
<p><strong>Supplementary Table 1:</strong></p> <p>A map between each of the genomes in mOTUs-db (3’747’151), the associated study and its metagenomic sample (in case of MAGs).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> STUDY → Unique mOTUs-db name of the study</code><br><code> IS_MAG → True if genome is a MAG, otherwise False </code><br><code> METAGENOMIC_SAMPLE → Unique name of the metagenomic sample or NA in case of non-MAG genome</code></p> <p>Example:</p> <p><code> GENOME STUDY IS_MAG METAGENOMIC_SAMPLE</code><br><code> ---------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_MAG_00000001 ACIN21-1 True ACIN21-1_SAMN05421555_METAG</code><br><code> RSGB23-1_GCA-006096615-V1_GENO_10000001 RSGB23-1 False NA</code></p> <p><strong>Supplementary Table 2:</strong></p> <p>A map between all non-MAG genomes (919’090) and their source (e.g. Refseq or JGI).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> SOURCE_SAMPLE_LINK → Link to the original location of this genome</code></p> <p>Example:</p> <p><code> #GENOME SOURCE_SAMPLE_LINK</code><br><code> --------------------------------------------------------------------------------------------------------</code><br><code> JGIG23-1_GA0055041_GENO_10000001 https://gold.jgi.doe.gov/analysis_project?id=Ga0055041</code><br><code> RSGB23-1_GCA-006717865-V1_GENO_10000001 https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_006717865.1</code></p> <p><strong>Supplementary Table 3:</strong></p> <p>A list of all metagenomic studies processed for the mOTUs-db, their number of samples, the number of reconstructed MAGs and the associated publication.</p> <p>Columns:</p> <p><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> BIOPROJECT --> Public identifier (NCBI/JGI) of metagenomic sequencing project</code><br><code> SAMPLES --> Number of metagenomic samples</code><br><code> MAGs --> Number of reconstructed MAGs</code><br><code> PUBLICATION --> Link to publication</code></p> <p>Example:</p> <p><code> STUDY BIOPROJECT SAMPLES MAGs PUBLICATION</code><br><code> -------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1 PRJEB44456 58 1,110 https://www.nature.com/articles/s42003-021-02112-2</code></p> <p><strong>Supplementary Table 4:</strong></p> <p>Mapping between mOTUs-db sample identifier, the associated biosample and the environment.</p> <p>Columns:</p> <p><code> SAMPLE --> Unique mOTUS-db sample identifier</code><br><code> BIOSAMPLE --> Public identifier (NCBI/JGI) of metagenomic sample</code><br><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> ENVIRONMENT --> Environment of metagenomic sample</code><br><code> SOURCE_SAMPLE_LINK --> Link to the original location of this sample</code></p> <p>Example:</p> <p><code> #SAMPLE BIOSAMPLE STUDY ENVIRONMENT SOURCE_SAMPLE_LINK</code><br><code> ---------------------------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_METAG SAMN05421555 ACIN21-1 marine https://www.ncbi.nlm.nih.gov/biosample/SAMN05421555/</code></p> <p><strong>Supplementary Table 5:</strong></p> <p>A list of environments covered in the mOTUs-db mapped to the respective NCBI taxonomy (if possible)</p> <p>Columns:</p> <p><code> TERM --> Unique environment name</code><br><code> NCBI TAXONOMY ID --> Link to the NCBI taxonomy</code></p> <p>Example:</p> <p><code> TERM NCBI TAXONOMY ID</code><br><code> ----------------------------------------------</code><br><code> activated sludge metagenome NCBI:txid942017</code><br><code> air metagenome NCBI:txid655179</code></p>
Linked collectors and determiners for: IberArthro: A database compiling taxonomic and distributional data on Ibero-Balearic arthropods.
Natural history specimen data linked to collectors and determiners held within, "IberArthro: A database compiling taxonomic and distributional data on Ibero-Balearic arthropods". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/73319645-c321-48ed-b923-b3782a203589">https://bionomia.net/dataset/73319645-c321-48ed-b923-b3782a203589</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/73319645-c321-48ed-b923-b3782a203589">https://gbif.org/dataset/73319645-c321-48ed-b923-b3782a203589</a>. Formatted as a Frictionless Data package.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.