Skip to main content
zenodoopen

Diamond formatted protein database for taxonomic classification

<p>This is a diamond formatted database (diamond version 0.9.22) built on December 14th 2018.</p> <p>The database contains a total of&nbsp;17,694,143 sequences:</p> <ul> <li>2,708,401 protein sequences from the <a href="https://bitbucket.org/dbeisser/taxmapper_supplement/src/master/">taxmapper database</a>&nbsp;(commit&nbsp;450d337), containing 121 unique taxa.</li> <li>14,976,193 protein sequences from 1055 unique fungal taxa (downloaded from&nbsp;<a href="https://genome.jgi.doe.gov/1000_fungi_project">JGI 1000 fungi project</a>&nbsp;on November 23 2018).</li> <li>9,549 protein sequences from the&nbsp;<em>Hygrophorus russula</em>&nbsp;genome obtained from Genbank (accession&nbsp;GCA_003314125.1) on November 28 2018.</li> </ul> <p>The&nbsp;<em>Hygrophorus russula</em>&nbsp; protein sequences were obtained by running Augustus (v. 3.2.3) gene caller on the genomic fasta file using the laccaria_bicolor gene model.</p> <p>Taxonomic information was built into the diamond database by running:</p> <pre><code class="language-bash">zcat fasta.gz | diamond makedb -d diamond -p 4 --taxonmap taxonmap.gz --taxonnodes nodes.dmp</code></pre> <p>The nodes.dmp file was obtained from the <a>taxdump.tar.gz</a>&nbsp;file on December 11 2018.</p>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
8
Reuse readiness
8
Engagement
0