Diamond formatted protein database for taxonomic classification
<p>This is a diamond formatted database (diamond version 0.9.22) built on December 14th 2018.</p> <p>The database contains a total of 17,694,143 sequences:</p> <ul> <li>2,708,401 protein sequences from the <a href="https://bitbucket.org/dbeisser/taxmapper_supplement/src/master/">taxmapper database</a> (commit 450d337), containing 121 unique taxa.</li> <li>14,976,193 protein sequences from 1055 unique fungal taxa (downloaded from <a href="https://genome.jgi.doe.gov/1000_fungi_project">JGI 1000 fungi project</a> on November 23 2018).</li> <li>9,549 protein sequences from the <em>Hygrophorus russula</em> genome obtained from Genbank (accession GCA_003314125.1) on November 28 2018.</li> </ul> <p>The <em>Hygrophorus russula</em> protein sequences were obtained by running Augustus (v. 3.2.3) gene caller on the genomic fasta file using the laccaria_bicolor gene model.</p> <p>Taxonomic information was built into the diamond database by running:</p> <pre><code class="language-bash">zcat fasta.gz | diamond makedb -d diamond -p 4 --taxonmap taxonmap.gz --taxonnodes nodes.dmp</code></pre> <p>The nodes.dmp file was obtained from the <a>taxdump.tar.gz</a> file on December 11 2018.</p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 8
- Engagement
- 0