Skip to main content
zenodoopen

Supplemental Datasets for: WhatsGNU: A Tool For Identifying Proteomic Novelty.

<p>Five precompressed databases&nbsp;are available for WhatsGNU:</p> <p>Ortholog Mode:</p> <ol> <li>WhatsGNU_TB_Ortholog.zip:&nbsp;<em>Mycobacterium tuberculosis</em> Version-07/09/2019 (compressed 26,794,006 proteins in 6563 genomes to 434,725 protein variants).</li> <li>WhatsGNU_Pa_Ortholog.zip:&nbsp;<em>Pseudomonas aeruginosa</em> Version-07/06/2019 (compressed 14,475,742 proteins in 4712 genomes to 1,288,892 protein variants).</li> <li>WhatsGNU_Sau_Ortholog.zip:&nbsp;<em>Staphylococcus aureus</em> Version-06/14/2019 (compressed 27,213,667 proteins in 10350 genomes to 571,848 protein variants).</li> </ol> <p>Big Data basic Mode:</p> <ol> <li><em>Senterica_Enterobase_basic_216642.pickle.gz: Salmonella enterica</em> Enterobase Version-08/29/2019 (compressed 975,262,506 proteins in 216,642 genomes to 5,056,335 protein variants).</li> <li><em>Sau_Staphopia_basic_43914.pickle: Staphylococcus aureus</em> Staphopia Version-06/27/2019 (compressed 115,178,200 proteins in 43,914 genomes to 2,228,761 protein variants).</li> </ol> <p>Figures_datasets.zip:&nbsp;The FASTA files and the WhatsGNU reports as input and output, respectively, which were used to produce panels in Figures 1 and 2 and supplementary figures 1, 2 and 3.</p>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
0
Engagement
0