Skip to main content
zenodoopen

Supplemental tables for 'Pan-microalgal dark proteome mapping via interpretable deep learning and synthetic chimeras'

<p>Distinguishing genuine microbial proteins from contaminants remains a major bottleneck in genomics, particularly for environmental and non-model organisms where conventional homology-based tools are slow, resource-intensive, and leave large fractions of the "dark proteome" unclassified. LA<sup>4</sup>SR offers a scalable, interpretable framework that classifies algal and bacterial proteins directly from translated sequence data, achieving near-complete recall while accelerating inference by ~ 10,000-fold relative to BLASTP. By revealing that internal sequence features alone can drive robust classification, LA4SR bypasses the need for complete gene models or perfect annotations&mdash;opening new opportunities for analyzing complex microbial communities and metagenomes. Interpretability methods further link emergent amino acid signatures to evolutionary and ecological features, highlighting the potential of language models not only to accelerate genomics workflows but also to uncover new biological insights.</p> <p>&nbsp;</p> <p><strong>Table S1 | External spreadsheet. </strong>This spreadsheet contains LA<sup>4</sup>SR performance metrics, technical performance estimations, and BLAST results and runtimes of genomes comprising the algal training data.<strong> </strong></p> <p><strong>Table S2 | External spreadsheet. </strong>Captum attributions for 100 sequences each of algal and bacterial origin obtained using the LayerIntegratedGradients function.</p> <p><strong>Table S3 | External spreadsheet. </strong>Influential motifs found with the DeepMotifMinerPro software introduced in this work (see Data S3).</p> <p><strong>Table S4 | External spreadsheet.</strong> LA4SR and Diamond BLAST results for data from new assemblies from seen species (Fig. S7), contaminated assemblies from unseen genera (Fig. S7), and clean assemblies from unseen genera (Fig. 8). For the LA4SR results for genome assemblies from unseen genera, the genomes were published after the model was trained, and the genera shown were not included in the training dataset. Newly sequenced genomes uploaded to NCBI SRA accession SUB14799921.</p> <p>&nbsp;</p> <p>&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0