Skip to main content
zenodoopen

Protein sequences matching TIGRFAM models

<p>A set of 411 large and 3,001 small TIGRFAM&nbsp;protein families originally used for benchmarking sequence clustering programs. Large families contain at least 20,000 sequences with a genus label, whereas small contain fewer than 20,000. Sequences were&nbsp;downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a>&nbsp;and renamed by their original name (accession number)&nbsp;followed by their semi-colon separated NCBI taxonomy of the originating organism. For example, the first sequence in TIGRFAM_large/TIGRFAM00005.fas.gz is named:</p><blockquote><p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p></blockquote>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0