Skip to main content
zenodoopen

TIGRFAM protein sequences named by taxonomy

<p>A set of 411 TIGRFAM&nbsp;protein families originally used for benchmarking sequence clustering programs. Sequences were&nbsp;downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a>&nbsp;and renamed by their original name (accession number)&nbsp;followed by their semi-colon separated NCBI taxonomy. For example, the first sequence in TIGRFAM00005.fas.gz is named:</p> <blockquote> <p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p> </blockquote>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics