Annotated sequences extracted from bacterial genomes
<p>Three files containing sequences extracted from 1,049,210 bacterial genomes available from GenBank (release 252). Protein coding sequences were annotated with IDTAXA (PMID: 34541527) using taxon-specific KEGG groups (Bacteria_Protein_subset.fas.gz). These annotations were transferred to their corresponding (nucleotide) coding sequences (Bacteria_Nucleotide_subset.fas.gz). Intergenic regions were extracted from each genome and annotated by FindNonCoding (PMID: 34636849) for their overlap with any of 25 common bacterial non-coding RNAs in Rfam (v14). Intergenic regions were required to be at least 100 nucleotides long and contain no ambiguities (Bacteria_Intergenic_subset.fas.gz). Each subset contains only distinct sequences randomly ordered.</p> <p><strong><em>Headers</em></strong></p> <p>Sequence headers contain the assembly accession followed by the annotation and separated by a "|" character. For example:</p> <p><strong>Bacteria_Intergenic_subset.fas.gz</strong></p> <p>>GCA_022121725.1|RF00000<br> ATGTTACCTTCTTGAGTGATACGGGATGAA[...]</p> <p><strong>Bacteria_Protein_subset.fas.gz</strong></p> <p>>GCA_014764685.1|K02049<br> MPRDLIRISGLEKTYADGSVHALSNIDLSIKD[...]</p> <p><strong>Bacteria_Nucleotide_subset.fas.gz</strong></p> <p>>GCA_015948525.1|K02197<br> GTGAACCTGCGACGTAAAAACCGGCTAYG[...]</p> <p><strong><em>Annotations</em></strong></p> <p>Protein and protein coding sequences are labeled with their KEGG group, starting with "K". Intergenic sequences are named by any overlapping Rfam families, starting with a "RF", and separated by commas when multiple are predicted. "RF00000" is a placeholder for the absence of any predicted RF families.</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0