Skip to main content
zenodoopen

Full-length and split homologs of human proteins in the gut microbiome

<p>These files were generated as part of the manuscript "Human xenobiotic metabolism proteins have full-length and split homologs in the gut microbiome" (submitted).</p> <p>The .tar file contains .ipc files that are tables of full-length (full_humcover3.ipc) and split homologs (part_humcover3.ipc) of human proteins in the gut microbiome, organized by alignment coverage threshold. For example, the directory `HumanUPR_0.67_src_20000_70` contains results obtained at a 67% alignment coverage threshold for the bacterial protein, and 70% for the human protein. Note that our pipeline collapses full-length alignments to the same UHGP-90 protein family into a single entry per species, with the number of genomes reported in the column nGenomes. Split homologs are not collapsed because genomic context is used to define them, and this context may differ across individual genomes.</p> <p>These files are in Arrow <a href="https://arrow.apache.org/docs/python/ipc.html#ipc">IPC</a> format, which provides compression and fast I/O for large tables. We recommend reading them using <a href="https://pola.rs/">pola.rs</a> or the <a href="https://arrow.apache.org/docs/r/">R Arrow</a> package. In particular, because the full-length homolog table is large, you may wish to work with it without loading it into memory, which can be accomplished using&nbsp;<a href="https://docs.pola.rs/api/python/dev/reference/api/polars.scan_ipc.html">scan_ipc</a> in pola.rs or <a href="https://arrow.apache.org/docs/r/reference/open_dataset.html">open_dataset</a> in R Arrow.</p> <p>We also provide gzipped .csv format datasets of full-length (pgkb_FH_drugs.csv.gz) and split (pgkb_SH_drugs.csv.gz) homologs, at the default 67% alignment coverage threshold for bacterial and 70% for human proteins, organized by their&nbsp;<a href="https://www.pharmgkb.org/">PharmGKB</a> annotations. For each drug annotated in PharmGKB as being metabolized by a human protein with full-length or split homologs, we provide the human protein(s) responsible, its xenobiotic enzyme class, the bacterial protein homolog(s), length and percent identity of the alignment, and either the specific genome (g, split homologs only) or the number of genomes (nGenomes, full homologs only). Xenobiotic enzyme classes are defined as in Figure 4 of the manuscript, with the additional classes "nucl" (nucleobase-containing metabolic proteins not annotated to any other class), "redox" (oxidoreductases not annotated to any other class), and "other" (all remaining proteins).</p>

ShareScore

52/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
4
Access
20
Reuse readiness
8
Engagement
8

Topics