Gene family data from the PhyloGenes (release version 1.2, phylogenes.org)
<p>The compressed file contains: </p> <p><br> 1. PhyloXML_files </p> <p>This folder has family trees in PhyloXML format, one file per family (e.g. <family_ID>.xml).</p> <p>The following information is provided for each node of a tree:<br> 1) leaf node:<br> branch length<br> name <gene_id><br> taxonomy scientific_name<br> sequence accession <UniProt ID></p> <p>2) non-leaf node:<br> branch length<br> events <duplication or speciation></p> <p><br> 2. phylogenes_csv.tar.xz</p> <p>This tar file has gene information of family members in CSV format, one file per family (e.g. <family_ID>.csv). </p> <p>A CSV file includes the following columns:<br> Uniprot ID<br> Gene <Gene name. If none then Gene ID><br> Gene ID<br> Gene name<br> Organism<br> Subfamily name</p> <p>Any columns displayed after 'Subfamily name' are 'Known functions'. Each 'Known function' is a GO molecular function term that is annotated to at least one member of the gene family AND that the annotation is supported by an experimental evidence. Number 1 or 0 indicates the presence or absence of a particular function in a gene.</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4