Data repository associated with 'A Functional Map of the Human Intrinsically Disordered Proteome'
<p><strong>ES_MAP.zip</strong></p> <ul> <li>a hierarchically clustered map of the human IDR-ome</li> <li>.cdt and .gtr files - outputs of Cluster3.0 software</li> <li>can be visualized using JavaTreeView (see Tutorial_ES.pdf)</li> </ul> <p><strong>TUTORIAL.zip</strong>, information on:</p> <ul> <li>visualization and analysis of the human IDR-ome map</li> <li>search for proteins of interest and exploratory analyses of clusters</li> <li>automatic export and analysis of exported clusters (code available at https://github.com/IPritisanac/ES_PW)</li> </ul> <p><strong>IDROME_SEQUENCES.zip</strong></p> <ul> <li>human proteome fasta file</li> <li>IDRome fasta file</li> <li>SPOT-Disorder v1.0 disorder boundaries <ul> <li>13 044 unique protein sequences with at least one IDR (>=30 amino acids)</li> <li>21 252 total unique human IDRs</li> </ul> </li> </ul> <p><strong>IDR_ALN.zip</strong></p> <ul> <li>alignments of IDR sequences across ENSEMBL orthologs</li> <li>19 459 IDR alignments</li> <li>UniProt ID and IDR boundaries for the human sequence are indicated in the name of the file</li> </ul> <p><strong>FAIDR_TSTATS.zip</strong></p> <ul> <li>hierarchical clustering of FAIDR t-statistics for 148 GO terms<br> <ul> <li>.cdt, .gtr files from Cluster3.0</li> <li>can be visualized using JavaTreeView</li> <li>reveals the most predictive molecular features for the top performing 148 models</li> </ul> </li> </ul> <p><strong>CLUSTERS_EXPLORE.zip</strong></p> <ul> <li>clusters obtained through exploratory analysis of the map provided in ES_MAP.zip</li> <li>93 exported clusters in .cdt file format</li> </ul> <p><strong>CLUSTERS_AUTO.zip</strong></p> <ul> <li>clusters extracted from the hierarchically clustered IDR-ome map at a range of distance thresholds (0.4 - 0.8) in .cdt file format</li> <li>distance refers to the uncentered correlation distance between vectors of Z-scores representing human IDRs</li> <li>clusters extracted at different distance thresholds are split into separate archives</li> <li>AUTO_GO_FEATS.xlsx - summary of GO-term overrepresentation and feature enrichment analyses; each distance threshold is in a separate sheet</li> </ul> <p><strong>FAIDR_HIGH_AUC_PPV_GO.zip</strong></p> <ul> <li>target files with annotations of 148 GO terms for which good quality FAIDR models could be obtained (AUC >= 0.7, PPV >= 0.4)</li> <li>file format: three columns; 1st: IDR ID (includes IDR boundaries); 2nd: protein UniProt ID; 3rd: annotation of the protein to a GO term (1 if known to be associated with the GO term, 0 if not)</li> </ul> <p> </p> <p> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4