Processed gene and clinical data
<p>The processed cancer dataset mentioned in the paper "Cox-Sage: Enhancing Cox proportional hazards model with interpretable graph neural networks for cancer prognosis," which is currently under review in Briefings in Bioinformatics. The gene expression data and clinical data of seven cancer types downloaded from TCGA (https://portal.gdc.cancer.gov/) were processed to retain only protein-coding genes. A patient similarity graph was constructed based on the similarity of clinical data. The data for each type of cancer consists of a `gene_expression.csv`, a `clinical.csv`, and an `adj_list.pkl`. In addition, the `prognostic_genes.zip` file contains the hazards contour plot of all prognostic genes identified in the study. And the `all_benchmarks_prediction_results.zip` file contains the hazards prediction results of all benchmarks that being reproduced.</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0