Skip to main content
zenodoopen

Processed gene and clinical data

<p>The processed cancer dataset mentioned in the paper "Cox-Sage: Enhancing Cox proportional hazards model with interpretable graph neural networks for cancer prognosis," which is currently under review in Briefings in Bioinformatics. The gene expression data and clinical data of seven cancer types downloaded from TCGA (https://portal.gdc.cancer.gov/) were processed to retain only protein-coding genes. A patient similarity graph was constructed based on the similarity of clinical data. The data for each type of cancer consists of a `gene_expression.csv`, a `clinical.csv`, and an `adj_list.pkl`. In addition, the `prognostic_genes.zip` file contains the hazards contour plot of all prognostic genes identified in the study. And the `all_benchmarks_prediction_results.zip` file contains the hazards prediction results of all benchmarks that being reproduced.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0