Skip to main content
zenodoopen

Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer

<p>The files in this data release represent&nbsp;processed data from the FORESEE study conducted by Cambridge Epigenetix Ltd, and reported in the preprint manuscript:&nbsp;&nbsp;&quot;Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer&quot; (<a href="https://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a>).&nbsp;</p> <p>&nbsp;</p> <p>As described in the manuscript, classifiers were trained and validated on genomic features extracted from sequencing datasets across&nbsp;cases and controls.&nbsp; Several classes of genomic features were constructed for training and validation data sets which are described below:</p> <p>&nbsp;</p> <p><strong>CRC_enhancer_znorm_training_matrix_v1.csv<br> CRC_enhancer_znorm_validation_matrix_v1.csv</strong></p> <p>Columns contain sample names, rows contain genomic features.&nbsp;</p> <p><em>Description of feature generation process.&nbsp;</em></p> <p>To calculate 5hmC levels at gene enhancers, we first calculated read counts using Bam readcounts v0.01. RPKM were calculated over candidate gene-enhancers downloaded from GeneCards v4.4.&nbsp;5hmC enrichment was computed as the log2 ratio between the hydroxymethylome library RPKM and the input library RPKM after the inclusion of pseudocounts. Feature scaled (z-score normalization) 5hmC levels&nbsp;of enhancers quantile-normalized over samples.</p> <p><br> <strong>CRC_cegxdelfi_znorm_training_matrix_v1.tsv<br> CRC_cegxdelfi_znorm_validation_matrix_v1.tsv</strong></p> <p>Columns contain sample names, rows contain genomic features.&nbsp;<br> &nbsp;</p> <p><em>Description of feature generation process.&nbsp;</em><br> We divided the genome into 100KB bins and quantified cfDNA fragment sizes per bin. We removed blacklisted regions, genomic gaps (UCSC table) and non-standard chromosomes a priori.&nbsp;We excluded outlier bins in fragment size, only retaining fragments between 100nt to 220nt length. Finally, we split the genome into 100KB bins (in total 26170 non-overlapping genomic regions)&nbsp;&nbsp;and calculated the following characteristics of fragment size distribution per genomic bin: number of short fragments (100-150nt), number of long fragments (151-220nt), ratio between short and&nbsp;&nbsp;long fragments and the total number of fragments. This approach generates 26170 features per metric and per sample. The last step is the averaging of the 100 KB bins into larger non-overlapping&nbsp;&nbsp;genomic regions of 5 MB (in total 512 bins).</p> <p><br> <strong>CRC_cegxnps_znorm_training_matrix_v1.tsv</strong></p> <p><strong>CRC_cegxnps_validation_matrix_v1.tsv</strong></p> <p>&nbsp; Columns contain sample names, rows contain genomic features. &nbsp;<br> <em>Description of feature generation process.&nbsp;</em><br> &nbsp; &nbsp; Further detail in the manuscript:&nbsp;<a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <p><br> <strong>FORESEE_sample_description.tsv</strong></p> <p>This file holds sample data for colorectal cancer and control samples described in <a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <ul> <li>The sample_name column&nbsp;links to the column names in the *_matrix.tsv files</li> <li>The columns denoted raw_file1 and raw_file2 link the sample metadata with the enhancer&nbsp;readcount files contained in the gh_readcount_training.tar and gh_readcount_validation.tar.</li> </ul> <p>The columns in the table are briefly described below:</p> <p><em>sample_name</em>:<em> </em>Sample identifier<br> <em>Title</em>: Composed of the the disease name, gender and sample_name<br> <em>Source_name</em>: Tissue source<br> <em>Organism</em>: Contains the term: &ldquo;Homo sapiens&rdquo;<br> <em>Characteristics_indication</em>: Disease indication&nbsp;<br> <em>Characteristics_stage</em>: Cancer stage where appropriate. Indicated by roman numerals (I,II,III,IV)<br> <em>Characteristics_gender</em>: Described as &ldquo;Female&rdquo; or &ldquo;Male&rdquo;<br> <em>Characteristics_ethnicity</em>: Ethnicity description<br> <em>Characteristics_age_at_collection</em>: Age value in years<br> <em>Molecule</em>: Contains the value &ldquo;cell free DNA&rdquo;<br> <em>Description</em>: Contains value: &ldquo;Training sample&rdquo; or &ldquo;Validation sample&rdquo;</p> <p><em>Processed_data_file</em>: Contains the term: &ldquo;CRC_enhancer_training_matrix&rdquo; or &ldquo;CRC_enhancer_validation_matrix&rdquo;. &nbsp;<br> <em>raw_file1</em>: Refers to the readcount file from the 5hmC capture library<br> <em>raw_file2</em>: Refers to the readcount file from the Input control (shallow sequenced) library</p> <p>&nbsp;</p> <p><strong>gh_readcount_training.tar<br> gh_readcount_validation.tar</strong></p> <p>These tar files include the raw read counts computed across enhancer regions for case and control data and are referenced in the FORESEE_sample_description.tsv file.</p> <p>&nbsp;</p> <p><strong>Manuscript Abstract</strong></p> <p>Our classifier discriminated CRC samples from controls with an area under the receiver operating characteristic curve (AUC) of 90% (sensitivity was 55% at 95% specificity). Performance was similar&nbsp;for&nbsp;early stage 1 (AUC 89%) and late stage 4 CRC (AUC 94%). Performance was independent of the proportion of tumor-DNA in the cell free DNA.&nbsp;&nbsp;</p> <p>We expanded the classifier to include information about cell free DNA fragment size and abundance across the genome. Overall performance was similar (AUC 91%), with gains in sensitivity (63% at 95% specificity).&nbsp;</p> <p>The 5-hydroxymethylcytosine signal&nbsp;allows detection of CRC, even&nbsp;in&nbsp;cell free DNA&nbsp;samples with undetectable tumor DNA.&nbsp;Including&nbsp;5-hydroxymethylcytosine in multi-analyte&nbsp;screening, will improve&nbsp;sensitivity&nbsp;for early-stage cancer.&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics