Skip to main content
zenodoopen

ecDNA machine learning modeling

<h3><strong>1. Today (2024-06-27), we discovered an issue with the labeling of sample groups in one of the supplementary figures (Supplementary Figure 14c) in our published article. We have corrected the figure and present it here, and we extend our apologies to all readers for any confusion this may have caused (although no report received).</strong></h3> <h3><strong>2. The source data of supplementary figure 13 in the accompanying article table has been found to have issues, which were identified as a result of improper Excel operation. Here, we have uploaded the correct data table</strong></h3> <p>--------------------------------------------------</p> <p>&nbsp;</p> <p>1.&nbsp;ecDNA_cargo_gene_modeling_data.csv.gz</p> <p>The dataset contains features from 386 TCGA tumors for modeling ecDNA cargo gene prediction. It was converted from R data format with the following&nbsp;code. NOTE: columns 'sample' and 'gene_id' are not used for actual modeling but for identifying, and sampling purposes.</p> <p>library(data.table)</p> <p>data = readRDS("~/../Downloads/ecDNA_cargo_gene_modeling_data.rds")</p> <p>colnames(data)[3] = "total_cn"</p> <p>data.table::fwrite(data, file = "~/../Downloads/ecDNA_cargo_gene_modeling_data.csv.gz", sep = ",")</p> <p>&nbsp;</p> <p>2.&nbsp;gcap_pcawg_WGS_result.tar.gz</p> <p>GCAP analysis results for PCAWG allele-specific copy number profiles derived from WGS.</p> <p>&nbsp;</p> <p>3.&nbsp;gcap_tcga_snp6_result.tar.gz</p> <p>GCAP analysis results for TCGA allele-specific copy number profiles derived from SNP6 array.</p> <p>&nbsp;</p> <p>4.&nbsp;gcap_Changkang_WES_result.tar.gz</p> <p>GCAP analysis results for SYSUCC Changkang&nbsp;allele-specific copy number profiles derived from tumor-normal paired WES.</p> <p>&nbsp;</p> <p>5.&nbsp;tcga_overlap_gene_wgs.rds,&nbsp;tcga_overlap_gene_snp.rds and&nbsp;tcga_overlap_gene_wes.rds</p> <p>These datasets contain TCGA gene-level copy number results in R data format from overlapping samples (dataset above). WGS from PCAWG, SNP array, and WES from GDC portal.</p> <p>&nbsp;</p> <p>6.&nbsp;cellline-batch1.zip &amp;&nbsp;cellline-batch1.zip</p> <p>&nbsp;</p> <p>GCAP results of cell line batch 1 and batch 2.</p> <p>&nbsp;</p> <p>7.&nbsp;AA_cellline_wgs.zip</p> <p>AA software results for cell line batch 1.</p> <p>&nbsp;</p> <p>8.&nbsp;Batch2_AA_summary.xlsx</p> <p>AA software results for cell line batch 2.</p> <p>&nbsp;</p> <p>9.&nbsp;FISH-for-supp-file.zip</p> <p>Extended raw FISH images from 12 CRC samples.</p> <p>&nbsp;</p> <p>10. SNU216.zip</p> <p>Extended AA and GCAP analysis on SNU216.</p> <p>&nbsp;</p> <p>11. aa_ffpe.zip and AA_summary_table_of_6_erbb2_ffpe_samples.xlsx</p> <p>Extended AA running files (all results) and result summary data for 6 GCAP predicted ERBB2 amp clinical samples.</p> <p>&nbsp;</p> <p>12. source data of fig.4</p> <p>&nbsp;</p> <p>13. source data of supp fig.2 subplots</p> <p>&nbsp;</p> <p>13. source data of supp fig.15</p> <p>&nbsp;</p> <p>14. GCAP result data objects for three ICB cohorts. Both gene-level and sample-level data included.</p> <p>&nbsp;</p> <p>15. PDX-P68: processed (AA and CNV) data of P68 from WGS and WES data.</p> <p>&nbsp;</p> <p>16. source data of supp fig.13</p> <p>&nbsp;</p> <p>17. updated supplementary figure 14</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0

Topics