Skip to main content
zenodoopen

Single Cell CPTAC Renal Cell Carcinoma

<p>These data include a subset of single-cell samples from the CPTAC Renal Cell Carcinoma data processed using the following steps:<br><br></p> <ol> <li><em>Loom</em>&nbsp;files were read into R and converted into&nbsp;<em>SingleCellExperiment</em>&nbsp;objects.</li> <li>Ensembl gene ID's were matched to HGNC symbols, chromosome name, starting position, and ending position via Biomart using the&nbsp;<em>scater</em>&nbsp;R package.</li> <li>Size factors were computed using the&nbsp;<em>scran</em>&nbsp;R package.</li> <li>UMAP dimensions were computed using the&nbsp;<em>scater</em>&nbsp;R package.</li> <li>Probes without matching HGNC symbols were removed.</li> <li>Where duplicate HGNC symbols were present, the gene with the maximum normalized range was retained.</li> <li>Cell types were inferred using the&nbsp;<em>scMRMA</em>&nbsp;R package.</li> <li>Cells with less than or equal to 1,000 features were removed.</li> <li>Cells with mitochondrial reads greater than or equal to 50% were removed.</li> <li>Expression data for podocytes and macrophages were saved separately for each sample.</li> <li>For podocytes and macrophages for each sample, expression data were projected onto the first 100 principal components using the&nbsp;<em>irlba</em>&nbsp;R package. The results are available from&nbsp;<strong>CPTAC_RCC_PCA.zip.</strong></li> <li>For macrophages for each sample, a differentiation trajectory was estimated using the&nbsp;<em>monocle3</em>&nbsp;R package, and plots were colored by combined expression of the M0 markers CSF1R, CD14, CD68, and CD11B, the M1 markers CD86, MARC0, CXCL9, CXCL10, CXCL11, NOS2, SOCS1, and CD64, and the M2 markers TGM2, CD23, ARG1, CCL22, CD163, and CD206 (from PMC8268869). Pseudotime starting points were annotated in&nbsp;<em>monocle3</em>&nbsp;using visual inspection of plots. Only samples in which a visible trajectory from M0 -&gt; M1 -&gt; M2 was evident using these markers were retained.</li> <li>For podocytes for each sample, a differentiation trajectory was estimated using the&nbsp;<em>monocle3</em>&nbsp;R package, and plots were colored by combined expression of the dedifferentiation markers DACH1 (from PMC5908116) and PTPRO (from PMID9639039. Pseudotime starting points were annotated in&nbsp;<em>monocle3</em>&nbsp;using visual inspection of plots. Only samples in which a visible trajectory of dedifferentiation was evident using these markers were retained. These pseudotime assignments and the macrophage assignments are available from&nbsp;<strong>CPTAC_RCC_pseudotime_all_cells.zip.</strong></li> <li>For podocytes and macrophages for each sample, 10-fold cross-validation matrices for expression, PCA, and pseudotime were generated across 5 random splits, for a total of 50 files per cell type, per sample. These files are available from&nbsp;<strong>CPTAC_RCC_expression_all_cells.zip.</strong></li> <li>Expression data were subset to include 63 randomly-selected podocytes and 63 randomly-selected macrophages to ensure balanced data. These expression data are available from&nbsp;<strong>CPTAC_RCC_expression.zip</strong>. Pseudotimes were also subset and are available in&nbsp;<strong>CPTAC_RCC_pseudotime.zip</strong>.</li> <li>Expression data were projected onto the first 100 principal components. These data are available from&nbsp;<strong>CPTAC_RCC_expression_dimReduced.zip</strong>.</li> </ol>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0