Skip to main content
zenodoopen

SCID Multiomics Post-Processed Data and Analysis

<p>In this repository are the post-processed datasets and analytical code for the SCID Multiomics&nbsp;paper.&nbsp;The repository is structured as an installable R package for dependency management and dataset&nbsp;loading; it does not export any functions.</p> <p><strong>Installation</strong></p> <p>The easiest way to install this is to download the repository and install using `devtools::install()`.&nbsp;This will allow the import of various datasets using the `data()` function, upon which many of the&nbsp;analysis scripts depend.</p> <p><strong>Datasets</strong></p> <p>In no particular order, the important datasets are described below:</p> <p>- <strong>intsites</strong>: summary statistics from (Wang et al, Blood, 2010) for timepoints used in this study<br> - <strong>tcr</strong>: Aggregate TCR data from Adaptive Biotechnology&#39;s ImmunoSeq pipeline.<br> - <strong>mb</strong>: Metadata for the microbiome sampling timepoints, as well as species data from Metaphlan (not used)<br> -&nbsp;<strong>agg.mb.kz</strong>: Kraken species data for the microbiome samples, after low-complexity filtering<br> - <strong>agg.vp.kz</strong>: Kraken species data for the virome samples, after low-complexity filtering<br> - <strong>card</strong>: Antibiotic resistance gene data from CARD<br> - <strong>subject_ids.csv</strong>: Provides a mapping from the original sample IDs used in the datasets to the ones used in the manuscript.</p> <p>The code for creating these datasets from the original data files are in the `data-raw` directory.</p> <p><strong>Analysis/Figures</strong></p> <p>The analysis code is broken apart by subject and is largely concerned with figure generation. The R<br> scripts are all located in the `inst` folder. To generate all figures, you should run each script in the<br> order specified by the `GenerateFigures.R` file.</p> <p>Figures are output to the `figures` directory, while tables are output to the `tables` directory.</p> <p>Please note: many of the figures used in the manuscript were aesthetically modified after generation (text size, color palette, orientation), precluding exact figure replication</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4