SCID Multiomics Post-Processed Data and Analysis
<p>In this repository are the post-processed datasets and analytical code for the SCID Multiomics paper. The repository is structured as an installable R package for dependency management and dataset loading; it does not export any functions.</p> <p><strong>Installation</strong></p> <p>The easiest way to install this is to download the repository and install using `devtools::install()`. This will allow the import of various datasets using the `data()` function, upon which many of the analysis scripts depend.</p> <p><strong>Datasets</strong></p> <p>In no particular order, the important datasets are described below:</p> <p>- <strong>intsites</strong>: summary statistics from (Wang et al, Blood, 2010) for timepoints used in this study<br> - <strong>tcr</strong>: Aggregate TCR data from Adaptive Biotechnology's ImmunoSeq pipeline.<br> - <strong>mb</strong>: Metadata for the microbiome sampling timepoints, as well as species data from Metaphlan (not used)<br> - <strong>agg.mb.kz</strong>: Kraken species data for the microbiome samples, after low-complexity filtering<br> - <strong>agg.vp.kz</strong>: Kraken species data for the virome samples, after low-complexity filtering<br> - <strong>card</strong>: Antibiotic resistance gene data from CARD<br> - <strong>subject_ids.csv</strong>: Provides a mapping from the original sample IDs used in the datasets to the ones used in the manuscript.</p> <p>The code for creating these datasets from the original data files are in the `data-raw` directory.</p> <p><strong>Analysis/Figures</strong></p> <p>The analysis code is broken apart by subject and is largely concerned with figure generation. The R<br> scripts are all located in the `inst` folder. To generate all figures, you should run each script in the<br> order specified by the `GenerateFigures.R` file.</p> <p>Figures are output to the `figures` directory, while tables are output to the `tables` directory.</p> <p>Please note: many of the figures used in the manuscript were aesthetically modified after generation (text size, color palette, orientation), precluding exact figure replication</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4