Skip to main content
zenodoopen

Synthetic gene expression data with underlying gene network

<p>This is the synthetic gene expression data along with the underlying gene network used in the simulation studies of Hu and&nbsp;Szymczak (2023) for evaluating network-guided random forest.</p> <p>In this dataset we consider the situation of 1000 genes and 1000 samples each for training and testing sets. Each file contains a list of 100 replications of the considered scenario which can be identified via the file name. In particular, we consider 6 different scenarios depending on the number of disease modules and how are&nbsp;the effects of disease genes&nbsp;distributed within the disease module. When there are&nbsp;disease genes, we also consider 3 different levels of effect sizes. The binary responses are then generated via a logistic regression model.&nbsp;More details on these scenarios and the data generation mechanism can be found in&nbsp;Hu and&nbsp;Szymczak (2023).</p> <p>The data is generated by the function <em>gen_data</em> in R package <em>networkRF</em> which can be accessed at&nbsp;https://github.com/imbs-hl/networkRF. To obtain the datasets with 3000 genes, which is the other part of the data used in the simulation studies of&nbsp;Hu and&nbsp;Szymczak (2023), simply modify the <em>num.var</em> argument of the function&nbsp;<em>gen_data.</em>&nbsp;More descriptions on the implementation and the format of the output can&nbsp;be found in the help page of the R package.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
8
Engagement
0

Topics