Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.7.1
Dataset results
5 results for “heterogeneous treatment effect”
Synthetic Data for Uplift Modeling and Heterogenous Treatment Effect with Known Counterfactuals and ITE
<p>This dataset is designed and simulated for evaluating uplift modeling. The data generation process is based on a logistic regression model - no real data is included or used for generating this dataset.</p> <p>This dataset has several signatures:</p> <ul> <li>It generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</li> <li>The true counterfactual outcomes under control and treatment are known for each user, as well as the true ITE (Individual treatment effect).</li> </ul> <p>This dataset consists of 50 trials (replicates with different random seeds), each trial with 20,000 samples and 36 features. The outcome variable is binary, which makes this dataset for classification problems. The samples are equally split for the control and treatment groups (10,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect.</p> <p>To simulate the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <p> Trial ID: 'trial_id'<br> Experiment group label: 'treatment_group_key'<br> Outcome variable (classification label): 'conversion'<br> Feature names: ['x1_informative',<br> 'x2_informative',<br> 'x3_informative',<br> 'x4_informative',<br> 'x5_informative',<br> 'x6_informative',<br> 'x7_informative',<br> 'x8_informative',<br> 'x9_informative',<br> 'x10_informative',<br> 'x11_irrelevant',<br> 'x12_irrelevant',<br> 'x13_irrelevant',<br> 'x14_irrelevant',<br> 'x15_irrelevant',<br> 'x16_irrelevant',<br> 'x17_irrelevant',<br> 'x18_irrelevant',<br> 'x19_irrelevant',<br> 'x20_irrelevant',<br> 'x21_irrelevant',<br> 'x22_irrelevant',<br> 'x23_irrelevant',<br> 'x24_irrelevant',<br> 'x25_irrelevant',<br> 'x26_irrelevant',<br> 'x27_irrelevant',<br> 'x28_irrelevant',<br> 'x29_irrelevant',<br> 'x30_irrelevant',<br> 'x31_uplift_increase',<br> 'x32_uplift_increase',<br> 'x33_uplift_increase',<br> 'x34_uplift_increase',<br> 'x35_uplift_increase',<br> 'x36_uplift_increase']<br> True underlying control conversion probability: 'control_conversion_prob'<br> True underlying treatment conversion probability: 'treatment1_conversion_prob'<br> True treatment effect: 'treatment1_true_effect'</p>
Polygenic modelling of treatment effect heterogeneity
<p>Xu, ZM, Burgess, S. Polygenic modelling of treatment effect heterogeneity. <em>Genetic Epidemiology</em>. 2020; 1– 12. <a href="https://doi.org/10.1002/gepi.22347">https://doi.org/10.1002/gepi.22347</a></p> <p>File contains: </p> <p>1. Summary statistics of genome-wide interaction study (between moderating variants and HMGCR pharmacomimetic score)</p> <p>2. Random forest of interaction tree (RFIT) objects. </p> <p>See README for detail</p>
Finding treatment effects in Alzheimer’s trials in the face of disease progression heterogeneity
Open the record for dataset details and reuse information.
Spatial transcriptomics reveals the pharmacological effects on tumors and TME structural heterogeneity under MASK vaccine treatment
GEO Series GSE248356. Mus musculus. 3 samples. Type: Expression profiling by high throughput sequencing.
Development and Prospective Validation of a Heterogeneous Treatment Effect-Based Decision Model for Transarterial Chemoembolization Combined With or Without Atezolizumab Plus Bevacizumab in Unresectab
ClinicalTrials.gov study NCT07109336. IPD Sharing: Not stated. Countries: 0. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.