Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7 results for “uplift modeling”

Learn how ShareScore rates datasets ↗
zenodo44/100

Synthetic Data for Uplift Modeling and Heterogenous Treatment Effect with Known Counterfactuals and ITE

<p>This dataset is designed and simulated for evaluating uplift modeling. The data generation process is based on a logistic regression model - no real data is included or used for generating this dataset.</p> <p>This dataset has several signatures:</p> <ul> <li>It generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</li> <li>The true counterfactual outcomes under control and treatment are known for each user, as well as the true ITE (Individual treatment effect).</li> </ul> <p>This dataset consists of 50 trials (replicates with different random seeds), each trial with 20,000 samples and 36 features. The outcome variable is binary, which makes this dataset for classification problems. The samples are equally split for the control and treatment groups (10,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect.</p> <p>To simulate the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <p>&nbsp;&nbsp;&nbsp; Trial ID: &#39;trial_id&#39;<br> &nbsp;&nbsp;&nbsp; Experiment group label: &#39;treatment_group_key&#39;<br> &nbsp;&nbsp;&nbsp; Outcome variable (classification label):&nbsp; &#39;conversion&#39;<br> &nbsp;&nbsp;&nbsp; Feature names: [&#39;x1_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x2_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x3_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x4_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x5_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x6_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x7_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x8_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x9_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x10_informative&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x11_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x12_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x13_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x14_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x15_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x16_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x17_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x18_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x19_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x20_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x21_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x22_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x23_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x24_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x25_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x26_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x27_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x28_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x29_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x30_irrelevant&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x31_uplift_increase&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x32_uplift_increase&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x33_uplift_increase&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x34_uplift_increase&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x35_uplift_increase&#39;,<br> &nbsp;&nbsp;&nbsp; &#39;x36_uplift_increase&#39;]<br> &nbsp;&nbsp;&nbsp; True underlying control conversion probability: &#39;control_conversion_prob&#39;<br> &nbsp;&nbsp;&nbsp; True underlying treatment conversion probability: &#39;treatment1_conversion_prob&#39;<br> &nbsp;&nbsp;&nbsp; True treatment effect:&nbsp; &#39;treatment1_true_effect&#39;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Processed model output of the climate simulation in the study: The effects of diachronous surface uplift of the European Alps on regional climate and the isotopic composition of precipitation (δ18Op) [Boateng et al.]

<p><strong>The geodynamic evolution of the Alps suggests that the Alps did not rise monotonically due to the different post-collisional processes such as slab break-off. However, understanding such subsurface dynamics would require adequate knowledge about its surface uplift history. Stable isotope paleoaltimetry methods are widely used to infer past surface elevation using geologic archives. However, its accurate interpretation relies on attributing the extracted isotopic signal from proxies to surface uplift despite other influences such as climate. To resolve this issue, topographic sensitivity experiments across the Alps are used to investigate the impacts of the diachronous surface uplift on regional climate and &delta;18Op. The Atmospheric General Circulation Model ECHAM5 with water isotope tracking capabilities (ECHAM5-wiso) is used to simulate the climate with varied topographic scenarios. We present the processed (long-term means) model output of the relevant climate variables (i.e &delta;18Op, near-surface temperature, precipitation amount, near-surface meridional and zonal winds, mean sea level pressure, and elevation) in response to the changes in topography. The file names are representative of the topographic scenarios used for the simulations. For example, the file &ldquo;W2E1.nc&rdquo; is the model output produced by a topographic scenario in which the topography across the west-central Alps was set to 200% of its modern height, and the Eastern Alps were kept at 100%. The &ldquo;CTL.nc&rdquo; file contains model output from the control simulation that uses present-day topography. The datasets for instance can be used to select far-field sampling points for the &delta;-&delta; paleoaltimetry method that are not significantly affected by the topographic changes.</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Processed model outputs for "South Asian summer monsoon enhanced by the uplift of Iranian Plateau in Middle Miocene"

<p>This dataset contains processed model outputs from modeling experiments performed in Zuo et al. (2024), including a set of 12 experiment with different CO2 concentrations and topography during the Middle Miocene. Due to space limitations, we provide the summer(JJA) mean climatology data for each experiment and some data that can be used to reproduce the figures in this paper. The raw data can be obtained by contacting the corresponding author.</p> <p><strong><span>Table 1. </span></strong><span>Simulations performed with CESM1.2 in this study.</span></p> <table> <tbody> <tr> <td> <p><span>experiment</span></p> </td> <td> <p><span>Geolography</span></p> </td> <td> <p><span>vegetation</span></p> </td> <td> <p><span>CO2</span></p> <p><span>(ppm)</span></p> </td> <td> <p><span>IP</span></p> </td> <td> <p><span>HM</span></p> </td> </tr> <tr> <td> <p><span>piControl</span></p> </td> <td> <p><span>Modern</span></p> </td> <td> <p><span>Modern</span></p> </td> <td> <p><span>280 </span></p> </td> <td> <p><span>Modern</span></p> </td> <td> <p><span>Modern</span></p> </td> </tr> <tr> <td> <p><span>MMIO</span></p> <p><span>(IP100HM80)</span></p> </td> <td> <p><span>M.Miocene*</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> </tr> <tr> <td> <p><span>IP0HM0</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>0</span></p> </td> <td> <p><span>0</span></p> </td> </tr> <tr> <td> <p><span>IP50HM0</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>50%</span></p> </td> <td> <p><span>0</span></p> </td> </tr> <tr> <td> <p><span>IP100HM0</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>100%</span></p> </td> <td> <p><span>0</span></p> </td> </tr> <tr> <td> <p><span>IP0HM100</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>0</span></p> </td> <td> <p><span>100%**</span></p> </td> </tr> <tr> <td> <p><span>IP50HM100</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>50%</span></p> </td> <td> <p><span>100%</span></p> </td> </tr> <tr> <td> <p><span>IP100HM100</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>400</span></p> </td> <td> <p><span>100%</span></p> </td> <td> <p><span>100%</span></p> </td> </tr> <tr> <td> <p><span>MMIO280</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>280</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> </tr> <tr> <td> <p><span>MMIO560</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>560</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> </tr> <tr> <td> <p><span>MMIO800</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>800</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> </tr> <tr> <td> <p><span>MMIO1000</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> <td> <p><span>1000</span></p> </td> <td> <p><span>M. Miocene</span></p> </td> <td> <p><span>M.Miocene</span></p> </td> </tr> </tbody> </table> <p><span>*M.Miocen</span><span>e</span><span>: Middle Miocene</span></p> <p><span>** 100% of the height of modern HM.</span></p>

opencc-by-4.0Jun 2024View details →
dryad32/100

High-resolution modelling of uplift landscapes can inform micro-siting of wind turbines for soaring raptors

<p>Collision risk of soaring birds is partly associated with updrafts to which they are attracted. To identify risk-enhancing landscape features, a micro-siting tool was developed to model orographic and thermal updraft velocities from high-resolution remote sensing data. The tool was applied to the island of Hitra, and validated using GPS-tracked white-tailed eagles (<i>Haliaeetus albicilla</i>). Resource selection functions predicted that eagles preferred ridges with high orographic uplift, especially at flight altitudes within the rotor-swept zone (40-110 m). Flight activity was negatively associated with the widely distributed areas with high thermal uplift at lower flight altitudes (&lt;110 m). Both the existing wind-power plant and planned extension are placed at locations rendering maximum orographic updraft velocities around the minimum sink rate for white-tailed eagles (0.75 m/s) but slightly higher thermal updraft velocities. The tool can contribute to improved micro-siting of wind turbines to reduce environmental impacts, especially for soaring raptors.</p>

opencc-zeroJul 2021View details →
dryad32/100

High-resolution modelling of uplift landscapes can inform micro-siting of wind turbines for soaring raptors

Open the record for dataset details and reuse information.

publicJul 2021View details →
zenodo28/100

Synthetic Data Set for Uplift Modeling (One Trial)

<p>This dataset is designed and simulated for evaluating uplift modeling and feature selection methods.</p> <p>This dataset contains 10,000 samples and 36 features (one trial).</p> <p>The samples are equally split for control and treatment group.</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect. To model the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <ul> <li>Experiment group label: &#39;treatment_group_key&#39;</li> <li>Feature names: [&#39;x1_informative&#39;,<br> &#39;x2_informative&#39;,<br> &#39;x3_informative&#39;,<br> &#39;x4_informative&#39;,<br> &#39;x5_informative&#39;,<br> &#39;x6_informative&#39;,<br> &#39;x7_informative&#39;,<br> &#39;x8_informative&#39;,<br> &#39;x9_informative&#39;,<br> &#39;x10_informative&#39;,<br> &#39;x11_irrelevant&#39;,<br> &#39;x12_irrelevant&#39;,<br> &#39;x13_irrelevant&#39;,<br> &#39;x14_irrelevant&#39;,<br> &#39;x15_irrelevant&#39;,<br> &#39;x16_irrelevant&#39;,<br> &#39;x17_irrelevant&#39;,<br> &#39;x18_irrelevant&#39;,<br> &#39;x19_irrelevant&#39;,<br> &#39;x20_irrelevant&#39;,<br> &#39;x21_irrelevant&#39;,<br> &#39;x22_irrelevant&#39;,<br> &#39;x23_irrelevant&#39;,<br> &#39;x24_irrelevant&#39;,<br> &#39;x25_irrelevant&#39;,<br> &#39;x26_irrelevant&#39;,<br> &#39;x27_irrelevant&#39;,<br> &#39;x28_irrelevant&#39;,<br> &#39;x29_irrelevant&#39;,<br> &#39;x30_irrelevant&#39;,<br> &#39;x31_uplift_increase&#39;,<br> &#39;x32_uplift_increase&#39;,<br> &#39;x33_uplift_increase&#39;,<br> &#39;x34_uplift_increase&#39;,<br> &#39;x35_uplift_increase&#39;,<br> &#39;x36_uplift_increase&#39;]</li> <li>Outcome variable: &nbsp;&#39;conversion&#39;</li> <li>True underlying control conversion probability: &#39;control_conversion_prob&#39;</li> <li>True underlying treatment conversion probability: &#39;treatment1_conversion_prob&#39;</li> <li>True treatment effect: &nbsp;&#39;treatment1_true_effect&#39;</li> <li>Note columns names with &#39;_transformed&#39; suffix are feature variables used in the intermediate steps during the data generation, that should be excluded for model training.</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo24/100

Synthetic Data Set for Uplift Modeling

<p>This dataset is designed and simulated for evaluating uplift modeling and feature selection methods. The main feature of this dataset is that it generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</p> <p>This dataset consists of 100 trials (replicates with different random seeds), each trial with 10,000 samples and 36 features. The outcome variable is binary, that makes this dataset for classification problem. The samples are equally split for control and treatment group (5,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect. To model the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <ul> <li>Trial ID: &#39;trial_id&#39;</li> <li>Experiment group label: &#39;treatment_group_key&#39;</li> <li>Outcome variable (classification label): &nbsp;&#39;conversion&#39;</li> <li>Feature names: [&#39;x1_informative&#39;,<br> &#39;x2_informative&#39;,<br> &#39;x3_informative&#39;,<br> &#39;x4_informative&#39;,<br> &#39;x5_informative&#39;,<br> &#39;x6_informative&#39;,<br> &#39;x7_informative&#39;,<br> &#39;x8_informative&#39;,<br> &#39;x9_informative&#39;,<br> &#39;x10_informative&#39;,<br> &#39;x11_irrelevant&#39;,<br> &#39;x12_irrelevant&#39;,<br> &#39;x13_irrelevant&#39;,<br> &#39;x14_irrelevant&#39;,<br> &#39;x15_irrelevant&#39;,<br> &#39;x16_irrelevant&#39;,<br> &#39;x17_irrelevant&#39;,<br> &#39;x18_irrelevant&#39;,<br> &#39;x19_irrelevant&#39;,<br> &#39;x20_irrelevant&#39;,<br> &#39;x21_irrelevant&#39;,<br> &#39;x22_irrelevant&#39;,<br> &#39;x23_irrelevant&#39;,<br> &#39;x24_irrelevant&#39;,<br> &#39;x25_irrelevant&#39;,<br> &#39;x26_irrelevant&#39;,<br> &#39;x27_irrelevant&#39;,<br> &#39;x28_irrelevant&#39;,<br> &#39;x29_irrelevant&#39;,<br> &#39;x30_irrelevant&#39;,<br> &#39;x31_uplift_increase&#39;,<br> &#39;x32_uplift_increase&#39;,<br> &#39;x33_uplift_increase&#39;,<br> &#39;x34_uplift_increase&#39;,<br> &#39;x35_uplift_increase&#39;,<br> &#39;x36_uplift_increase&#39;]</li> <li>True underlying control conversion probability: &#39;control_conversion_prob&#39;</li> <li>True underlying treatment conversion probability: &#39;treatment1_conversion_prob&#39;</li> <li>True treatment effect: &nbsp;&#39;treatment1_true_effect&#39;</li> </ul>

opencc-by-4.0Feb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record