Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
28
datasets available to search
ShareScore release 0.9.0
Dataset results
28 results for “counterfactual”
Water levels at tide gauges from: Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution
<p>Data to reproduce the analysis of the Hourly Coastal water levels with Counterfactual (HCC) dataset, presented in the publication "<strong>Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution</strong>" published in Earth System Science Data (ESSD). </p><p>Note that in this repository, water levels are only provided tide gauge locations which were used for the analysis presented in the paper. The full Hourly Coastal water levels with Counterfactual (HCC) dataset is published in the <a href="https://doi.org/10.48364/ISIMIP.749905">ISIMIP repository</a>.</p><h2>File Descriptions</h2><h4>HCC_analysis_and_plots.ipynb</h4><p>This jupyter-notebook contains all scripts to produce the plots presented in the paper. Make sure that all necessary python packages are installed. The script assumes all netCDF files from this repository to be stored in a sub-directory called "data".</p><h3>hcc_gesla3_99pctl_surge_2011_2015.nc</h3><p>Extreme surge levels from 2011-2015 at 999 GESLA-3 tide gauge stations with at least 90 percent of data in the considered period. As astronomical tides are removed from the modeled and observed water levels to yield the surge component. The file also contains monthly relative water levels and monthly geocentric water levels from 1900-2015 from the HCC dataset.</p><h4>Variables:</h4><ul><li><i>observed_99pctl_surge_level_anomaly</i> -- 99th percentile of daily maximum surge level anomalies from 2011-2015</li><li><i>hcc_99pctl_surge_level_anomaly -- </i>HCC surge level anomalies at the same time steps as <i>observed_99pctl_surge_level_anomaly</i></li><li><i>hcc_counterfactual_99pctl_surge_level_anomaly</i> -- HCC counterfactual surge levels at the same time steps as <i>observed_99pctl_surge_level_anomaly</i></li><li><i>hcc_water_level_monthly</i> – Monthly relative water level from 1900-2015</li><li><i>hcc_geocentric_water_level_monthly</i> – Monthly geocentric water level from 1900-2015</li></ul><h3>hcc_hr_psmsl_water_level_monthly_1900_2015.nc</h3><p>Monthly water levels at 663 PSMSL tide gauge stations of at least 20 year length and with at least 30 percent data coverage in the 1993-2012 period. The file contains data from the HCC, HR and PSMSL datasets. To align PSMSL and HR with HCC, the 1993-2012 average from PSMSL and HR is removed from each of those datasets respectively and the 1993-2012 average of HCC is added. The average is calculated only over all time steps where the associated observational record has valid data.</p><h4>Variables:</h4><ul><li><i>hcc_water_level_monthly</i> – Monthly relative water level from the HCC dataset</li><li><i>hr_aligned_water_level_monthly</i> -- Monthly relative water level from the HR dataset, aligned with <i>hcc_water_level_monthly</i></li><li><i>psmsl_aligned_water_level_monthly</i> -- Monthly relative water level from the PSMSL database, aligned with <i>hcc_water_level_monthly</i></li></ul><h3>hcc_codec_hr_gesla3_water_level_hourly_monthly_1979_2015.nc</h3><p>Hourly water levels at 1040 GESLA-3 tide gauge stations which have at least 30 percent of valid observations between 1979 and 2015. The file contains data from the HCC, CoDEC, HR and GESLA-3 datasets. The different records are not vertically aligned.</p><h4>Variables:</h4><ul><li><i>gesla3_water_level_hourly</i> -- Hourly relative water level from the GESLA3 database</li><li><i>hcc_water_level_hourly</i> -- Hourly relative water level from the HCC dataset</li><li><i>codec_water_level_hourly</i> -- Hourly relative water level from the CoDEC dataset</li><li><i>hr_water_level_monthly</i> -- Monthly relative water level from the HR dataset</li></ul><h3> </h3><h3>hcc_gesla3_water_level_hourly_2011_2015.nc</h3><p>Water levels from the HCC and GESLA-3 datasets, only for tide gauge stations with a complete record in the period 2011-2015 and associated HCC grid points.</p><h4>Variables:</h4><ul><li><i>gesla3_water_level_hourly</i> -- Hourly relative water level from the GESLA3 database</li><li><i>hcc_water_level_hourly</i> -- Hourly relative water level from the HCC dataset</li></ul><h3>slr_ds_psmsl_selected.nc</h3><p>Linear estimates of relative sea level rise from 1900 to 2015. Data is provided at 663 PSMSL tide gauge stations of at least 20 year length and with at least 30 percent data coverage in the 1993-2012 period. Estimates are calculated for the HCC, HR and PSMSL datasets.</p><h4>Variables:</h4><ul><li><i>psmsl_rslr, psmsl_rslr_lower, psmsl_rslr_upper</i> -- Relative sea level rise for PSMSL with lower and upper bounds for a 95 percent confidence interval</li><li><i>hcc_long_rslr, hcc_long_rslr_lower, hcc_long_rslr_upper </i>-- Relative sea level rise for HCC with lower and upper bounds for a 95 percent confidence interval</li><li><i>hr_rslr, hr_rslr_lower, hr_rslr_upper</i> -- Relative sea level rise for HR with lower and upper bounds for a 95 percent confidence interval</li></ul><h3>reg_mask_xr.nc</h3><p>Split of the world into 7 ocean basins: Indian Ocean - South Pacific, Northwest Pacific, East Pacific, South Atlantic, Subtropical North Atlantic, Subpolar North Atlantic West and Subpolar North Atlantic East.</p><h4>Variables:</h4><p><i>reg_mask</i> – Float value, representing the ocean basins</p><p> </p>
Synthetic Data for Uplift Modeling and Heterogenous Treatment Effect with Known Counterfactuals and ITE
<p>This dataset is designed and simulated for evaluating uplift modeling. The data generation process is based on a logistic regression model - no real data is included or used for generating this dataset.</p> <p>This dataset has several signatures:</p> <ul> <li>It generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</li> <li>The true counterfactual outcomes under control and treatment are known for each user, as well as the true ITE (Individual treatment effect).</li> </ul> <p>This dataset consists of 50 trials (replicates with different random seeds), each trial with 20,000 samples and 36 features. The outcome variable is binary, which makes this dataset for classification problems. The samples are equally split for the control and treatment groups (10,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect.</p> <p>To simulate the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <p> Trial ID: 'trial_id'<br> Experiment group label: 'treatment_group_key'<br> Outcome variable (classification label): 'conversion'<br> Feature names: ['x1_informative',<br> 'x2_informative',<br> 'x3_informative',<br> 'x4_informative',<br> 'x5_informative',<br> 'x6_informative',<br> 'x7_informative',<br> 'x8_informative',<br> 'x9_informative',<br> 'x10_informative',<br> 'x11_irrelevant',<br> 'x12_irrelevant',<br> 'x13_irrelevant',<br> 'x14_irrelevant',<br> 'x15_irrelevant',<br> 'x16_irrelevant',<br> 'x17_irrelevant',<br> 'x18_irrelevant',<br> 'x19_irrelevant',<br> 'x20_irrelevant',<br> 'x21_irrelevant',<br> 'x22_irrelevant',<br> 'x23_irrelevant',<br> 'x24_irrelevant',<br> 'x25_irrelevant',<br> 'x26_irrelevant',<br> 'x27_irrelevant',<br> 'x28_irrelevant',<br> 'x29_irrelevant',<br> 'x30_irrelevant',<br> 'x31_uplift_increase',<br> 'x32_uplift_increase',<br> 'x33_uplift_increase',<br> 'x34_uplift_increase',<br> 'x35_uplift_increase',<br> 'x36_uplift_increase']<br> True underlying control conversion probability: 'control_conversion_prob'<br> True underlying treatment conversion probability: 'treatment1_conversion_prob'<br> True treatment effect: 'treatment1_true_effect'</p>
The SPIN covid19 RMRIO dataset: Global trade network data for the years 2016-2026 reflecting macroeconomic effects of the covid19 pandemic - C. Data for 2020 - 2026 - Counterfactual scenario
<p>The SPIN covid19 RMRIO dataset is a time series of MRIO tables covering years from 2016-2026 on a yearly basis. The dataset covers 163 sectors in 155 countries.</p> <p>This repository includes data for years from 2020 to 2026 (<em>counterfactual</em> scenario).<br> Code, method material and data for years 2016-2019 are stored in the following repository: <a href="http://doi.org/10.5281/zenodo.5713811">10.5281/zenodo.5713811</a><br> Data for the <em>covid</em> scenario are stored in the following repository: <a href="https://doi.org/10.5281/zenodo.5713825">10.5281/zenodo.5713825</a></p> <p>Tables are generated using the <a href="https://github.com/TBeaufils/SPIN">SPIN method</a>, based on the <a href="https://doi.org/10.5281/ZENODO.3993659">RMRIO tables</a> for the year 2015, GDP, imports and exports data from the <a href="https://data.imf.org/?sk=4c514d48-b6ba-49ed-8ab9-52b0c1a0179b">International Financial Statistics</a> (IFS) and the World Economic Outlooks (WEO) of <a href="https://www.imf.org/en/Publications/WEO/weo-database/2019/October">October 2019</a> and <a href="https://www.imf.org/en/Publications/WEO/weo-database/2021/April">April 2021</a>.</p> <p>The<em> counterfactual</em> scenario is in line with October 2019 WEO's data and simulates the global economy without Covid 19.</p> <p>All tables are labelled in 2015 US$ and valued in basic prices.</p>
Input data for: Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution
<p>This Zenodo archive contains essential input datasets utilized in our <a href="https://doi.org/10.5194/essd-2023-112">research study</a> titled "Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution". </p><p>This archive contains only input data. The Hourly Coastal water levels with Counterfactual (HCC) dataset is published in the <a href="https://doi.org/10.48364/ISIMIP.749905">ISIMIP repository</a>.</p><p><strong>Datasets Included</strong>:</p><p><strong>CoDEC (Coastal Dataset for the Evaluation of Climate Impact)</strong>:</p><ul><li>This dataset is described in<a href="https://doi.org/10.3389/fmars.2020.00263"> Muis et al. (2020)</a></li><li><strong>cf_esl folder</strong>: Contains data representing total CoDEC water levels. Individual NetCDF files store data for each grid point.</li><li><strong>cf_tides folder</strong>: This folder holds data related to tidal elevation.</li><li><strong>coor_coastal.nc</strong>: A NetCDF file featuring the spatial grid utilized in CoDEC. This dataset comprises only coastal grid points.</li></ul><ol><li><strong>HR (Hybrid Reconstructions)</strong>:<ul><li><strong>HybridRec_Upd0422.mat</strong>: This file contains data from the Hybrid Reconstructions dataset (<a href="https://doi.org/10.1038/s41558-019-0531-8">Dangendorf et al 2019</a>), aligned to the CoDEC grid, and includes satellite altimetry integral to producing the Hybrid Reconstructions dataset. Each row corresponds to one grid point on the CoDEC grid. For ease of use in our applications, we offer a preprocessing script in our <a href="https://doi.org/10.5281/zenodo.7771501">source code</a> named split_hr_dataset_to_stations.py.</li></ul></li></ol><p>We here provide the specific versions of HR and CoDEC that are used in our study to ensure accurate replication.</p>
Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System
Open the record for dataset details and reuse information.
Research data: Counterfactual assessment of protected area avoided deforestation in Cambodia version 4
<p>This dataset includes the data, the R scripts used for analysis and results that are the basis of the journal article: Black, B., Anthony, B. In review. Counterfactual assessment of protected area avoided deforestation in Cambodia: Trends in effectiveness, spillover effects and the influence of establishment date. Global Ecology and Conservation.</p> <p>Each folder includes a specific readme file in .txt format which includes metadata and instructions for reproducing the research.</p> <p> </p>
Data from ATTRICI 1.1 - counterfactual climate for impact attribution
<p>Data as produced and presented in the publication</p> <pre><strong>ATTRICI v1.1 - counterfactual climate for impact attribution</strong></pre> <p>in Geoscientific Model Development.</p> <p>Abstract:</p> <p>Attribution in its general definition aims to quantify drivers of change in a system. According to IPCC WGII a change in a natural, human or managed system is attributed to climate change by quantifying the difference between the observed state of the system and a counterfactual baseline that characterizes the system’s behavior in the absence of climate change, where “climate change refers to any long-term trend in climate, irrespective of its cause". Impact attribution following this definition remains a challenge because the counterfactual baseline cannot be observed. Process-based and empirical impact models can fill this gap as they allow to simulate the counterfactual climate impact baseline. In those simulations, the models are forced by observed direct (human) drivers such as land use changes, changes in water or agricultural management but a counterfactual climate without long-term changes. We here present ATTRICI (ATTRIbuting Climate Impacts), an approach to construct the required counterfactual stationary climate data from observational (factual) climate data. Our method identifies the long-term shifts in the considered daily climate variables that are correlated to global mean temperature change assuming a smooth annual cycle of the associated scaling coefficients for each day of the year. The produced counterfactual climate datasets are used as forcing data within the impact attribution set-up of the Inter-Sectoral Impact Model Intercomparison Project (ISIMIP3a). Our method preserves the internal variability of the observed data in the sense that factual and counterfactual data for a given day have the same rank in their respective statistical distributions. The associated impact model simulations allow for quantifying the contribution of climate change to observed long-term changes in impact indicators and for quantifying the contribution of the observed trend in climate to the magnitude of individual impact events. Attribution of climate impacts to anthropogenic forcing would need an additional step separating anthropogenic climate forcing from other sources of climate trends, which is not covered by our method.</p>
Spreadsheets to model counterfactual tropical forest losses (1990-2019) for Brazil, Democratic Republic of Congo and Indonesia
<p>18 spreadsheets used to simulate the counterfactual forest losses underlying the publication:<br> "Trends in tropical forest loss and the social value of emission reductions"<br> by Thomas Knoke, Nick Hanley, Rosa Maria Roman-Cuesta, Ben Groom, Frank Venmans and Carola Paul published in Nature Sustainability (DOI: 10.1038/s41893-023-01175-9)</p> <p>Each spreadsheet covers a period of five years. The dynamic robust multifunctional land-use allocation model is based on</p> <p>Knoke, T. et al. Accounting for multiple ecosystem services in a simulation of land-use<br> decisions: Does it reduce tropical deforestation? Glob. Chang. Biol. 26, 2403–2420; 10.1111/gcb.15003 (2020).</p> <p>For details see linked publication and README file</p> <p> </p>
SemEval-2020 Task 5: Modelling Causal Reasoning in Language: Detecting Counterfactuals
<p><strong>SemEval-2020 Task 5</strong></p> <p> </p> <p><strong>Subtask-1:</strong> Recognizing Counterfactual Statements (RCS) -- Determine whether a given sentence is counterfactual or not.</p> <p><strong>Subtask-2: </strong>Detecting Antecedent and Consequent (DAC) -- Extract the antecedent and consequent part in a given counterfactual sentence.</p> <p> </p> <p>The released dataset consists of train/test data of both subtask-1 and subtask-2. In our competition, participants could only use the corresponding dataset in each subtask.</p> <p> </p> <p><strong>Task 5 Codalab Website:</strong> <a href="https://competitions.codalab.org/competitions/21691">https://competitions.codalab.org/competitions/21691</a></p>
Data for "Do the Laws of Physics Prohibit Counterfactual Communication"
<p>Data (and code used to analyse the data) used in the weak measurement experiment in the paper "Do the Laws of Physics Prohibit Counterfactual Communication" (arXiv:<a href="https://arxiv.org/abs/1806.01257">1806.01257</a>).</p>
Data and Code for "Informative risk analyses of radiative forcing geoengineering require proper counterfactuals"
<p>This repository contains data and code necessary to generate the figures presented in the maunscript "Data and Code for “Informative risk analyses of radiative forcing geoengineering require proper counterfactuals”, submitted for publication to the Nature journal <em>Communications Earth & Environment</em>.</p>
TRACE-Omicron Policy Counterfactuals Simulation Data
<p>This repository contains the simulation data for the Policy Counterfactuals in the study in "TRACE-Omicron: Policy Counterfactuals to Inform Mitigation of COVID-19 Spread in the United States" published in Advanced Theory and Simulations (doi:10.1002/adts.202300147)</p>
TRACE-Omicron: Epidemiological Counterfactuals; Tractable Strain
<p>This repository contains the simulation data corresponding to the Epidemiological Counterfactuals under the Tractable Strain baseline scenario in "TRACE-Omicron: Policy Counterfactuals to Inform Mitigation of COVID-19 Spread in the United States" published in Advanced Theory and Simulations (doi:10.1002/adts.202300147)</p>
TRACE-Omicron: Epidemiological Counterfactuals; Best Calibration Fit
<p>This repository contains the simulation data corresponding to the Epidemiological Counterfactuals under the Best Calibration Fit baseline scenario in "TRACE-Omicron: Policy Counterfactuals to Inform Mitigation of COVID-19 Spread in the United States" published in Advanced Theory and Simulations (doi:10.1002/adts.202300147)</p>
TRACE-Omicron: Epidemiological Counterfactuals; High Immune Escape
<p>This repository contains the simulation data corresponding to the Epidemiological Counterfactuals under the High Immune Escape baseline scenario in "TRACE-Omicron: Policy Counterfactuals to Inform Mitigation of COVID-19 Spread in the United States" published in Advanced Theory and Simulations (doi:10.1002/adts.202300147)</p>
Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System
Open the record for dataset details and reuse information.
Studying Therapy Effects and Disease Outcomes in Silico using Artificial Counterfactual Tissue Samples
<p>Counterfactual samples created by our generative method the CF-HistoGAN introduced in https://doi.org/10.48550/arXiv.2302.03120. This model was trained on and transformed data from 2 different datasets: the colorectal cancer (CRC) dataset by Schürch et al. https://doi.org/10.1016/j.cell. 2020.07.005 and the cutaneous T cell lymphoma (CTCL) by (Phillips et al https://doi.org/10.1038/s41467-021-26974-6.</p>
Counterfactual Strategies, Physical Activity, and Wearable Trackers
ClinicalTrials.gov study NCT05192226. IPD Sharing: YES. Countries: 1. Publications: 12.
CoPhy: Counterfactual Learning of Physical Dynamics (Benchmark Dataset)
<p> </p> <p>Benchmark website: https://projet.liris.cnrs.fr/cophy/</p> <p>Understanding causes and effects in mechanical systems is an essential component of reasoning in the physical world. This work poses a new problem of counterfactual learning of object mechanics from visual input. We develop the COPHY benchmark to assess the capacity of the state-of-the-art models for causal physical reasoning in a synthetic 3D environment. Having observed a mechanical experiment that involves, for example, a falling tower of blocks, a set of bouncing balls or colliding objects, we require to learn to predict how its outcome is affected by an arbitrary intervention on its initial conditions, such as displacing one of the objects in the scene.</p> <p>The main objective for the creation of our benchmark is (a) to focus specifically on evaluating capabilities of state of the art models for performing counterfactual reasoning, (b) to be unbiased in terms of distributions of parameters to be estimated and balanced with respect to possible outcomes, and (c) to have sufficient variety in terms of scenarios<br> and latent physical characteristics of the scene that are not visually observed and therefore can act<br> as confounders.</p> <p>If you use this benchmark, you need to cite the following paper:</p> <p>Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, Christian Wolf. COPHY: Counterfactual Learning of Physical Dynamics. pre-print arXiv:1909.12000, 2019.</p>
Dataset - What enabled the forest transition? A socio-ecological counterfactual assessment for Austria, 1830-1910
<p>This excel file contains data on forest change, food and energy provision, as well as greenhouse gas emissions from agriculture, forestry and energy use in Austria 1830-1910, and assessments of five counterfactual scenarios in the absence of specific forest relief processes. The data were used to create figures 1a and b, 2a, b, c, and d, 3a, b, c, and d, and 4a and b in the manuscript "What enabled the forest transition? A socio-ecological counterfactual assessment for Austria, 1830-1910" submitted to Global Biogeochemical Cycles in July 2020 by Simone Gingrich, Christian Lauk, Fridolin Krausmann, Karl-Heinz Erb, Julia Le Noë.</p> <p> </p> <p>Dr. Simone Gingrich<br> Institute of Social Ecology (SEC)<br> Department of Economics and Social Sciences (WiSo)</p> <p>University of Natural Resources & Life Sciences, Vienna (BOKU)</p> <p><br> Post: Schottenfeldgasse 29, 1070 Vienna, Austria<br> Fon: <a>+43 1 47654-73724</a><br> Web: <a href="http://www.wiso.boku.ac.at/wiso/sec/">http://www.wiso.boku.ac.at/wiso/sec/</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.