Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
376
datasets available to search
ShareScore release 0.9.0
Dataset results
376 results for “Causality”
MAL02 Causal loop diagrams for Soutwest Messinia (Greece)
<p>This dataset includes the causal loop diagrams (CLDs) developed by the H2020 COASTAL project’s MAL #2 for the Social Ecological Land Sea System of Soutwest Messinia. These CLDs were developed as a result of 6 consecutive stakeholder workshops representing the different interest groups in the area (Agriculture, olive oil industry, tourism, fisheries, public administration, NGOs and Research Institutes) aiming in describing the functioning of the territory in a systemic way, highlighting its main components and interactions among them. The CLD's were created using Vensim PLE software and lead to the creation of a Systems Dynamic model for the area.</p>
Study: Layout of Causal Graphs
<p>This repository contains the material and obtained data of an eye tracking study on the topic "Layout of Causal Graphs".</p> <p>For more information, please feel free to contact <lisa.grabinger@oth-regensburg.de>.</p>
Microdata used to construct the Causal Diagrams to model investment decisions related to the energy transition
<ul> <li><strong>Name</strong>: Microdata used to construct the Causal Diagrams to model investment decisions related to the energy transition</li> <li><strong>Summary</strong>: This dataset contains answers from a panel of experts to build a) a taxonomy of determinants that explain the investment decision making on assets related to the energy transition, b) the individual contributions when sorting the taxonomy of determinantes on the different stages of the transtheoretical model for different archetypes of persons and c) the causal diagrams agreed between the different groups of experts.</li> <li><strong>License</strong>: cc-BY-SA</li> <li><strong>Acknowledge</strong>: These data have been collected in the framework of the WHY project. This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 891943.</li> <li><strong>Disclaimer</strong>: The sole responsibility for the content of this publication lies with the authors. It does not necessarily reflect the opinion of the Executive Agency for Small and Medium-sized Enterprises (EASME) or the European commission (Ec). EASME or the Ec are not responsible for any use that may be made of the information contained therein.</li> <li><strong>Collection Date</strong>: 22/07/2022</li> <li><strong>Publication Date</strong>: 01/06/2024</li> <li><strong>DOI</strong>: 10.5281/zenodo.11234441</li> <li><strong>Other repositories:</strong></li> <li><strong>Author</strong>: University of Deusto</li> <li><strong>Objective of collection</strong>: This data was originally collected to build a set of causal diagrams of the .</li> <li><strong>Description:</strong> <br> <ul> <li><strong>Scenarios: </strong>This dataset contains the description of 20 different scenarios used in this research activity. </li> <li><strong>File 1 - individual reasons to be coded<br></strong>This dataset compiles the reasons given by experts of different panels of the Intrinsic and Extrinsic Determinants, and the Barriers and potential Rebound effects of citizens towards a set of 20 different scenarios. The file contains the following sheets:<br> <ul> <li><strong>Methodology</strong>: Methodology followed by the coders.</li> <li><strong>Help</strong>: Short summary of the Social Cognitive Theor and Self Determination Theory used for coding. </li> <li><strong>Glossary</strong>: Glossary of terms build by the experts coding the answers. </li> <li><strong>Appliances/Flexibility/Buildings/Mobility</strong>: The contributions of each expert, the code provided by the two researchers and the consensus achived. </li> <li><strong>Summary</strong>: Assesment of the results.</li> </ul> </li> <li><strong>File 2 - individual microdata to sort determinants into causal threads from experts</strong>This dataset includes the individual sortings made by the experts of the taxonomy of determinantes into each one of the stages of the transtheoretical model. The file includes one sheet per expert where he/she has sort each determinant for each arquetype into the stage he/she thinks is more relevant to advance to the next step of the TTM. </li> <li><strong>File 3 - collective microdata to sort determinants into causal threads from EU and LATAM experts</strong> <p>This dataset compiles the results, stage by stage, of the consensus reached by each panel regarding the determining factors that make up each of the archetypes in the contexts of Europe (EU) and Latin America (LATAM). And in which stage of the change of the Transtheoretical Model (TTM) the factors should appears.</p> <ul> <li> <p><strong>Stage 1</strong>: The panels reached a consensus on the factors that describe each of the archetypes in their context. In the case of Latin America, for the panels of some countries, the existence of all eight archetypes was not evident. The number of archetypes analysed by each panel is indicated in parentheses in the following list:</p> <ul> <li> <p><strong>European panels</strong>: Group – F (8), Group–A (8). Group–FF (8), Group–M (4)</p> </li> <li> <p><strong>Latin America panels</strong>: Group-MX (5), Group-CO (8), Group-CL (7), Group-SV (7)</p> </li> </ul> </li> </ul> <ul> <li> <p><strong>Stage 2</strong>: For each of the eight archetypes, the results of the consensus for each panel are consolidated in the tabs indicated in the list below. The column on the far right shows the weights (percentage) of each factor in each stage of the TTM: Archetype-EarlyAdopter, Archetype-Uninterested, Archetype-HomoEconomicus, Archetype-Fearful, Archetype-Stubborn, Archetype-Influencer, Archetype-Careful and Archetype-Activist.</p> </li> <li> <p><strong>Stage3</strong>: In the "<em>Archetypes - Consensus Results</em>" tab, the weights of the factors for each archetype are consolidated. The far-right column calculates the average weight of each factor at each stage of the TTM (Transtheoretical Model of Change).</p> </li> <li> <p><strong>Stage 4</strong>. In the “EU vs Latam - split context” sheet, it is presented a comparative assessment between the European and Latin American results. The comparison has four tables:</p> <ul> <li> <p><em>Table (s)</em>: Difference and Agreements between both context: European & Latin American Archetypes. The table highlights the regions of determinants that mark the differences between both contexts for each archetype. If a determinant is identified by both contexts (EU, Latam), it is considered an agreement and allocated to the early TTM stage. The remaining determinants highlight the differences between the two contexts. European (-1) & Latin American (1) Archetypes FINAL Consensus (0) on TTM Stages.</p> </li> <li> <p><em>Table (t)</em>: This table shows the difference (E, L) and agreements (X) between both context: European (E) & Latin American (L) Archetypes.</p> </li> <li> <p><em>Table (t.1)</em>: This table shows just the <strong>agreements</strong> (X) between both context: European & Latin American Archetypes.</p> </li> <li> <p><em>Table (t.2)</em>: Show the difference between both context: European (E) & Latin American Archetypes (L).</p> </li> <li> <p><em>Table (t.3)</em>: This table shows the differences (E, L) and agreements (X) between both contexts: European (E) & Latin American (L) archetypes. In this table, the main regions of factors for each archetype are coloured to highlight the set of factors that make the main differences.</p> </li> </ul> </li> </ul> </li> </ul> </li> <li><strong>5 star</strong>: ⭐⭐⭐</li> <li><strong>Preprocessing steps:</strong> Data transcription from written documents and oral discussions.</li> <li><strong>Reuse:</strong> NA</li> <li><strong>Update policy:</strong> No more updates are planned.</li> <li><strong>Ethics and legal aspects:</strong> Names of the persons involved have been removed. </li> <li><strong>Technical aspects</strong>: </li> <li><strong>Other:</strong></li> </ul>
CausalBench A Comprehensive Benchmark for Evaluating Causal Reasoning Capabilities of Large Language Models
<p>CausalBench is a comprehensive benchmark dataset designed to evaluate the causal reasoning capabilities of large language models. The primary uses of this dataset include, but are not limited to:</p> <p>- Testing the performance of large language models on causal reasoning tasks</p> <p>- Serving as a benchmark dataset for causal reasoning research</p> <p>- Improving and developing new causal reasoning algorithms and models</p>
FIGURE 3 in Reassessing the causal connection between satDNA dynamics and chromosomal evolution in Ctenomys (Rodentia, Ctenomyidae): Unveiling the overlooked importance of the Y chromosome
FIGURE 3 Ancestral RPCS copy number reconstruction. RPCS copy number was mapped along the mtDNA phylogeny using the phytools function anc.ML, for males and females of the Ctenomys Corrientes group separately. The projection of the reconstruction onto the edges of the tree was made with the function contMap. RPCS copy number is expressed as thousands of copies. The scale of the tree is expressed in substitutions per site. Letters A–D correspond to the four main clades of the group.
FIGURE 2 in Reassessing the causal connection between satDNA dynamics and chromosomal evolution in Ctenomys (Rodentia, Ctenomyidae): Unveiling the overlooked importance of the Y chromosome
FIGURE 2 Geographic distribution of mean RPCS copy number. The Corrientes group lineages are surrounded by dashed lines. Locality numbers are: 1 – San Alonso, 2 – Loreto, 3 – Contreras_Cué, 4 – Estancia La Tacuarita, 5 – Saladas Sur, 6 – Saladas, 7 – Santa Rosa, 8 – San Roque, 9 – Estancia San Luis, 10 – Pago Alegre, 11 – Mbarigüí, 12 – Paraje Angostura, 13 – Goya, 14 – Chavarría, 15 – Colonia 3 de abril, 16 – Rincón de Ambrosio.
FIGURE 1 in Reassessing the causal connection between satDNA dynamics and chromosomal evolution in Ctenomys (Rodentia, Ctenomyidae): Unveiling the overlooked importance of the Y chromosome
FIGURE 1 Differences in RPCS copy number in males and females. Scatter plot showing differences in RPCS copy numbers between males and females of the Ctenomys Corrientes group, expressed as thousands of copies. Clades A-D correspond to the four different clades of the phylogeny (figs 3 and 4). A smoothing function was applied with the package ggplot2.
FIGURE 4 in Reassessing the causal connection between satDNA dynamics and chromosomal evolution in Ctenomys (Rodentia, Ctenomyidae): Unveiling the overlooked importance of the Y chromosome
FIGURE 4 Ancestral reconstruction of diploid numbers (2n) and main RPCS reductions and amplifications in the Ctenomys Corrientes group. Ancestral diploid numbers were inferred with the ChromEvol model implemented in RevBayes, over the mtDNA Bayesian phylogeny of the Corrientes group. Numbers in internal nodes/ terminals represent inferred/observed 2n. Colored circles depict 2n (size) and posterior probability of the inferred value (color). Red and green branches depict significant reductions and amplifications in diploid numbers, respectively. Smaller equal-sized black circles show well-supported nodes (posterior probability> 0.75). Black arrowheads denote a marked increase/decrease in RPCS copy numbers (inferred from females). The scale bar is expressed in substitutions per site.
Figure 2. Causal loop diagram for the problem situation-An Efficient Expert System Generator for Qualitative Feed-Back Loop Analysis
<p>The problem situation could be represented in the following form of a causal loop diagram<br> as shown in Figure 2.</p>
Dataset Artifact for paper "Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?"
<p>Artifacts for the paper titled <strong><em>Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?</em></strong>.</p> <p>This artifact repository contains 9 compressed folders, as follows: </p> <table> <tbody> <tr> <td><strong>ID</strong></td> <td><strong>File Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>1</td> <td>syn_circa.zip</td> <td>CIRCA10, and CIRCA50 datasets for Causal Discovery</td> </tr> <tr> <td>2</td> <td>syn_rcd.zip</td> <td>RCD10, and RCD50 datasets for Causal Discovery</td> </tr> <tr> <td>3</td> <td>syn_causil.zip</td> <td>CausIL10, and CausIL50 datasets for Causal Discovery</td> </tr> <tr> <td>4</td> <td>rca_circa.zip</td> <td>CIRCA10, and CIRCA50 datasets for RCA</td> </tr> <tr> <td>5</td> <td>rca_rcd.zip</td> <td>RCD10, and RCD50 datasets for RCA</td> </tr> <tr> <td>6</td> <td>online-boutique.zip</td> <td>Online Boutique dataset for RCA</td> </tr> <tr> <td>7</td> <td>sock-shop-1.zip</td> <td>Sock Shop 1 dataset for RCA</td> </tr> <tr> <td>8</td> <td>sock-shop-2.zip</td> <td>Sock Shop 2 dataset for RCA</td> </tr> <tr> <td>9</td> <td>train-ticket.zip</td> <td>Train Ticket dataset for RCA</td> </tr> </tbody> </table> <p>Each zip file contains the generated/collected data from the corresponding data generator or microservice benchmark systems (e.g., online-boutique.zip contains metrics data collected from the Online Boutique system). </p> <p><strong>Details about the generation of our datasets</strong></p> <p><em>1. Synthetic datasets</em></p> <p>We use three different synthetic data generators from three previous RCA studies [15, 25, 28] to create the synthetic datasets: CIRCA, RCD, and CausIL data generators. Their mechanisms are as follows:<br><br>1. CIRCA datagenerator [28] generates a random causal directed acyclic graph (DAG) based on a given number of nodes and edges. <span>From this DAG, time series data for each node is generated using a </span><span>vector auto-regression (VAR) model. A fault is injected into a node </span><span>by altering the noise term in the VAR model for two timestamps. <br></span><span><br>2. RCD data generator [25] uses the pyAgrum package [3] to generate </span><span>a random DAG based on a given number of nodes, subsequently </span><span>generating discrete time series data for each node, with values ranging from 0 to 5. A fault is introduced into a node by changing its </span><span>conditional probability distribution.<br><br>3. CausIL data generator [15] generates causal graphs and time series data that simulate </span><span>the behavior of microservice systems. It first constructs a DAG of </span><span>services and metrics based on domain knowledge, then generates </span><span>metric data for each node of the DAG using regressors trained on </span><span>real metrics data. Unlike the CIRCA and RCD data generators, the </span><span>CausIL data generator does not have the capability to inject faults.<br><br></span>To create our synthetic datasets, we first generate 10 DAGs whose nodes range from 10 to 50 for each of the synthetic data generators. Next, we generate fault-free datasets using these DAGs with different seedings, resulting in 100 cases for the CIRCA and RCD generators and 10 cases for the CausIL generator. We then create faulty datasets by introducing ten faults into each DAG and generating the corresponding faulty data, yielding 100 cases for the CIRCA and RCD data generators. The fault-free datasets (e.g. `syn_rcd`, `syn_circa`) are used to evaluate causal discovery methods, while the faulty datasets (e.g. `rca_rcd`, `rca_circa`) are used to assess RCA methods. </p> <p><em>2. Data collected from benchmark microservice systems </em></p> <p>We deploy three popular benchmark microservice systems: Sock Shop [6], Online Boutique [4], and Train Ticket [8], on a four-node Kubernetes cluster hosted by AWS. Next, we use the Istio service mesh [2] with Prometheus [5] and cAdvisor [1] to monitor and collect resource-level and service-level metrics of all services, as in previous works [ 25 , 39, 59 ]. To generate traffic, we use the load generators provided by these systems and customise them to explore all services with 100 to 200 users concurrently. We then introduce five common faults (CPU hog, memory leak, disk IO stress, network delay, and packet loss) into five different services within each system. Finally, we collect metrics data before and after the fault injection operation. An overview of our setup is presented in the Figure below.</p> <p></p> <p><strong>Code</strong></p> <p>The code to reproduce the experimental results in the paper is available at <a href="https://github.com/phamquiluan/RCAEval">https://github.com/phamquiluan/RCAEval</a>.</p> <p><strong>References</strong></p> <p>As in our paper.</p>
Supplementary materials for paper "Untangling the drivers of change and policy impact in coastal wetland area in the Yangtze Estuary using causal inference"
<h1>Annual coastal wetland vegetation maps of the Yangtze Estuary from 1986 to 2021</h1> <p> </p> <h2><strong>Basic information</strong></h2> <p>Using remotely sensed data from Google Earth Engine, we generated a 30 m resolution annual dataset of Yangtze Estuary wetland vegetation for the period 1986-2021. This dataset includes three dominant vegetation types (<em>Spartina alterniflora</em>, <em>Phragmites australis</em>, and <em>Scirpus mariqueter</em>) and tidal flat areas. We combined fieldwork data and high-resolution images for accuracy assessment, achieving an overall accuracy exceeding 80% in different years. Detailed information about the mapping methods can be found in our paper and accompanying supplementary materials.</p> <h2><strong>Notes:</strong></h2> <p>In the image classification scheme: 0-Tidal flats, 1-<em>Spartina alterniflora, </em>2-<em>Phragmites australis, </em>3-<em>Scirpus mariqueter.</em></p> <h2><strong>Usage Policy:</strong></h2> <p>This dataset is a collaborative effort between East China Normal University and Deakin University. If you plan to use our data in <strong>a scientific analysis paper or other research work</strong>, we strongly recommend contacting us in advance to seek our opinions, <strong>citing the unique DOI of this dataset</strong>, and considering acknowledging our contributions or including us as co-authors.</p>
Supplement to "Endogenous DHEAS is causally linked with lumbar spine bone mineral density and forearm fractures in women - A mendelian randomization study"
<p>Supplemental tables to "Endogenous DHEAS is causally linked with lumbar spine bone mineral density and forearm fractures in women - A mendelian randomization study" by Johan Quester, Maria Nethander, Anna Eriksson and Claes Ohlsson.</p>
Clausal Causal Markers in the Languages of Europe: A Database
<p>The present typological database accumulates clausal causal markers in the languages of Europe. Clausal causal markers are used in polypredicative constructions (cf. <em>I fell down, because the floor was slippery</em>). The database contains more than 100 markers from Russian, Ukrainian, Belarusian, Polish, Czech, Slovak, Serbian, Bulgarian, Macedonian, Latvian, Lithuanian, German, English, Dutch, Danish, Swedish, Norwegian, Portuguese, Spanish, French, Italian, Catalan, Modern Greek, Albanian, Estonian, Finnish, Hungarian, Basque, Chuvash, and Georgian (currently 30 languages in total). The data were collected from grammars and language corpora and via elicitation.</p> <p>The typological parameters taken into consideration are:</p> <p>1) semantic varieties of causal meanings: direct causal meaning (<em>I fell down, because the floor was slippery</em>); motive (<em>As it is your birthday today, I have bought you a present</em>); inferential (<em>Peter is at home, because he is not at work</em>); illocutionary-imperative (<em>Wrap it up with your packing, or (= because) you'll be late for your train</em>); illocutionary-question<em> </em>(<em>How much are the cucumbers, because I want to take a couple of kilos?</em>); logic (<em>It is a triangle, because it has three angles</em>);</p> <p>2) assessment (positive vs. negative);</p> <p>3) reality of the cause;</p> <p>4) stylistic features;</p> <p>5) morphosyntactic features (position of the marker within the clause, preposition vs. postposition of the causal clause; occurrence in emphatic constructions and under negation, etc.)</p> <p>6) diachronic information and polysemy patterns.</p> <p>Further updates and expansion of the database are possible.</p> <p>The database was created at the Institute for Linguistic Studies, RAS with support from the Russian Science Foundation grant # 18-18-00472 “Causal Constructions in World Languages (Semantics and Typology)”.</p>
Webis Causal Question Answering 2022
<p>The Webis Causal Question Answering 2022 (Webis-CausalQA-22) corpus comprises 1.1M causal question-answer pairs collected from the public QA datasets. This dataset was developed to support the development of tailored approaches that can answer causal questions.</p> <p>Overview:</p> <p>The directory "input" contains the train and validation splits (used for evaluation), the directory "output" contains the evaluation results, and the directory "models" includes the fine-tuned checkpoints.</p> <p> </p> <p> </p>
CausalOrca: An ORCA-based Diagnostic Dataset for Causally-aware Multi-agent Trajectory Prediction
<p>CausalOrca is a synthetic diagnostic dataset created through controlled simulations. It is designed to provide annotations of ground-truth causal effects and fine-grained agent categories for social interactions in multi-agent scenarios. The dataset is constructed using a modified RVO2 simulator and incorporates the ORCA optimization-based collision avoidance algorithm known for crowd simulation. With full control over scene configurations, the dataset enables the collection of motion behaviors in paired scenes before and after agent removal, generating a large set of counterfactual pairs with annotations of ground-truth causal effects. CausalOrca can serve as a valuable resource for studying and developing causally-aware neural representations of social interactions and trajectory prediction models. Please see the <a href="https://github.com/rebuttal-anonymous/causalorca">GitHub repository</a> for a more detailed description of the dataset, including dataset statistics and documentation on how to use, visualize, and generate the data.</p>
BISCUIT: Causal Representation Learning from Binary Interactions
<p>This repository contains the datasets from the paper "BISCUIT: Causal Representation Learning from Binary Interactions" (<a href="https://phlippe.github.io/BISCUIT/">link</a>) by Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano, Taco Cohen, Efstratios Gavves. </p> <p><strong>iTHOR Embodied AI </strong>- The Embodied AI dataset is generated with the iTHOR simulator. We use the default kitchen environment, FloorPlan10, and position the robot in front of the kitchen counter. The robot interacts with different objects in the environment, including a Microwave, cabinet, stove, and an egg. For more details on the dataset, see <a href="https://github.com/phlippe/BISCUIT">our GitHub repository</a> and the appendix of our paper.</p> <p><strong>CausalWorld </strong>- The CausalWorld environment implements a tri-finger robot which can interact with a cube in the center of a stage. We additionally introduce causal variables for the friction parameters of the stage, floor and cube, as well as color changes. For more details on the dataset as well as the code to generate this dataset, see <a href="https://github.com/phlippe/BISCUIT">our GitHub repository</a> and the appendix of our paper.</p> <p><strong>Voronoi</strong> - The Voronoi benchmark allows for creating causal systems with arbitrary number of causal variables and causal graphs. We provide datasets with 6 and 9 variables, as well as systems with minimal number of interactions. For details, see <a href="https://github.com/phlippe/BISCUIT">our GitHub repository</a> and our paper.</p>
Disgust and Politics: Investigating the Causal Mechanism Between Pathogen-Avoidance and Conservatism
<p>Data for the bachelor thesis of L. Y. Kogelheide. (.csv format)</p> <p>FEE_FEE are answers from the disgust sensitivity questionnaire</p> <p>Path are answers for the Disgust Image Set - Experimental Condition</p> <p>No_Path are answers for the Disgust Image Set - Control Condition</p> <p>Traditionalism_T are answers for the traditionalism questionnaire</p> <p>SDO_SDO are answers for the social dominance orientation questionnaire</p> <p>HC is honesty check</p>
Resting-state fMRI data for locating causal hubs of memory consolidation in spontaneous brain network
<p>The mouse fMRI data for the paper "<strong>Locating causal hubs of memory consolidation in spontaneous brain network in male mice</strong>"<strong> </strong>published in <strong>Nature Communications </strong>(DOI: 10.1038/s41467-023-41024-z)<strong>. </strong>This includes longitudinal resting-state fMRI data in mice after behavioural training for 1-Day or 5-Day Active Place Avoidance (APA) task, acquired at post-training day 1 and day 8. Due to the large datasets, each group has been packed into several 2GB zip files. They need to be downloaded into the same folder and unpacked together (e.g. 1-Day APA Post training day 1 has five zip files starting with "1DAPA_PostDay1"). The structural and EPI templates and the ROI labels in the AMBMC atlas space are provided in the AMBMC_label.zip. </p>
Causal evidence for social group sizes from Wikipedia editing data
Open the record for dataset details and reuse information.
Data from: Multisensory perceptual and causal inference is largely preserved in medicated post-acute individuals with schizophrenia
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.