Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

376

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

376 results for “Causality”

Learn how ShareScore rates datasets ↗
zenodo48/100

CauseNet: Towards a Causality Graph Extracted from the Web

<p>Causal knowledge is seen as one of the key ingredients to advance artificial intelligence. Yet, few knowledge bases comprise causal knowledge to date, possibly due to significant efforts required for validation. Notwithstanding this challenge, we compile CauseNet, a large-scale knowledge base of <em>claimed&nbsp;</em>causal relations between causal concepts. By extraction from different semi- and unstructured web sources, we collect more than 11 million causal relations with an estimated extraction precision of 83% and construct the first large-scale and open-domain causality graph. We analyze the graph to gain insights about causal beliefs expressed on the web and we demonstrate its benefits in basic causal question answering. Future work may use the graph for causal reasoning, computational argumentation, multi-hop question answering, and more.</p> <p>When using the data, please make sure to refer to it as follows:</p> <pre><code>@inproceedings{heindorf2020causenet, author = {Stefan Heindorf and Yan Scholten and Henning Wachsmuth and Axel-Cyrille Ngonga Ngomo and Martin Potthast}, title = {CauseNet: Towards a Causality Graph Extracted from the Web}, booktitle = {{CIKM}}, pages = {3023--3030}, publisher = {{ACM}}, year = {2020} }</code></pre>

opencc-by-4.0Oct 2020View details →
zenodo48/100

Data for paper "Stratocumulus adjustments to aerosol perturbations disentangled with a causal approach"

<p>Timeseries data used for the causal effect estimation of the paper &quot;Stratocumulus adjustments to aerosol perturbations disentangled with a causal approach&quot;.</p> <p>This dataset contains several cloud parameters and meteorological co-variates corresponding to the evolution of the South-East Atlantic stratocumulus deck for the time period January 2016 to December 2017 and the spatial domain [lon1,lon2,lat1,lat2]=[0, 10, -20, -10].&nbsp;</p> <p>The processing code used to generate the timeseries data, as well as the analysis code are uploaded separately on Zenodo. The input raw satellite and reanalysis data for the processing code are from EUMETSAT (Copyright (c) (2020) EUMETSAT), NASA&nbsp;and COPERNICUS data (generated using Copernicus Climate Change Service information [2022]).&nbsp;</p> <p>&nbsp;</p> <p>Citations for the raw data sources:&nbsp;</p> <p>Finkensieper, S., Meirink, J.-F., van Zadelhoff, G.-J., Hanschmann, T., Benas, N., Stengel, M., Fuchs, P., Hollmann, R., Kaiser, J., Werscheck, M.: CLAAS-2.1: CM SAF CLoud property dAtAset using SEVIRI - Edition 2.1. Satellite Application Facility on Climate Monitoring (2020). <a href="https://doi.org/10.5676/EUM_SAF_CM/CLAAS/V002_01">https://doi.org/10.5676/EUM_SAF_CM/CLAAS/V002_01</a></p> <p>Huffman, G.J., Stocker, E.F., Bolvin, D.T., Nelkin, E.J., Tan, J.: GPMIMERG Final Precipitation L3 Half Hourly 0.1 degree x 0.1 degree V06. MD, Goddard Earth Sciences Data and Information Services Center (GES DISC) (2019). <a href="https://doi.org/10.5067/GPM/IMERG/3B-HH/06">https://doi.org/10.5067/GPM/IMERG/3B-HH/06</a>.</p> <p>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horanyi, A., Munoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., Th ́ebaut, J.-N.: ERA5 hourly data on single levels from 1959 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS) (2018). <a href="https://doi.org/10.24381/cds.adbb2d47">https://doi.org/10.24381/cds.adbb2d47</a></p> <p>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horanyi, A., Munoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., Th ́ebaut, J.-N.: ERA5 hourly data on pressure levels from 1959 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS) (2018). <a href="https://doi.org/10.24381/cds.bd0915c6">https://doi.org/10.24381/cds.bd0915c6</a></p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Replication package and appendixes for Causal inference of server- and client-side code smells in web apps evolution

<p>-Analysis&nbsp;<br>--R scripts used to make the analisys, divided by folders<br>--Data folders used in the questions</p> <p>-Appendixes - used in the article to shwo extra tables and plots</p> <p>-data folders - Aggregation of data, each app has two files, CSV and xls</p> <p>-separated data folders - 5 files for each app, with lines corresponding to the each released official version<br>--serversmells<br>--clientsmells<br>--javascriptsmells<br>--Cloc(metrics)<br>--version (all oficial releases)</p> <p>-issues_bugs<br>--data -issues by app by release&nbsp;<br>--data_bugs_more - the same but only bugs, by app by release<br>--scripts - scrips used to aggregate issues (from daily issues to by release) anf the same for bugs</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Improving causality perception judgments in schizophrenia spectrum disorder via transcranial direct current stimulation - Dataset

<p>Raw data related to the publication:</p> <p>Sch&uuml;lke, R., Schmitter, C. V., &amp; Straube, B. (2023). Improving causality perception judgments in schizophrenia spectrum disorder via transcranial direct current stimulation. <em>Journal of Psychiatry and Neuroscience</em>, <em>48</em>(4), E245&ndash;E254. <a href="https://doi.org/10.1503/jpn.220184">https://doi.org/10.1503/jpn.220184</a></p> <p>Variables:</p> <ul> <li>Subject</li> <li>Condition &ndash; Stimulation condition; parietal (left parietal cathodal, right parietal anodal [LPC-RPA]), frontoparietal (left frontal cathodal, right parietal anodal [LFC-&shy;RPA]), frontal (left frontal cathodal, right frontal anodal [LFC&shy;-RFA])</li> <li>Timepoint &ndash; Before/After (stimulation)</li> <li>Angle &ndash; in degrees</li> <li>Angle_scaled &ndash; mean-centered and scaled Angle</li> <li>Delay_ms &ndash; in milliseconds</li> <li>Delay_ms_scaled &ndash; mean-centered and scaled Delayed_ms</li> <li>Causality &ndash; causal/non-causal (judgment)</li> <li>RT &ndash; reaction time in milliseconds</li> </ul> <p>In the original version of the data, the data had been incorrectly labelled: The data actually corresponding to the LFC-RPA condition had been incorrectly labelled as LPC-RPA, and the data actually corresponding to the LPC-RPA condition had been incorrectly labelled as LFC-RPA. This has been corrected with the 04/2024 version of the dataset.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Annotation of the the assembled genome of Fusarium oxysporum f. sp. albedinis strain 133, the causal agent of date palm dieback.

<p>Annotation of&nbsp;the the assembled genome of <em>Fusarium oxysporum f. sp. albedinis</em> strain 133 (Khayi et al., 2020). Gene prediction and annotation were carried out using funnotate pipeline v1.8.1 (Stajich, 2020), which&nbsp;includes masking, ab initio gene-prediction training, using Augustus and Genmark, with the EST dataset&nbsp;reported to the Ganoderma mycocosm repository, gene prediction, and the assignment of functional&nbsp;annotation to protein-coding gene models.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

MediCause Dataset of Causal Sentences with Annotated Entities

<p>The MediCause dataset contains 1202 causal sentences from medical publications where the entities involved in the causal relations have been annotated according to the MediCause ontological model for causal relations. The entities are annotated according to the Inside-Outside-Beginning (IOB) format. The labels used for the annotation are B-C (Cause), B-VC (Causal Variable), B-CS (Beginning Causal Specifier), I-CS (Inside Causal Specifier), B-CON (Beginning Connective), I-CON (Inside Connective), B-EF (Effect), B-VE (Effect Variable), B-ES (Beginning Effect Specifier), I-ES (nside Effect Specifier), O (Outside).</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

CITRIS - Causal Representation Learning Datasets

<p>This repository contains the datasets from the paper &quot;CITRIS: Causal Identifiability from Temporal Intervened Sequences&quot; (<a href="https://arxiv.org/abs/2202.03169">link</a>) by Phillip Lippe,&nbsp;Sara Magliacane,&nbsp;Sindy L&ouml;we,&nbsp;Yuki M. Asano,&nbsp;Taco Cohen,&nbsp;Efstratios Gavves.</p> <p><strong>Temporal Causal3DIdent</strong>&nbsp;-&nbsp;The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains&nbsp;7 (multidimensional) causal factors. The 7 shapes used are&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>,&nbsp;<a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>,&nbsp;<a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>,&nbsp;<a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>, <a href="https://github.com/brendel-group/cl-ica">Teapot</a>.&nbsp;We provide two versions of the dataset: one that only contains images of the Teapot, and one that uses all 7 shapes. For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Interventional Pong</strong>&nbsp;- The Interventional Pong environment is inspired by the game dynamics of Pong, where both paddles follow the policy of moving towards the ball, and the ball has slightly random movements. This dataset considers the 5 causal variables paddle left, paddle right, the ball position, the ball velocity, and the score. For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Ball-in-Boxes</strong> -&nbsp;The Ball-in-Boxes is a simple dataset for showcasing the concept of the minimal causal variables. The system consists of a ball which randomly moves within a box, but only under an intervention can swap between the two boxes. Thereby, the intervention does not affect the x-position in the box. Thus, one can only discover the box assignment as a causal variable, and not whether the inner x-position also belongs to it.&nbsp;For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

iCITRIS - Causal Representation Learning Datasets

<p>This repository contains the datasets from the paper &quot;iCITRIS: Causal Representation Learning for Instantaneous Temporal Effects&quot;&nbsp;(<a href="http://arxiv.org/abs/2206.06169">link</a>) by Phillip Lippe,&nbsp;Sara Magliacane,&nbsp;Sindy L&ouml;we,&nbsp;Yuki M. Asano,&nbsp;Taco Cohen,&nbsp;Efstratios Gavves.&nbsp;</p> <p><strong>Instantaneous Temporal Causal3Ident&nbsp;</strong>-&nbsp;The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains&nbsp;7 (multidimensional) causal factors with instantaneous and temporal causal relations between them. The 7 shapes used are&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>,&nbsp;<a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>,&nbsp;<a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>,&nbsp;<a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>,&nbsp;<a href="https://github.com/brendel-group/cl-ica">Teapot</a>.&nbsp;For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Causal Pinball </strong>- The Causal Pinball environment implements the simplified, real-world game dynamics of Pinball. This dataset considers 5 causal variables with instantaneous effects: the paddle position left, the paddle position right, the ball (velocity and position), the state of all bumpers, and the score.&nbsp;For more details on the dataset as well as the code to generate this dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

MAL04 Causal loop diagrams for the Charente River basin and its coastal zone (France)

<p>This dataset includes the causal loop diagrams (CLDs) developed by the H2020 COASTAL project&rsquo;s MAL #4 for the Charente River basin and its coastal zone. These CLDs represent the functioning of the territory in a systemic way, highlighting its main components and interactions among them. The CLDs are the result of multiple sectoral and multi-actor workshops during which stakeholders from different sectors discussed and collaborated to establish a common vision of the land-sea system. The CLDs concern the whole territory and some specific sectors.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Supplementary Data from, "Causal health impacts of power plant emission controls under modeled and uncertain physical process interference."

<p>These data are used to conduct the analysis in, "<a href="https://arxiv.org/abs/2306.05665">Causal health impacts of power plant emission controls under modeled and uncertain physical process interference</a>," by Wikle and Zigler (2024), to appear in <em>Annals of Applied</em> Statistics. This is purely for archival purposes to facilitate access to and replication of the aforementioned analysis. Data were obtained from the following sources:</p> <ol> <li>&nbsp;U.S. Emissions Data [<a href="https://ampd.epa.gov/ampd">U.S. EPA, Air markets program data (AMPD)</a>] <ul> <li>AMPD_Unit_with_Sulfur_Content_and_Regulations_with_Facility_Attributes.csv</li> </ul> </li> <li>&nbsp;US Census 2016 American Community Survey [<a href="https://www.census.gov/programs-surveys/acs">US Census Bureau ACS</a>] <ul> <li>Census_2016_TxZCTA.RDS</li> <li><em>Note: data were obtained using the r package &lsquo;<a href="https://walker-data.com/tidycensus/">tidycensus</a>&rsquo;.</em></li> </ul> </li> <li>&nbsp;Daymet Annual Climate Summaries [<a href="https://daac.ornl.gov/DAYMET/guides/Daymet_V4_Annual_Climatology.html">Daymet Version 4</a>] <ul> <li>daymet_v4_prcp_annttl_na_2016.nc</li> <li>daymet_v4_tmax_annavg_na_2016.nc</li> <li>daymet_v4_tmin_annavg_na_2016.nc</li> <li>daymet_v4_vp_annavg_na_2016.nc</li> </ul> </li> <li>&nbsp;SO<sub>4</sub> and Black Carbon Concentrations [<a href="https://sites.wustl.edu/acag/datasets/surface-pm2-5/#V4.NA.03">Randall Martin Atmospheric Composition Analysis Group, North American Regional Estimates, version V4.NA.02</a>] <ul> <li>GWRwSPEC_BC_NA_201601_201612.nc</li> <li>GWRwSPEC_SO4_NA_201601_201612.nc</li> </ul> </li> <li>&nbsp;HyADS Coal-Attributed PM2.5 Concentrations [<a href="https://doi.org/10.1097/EDE.0000000000001024">Henneman et al. (2019)</a>] <ul> <li>HyADS_grids_pm25_byunit_2016.fst</li> <li>HyADS_grids_pm25_total_2016.fst</li> </ul> </li> <li>&nbsp;Mexico Emissions Data [<a href="https://www.epa.gov/air-emissions-modeling/2014-2016-version-7-air-emissions-modeling-platforms">National Emissions Inventory Collaborative, 2016v1 emissions modeling platform</a>] <ul> <li>Mexico_2016_point_interpolated_02mar2018_v0.csv</li> </ul> </li> <li>&nbsp;North American Regional Reanalysis Meteorological Data [<a href="https://psl.noaa.gov/data/gridded/data.narr.monolevel.html">NOAA</a>] <ul> <li>rhum.2m.mon.mean.nc</li> <li>uwnd.10m.mon.mean.nc</li> <li>vwnd.10m.mon.mean.nc</li> </ul> </li> <li>&nbsp;Cigarette smoking data [<a href="https://doi.org/10.1186/1478-7954-12-5">Dwyer-Lindgren et al. (2014)</a>] <ul> <li>smokedatwithfips_1996-2012.csv</li> </ul> </li> <li>&nbsp;Synthetic pediatric asthma data [<em>Note:<strong> synthetic data!</strong> Simulated to match the format, but not the observations, from the <a href="https://www.dshs.texas.gov/texas-health-care-information-collection">Texas Health Care Information Collection (THCIC), Texas DSHS</a></em>] <ul> <li>synth-ped-asthma-data.csv</li> </ul> </li> <li>&nbsp;Texas state shape file [<a href="https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html">US Census</a>] <ul> <li>texas-state-sf.RDS</li> </ul> </li> <li>&nbsp;US ZIPcode-to-county data crosswalk [<a href="https://mcdc.missouri.edu/applications/geocorr2014.html">Missouri Census Data Center</a>] <ul> <li>tx-zip-to-county.csv</li> </ul> </li> </ol> <p>Code and supplementary material from this analysis, as well as more detailed data descriptions, are available at: <a href="https://github.com/nbwikle/estimating-interference">https://github.com/nbwikle/estimating-interference</a></p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Online Appendix for PhD Thesis Titled "Dissecting Causal Relationships and Molecular Mechanisms in Disease using Genetic Risk Profiles"

<p>This repository contains 23 tables and two figures, which are too big to be included in the Appendix section of my thesis document.</p> <p>The second version includes additional summary statistics of metabolite-PGS associations which can be found at http://mrcieu.mrsoftware.org/metabolites_PRS_atlas/.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Attack of the clones: population genetics reveals clonality of Colletotrichum lupini, the causal agent of lupin anthracnose

<p><em>Colletotrichum lupini</em>, causing lupin anthracnose, is one of the worst pathogens to lupin cultivation worldwide. Understanding its population structure and evolutionary potential is crucial to design successful disease management strategies. The objective of this study was to employ population genetics to investigate the diversity, evolutionary dynamics and molecular basis of host interaction of this notorious lupin pathogen. A collection of globally representative <em>C. lupini </em>isolates was genotyped through triple digest restriction-site associated DNA sequencing (3D-RADseq), resulting in a dataset of unparalleled resolution. Phylogenetic and structural analysis could distinguish four (I &ndash; IV) independent lineages. The strong population structure, low recombination and slow linkage decay strongly suggests that <em>C. lupini</em> reproduces clonally. Different morphologies and virulence patterns on white (<em>Lupinus albus</em>) and Andean lupin (<em>L. mutabilis</em>) were observed between and within clonal lineages. Isolates belonging to lineage II were shown to have a mini-chromosome which was also partly present in lineage III and IV, but not in lineage I isolates. Variation in the presence of this mini-chromosome could indicate a role in host interaction. All four lineages were present in the South American Andes region, which is concluded to be the center of origin of this species. Only members of lineage II have been found outside South America since the 1990s, indicating it as the current pandemic population. As a seed-borne pathogen, <em>C. lupini</em> has mainly spread through infected but symptomless seeds, stressing the importance of phytosanitary measures to prevent future outbreaks of strains that are yet confined to South America.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

WikiCausal Corpus for Evaluation of Causal Knowledge Graph Construction

<p>Documentation on the data format and how it can be used can be found on: <a href="https://github.com/IBM/wikicausal">https://github.com/IBM/wikicausal</a> as well as our paper:</p> <pre><code>@unpublished{, author = {Oktie Hassanzadeh and Mark Feblowitz}, title = {{WikiCausal}: Corpus and Evaluation Framework for Causal Knowledge Graph Construction}, year = {2023}, doi = {10.5281/zenodo.7897996} }</code></pre> <pre>Corpus derived from Wikipedia and Wikidata. Refer to Wikipedia and Wikidata <a href="https://en.wikipedia.org/wiki/Wikipedia:Copyrights">license and terms of use</a> for more details:</pre> <ul> <li><strong>Permission is granted</strong> to copy, distribute and/or modify Wikipedia&#39;s text under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License and, <em>unless otherwise noted</em>, the GNU Free Documentation License, unversioned, with no invariant sections, front-cover texts, or back-cover texts.</li> <li>A copy of the Creative Commons Attribution-ShareAlike 3.0 Unported License is included in the section entitled &quot;<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_Creative_Commons_Attribution-ShareAlike_3.0_Unported_License">Wikipedia:Text of Creative Commons Attribution-ShareAlike 3.0 Unported License</a>&quot;</li> <li>A copy of the GNU Free Documentation License is included in the section entitled &quot;<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_the_GNU_Free_Documentation_License">GNU Free Documentation License</a>&quot;.</li> <li>Content on Wikipedia is covered by <a href="https://en.wikipedia.org/wiki/Wikipedia:General_disclaimer">disclaimers</a>.</li> </ul> <pre>THIS DATA IS PROVIDED &quot;AS IS&quot;, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</pre>

opencc-by-3.0May 2023View details →
zenodo44/100

Modeling islet enhancers using deep learning identifies candidate causal variants at loci associated with T2D and glycemic traits

<p>Genetic association studies have identified hundreds of independent genetic signals associated with type 2 diabetes (T2D) and related traits. Despite these successes, the identification of specific causal variants underlying a genetic association signal remains challenging. In this study, we describe a deep learning method to analyze the impact of sequence variants on enhancers. Focusing on pancreatic islets, a relevant T2D tissue, we show that our model learns islet-specific transcription factor (TF) regulatory patterns and can be used to prioritize candidate causal variants. At 101 genetic signals associated with T2D and related glycemic traits where multiple variants occur in linkage disequilibrium, our method nominates a single causal variant for each association signal, including three variants previously shown to alter reporter activity in islet-relevant cell types. For another signal associated with blood glucose levels, we biochemically test all candidate causal variants from statistical fine-mapping using a pancreatic islet beta cell line and show biochemical evidence of allelic effects on TF binding for the model-prioritized variant. To aid in future research, we publicly distribute our model and islet enhancer perturbation scores across ~67 million variants. We anticipate that deep learning methods like the one presented in this study will enhance the prioritization of candidate causal variants for functional studies.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Gauging Size Resolved Ambient Particulate Matter Concentration Solely Using Biometric Observations: A Machine Learning and Causal Approach

<p>Notebook and data to accompany the (unpublished) paper titled "Gauging Size Resolved Ambient Particulate Matter Concentration Solely Using Biometric Observations: A Machine Learning and Causal Approach". This work expands a previous study, relating particulate matter concentrations and short-term biometric features across multiple participants.&nbsp;</p><p>Github link: https://github.com/mi3nts/DUEDARE_multiple_participants</p>

opencc-by-4.0Nov 2023View details →
dryad40/100

Causal evidence for social group sizes from Wikipedia editing data

<p>Human communities have self-organizing properties in which specific Dunbar Numbers may be invoked to explain group attachments.  By analyzing Wikipedia editing histories across a wide range of subject pages, we show that there is an emergent coherence in the size of transient groups formed to edit the content of subject texts, with two peaks averaging at around $N=8$ for the size corresponding to maximal  contention, and at around $N=4$ as a regular team. These values are consistent with the observed sizes of conversational groups, as well as the hierarchical structuring of Dunbar graphs.  We use the Promise Theory model of bipartite trust to derive a scaling law that  fits the data and may apply to all group size distributions, when based on attraction to a seeded group process.  In addition to  providing further evidence that even spontaneous communities of strangers are self-organizing, the results have important implications for the governance of the Wikipedia commons and for the security of all online social platforms and associations.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Summary statistics from "Sex-Specific Causal Relations between Steroid Hormones and Obesity—A Mendelian Randomization Study"

<p>GWAMA summary statistics of four steroid hormone levels and one steroid hormone ratio using fixed-effect model.</p> <p>When using this data, please cite: Pott J, Horn K, Zeidler R, et al.. Sex-Specific Causal Relations between Steroid Hormones and Obesity - A Mendelian Randomization Study. <em>Metabolites</em> <strong>2021</strong>, <em>11</em>, 738. https://doi.org/10.3390/metabo11110738</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Global Soil Moisture-Air Temperature Interactions from Linear and Nonlinear Granger Causalities

<p>These datasets were generated to assess linear and nonlinear Granger causalities in the submitted manuscript, Global Soil Moisture-Air Temperature Interactions from Linear and Nonlinear Granger Causalities by Bhatti et al. submitted to AGU-GRL. Nonlinear GC here is achieved with the Kernel Granger causality by Marinazzo et al. (2008). The data was used to develop theoretical experiments that help validate the strengths and limitations of both the linear Granger causality and the Kernel Granger causality before applying to real world datasets</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Interactive Causal Structure Discovery with Hyytiälä measurements (experiment code and data)

<p>This archive contains code and data required to reproduce the results presented in the following two papers.</p> <p>Interactive Causal Structure Discovery in Earth System Sciences<br> published in Proceedings of The KDD&#39;21 Workshop on Causal Discovery, 2021.</p> <p>Technical note: incorporating expert domain knowledge into causal structure discovery workflows<br> published in Biogeosciences, 2022</p> <p>The archive contains a README markdown document detailing the contents and how to run the experiments.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

General causal loop diagram of main land-sea interactions for the Belgian coastal zone.

<p>This is a polished&nbsp;diagram of the main causal land-sea interactions, based on aggregation of the detailed mind maps of land-sea interactions identified by the coastal and rural stakeholders who participated in the Belgian Multi-Actor Lab of the EU-funded H2020 project COASTAL (https://h2020-coastal.eu). The diagrams are used for designing System Dynamics models and evidence-based business road maps.&nbsp;&nbsp;</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record