Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
376
datasets available to search
ShareScore release 0.9.0
Dataset results
376 results for “Causality”
Data for: How do we measure and increase systems thinking? Comparing self-reported and performative metrics in response to building causal loop models
Open the record for dataset details and reuse information.
CLEVRER-Humans: Describing physical and causal events the human way
Open the record for dataset details and reuse information.
Data from: Flight power muscles have a coordinated, causal role in controlling hawkmoth pitch turns
Open the record for dataset details and reuse information.
Dogs’ looking times and pupil dilation response reveal expectations about contact causality
Open the record for dataset details and reuse information.
Data from: An integrative approach to prioritize candidate causal genes for complex traits in cattle
Open the record for dataset details and reuse information.
Modelling trait heterogeneity and inferring causal links in the macroevolution of growth habit in eudicot angiosperms
Open the record for dataset details and reuse information.
Causal inference and risk prediction of gestational diabetes mellitus based on case-control study and Mendel randomization
Open the record for dataset details and reuse information.
The extended ‘common cause’: causal links between punctuated evolution and sedimentary processes
Open the record for dataset details and reuse information.
Data for: Causal mechanisms for negative impacts of energy development inform management triggers for sagebrush birds
Open the record for dataset details and reuse information.
Past Causalities and Event Categories for Connecting Similar Past and Present Causalities
<p>This dataset includes past causalities and their categories to connect similar past and present causalities. We report how to use this dataset in the following papers.</p> <p><em>Ryohei Ikejiri, Yasunobu Sumikawa: "Developing world history lessons to foster authentic social participation by searching for historical causation in relation to current issues dominating the news". Journal of Educational Research on Social Studies 84, 37–48 (2016). (in Japanese).</em></p> <p><em>Yasunobu Sumikawa and Ryohei Ikejiri, "Mining Historical Social Issues", Intelligent Decision Technologies, Smart Innovation, IDT'15, Systems and Technologies, Vol. 39, Springer, pp. 587--597, 2015.</em></p> <p>This dataset is based on some textbooks that are popular ones in Japanese high-school. We first collect past causalities by referencing the textbooks. We then select the causalities if they can be useful for considering solutions for present social issues. To enhance the analogy, we describe each causality in three kinds of texts: background including problems, solution ways, and their results. From the selected causalities and an Encyclopedia of Historiography, we define categories for them. Finally, the created dataset contains 138 past causalities and 13 categories. Each past causality has more than one categories.</p> <p>To help training machine learning models, this dataset additionally provides 900 past event data in past_events_wikipedia.tsv. The event data were collected from Wikipedia, and then were assigned one or more categories from the above 13 ones. We have confirmed that SVM-RBF equipped with the above all categorized data obtained 73.6% precision, 55.8% recall and 63.5% F1 score</p> <p> </p> <p><strong>File contents</strong>:</p> <ul> <li>Past causality data <ol> <li>historical_causalities_data.tsv: Detail of stored causalities.</li> <li>historical_causalities_regions.tsv: Regions where the causalities happened.</li> <li>historical_causalities_categories.tsv: Categories of the causalities.</li> </ol> </li> <li>Past event data <ol> <li>past_events_wikipedia.tsv: Descriptions of past events stored in Wikipedia. This file is useful for training machine learning model such as SVM.</li> </ol> </li> <li>Statistics (Statistics.tsv) <p> Results of statistical analyses for the dataset. We used Calinski and Harabaz method, mutual information, Jaccard Index, TF-IDF+JS divergence, and Meta-data Similarity that counts how many common categories two causalities share in order to measure qualities of the dataset.</p> </li> </ul> <p><strong>Grants</strong>: JSPS KAKENHI Grant Number 26750076, 17K12792, and 19K20631</p>
Data from: QTG-Finder2: a generalized machine-learning algorithm for prioritizing QTL causal genes in plants
Linkage mapping has been widely used to identify quantitative trait loci (QTL) in many plants and usually requires a time-consuming and labor-intensive fine mapping process to find the causal gene underlying the QTL. Previously, we described QTG-Finder, a machine-learning algorithm to rationally prioritize candidate causal genes in QTLs. While it showed good performance, QTG-Finder could only be used in Arabidopsis and rice because of the limited number of known causal genes in other species. Here we tested the feasibility of enabling QTG-Finder to work on species that have few or no known causal genes by using orthologs of known causal genes as training set. The model trained with orthologs could recall about 64% of Arabidopsis and 83% of rice causal genes when the top 20% ranked genes were considered, which is similar to the performance of models trained with known causal genes. The average precision was 0.027 for Arabidopsis and 0.029 for rice. We further extended the algorithm to include polymorphisms in conserved non-coding sequences and gene presence/absence variation as additional features. Using this algorithm, QTG-Finder2, we trained and cross-validated Sorghum bicolor and Setaria viridis models. The S. bicolor model was validated by causal genes curated from the literature and could recall 70% of causal genes when the top 20% ranked genes were considered. In addition, we applied the S. viridis model and public transcriptome data to prioritize a plant height QTL and identified 13 candidate genes. QTL-Finder2 can accelerate the discovery of causal genes in any plant species and facilitate agricultural trait improvement.
Simulated dataset from 'Quantifying the causal pathways contributing to natural selection'
<p>This dataset (Antechinus.csv) relates to the worked example in the appendix of the paper:</p> <p>Henshaw JM, Morrissey MB, Jones AG (2020). Quantifying the causal pathways contributing to natural selection. Evolution (doi:10.1111/evo.14091)</p> <p>The worked example concerns a hypothetical study of female antechinus. These are small carnivorous marsupials that use torpor to reduce energy consumption from late summer to early winter. They reproduce once per year in late winter or early spring, following which most individuals die. We suppose that researchers tracked female antechinus from mid summer to the end of the breeding season. They recorded the animals' body size, their date of last torpor, whether they survived to breed, their number of mates, and their fecundity. I simulated the dataset resulting from this hypothetical study in Wolfram Mathematica (see Methods below). In the above paper, we analyse the causal structure of natural selection in this dataset (the R code for the causal analysis is included here as 'AntechinusAnalysis.R').</p> <p>The variables in the dataset are: body size in unspecified units (BodySize); the date of last torpor (TorporDate), standardized as the number of days before/after an unspecificed reference date; whether the individual survived to breed, given as a binary variable (Survival); an individual's number of mates (Mates); and her fecundity (Fecundity).</p>
Analysis and Figures from "Causal network inference from gene transcriptional time-series response to glucocorticoids"
<p>Gene regulatory network inference is essential to uncover complex relationships among gene pathways and inform downstream experiments, ultimately enabling regulatory network re-engineering. Network inference from transcriptional time-series data requires accurate, interpretable, and efficient determination of causal relationships among thousands of genes. Here, we develop Bootstrap Elastic net regression from Time Series (BETS), a statistical framework based on Granger causality for the recovery of a directed gene network from transcriptional time-series data. BETS uses elastic net regression and stability selection from bootstrapped samples to infer causal relationships among genes. BETS is highly parallelized, enabling efficient analysis of large transcriptional data sets. We show competitive accuracy on a community benchmark, the DREAM4 100-gene network inference challenge, where BETS is one of the fastest among methods of similar performance and additionally infers whether the causal effects are activating or inhibitory. We apply BETS to transcriptional time-series data of 2,768 differentially-expressed genes from A549 cells exposed to glucocorticoids over a period of 12 hours. We identify a network of 2,768 genes and 31,945 directed edges (FDR <= 0.2). We validate inferred causal network edges using two external data sources: overexpression experiments on the same glucocorticoid system, and genetic variants associated with inferred edges in primary lung tissue in the Genotype-Tissue Expression (GTEx) v6 project. BETS is available as an open source software package at https://github.com/lujonathanh/BETS</p> <p>This upload documents the analysis and figure files that support each numerical claim of the manuscript. Full Progeny.xlsx lists out the relevant code and files for each numerical claim of the manuscript, assuming the home folder of port-from-della</p>
The causal relationship k
<p><strong>Published at: </strong></p> <p>MATEC Web of Conferences: ISSN: 2261-236X.<br> ( https://www.matec-conferences.org/ )<br> <a href="https://doi.org/10.1051/matecconf/202133609032">https://doi.org/10.1051/matecconf/202133609032 </a></p> <p>MATEC Web Conf.</p> <p><strong>Volume </strong>336, 2021<br> Article Number: 09032</p> <p>Number of page(s): 28</p> <p>2020 2<sup>nd</sup> International Conference on Computer Science Communication and Network Security (CSCNS2020)<br> <a href="https://doi.org/10.1051/matecconf/202133609032">https://doi.org/10.1051/matecconf/202133609032 </a></p> <p><strong>Published online: </strong>15 February 2021</p> <p>see also:</p> <p><strong>2020 2nd International Conference on Computer Science, Communication and Network Security (CSCNS2020</strong></p> <p>Sanya, China<br> <strong>December 22-23, 2020</strong><br> http://www.cscns.org/<br> Contact via email:<br> submit@cscns.org<br> or<br> cscns2020@sina.com</p> <p>http://www.cscns.org/keynote.html</p>
Data from: Development of a protocol for environmental impact studies using causal modelling
1- The global issue of water scarcity caused by climate change and human utilisation highlights the importance of an efficient assessment of water quality in freshwater systems. One of the challenges facing water management in environmental impact studies is the difficulty of inferring causality in complex systems. Traditional water assessment methods are inadequate because they are challenged to separate natural variation from the effect of human activities. 2- Knowing the causal structure of a complex ecosystem will enable managers to identify key anthropogenic, climate and flow drivers of water quality, and make informed decisions about interventions that improve water and environmental quality. In this study, I show how causal modelling can facilitate decision making for water treatment plant managers to improve their environmental management of this valuable resource. 3- Models built using causal modelling techniques, including structural equation modelling and the principles of Bayesian Networks, were utilised for management decision making purposes. The discharge load values were manipulated in the models to predict the effect of a potential intervention, e.g. treatment plant upgrade, on the values of the water quality variables in the creek. That is, water quality variables were predicted when an imaginary or counterfactual situation was imposed on the models. 4- This study showed that there would not be any observable effect of effluent on macroinvertebrate communities if the discharge loads of chlorophyll a, total organic carbon, total phosphorus, nitrate, and conductivity were reduced to 0.1 of observed values. The concentrations of environmental variables in the creek would return to their baseline levels when their corresponding discharge loads in the effluent were halved or divided by 10. 5- Based on the findings of this study, managers in the field of environmental impact studies can predict the response of a system in the presence of potential interventions under complex and uncertain conditions. The implementation of such techniques offers great promise in the wider field of environmental management where accounting for multiple factors structuring ecosystems in necessary to adequately represent causality.
Data from: Timing manipulations reveal the lack of a causal link across timing of annual-cycle stages in a long-distance migrant
Organisms need to time their annual-cycle stages, like breeding and migration, to occur at the right time of the year. Climate change has shifted the timing of annual-cycle stages at different rates, thereby tightening or lifting time constraints of these annual-cycle stages, a rarely studied consequence of climate change. The degree to which these constraints are affected by climate change depends on whether consecutive stages are causally linked (I) or whether the timing of each stage is independent of other stages (II). Under (I), a change in timing in one stage has knock-on timing effects on subsequent stages, whereas under (II) a shift in the timing of one stage affects the degree of overlap with previous and subsequent stages. For testing this we combined field manipulations, captivity measurements and geolocation data. We advanced and delayed hatching dates in pied flycatchers (Ficedula hypoleuca) and measured how the timing of subsequent stages (male moult and migration) were affected. There was no causal effect of manipulated hatching dates on the onset of moult and departure to Africa. Thus, advancing hatching dates reduced the male moult-breeding overlap with no effect on the moult-migration interval. Interestingly, the wintering location of delayed males was more westwards, suggesting that delaying the termination of breeding carries-over to winter location. Because we found no causal linkage of the timing of annual-cycle stages, climate change can shift these stages at different rates, with the risk that the time available for some become so short that this will have major fitness consequences.
Data from: The causal relationship between sexual selection and sexual size dimorphism in marine gastropods
Sexual size dimorphism is widespread among dioecious species but its underlying driving forces are often complex. A review of sexual size dimorphism in marine gastropods revealed two common patterns: firstly, sexual size dimorphism, with females being larger than males, and secondly females being larger than males in mating pairs; both of which suggest sexual selection as being causally related with sexual size dimorphism. To test this hypothesis, we initially investigated mechanisms driving sexual selection on size in three congeneric marine gastropods with different degrees of sexual size dimorphism, and, secondly, the correlation between male/female sexual selection and sexual size dimorphism across several marine gastropod species. Male mate choice via mucus trail following (as evidence of sexual selection) was found during the mating process in all three congeneric species, despite the fact that not all species showed sexual size dimorphism. There was also a significant and strong negative correlation between female sexual selection and sexual size dimorphism across 16 cases from seven marine gastropod species. These results suggest that sexual selection does not drive sexual size dimorphism. There was, however, evidence of males utilizing a similar mechanism to choose mates (i.e. selecting a female slightly larger than own size) which may be widespread among gastropods, and in tandem with present variability in sexual size dimorphism among species, provide a plausible explanation of the observed mating patterns in marine gastropods.
Processed data from "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"
<p>This is the processed data from our manscript "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"</p>
Processed data from "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"
<p>This is the processed data from our manscript "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"</p>
Processed data from "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"
<p>This is the processed data from our manscript "Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.