Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
384
datasets available to search
ShareScore release 0.7.1
Dataset results
384 results for “risk model”
Data from: In vitro to in vivo extrapolation from three-dimensional hiPSC-derived cardiac microtissues and physiologically based pharmacokinetic modeling to inform next-generation arrythmia risk assessment
Open the record for dataset details and reuse information.
Spatiophylogenetic modelling of extinction risk reveals evolutionary distinctiveness and brief flowering period as threats in a hotspot plant genus
Comparative models used to predict species threat status can help to identify diagnostic features of species at risk. Such models often combine variables measured at the species level with spatial variables, causing multiple statistical challenges, including phylogenetic and spatial non-independence. We present a novel Bayesian approach for modelling threat status that simultaneously deals with both forms of non-independence and estimates their relative contribution, and we apply the approach to modelling threat status in the Australian plant genus Hakea. We find that after phylogenetic and spatial effects are accounted for, species with greater evolutionary distinctiveness and a shorter annual flowering period are more likely to be threatened. The model allows us to combine information on evolutionary history, species biology, and spatial data, to calculate latent extinction risk (potential for non-threatened species to become threatened), estimate the most important drivers of risk for individual species, and map spatial patterns in the effects of different predictors on extinction risk. This could be of value for proactive conservation decision-making based on the early identification of species and regions of potential conservation concern.
Feature attention graph neural network for estimating brain age and identifying important neural connections in mouse models of genetic risk for Alzheimer's disease
<p>Connectome, traits and behavior data for APOE234 mice.</p> <ul> <li>1. connectome.zip: mouse brain structural connectivity matrices from diffusion MRI.</li> <li>2. FAGNN_Phenotype.csv: a sheet of trait information of mice used in the study.</li> </ul> <p>columns: winding numbers, total distance, normalized NE time, normalized NE distance, normalized NW time, normalized NW distance, normalized SE time, normalized SE distance, normlaized SW time, normalized SW distance, island latency to first entry, island entries, normalized thigmataxis time, and normalized thigmotaxis distance</p> <div>rows: 4 trials for each day from day 1 to day 5 with 1 probing test each at day 5 and day 8</div> <ul> <li>3. mouse_anatomy.csv: brain region information regarding the connectivity matrix.</li> <li>4. behavior.zip: behavioral data for each mouse from Morris Water Maze experiments.</li> </ul>
Assessing Heavy Metal Contamination in Agricultural Soils: A Predictive Model Integrating GIS Tools and Probability-Risk Matrix – Case Study: Guarda Region, Portugal
<p>In these files we can find the final risk map of heavy metal contamination for the guarding area in Portugal obtained according to the methodology explained in the paper "Assessing Heavy Metal Contamination in Agricultural Soils: A Predictive Model Instegrating GIS Tools and Probability-Risk Matrix - Case Study: Guarda Region (Portugal)</p> <p>Final Risk Equal.tiff: GeoTiff with a pixel size of 30m. EPSG:3763 - ETRS89 / Portugal TM06</p> <p>Also attached is the symbolisation for the image in .qml (Quantum GIS Layer Style File) format.</p> <p>A file called RISK RECLASS is also available, where you can find the risk classification maps for each of the studied factors: </p> <ul> <li>Proximity to roads</li> <li>Proximity to industrial areas</li> <li>Ph</li> <li>Soil organic content</li> <li>Slope</li> <li>Soil texture</li> <li>Mining extraction areas </li> <li>Drainage</li> </ul> <p>finally a DATABASE file where the data of the 360 points for the calculation of the risk maps can be found. </p>
Data for "A stochastic model of geomorphic risk due to episodic river aggradation and degradation"
<p>The code and the dataset can be read/run by using Matlab. The description as follows:<br>1. Dataset of riverbed measurement (long profile and water level gauge data), carbon dating data, and rainfall record in the Laonong River (Taiwan). The dataset are used for the model calibration and the model application. <br>2. The developed riverbed stochastic processing model and the maximum likelihood calibration model. </p> <p>Note: this new version includes the corrected Monte Carlo simulation code and a required Matlab function (fminsearchbnd.m) that was missing in the first version.</p>
Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"
<h1>Replication Package for the Paper: “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis”</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.”</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>
Brain Transcriptome Single-cell (BTS) Atlas: Anndata, Seurat Object, CellTypist model, and Disorder Risk Geneplot
<p>Brain Transcriptome Single-cell Atlas (BTS) Anndata, Seurat object, and Celltypist model for further use of the atlas. The Celltypist model can be utilized to accurately annotate cell types in new datasets based on the atlas. Plots illustrating the expression profile for 3,380 neurological disorder risk genes across the atlas are also uploaded. Further availability for the data can be requested by the corresponding author.<br><br>This dataset is published in Kim, S., Lee, J., Koh, I.G. <em>et al.</em> An integrative single-cell atlas for exploring the cellular and temporal specificity of genes related to neurological disorders during human brain development. <em>Exp Mol Med</em> <strong>56</strong>, 2271–2282 (2024). https://doi.org/10.1038/s12276-024-01328-6</p>
Quantitative assessment of the reservoir-induced and urbanization-induced impact on multivariate flood risk via a nonstationary vine Copula model
<p>Here we show the results of the characteristics of the floods at Huayuankou, Lanzhou and Toudaoguai gauges selected by AMS and POT mentod, respectively. Besides that, the inormation about the reservoirs and the imprevious layer in the control catchment of each station is also uploaded.</p>
Optimization of cancer risk assessment model for PM2.5-bound PAHs Application in Shanxi of China SM
<p><span>To optimize the models for cancer risk assessment of PM<sub>2.5</sub></span><span>-bound polycyclic aromatic hydrocarbons (PAHs), six models of three types utilized with high frequency in recent years, including the USEPA recommendation (Model I), inhalation carcinogen unit risk (Models </span><span>II</span><span>A–IID), and three exposure pathways (inhalation, dermal, and oral) (Model </span><span>III), were selected to calculate potential cancer risk using the benzo[a]pyrene toxicity equivalent method. The results indicated no significant differences in the risk values between Models IIA and III, and no significant differences were found among Models </span><span>IIB</span><span>, </span><span>IIC</span><span>, and </span><span>IID</span><span>. </span><span>However, there were significant differences between Models I and </span><span>II</span><span>, and between Models I and </span><span>III. </span><span>Furthermore, the optimization of Model I indicated that the population exposure parameters for each country should be used and age differences should not be disregarded. However, no significant differences were observed with respect to gender.</span><span> In conclusion, Model I was superior to the other models; when age differences were not considered, Model IID (the USEPA's integrated risk information system) was preferred to replace Models </span><span>IIB</span><span> and </span><span>IIC</span><span>. Models IIA and </span><span>III may overestimate the risk of carcinogens, </span><span>and their parameters need to be reoptimized to accurately assess cancer risks.</span></p>
On modelling airborne infection risk
<div> <div> <div> <p>Airborne infection risk analysis is usually performed for enclosed spaces where susceptible indi- viduals are exposed to infectious airborne respiratory droplets by inhalation. It is usually based on exponential, dose-response models of which a widely used variant is the Wells-Riley (WR) model. We revisit this infection-risk estimate and extend it to the population level. We use an epidemiolog- ical model where the mode of pathogen transmission, airborne or contact, is explicitly considered. We illustrate the link between epidemiological models and the WR and the Gammaitoni and Nucci models. We argue that airborne infection quanta are, up to an overall density, airborne infectious respiratory droplets modified by a parameter that depends on biological properties of the pathogen, physical properties of the droplet, and behavioural parameters of the individual. We calculate the time-dependent risk to be infected for two scenarios. We show how the epidemic infection risk de- pends on the viral latent period and the event time, the time infection occurs. Infection risk follows the dynamics of the infected population. As the latency period decreases, infection risk increases. The longer a susceptible is present in the epidemic, the higher its risk of infection for equal exposure time to the mode of transmission is.</p> </div> </div> </div>
Eutrophication Risk Index (ERI) for the Cerrado and Caatinga: Modeling Scenarios for 2030 and 2040
<p><span>To assist in mapping Water Pollution Risk (WPR), we developed the Eutrophication Risk Index (ERI). This index is designed to assess and predict the vulnerability of water bodies to eutrophication, a process driven by excessive nutrient accumulation—particularly nitrogen and phosphorus—resulting in uncontrolled algal growth. The ERI helps identify at-risk areas and supports the development of more effective mitigation strategies aimed at preserving water quality and sustaining aquatic ecosystems.</span></p> <p><span>The proposed Eutrophication Risk Index (ERI) specifically accounts for human pressures on aquatic ecosystems. The ERI is determined by the phosphorus contribution to aquatic environments, derived from urban effluents and the excess nutrients (phosphorus and nitrogen) applied to the soil, measured in tons per hectare per year (Ton ha⁻¹ year⁻¹). To calculate the ERI for the Cerrado and Caatinga, we employed an equation with two main components: one concerning nutrient loss from agricultural systems and the other related to nutrient loss in wastewater.</span></p> <p><strong><span>Nutrient Loss in Agricultural Areas:</span></strong><span><br>Nutrient loss from agricultural areas was estimated using a spatially explicit soil nutrient balance model, incorporating secondary data sources and land use and land cover maps of the study area. For this analysis, we assumed that the nutrient balance in the soil is the difference between total inputs (IN) and total outputs (OUT), where IN includes chemical and organic fertilizers and OUT represents agricultural products. A positive nutrient balance, or surplus, indicates potential nutrient loss that could impact adjacent ecosystems. In our calculations, we also considered phosphorus saturation levels as a risk factor for phosphorus loss, alongside soil types.</span></p> <p><strong><span>Nutrient Loss in Wastewater:</span></strong><span><br>Nutrient loss in wastewater was based on data from the National Water and Sanitation Agency. This method considers the nutrient content in untreated wastewater and in effluents from wastewater treatment plants. We assumed a constant treatment efficiency of 30%, although this value may be optimistic given the primary effluent treatment processes in Brazil. For future assessments, local data on sewage treatment plants could be incorporated into the calculations if available during the project's execution.</span></p> <p><span>This dataset includes empirical data and model simulations developed under the NEXUS project (</span><a href="https://nexus.ccst.inpe.br/" target="_new"><span>https://nexus.ccst.inpe.br/</span></a><span>), which analyzed the interrelationship and challenges of agricultural production, energy, and water resource use in the Caatinga and Cerrado regions. Conducted between 2018 and 2024, the NEXUS project employed a participatory multiscale approach, combining qualitative and quantitative methods from natural and social sciences. Over its six-year duration, the project engaged more than one hundred stakeholders from various sectors, producing diagnostics and scenarios for sustainable futures in these biomes.</span></p> <p><strong><span>Scenarios Descriptions:</span></strong><span><br>The “Green Transition” scenario aligns with the dominant sustainability narrative in the business sector, focusing on efficiency gains and technological solutions (e.g., low-carbon agriculture, energy transition led by large corporations) to address environmental challenges. This scenario envisions agricultural production concentrated in highly productive areas, facilitating the restoration of natural vegetation and fostering an increasingly urban future.</span></p> <p><span>Conversely, the “Lives in Balance” scenario reflects the aspirations and struggles of social movements and traditional communities for recognition and the coexistence of diverse ways of life. It advocates transforming production systems, particularly through decentralized food and energy production, and emphasizes strengthening family farming and agroecological systems.</span></p> <p> </p> <p><strong><span>Acknowledgements</span></strong><span><br>The authors would like to thank the NEXUS Project, funded by the São Paulo Research Foundation – FAPESP (grants 2022/00917-0 and 2017/22269-2), and the Coordination for the Improvement of Higher Education Personnel (CAPES) for their support to Marcela Miranda through the National Postdoctoral Program (grants 88882.317530/2019-1 and 1732909/2017-2).</span></p>
Predictive modelling of brain metastasis risk and non-invasive biomarker detection using DNA methylation signatures
<p>Methylated cell-free DNA was sequenced for 123 BM plasma and compared to plasma methylomes from 107 gliomas, central nervous system (CNS) lymphomas (CNSL), and non-CNS tumor controls. Plasma methylome-based classifiers of BM from other entities were built in fifty 80% discovery set iterations of 92/123 BM samples. External publicly-available tissue methylation data on 442 LUAD, 85 BM, and 146 glioma/CNSL/control samples were acquired for validation and the remaining 31/123 BM plasma samples were used for additional validation.</p>
Monitoring of postpartum body condition at the cow and herd levels: assessing explanatory and predictive power of disease risk models
<p>Objectives</p> <p>1- To define the herd threshold for cows with poor body condition based on its predictive capacity for disease risk at the herd level, and</p> <p>2- to estimate the impact measures on disease rates due to body condition indicators in transition period.</p> <p>Two commercial grazing dairy herds (Herd A=5.034 and herd B=7.965 lactations) from Argentinean Pampa region were used to perform a longitudinal retrospective study during a 4-year period (2014 –2017).Health, reproductive and body condition score (BCS) records were gathered. The BCS (5-point scale) was performed at calving and at the time of reproductive release. The difference between both measures of BCS was used to assess the body condition loss (∆BCS). All the cows not bred by 70 DIM were checked for anestrus.Calving cohorts of 21-day were defined at each herd and parity group through the entire study period. The frequency of cows with BCS<3 or ∆BC>-0.5 at each cohort were calculated and used to define quartiles through whole study period. Quartiles were used, one at a time, as threshold to dichotomize the cohorts to predict the risk that a cohort has a frequency of anestrus over the median.The higher AUC was used as selection criterium to determine the herd level threshold at each HERD and PARITY level. The population attributable fraction (AFP) of anestrus rate to body condition indicators at each cohort was calculated, for every HERD and PARITY level. </p>
Geospatial based model for malaria risk prediction in Kilombero Valley, south-eastern Tanzania
<div> <p><strong>Background</strong>: Malaria continues to pose a major public health challenge in tropical regions. Despite significant efforts to control malaria in Tanzania, there are still residual transmission cases. Unfortunately, little is known about where these residual malaria transmission cases occur and how they spread. In Tanzania, for example, the transmission is heterogeneously distributed. In order to effectively control and prevent the spread of malaria, it is essential to understand the spatial distribution and transmission patterns of the disease. This study seeks to predict areas that are at high risk of malaria transmission so that intervention measures can be developed to accelerate malaria elimination efforts.</p> </div> <p><strong>Methods</strong>: This study employs a geospatial-based model to predict and map out malaria risk area in Kilombero Valley. Environmental factors related to malaria transmission were considered and assigned valuable weights in the Analytic Hierarchy Process (AHP), an online system using a pairwise comparison technique. The malaria hazard map was generated by a weighted overlay of the altitude, slope, curvature, aspect, rainfall distribution, and distance to streams in Geographic Information Systems (GIS). Finally, the risk map was created by overlaying components of malaria risk including hazards, elements at risk, and vulnerability.</p> <p><strong>Results</strong>: The study demonstrates that the majority of the study area falls under the moderate-risk level (61%), followed by the low-risk level (31%), while the high-malaria risk area covers a small area, which occupies only 8% of the total area.</p> <p><strong>Conclusion</strong>: The findings of this study are crucial for developing spatially targeted interventions against malaria transmission in residual transmission settings. Predicted areas prone to malaria risk provide information that will inform decision-makers and policymakers for proper planning, monitoring, and deployment of interventions.</p>
A global hybrid tropical cyclone risk model based upon statistical and coupled climate models - Supporting figures and data
<p><strong>Introduction</strong></p><p>This contribution consists of 1) supporting figures and 2) supporting data for the submitted manuscript "A Global Hybrid Tropical Cyclone Risk Model based upon Statistical and Coupled Climate Models." Supporting figures are presented in two interactive HTML documents. The supporting datasets contain tropical cyclone event sets, catalogs, and an example analysis that plots summaries of the simulated catalogs and compares them to historical observations. All files are provided for the 400 ensemble members (event sets) that represent climate model years 1981-2020 (40 years) from the first 10 CESM-LE members.</p><p><strong>Contents</strong></p><p>./Catalog/sim_650</p><p>Tropical cyclone annual catalog for each basin based on the CESM-LE distribution of ENSO phases. 650 simulations are provided for each of 400 event sets. Each file contains the catalog for a single basin and is written as catalog_tc_(BASIN)_sim650_my400_nbinom_condmeanensojma.csv., where BASIN can be NA, EP, WP, NI, SI, SP.</p><p>./Documentation</p><p>TCMODEL_UQAM_EXAMPLE.html: Analysis script showing example of use of the tropical cyclone catalogs and comparison to historical observations.</p><p>UQAM_TC_Model_Data_Supplement_Dictionary.xlsx: Data dictionary of all data supplement file contents.</p><p>./IBTRACS</p><p>Summary of IBTrACS data required in the analysis script.</p><p>./TrajectoryBanks</p><p>Contains subdirectories for each basin (EP, NAT, NI, SI, SP, WP)</p><p>Each subdirectory contains several files summarizing the event sets, or banks, of tropical cyclone trajectories.</p><p><strong>Versions</strong></p><p>Version 1.0.1: Updated TCMODEL_UQAM_SUPPORTING_FIGURES.zip for revisions to submitted manuscript.</p><p>Version 1.0.0: Original version.</p><p> </p>
Implementation of the Care Ecosystem Training Model for Individuals With Dementia in a High-risk, Integrated Care Management
ClinicalTrials.gov study NCT04556097. IPD Sharing: NO. Countries: 1. Publications: 1.
Coronary Imaging and Metabolic Indicators-Based Risk Prediction Model for Coronary Artery Disease(CMI-RiskCAD)
ClinicalTrials.gov study NCT07353762. IPD Sharing: NO. Countries: 1. Publications: 1.
Development and Evaluation of High Risk Group Prediction Model in T1 Stage Renal Cell Cancer Using Molecular Biomarkers
ClinicalTrials.gov study NCT03694912. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.
Risk Factors and Prediction Model of Cancer-associated Venous Thromboembolism
ClinicalTrials.gov study NCT05729464. IPD Sharing: UNDECIDED. Countries: 1. Publications: 11.
Longitudinal Multimarker Risk Models for Very Elderly Patients With Heart Failure and Preserved Ejection Fraction
ClinicalTrials.gov study NCT05992558. IPD Sharing: YES. Countries: 1. Publications: 8.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.