Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
How to select predictive models for decision making or causal inference? Experiments data
<p>This is the full result data for the experiments of the paper : Doutreligne, M., & Varoquaux, G. (2023). How to select predictive models for decision making or causal inference?, https://hal.science/hal-03946902. <br><br>The code repository is : https://github.com/soda-inria/caussim/tree/main</p> <p>The files in this dataset are the one for the most computationnally costly experiments. There is one folder for each of the four datasets used in the paper. Then, one folder for each of the experimental setup. The files required for the main figure (Fig.3) of the paper are the one labelled #fig3 in the following descriptions.</p> <p>Details on the files : </p> <p>.<br>├── acic_2016_save<br>│ ├── acic_2016__nuisance_non_linear__candidates_hist_gradient_boosting__dgp_1-77__rs_1-5<br>│ │ └── run_logs.csv: results for the experiment with non linear models for both the nuisances and the candidates<br>│ ├── acic_2016__nuisance_non_linear__candidates_ridge__dgp_1-77__rs_1-10<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates<br>│ └── acic_2016__stacked_regressor__dgp_1-77__seed_1-10<br>│ └── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and non linear models for the candidates #fig3<br>├── acic_2018_save<br>│ └── acic_2018__nuisance_non_linear__candidates_hist_gradient_boosting__first_uid_432<br>│ └── run_logs.csv results for the experiment with stacked models (linear and non linear) for the nuisances models and non linear models for the candidates #fig3<br>├── caussim_save<br>│ ├── caussim__linear_regressor__test_size_5000__n_datasets_1000<br>│ │ ├── run_logs.csv: results for the experiment with stacked models for the nuisances models and linear models for the candidates <br>│ │ └── simu.yaml: configuration file of the experiment<br>│ ├── caussim__nuisance_non_linear__candidates_ridge__overlap_01-247_join_nuisance_train_set<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates, joined sets for the nuisances and the candidates<br>│ ├── caussim__nuisance_non_linear__candidates_ridge__overlap_01-247_separated_nuisance_train_set<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates, separated sets for the nuisances and the candidates<br>│ └── caussim__stacked_regressor__test_size_5000__n_datasets_1000<br>│ ├── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and linear models for the candidates #fig3<br>│ └── simu.yaml: configuration file of the experiment<br>└── twins_save<br> └── twins__stacked_regressor__rs_1-10__overlap_0.1-3<br> └── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and non linear models for the candidates #fig3</p>
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
<p>Despite their success, large language models (LLMs) face the critical challenge of hallucinations, generating plausible but incorrect content. While much research has focused on hallucinations in multiple modalities including images and natural language text, less attention has been given to hallucinations in source code, which leads to incorrect and vulnerable code that causes significant financial loss. To pave the way for research in LLMs' hallucinations in code, we introduce Collu-Bench, a benchmark for predicting code hallucinations of LLMs across code generation (CG) and automated program repair (APR) tasks. Collu-Bench includes 13,234 code hallucination instances collected from five datasets and 11 diverse LLMs, ranging from open-source models to commercial ones. <br>To better understand and predict code hallucinations, Collu-Bench provides detailed features such as the per-step log probabilities of LLMs' output, token types, and the execution feedback of LLMs' generated code for in-depth analysis. In addition, we conduct experiments to predict hallucination on Collu-Bench, using both traditional machine learning techniques and neural networks, which achieves 22.03 -- 33.15% accuracy. Our experiments draw insightful findings of code hallucination patterns, reveal the challenge of accurately localizing LLMs' hallucinations, and highlight the need for more sophisticated techniques.</p>
Supplementary files: Machine Learning Insights into Türkiye's Climate Variability: Predictive Modelling and Spatial Analysis
<p>This dataset and python code were used in the study titled "Machine Learning Insights into Türkiye's Climate Variability: Predictive Modelling and Spatial Analysis".</p>
Temperature data corresponding to "Hybrid Phenology Modeling for Predicting Temperature Effects on Tree Dormancy"
<p>MERRA2 Temperature data corresponding to "Hybrid Phenology Modeling for Predicting Temperature Effects on Tree Dormancy"</p>
Dataset for the publication of "WRF model parameter calibration to improve the prediction of tropicalcyclones over the Bay of Bengal using Machine Learning-basedMultiobjective Optimization"
<p>The dataset consists of the modified WRF model software, that can be extracted and used in any Linux system with preinstalled required software.</p> <p>The namelists_file.zip consists of the namelist.input files that are used for the default and calibration simulations with different driving data namely, FNL files at 1deg with two nested domains, ERA files at 1deg with two nested domains, ERA files at 0.25deg with a single domain, and the ERA files at 0.25deg with two nested domains.</p>
Supplementary Information for Consonance-emerging Hebbian Learning neural network model predicts discreteness of musical scales and the Natural Just Intonation scale
<p><strong>The following phenomena and features are apparent in music and auditory perception in general: the discreteness of the tones in musical scales</strong> [1]<strong>, the prevalence of the tonal frequency span of one semitone (100 cents) in musical scales across cultures </strong>[1]<strong>, the list of tonal intervals ordered by consonance [2], and the musical performers’ preference of the Natural Just-Intonation scale [3] (A). However, researchers still have no agreement about the causes and the emergence of said phenomena (A). Here we show that the consonance-pattern emerging neural network model introduced in our previous study [4], predicts and yields all the said phenomena (A) with a precision of 1/100<sup>th</sup> of a semitone (1 cent). This precision is beyond the resolution of human hearing </strong>[5], [6], [7]. <strong>Since the Hebbian learning paradigm and harmonicity are the main features of our model, we propose that they are sufficient conditions for any system to yield the said phenomena (A). Therefore, they have a crucial role in processing pitch, consonance, and music perception in general. As a consequence, we additionally propose that the mentioned phenomena (A) are a balanced result of the joint workings of the Hebbian paradigm (nurture and cultural exposure) and harmonicity (auditory physics and biology).</strong></p>
Comparative Accuracies of Models for Drag Prediction During Geomagnetically Disturbed Periods: A First Principles Model versus Empirical Models
<p>This dataset contains observational and TIEGCM simulation data used for a manuscript that is being submitted to the <em>Space Weather</em> journal. The abstract for the study follows: We examine the accuracy of density prediction by the first principals model Thermosphere Ionsosphere Electrodynamics General Circulation Model (TIEGCM) developed by the National Center for Atmospheric Research and compare it to the accuracy of three empirical models: Jacchia 71, the Naval Research Laboratory Mass Spectrometer Incoherent Scatter Extended 2000 (NRLMSIS), Jacchia 1971 and Jacchia-Bowman 2008. Comparisons are made for three large storms: the October 2003 storm, the March 2013 storm, and the March 2015 storm. To evaluate the accuracy of these models we use tracking data for nine space objects in low earth orbit (three for each storm). Additionally, and evaluate the accuracy of the TIEGCM and NRLMSIS with data from high precision accelerometers on the Challenging Minisatellite Payload (CHAMP) and Gravity field and Circulation Explorer (GOCE) satellites. The goal is to assess the use of a first principles model as a potential tool for forecasting satellite drag during large magnetic storms. We find that the TIEGCM accuracy is substantially better than for the Jacchia 71 and NRLMSIS models. The accuracies of the TIEGCM and JB2008 models are similar, but overall the TIEGCM is more accurate. We found smaller mean percentage differences for TIEGCM versus CHAMP than for NRLMIS for the Halloween Storm and smaller differences than results published for JB2008 and the assimilative model HASDM. The empirical models are at present more practical for operational purposes, but the first principles TIEGCM was developed as a research model and with a greater focus on operational use offers the potential for improved utility during stressing conditions.</p>
A machine learning based prediction model for life expectancy
<p>The social and financial systems of many nations throughout the world are significantly impacted by life expectancy (LE) models. Numerous studies have pointed out the crucial effects that life expectancy projections will have on societal issues and the administration of the global healthcare system. The computation of life expectancy has primarily entailed building an ordinary life table. However, the life table is limited by its long duration, the assumption of homogeneity of cohorts and censoring. As a result, a robust and more accurate approach is inevitable. In this study, a supervised machine learning model for estimating life expectancy rates is developed. The model takes into consideration health, socioeconomic, and behavioral characteristics by using the eXtreme Gradient Boosting (XGBoost) algorithm to data from 193 UN member states. The effectiveness of the model's prediction is compared to that of the Random Forest (RF) and Artificial Neural Network (ANN) regressors utilized in earlier research. XGBoost attains an MAE and an RMSE of 1.554 and 2.402, respectively outperforming the RF and ANN models that achieved MAE and RMSE values of 7.938 and 11.304, and 3.86 and 5.002, respectively. The overall results of this study support XGBoost as a reliable and efficient model for estimating life expectancy.</p>
Including a spatial predictive process in band recovery models improves inference for Lincoln estimates of animal abundance
<p>Abundance estimation is a critical component of conservation planning, particularly for exploited species where managers set regulations to restrict harvest based on current population size. An increasingly common approach for abundance estimation is through integrated population modeling (IPM), which uses multiple data sources in a joint likelihood to estimate abundance and additional demographic parameters. Lincoln estimators are one commonly used IPM component for harvested species, which combine information on the rate and the total number of individuals harvested within an integrated band-recovery framework to estimate abundance at large scales.</p> <p>A major assumption of the Lincoln estimator is that banding and recoveries are representative of the whole population, which may be violated if major sources of spatial heterogeneity in survival or harvest rates are not incorporated into the model. We developed an approach to account for spatial variation in harvest rates using a spatial predictive process, which we incorporated into a Lincoln estimator IPM.</p> <p>We simulated data under different configurations of sample sizes, harvest rates, and sources of spatial heterogeneity in harvest rate to assess potential model bias in parameter estimates. We then applied the model to data collected from a field study of wild turkeys (<em>Meleagris gallapavo</em>) to estimate local and statewide abundance in Maine, USA.</p> <p>We found that the band recovery model that incorporated a spatial predictive process consistently provided estimates of adult and juvenile abundance with low bias across a variety of spatial configurations of harvest rate and sampling intensities. When applied to data collected on wild turkeys, a model that did not incorporate spatial heterogeneity underestimated the harvest rate in some sub-regions. Consistent with simulation results, this led to over-estimation of both local and statewide abundance.</p> <p>Our work demonstrates that a spatial predictive process is a viable mechanism to account for spatial variation in harvest rates and limit bias in abundance estimates. This approach could be extended to large-scale band recovery datasets and has applicability for the estimation of population parameters in other ecological models as well.</p>
Predictive modelling of brain metastasis risk and non-invasive biomarker detection using DNA methylation signatures
<p>Methylated cell-free DNA was sequenced for 123 BM plasma and compared to plasma methylomes from 107 gliomas, central nervous system (CNS) lymphomas (CNSL), and non-CNS tumor controls. Plasma methylome-based classifiers of BM from other entities were built in fifty 80% discovery set iterations of 92/123 BM samples. External publicly-available tissue methylation data on 442 LUAD, 85 BM, and 146 glioma/CNSL/control samples were acquired for validation and the remaining 31/123 BM plasma samples were used for additional validation.</p>
Bayesian species distribution models integrate presence-only and presence-absence data to predict deer distribution and relative abundance
<p>Using geospatial data of wildlife presence to predict a species distribution across a geographic area is among the most common tools in management and conservation. The collection of high-quality presence-absence data through structured surveys is, however, expensive, and managers usually have access to larger amounts of low-quality presence-only data collected by citizen scientists, opportunistic observations, and culling returns for game species. Integrated Species Distribution Models (ISDMs) have been developed to make the most of the data available by combining the higher-quality, but usually scarcer and more spatially restricted presence-absence data, with the lower quality, unstructured, but usually more extensive presence-only datasets. Joint-likelihood ISDMs can be run in a Bayesian context using INLA (Integrated Nested Laplace Approximation) methods that allow the addition of a spatially structured random effect to account for data spatial autocorrelation. Here, we apply this innovative approach to fit ISDMs to empirical data, using presence-absence and presence-only data for the three prevalent deer species in Ireland: red, fallow and sika deer. We collated all deer data available for the past 15 years and fitted models predicting distribution and relative abundance at a 25 km<sup>2</sup> resolution across the island. Models' predictions were associated to spatial estimates of uncertainty, allowing us to assess the quality of the model and the effect that data scarcity has on the certainty of predictions. Furthermore, we checked the performance of the three species-specific models using two datasets, independent deer hunting returns and deer densities based on faecal pellet counts. Our work clearly demonstrates the applicability of spatially-explicit ISDMs to empirical data in a Bayesian context, providing a blueprint for managers to exploit unexplored and seemingly unusable data that can, when modelled with the proper tools, serve to inform management and conservation policies.</p>
Saved model and preprocessed data for "CRMnet:a deep learning model for predicting gene expression from large regulatory sequence datasets"
<p>Saved TUNet model and preprocessed training data for "CRMnet: a deep learning model for predicting gene expression from large regulatory sequence datasets"</p> <p>To load the trained model:</p> <pre><code class="language-python">import tensorflow as tf tf.keras.models.load_model("path to the model folder")</code></pre> <p>for more information please find our repository: https://github.com/jiayuwen/CRMnet</p>
Predictive Search Model of Flocking for Quadcopter Swarm in the Presence of Static and Dynamic Obstacles
<p>The folder includes experimental data for the paper titled "Predictive Search Model of Flocking for Quadcopter Swarm in the Presence of Static and Dynamic Obstacles".</p> <p>In the paper, we present a Predictive Search Model (PSM) for flocking with Heading and Speed Shared (HSS) and Heading and Speed Unshared (HSU) prediction methods. We compare the performance of PSM with Potential Field Model (PFM) in the presence of static and dynamic obstacles in simulation. Also, we validate the performance of PSM with a quadcopter swarm indoors.</p> <p>The 'simulation experiments' folder includes simulation experiment data and MATLAB scripts that can simulate the experiments and provide plots for analysis.</p> <p>The 'quadcopter experiments' folder includes quadcopter experiment data and MATLAB scripts that can simulate the experiments and provide plots for analysis.</p>
Predicting and modeling protein-protein interactions in E. coli envelopome
<p>Structural models and Supplementary data described in the reference:</p> <p>Deep learning-driven insights into super protein complexes for outer membrane protein biogenesis in bacteria.</p> <p>Mu Gao, Davi Nakajima An, and Jeffrey Skolnick<em>. eLife</em>, 2022. <strong>11</strong>: p. e82885.</p> <p>List of files:</p> <ul> <li>af2c_fea_220331.tar -- A tarball of input features of full E coli proteome to AF2Complex version 1.3.0. Feature files are pickled and gzipped, which AF2Complex v1.3 takes as input directly.</li> </ul> <p>Results of an application to E coli envelopome on four query proteins from the outer membrane biogenesis pathway.</p> <ul> <li>Supplementary Data.xlsx -- Virtual PPI screening results of PpiD, YfgM, SurA, and BamA.</li> <li>screening_top1_models.zip -- Compressed top 1 dimeric models of top hits from the PPI screening. Note that these models are unrelaxed.</li> <li>predicted structural models.zip -- Compressed structural models of supercomplexes formed in the OMP biogenesis pathway described in the reference.</li> </ul>
A big data–model integration approach for predicting epizootics and population recovery in a keystone species
<p>Infectious diseases pose a significant threat to global health and biodiversity. Yet, predicting the spatiotemporal dynamics of wildlife epizootics remains challenging. Disease outbreaks result from complex non-linear interactions among a large collection of variables that rarely adhere to the assumptions of parametric regression modeling. We adopted a non-parametric machine learning approach to model wildlife epizootics and population recovery, using the disease system of colonial black-tailed prairie dogs (BTPD, <em>Cynomys ludovicianus</em>) and sylvatic plague as an example. We synthesized colony data between 2001–2020 from eight USDA Forest Service National Grasslands across the range of BTPD in central North America. We then modeled extinctions due to plague and colony recovery of BTPD in relation to complex interactions among climate, topoedaphic variables, colony characteristics, and disease history. Extinctions due to plague occurred more frequently when BTPD colonies were spatially clustered, in closer proximity to colonies decimated by plague during the previous year, following cooler than average temperatures the previous summer, and when wetter winter/springs were preceded by drier summer/falls. Rigorous cross-validations and spatial predictions indicated that our final models predicted plague outbreaks and colony recovery in BTPD with high accuracy (e.g., AUC generally > 0.80). Thus, these spatially-explicit models can reliably predict the spatial and temporal dynamics of wildlife epizootics and subsequent population recovery in a highly complex host-pathogen system. Our models can be used to support strategic management planning (e.g., plague mitigation) to optimize benefits of this keystone species to associated wildlife communities and ecosystem functioning. This optimization can reduce conflicts among different landowners and resource managers, as well as economic losses to the ranching industry. More broadly, our big data–model integration approach provides a general framework for spatially-explicit forecasting of disease-induced population fluctuations, for use in natural resource management decision-making.</p>
Data Set for Predicting the Performance of ATL Model Transformations
<p>Model transformation languages are special-purpose languages, which are designed to define transformations as comfortably as possible, i.e., often in a declarative way. With the increasing use of transformations in various domains, the complexity and size of input models are also increasing. However, developers often lack suitable models for performance testing. We have therefore conducted experiments in which we predict the performance of model transformations based on characteristics of input models using machine learning approaches. This dataset contains our raw and processed input data, the scripts necessary to repeat our experiments, and the results we obtained.</p> <p>Our input data consists of the time measurements for six different transformations defined in the Atlas Transformation Language (ATL), as well as the collected characteristics of the real-world input models that were transformed. We provide the script that implements our experiments. We predict the execution time of ATL transformations using the machine learning approaches linear regression, random forests and support vector regression using a radial basis function kernel. We also investigate different sets of characteristics of input models as input for the machine learning approaches. These are described in detail in the provided documentation.pdf. The results of the experiments are provided as raw data in individual cvs files. Additionally, we calculated the mean absolute percentage error in % and the 95th percentile of the absolute percentage error in % for each experiment and provide these results. Furthermore, we provide our Eclipse plugin, which collects the characteristics for a set of given models, the Java projects used to measure the execution time of the transformations, and other supporting scripts, e.g. for the analysis of the results.</p> <p>A short introduction with a quick start guide can be found in README.md and a detailed documentation in documentaion.pdf.</p>
Raw Data for the article: Mortality after transjugular intrahepatic portosystemic shunt in older adult patients with cirrhosis: A validated prediction model
<p><strong>Background and aims: </strong>Implantation of a transjugular intrahepatic portosystemic shunt (TIPS) improves survival in patients with cirrhosis with refractory ascites and portal hypertensive bleeding. However, the indication for TIPS in older adult patients (greater than or equal to 70 years) is debated, and a specific prediction model developed in this particular setting is lacking. The aim of this study was to develop and validate a multivariable model for an accurate prediction of mortality in older adults.</p> <p><strong>Approach and results: </strong>We prospectively enrolled 411 consecutive patients observed at four referral centers with de novo TIPS implantation for refractory ascites or secondary prophylaxis of variceal bleeding (derivation cohort) and an external cohort of 415 patients with similar indications for TIPS (validation cohort). Older adult patients in the two cohorts were 99 and 76, respectively. A cause-specific Cox competing risks model was used to predict liver-related mortality, with orthotopic liver transplant and death for extrahepatic causes as competing events. Age, alcoholic etiology, creatinine levels, and international normalized ratio in the overall cohort, and creatinine and sodium levels in older adults were independent risk factors for liver-related death by multivariable analysis.</p> <p><strong>Conclusions: </strong>After TIPS implantation, mortality is increased by aging, but TIPS placement should not be precluded in patients older than 70 years. In older adults, creatinine and sodium levels are useful predictors for decision making. Further efforts to update the prediction model with larger sample size are warranted.</p>
Dataset from Maith, O., Baladron, J., Einhäuser, W., & Hamker, F. H. (2023). Exploration behavior after reversals is predicted by STN-GPe synaptic plasticity in a basal ganglia model. Submitted to iScience.
<p>This dataset contains all analyzed data from the study "Maith, O., Baladron, J., Einhäuser, W., & Hamker, F. H. (2023). Exploration behavior after reversals is predicted by STN-GPe synaptic plasticity in a basal ganglia model. Submitted to iScience.". It includes the behavioral data of 20 human participants (folder "psychExp") and of simulations of a neuro-computational basal ganglia model (folder "simulations") of the study.</p> <p>To replicate the results of the study, the dataset can be analyzed using the code provided separately under the following identifier: https://doi.org/10.5281/zenodo.6555886. The dataset is organized in the directory structure required for this purpose.</p> <p>For the human participants, only preprocessed eye-tracking and general behavioral data (.mat files) and the final analyzed behavioral data (output files) generated with the script "get_vps_outputs.m" (folder psychExp/..../3_srcAna/) are available. For more information about preprocessing steps as well as raw data of the eye-tracking experiment, please contact us by email (click <a href="https://www.tu-chemnitz.de/urz/mail/adrx.html?1-d29sZmdhbmcuZWluaGFldXNlci10cmV5ZXJAcGh5c2lrLg==">here</a>).</p>
SPIRIT Checklist & Model Consent for 'Predicting Acute and Post-Recovery Outcomes in Cerebral Malaria and Other Comas by Optical Coherence Tomography (OCT in CM) – A protocol for an observational cohort study of Malawian children'
<p>This dataset contains the SPIRIT checklist (adapted to a observational trial) and model consent forms for the OCT in CM study protocol. The protocol will be submitted as a paper to Wellcome Open Research.</p>
Data from: Prediction of three years of annual rain attenuation statistics at Ka-band in French Guiana using the Numerical Weather Prediction model WRF
<p><span>This study highlights the interest in using an Atmospheric Numerical Simulator (ANS) relying on a high-resolution weather forecast model coupled with an ElectroMagnetic Module (EMM) to compute Ka-band rain attenuation statistics in an equatorial region. An optimization of the parametrisation of the Weather Research and Forecasting meteorological model (WRF) is carried out using measurements collected from a propagation experiment carried out by CNES and ONERA near Kourou in French Guiana. Both simulated and experimental annual Complementary Cumulative Distribution Functions (CCDF) of rain attenuation are presented in this dataset.</span></p> <p>More specifically, this dataset includes the statistical distribution from both the WRF-EMM model and from the propagation experiment for the years 2017, 2018, 2020 and the whole three-year period.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.