Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Disentangling Sources of Uncertainty in CLM5 Model Predictions: Water, Energy, and Carbon Fluxes at European Observation Sites
<p>The datasets include:</p> <ul> <li>EC data from Europement measurement sites in <a href="https://www.icos-cp.eu/data-products/2G60-ZHAK">ICOS</a>, <a href="https://fluxnet.org/login/?redirect_to=/data/download-data/">FLUXNETS</a>, and <a href="https://doi.org/10.34731/x9s3-Kr48">COSMOS-Europe</a>.</li> <li>Ensemble simulation data used for analysis</li> </ul> <p>The atmospheric forcings used in driving the model were all local measurements pre-processed using the script in GitHub repository <a href="https://github.com/FedoAIworld/CLM5-Disentangling-Uncertainty/tree/main/00_create_forcing_ds">CLM5-Disentangling-Uncertainty</a>.</p>
Dataset and machine learning models for seismic response predictions of small-to-medium continuous girder bridges
<p>This upload includes the dataset and machine learning models (based on Matlab platform) for longitudinal seismic response predictions of multi-span highway girder bridges, which have a typical span length of 30 m supported by reinforced concrete (RC) bridge bents and abutments through spherical steel bearings. The input variables (features) are five structural parameters of studied bridges and seven intensity measures of earthquakes. The output variables (labels) are peak column drifts and peak bearing deformations. The dataset is developed by conducting a total number of 720 nonlinear time-history analyses considering the uncertainty of bridges and earthquakes. Machine learning models are developed using two popular machine learning algorithms named artificial neural network (ANN) and support vector regression (SVR).</p>
FESOM-REcoM model data: Predicting future distribution of Antarctic toothfish (Dissostichus mawsoni), with implications for Marine Protected Areas in the Southern Ocean
<p>This data set belongs to <strong>"Predicting future distribution of Antarctic toothfish (<em>Dissostichus mawsoni</em>), with implications for Marine Protected Areas in the Southern Ocean"</strong> by <a>Rebecca Konijnenberg, Cara Nissen, Casper Kraan, Jilda Cavacco, Peter Yates, Philippe Ziegler, and Katharina Teschke (in preparation). </a></p> <p>Contact for data set: cara.nissen@colorado.edu</p> <p>The data provided here are post-processed from the raw FESOM1.4-REcoM2 model output which can be obtained <br>at the World Data Center for Climate (WDCC): <a href="https://www.wdc-climate.de/ui/project?acronym=HighRes_highLat_SO">https://www.wdc-climate.de/ui/project?acronym=HighRes_highLat_SO</a></p> <p>simA (historical simulation): <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_hist_vA_vC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_hist_vA_vC</a><br>simA-ssp126: <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s126_vA_vC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s126_vA_vC</a><br>simA-ssp245: <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s245_vA_vC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s245_vA_vC</a><br>simA-ssp370: <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s370_vA_vC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s370_vA_vC</a><br>simA-ssp585: <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s585_vA_vC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_A_s585_vA_vC</a><br>simB (control simulation): <a href="https://doi.org/10.26050/WDCC/FESOM14-REcoM2_B_1921_cA_cC">https://doi.org/10.26050/WDCC/FESOM14-REcoM2_B_1921_cA_cC</a></p> <p>Here, we provide decadal monthly climatologies of the following variables: sea-ice fraction, surface small-phytoplankton chlorophyll, surface diatom chlorophyll, as well as potential temperature, practical salinity and oxygen concentrations (all at the surface, at the bottom, at 100m, at 250m, at 500m, at 1000m, and at 1500m). For example, "tos" is temperature in ocean at surface", while "tob" is "temperature in ocean at bottom" (with <em>bottom</em> here being the deepest available model grid cell at each location).</p> <p>We also provide the script used to re-grid the original model output to the regular 0.25° x 0.0625° mesh used in this study. </p>
SBC Predictive Model
<p>Data and code used for the development of an <strong>ethnoarchaeological inductive predictive model</strong> in the central italian Alps (Upper Brembo Valley, province of Bergamo, Italy). SBC is for "Sorgenti del Brembo di Carona" (Source of the Carona branch of the River Brembo).</p> <p>The position of modern summer farms, along with a sample of random points, and physical landscape data, are used to calculate an inductive predictive model centered on pastoral activities. The calculation is done mainly with GRASS GIS and R/RStudio. Some data are pre-processed using QGIS.</p> <p>The model creation and output is <strong>discussed in:</strong> CROCE, E., CARRER, F. & ANGELUCCI, D.E., 2025, Ethnoarchaeological Inductive Predictive Model: A Field Test in the Italian Alps<em>,</em> <em>Journal of Archaeological Method and Theory</em> 32, 43. <a href="https://doi.org/10.1007/s10816-025-09712-w">https://doi.org/10.1007/s10816-025-09712-w</a></p> <p>The inductive predictive model was <strong>originally created as part of a PhD Thesis</strong>: <em>CROCE E., 2022, Archeologia d'alta quota alle sorgenti del Brembo, </em>PhD Thesis, A.A. 2020/2021, Università di Trento<em>.</em> <a href="https://dx.doi.org/10.15168/11572_350299" rel="nofollow">https://dx.doi.org/10.15168/11572_350299</a></p> <p>The methodology is <strong>based upon the previous work of F. Carrer</strong> (which co-authored also the present model): <em>CARRER F., 2013, </em>An ethnoarchaeological inductive model for predicting archaeological site location: a case-study of pastoral settlement patterns in the Val di Fiemme and Val di Sole (Trentino, Italian Alps)<em>, Journal of Anthropological Archaeology, 32, pp. 54-62.</em> <a href="https://doi.org/10.1016/j.jaa.2012.10.001" rel="nofollow">https://doi.org/10.1016/j.jaa.2012.10.001</a></p> <p> </p> <p> </p>
Datasets for understanding the importance of conformation in property prediction models
<p>Descriptor and conformer data sets for molecular property and reaction selectivity prediction tasks. The PQC data set was created based on a part of the PubChemQC PM6 dataset (J. Chem. Inf. Model. 2020, 60, 12, 5891–5899), which contains two- and three-dimensional descriptors and conformers. The APTC data sets are based on the data sets for asymmetric phase transfer catalysts with enantio-selectivity (<a href="https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets" target="_blank" rel="noopener">https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets</a>). The melting point data set was created from the Jean-Claude Bradley Double Plus Good (Highly Curated and Validated) Melting Points Dataset (<a href="https://doi.org/10.6084/m9.figshare.1031638.v1">https://doi.org/10.6084/m9.figshare.1031638.v1</a>).</p> <p>They contained descriptors and conformers to train and validate machine learning models.</p> <p>Detailed explanations on how to use these datasets are found in the Github repository: <a href="https://github.com/YuHamakawa/Conformation-Importance-ML-Models">https://github.com/YuHamakawa/Conformation-Importance-ML-Models</a>. </p> <p> </p> <p> </p>
Electronic Supplement to: A Predictive Model for Divalent Element Partitioning between Clinopyroxene and Basaltic Melt and a Europium-in-Plagioclase-Clinopyroxene Oxybarometer for Cumulate Rocks
<p>Contents</p> <p>Supplementary Figures</p> <p>Supplementary Calculator Spreadsheet</p> <ul> <li>A calculator for divalent element partitioning between the clinopyroxene M2 site and silicate melt</li> <li>An fO2-, temperature- and composition- dependent clinopyroxene-melt Eu partition coefficient calculator for many samples, each at a single fO2</li> <li>An fO2-, temperature- and composition- dependent clinopyroxene-melt Eu partition coefficient calculator for a single sample at many fO2s</li> <li>A Eu-in-clinopyroxene-melt oxybarometer</li> <li>A Eu-in-plagioclase-clinopyroxene oxybarometer</li> </ul> <p>Supplementary Code</p> <ul> <li>A Eu-in-plagioclase-clinopyroxene oxybarometer</li> <li>A Monte Carlo-based fO2 uncertainty calculator</li> </ul>
Online water quality monitoring data from full scale CS#3 DWDN for the DBP prediction model
<p>Online water quality data though the drinking water distribution network. More than 1 year of data.</p> <p>SCADA data source.</p> <p>Provide water quality of the whole system at selected locations.</p>
Figure 4 in Calling phenology of anurans in a tropical rainforest in South Mexico: testing predictive models
Figure 4. Best models of each general structure. Model 6 (M6) for linear and 13 (M13) for sinusoidal structure. I = relative calling intensity, t = time. The solid lines represent the predictions of the models; the points represent the data acquired from the ARS.
Figure 3 in Calling phenology of anurans in a tropical rainforest in South Mexico: testing predictive models
Figure 3. Rose-diagram of anuran assemblage relative calling intensity. The length and colour (light = low, dark = high) of the bars from the centre indicate the intensity of vocalisation for the temporal section. A Rayleigh z test for unimodal orientation was performed, resulting in significant evidence of non-uniform distribution of vocalisations across the year (P = 0.0303).
Figure 1 in Calling phenology of anurans in a tropical rainforest in South Mexico: testing predictive models
Figure 1. Study site location and land use, natural protected area of Nahá, Ocosingo, Chiapas, México.
Figure 2. a in Calling phenology of anurans in a tropical rainforest in South Mexico: testing predictive models
Figure 2. a) Calling anuran species recorded at temporary sections (7 months). The horizontal axis corresponds to temporary sections and the vertical to species. b) Average temperature T � (° C) (continuous line) and accumulated rainfall Ra (dashed line) by temporary section (32 in 7 months). Both graphs share the horizontal axis, where breaks indicate recording gaps, and the numbers are the days of the months covered.
Validation of the predictive accuracy of health-state utility values based on the Lloyd model for metastatic or recurrent breast cancer in Japan
<p>Although there is a lack of data on health-state utility values (HSUVs) for calculating quality-adjusted life years in Japan, Cost-utility analysis has been introduced by the Japanese government to inform decision-making in the medical field since 2016. This study aimed to determine whether the Lloyd model which was a predictive model of HSUVs for metastatic breast cancer (MBC) patients in the United Kingdom can accurately predict actual HSUVs for Japanese patients with MBC. The prospective observational study, followed by the validation study of the clinical predictive model.<b> </b>Forty-four Japanese patients with MBC were studied at 336 survey points. This study consisted of two phases. In the first phase, we constructed a database of clinical data prospectively and HSUVs for Japanese patients with MBC to evaluate the predictive accuracy of HSUVs calculated using the Lloyd model. In the second phase, Bland-Altman analysis was used to determine how accurately predicted HSUVs (based on the Lloyd model) correlated with actual HSUVs obtained using the EuroQol 5-Dimension 5-Level questionnaire, a preference-based measure of HSUVs in patients with MBC. In the Bland-Altman analysis, the mean difference between HSUVs estimated by the Lloyd model and actual HSUVs, or systematic error, was -0.106. The precision was 0.165. The 95% limits of agreement ranged from -0.436 to 0.225. The t value was 4.6972, which was greater than the t value with 2 degrees of freedom at the 5% significance level (p=0.425). There were acceptable degrees of fixed and proportional errors associated with the prediction of HSUVs based on the Lloyd model for Japanese patients with MBC. We recommend that sensitivity analysis be performed when conducting cost-effectiveness analyses with HSUVs calculated using the Lloyd model.</p>
Molecular modelling of novel ADCY3 variant predicts a molecular target for tackling obesity
<p>The present study aimed to identify <em>ADCY3</em> genetic variants in severely obese young patients of Greek-Cypriot origin by genomic sequencing. Total genomic DNA samples were isolated from peripheral whole blood using the Gentra Puregene Blood kit (Qiagen GmbH). DNA sequencing was performed with 100 ng genomic DNA, which was amplified using primers designed by Primer3 software ver. 0.4.0 (http://frodo.wi.mit.edu/). PCR mixtures were prepared using the Taq DNA Polymerase Kit (Qiagen GmbH); they had a final volume of 20 μl and contained 2 μl PCR buffer (10X), 2 μl Q Solution (5X), 2 μl dNTPs (2mM), 0.3 μl of each primer (10μM), 0.2 μl Taq polymerase (5U/μl) and 100 ng genomic DNA. Amplification was performed with an initial denaturing temperature at 95˚C for 5 min, followed by 30 cycles of denaturation (95˚C, 45 sec), annealing (57˚C, 60 sec) and extension (72˚C, 45 sec), with a final extension at 72˚C for 5 min. The <em>ADCY3</em> gene primers covered all exons. The PCR products were analysed on an Applied Biosystems 3130xl Genetic Analyzer and the results were analysed using Sequencing Analysis R 5.3 software (Applied Biosystems; Thermo Fisher Scientific, Inc.). The results revealed a total of five variants in patients, four of which were previously reported. A novel variant was identified in two patients (6%). The novel variant involves a heterozygous c.349T>A change in exon 1 of the gene locus, leading to a missense p.Leu117Met substitution.</p>
Predicted HHV values for woody biomass samples from USDA-AFRI project using the best performing models.
<p>Predicted HHV values for samples from the USDA-AFRI project using the best-performing models from the cross-validation process.</p>
Predicting the execution time of COSMO weather forecast models
<p>This data set is the work of E. Di Giacomo Master Thesis at the University of Bologna.</p> <p>Predicting the execution time of a numerical weather forecast model is a complex task. Generally, these models simulate the evolution of atmospheric weather and they are typically used for the production of weather forecasts, one or multiple times a day. Given their computational complexity, they require large computing capabilities, such as High Performance Computing systems. In these systems, job scheduling and resource allocation are carefully managed to optimize the usage of the finite and expensive hardware resources; in particular, several allocation and related pricing decisions are based on estimates of the duration of the application submitted, such as the execution time of weather forecast models.<br> A reliable prediction of execution time allows for a better management of the overall system, an improved planning of the model execution, as well as the identification of possible anomalies during the execution, thus providing great benefits to both system administrators and users.</p> <p>This data set regards a particular weather forecast model, namely the COSMO model, the weather forecasting model used at the the Hydro-Meteo-Climate Structure of Arpae Emilia-Romagna. The data set contains many execution times of the COSMO meteorological model run under a variety of different scientific parameters and parallelization levels.</p>
A new model of forelimb ecomorphology for predicting the ancient habitats of fossil turtles
<p>Various morphological proxies have been used to infer habitat preferences among fossil turtles and their early ancestors, but most are tightly linked to phylogeny, thereby minimizing their predictive power. One particularly widely used model incorporates linear measurements of the forelimb (humerus + ulna + manus) but, in addition to the issue of phylogenetic correlation, it does not estimate the likelihood of habitat assignment. Here, we introduce a new model that uses intramanual measurements (digit III metacarpal + non-ungual phalanges + ungual) to statistically estimate habitat likelihood, and that has greater predictive strength than prior estimators. Application of the model supports the hypothesis that stem-turtles were primarily terrestrial in nature, and recovers the nanhsiungchelyid <i>Basilemys</i> (a fossil crown-group turtle) as having lived primarily on land, despite some prior claims to the contrary.</p>
Salmon louse infestation levels on sea trout can be predicted from a hydrodynamic lice dispersal model
<p>The abundance of the parasitic salmon louse has increased with the growth in aquaculture of salmonids in open net pens. This represents a threat to wild salmonid populations as well as a key limiting factor for salmon farming. The Norwegian 'traffic light' management system for salmon farming aims to increase aquaculture production while securing sustainable wild salmonid populations. However, this system is at present solely focusing on mortality in wild Atlantic salmon, while responses of sea trout with different ecological characteristics are not included.</p> <p>We analyze lice counts on sea trout from surveillance data and use Bayesian statistical models to relate the observed lice infestations to the environmental lice infestation pressure, salinity, and current speed. These models can be used in risk assessment to predict when and where lice numbers surpass threshold levels for expected serious health effects in wild sea trout.</p> <p>We find that in production areas with the highest density of salmon farms (West coast), more than 50 % of the sea trout experienced lice infestations above levels of expected serious health effects.</p> <p>We also observed high lice infestations on sea trout in areas with salinities below louse tolerance levels, indicating that the fish had been infested elsewhere but were returning to low-saline waters to avoid lice or delouse. This behavioural response may over time disrupt anadromy in sea trout.</p> <p>The observed infestations on sea trout can be explained by the hydrodynamic lice dispersal model, which provides continuous estimates of lice exposure along the whole Norwegian coast. These estimates, which are used in Atlantic salmon research and management, can also be used for sea trout.</p> <p>Synthesis and policy implications: Wild sea trout, spending its entire feeding migration in fjords and coastal areas, is at higher risk than Atlantic salmon to lice infestations from aquaculture. The observed high levels of lice infestation on sea trout question the environmental sustainability of the current aquaculture industry in areas with intensive farming. We discuss the complex responses of sea trout to salmon lice and how the 'traffic light' management system may include data on this species.</p>
Code for: Comparison and interpretability of machine learning models to predict severity of chest injury
<p><span><span><span><span><span><span><span><span><span><span><span><b>Objective:</b> Trauma quality improvement programs and registries improve care and outcomes for injured patients. Designated trauma centers calculate injury scores using dedicated trauma registrars; however, many injuries arrive at non-trauma centers, leaving a substantial amount of data uncaptured. We propose automated methods to identify severe chest injury using machine learning (ML) and natural language processing (NLP) methods from the electronic health record (EHR) for quality reporting.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Materials and Methods:</b> A level I trauma center was queried for patients presenting after injury between 2014 and 2018. Prediction modeling was performed to classify severe chest injury using a reference dataset labeled by certified registrars. Clinical documents from trauma encounters were processed into concept unique identifiers for inputs to ML models: logistic regression with elastic net regularization (EN), extreme gradient boosted machines (XGB), and convolutional neural networks (CNN). The optimal model was identified by examining predictive and face validity metrics using global explanations.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Results:</b> Of 8,952 encounters, 542 (6.1%) had a severe chest injury. CNN and EN had the highest discrimination, with an area under the receiver operating characteristic curve of 0.93 and calibration slopes between 0.88 and 0.97. CNN had better performance across risk thresholds with fewer discordant cases. Examination of global explanations demonstrated the CNN model had better face validity, with top features including "contusion of lung" and "hemopneumothorax." </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Discussion: </b>The CNN model featured optimal discrimination, calibration, and clinically relevant features selected. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Conclusion:</b> NLP and ML methods to populate trauma registries for quality analyses are feasible.</span></span></span></span></span></span></span></span></span></span></span></p>
Modeling management strategies for chronic disease in wildlife: predictions for the control of respiratory disease in bighorn sheep
<p>1. Controlling persistent infectious disease in wildlife populations is an on-going challenge for wildlife managers and conservationists worldwide.</p> <p>2. Here, we develop a dynamic pathogen transmission model capturing key features of M. ovipneumoniae infection, a major cause of population declines in North American bighorn sheep (Ovis canadensis). We explore the effects of model assumptions and parameter values on disease dynamics, including density versus frequency dependent transmission, the inclusion of a carrier class versus a longer infectious period, host survival rates, disease-induced mortality and recovery rates, and the epidemic growth rate.</p> <p>3. We compare the effectiveness of a suite of management actions following an epidemic, including test-and-remove, depopulation-and-reintroduction, range expansion, herd augmentation, and density reduction.</p> <p>4. Our results suggest that test-and-remove, depopulation-and-reintroduction, and range expansion have the potential to facilitate recovery of persistently infected bighorn sheep herds post-epidemic. By contrast, augmentation could lead to worse outcomes than those expected in the absence of management. Management that improves host survival or reduces disease-induced mortality are also likely to improve population size and persistence of chronically infected herds.</p> <p>5. Dynamic transmission models like the one employed here offer a structured, logical approach towards exploring hypotheses and can serve as a basis for planning field experiments and adaptive management. Models should be used iteratively with the field empirical approaches to triangulate on better approaches to wildlife management.</p>
Machine learning models predict calculation outcomes with the transferability necessary for computational catalysis
<p>data files, including ML models of dynamic classifiers, trajectories of electronic structure and geometric features, optimized geometries, and final csv files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.