Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Protein Condensate Atlas from predictive models of heteromolecular condensate composition
Open the record for dataset details and reuse information.
Dataset related to article "ECoG spiking activity and signal dimension are early predictive measures of epileptogenesis in a translational mouse model of traumatic brain injury"
<p>Excel folder with raw data related to the paper 10.1016/j.nbd.2023.106251</p>
Random Forest Cloud Model for Predicting Liquid Cloud Microphysical Properties from A-Train Data
<p>Code for creating and analyzing the performance of a random forest model to predict cloud optical depth and cloud top effective radius from A-train satellite observations. Because of storage limitations, this directory does not include full satellite dataset, but the CloudSat data is available from the CloudSat Data Processing center (https://www.cloudsat.cira.colostate.edu/) and the CALIPSO data from NASA's Atmospheric Science Data Center (https://asdc.larc.nasa.gov/project/CALIPSO). </p>
The detailed results of model comparison between basline models and pepMTL in RT, CCS, MS/MS prediction
<p>This is a record file of the running results of the pepMTL model in the article "pepMTL: a synchronous multi-properties predictor for peptides enabled by multi-task framework and pre-trained protein language model", as well as the various benchmark models in RT, CCS, and MS/MS aspects. By carefully tiling the results of the run and saving the tiling results in the xlsx file.</p> <p>这是文章 "pepMTL: a synchronous multi-properties predictor for peptides enabled by multi-task framework and pre-trained protein language model"中pepMTL模型以及RT、CCS、MS/MS方面各个基准模型的运行结果的记录文件。通过对运行结果进行仔细整理,并将整理的结果保存在了xlsx文件中。</p> <p> </p>
Laboratory water quality monitoring data from full scale CS#3 DWDN for the DBP prediction model
<p>Lab measurements of water quality based on manual sampling including ordinary monitoring and special monitoring performed in SafeCREW project.</p> <p>Laboratory LIMS database.</p> <p>Provide water quality measurements based on the planned monitoring schedule. Schedule tailored to SafeCREW project needs plus usual legal water sampling.</p>
Datalakes Observational model predictions (2019)
<p>Hydrodynamic model predictions for Lake Geneva generated using the SPUX Bayesian inference and uncertainty quantification package (https://doi.org/10.5281/zenodo.5638312) and the MITgcm hydrodynamics package (https://doi.org/10.5281/zenodo.5634041). The datasets contain the mean temperature/velocity predictions, as well as their respective spreads. The source repository for this dataset with predictions extended to September 2021 can be found at https://renkulab.io/gitlab/datalakes/hydrodynamic-model.</p> <p>The extended dataset is visualized on the Datalakes website: https://www.datalakes-eawag.ch/datadetail/17</p>
Cluster investigation results from model-predicted mountain lion feeding sites and non-predicted sites in southwest Wyoming
<p class="MsoNormal">Global positioning system (GPS) receivers allow researchers to collect location data that provide information about fine-scale animal movements. For large carnivores, these data are routinely processed to identify clusters of GPS locations which are investigated to validate feeding sites, estimate prey species composition, and model the likelihood of predation events based on characteristics of GPS location data within clusters. Although developing predation models entails a high level of field effort, researchers are apprehensive in applying system-specific models to other systems both spatially and temporally. Our objectives were to apply and compare multiple predation models to predict and identify feeding events outside of the geographic areas where models were developed. Using our multi-model approach, we identified feeding sites and estimated kill rates and prey composition of mountain lions (<em>Puma concolor</em>) in southwest Wyoming, USA from 2017 to 2019. Our approach increased our field efficiency by reducing potential field site investigations by 63%. We provide results of estimated diet composition in our study area where mountain lion prey included relatively high proportions of pronghorn (12.3%) and smaller mammals, particularly coyotes (7.6%). We also provide a comparison of the predation models developed across unique ecoregions of North America, and how they performed when applied to our area. We believe a similar approach could be adopted for other large carnivore populations where multiple models have been developed to characterize feeding events.</p>
Figure 4 from: Tachkov K, Mitov K, Savova A (2019) Predicting the outcomes and costs for a cohort of 426 patients with Chronic Obstructive Pulmonary Disease (COPD) in Bulgaria through a Markov model. Pharmacia 66(2): 53-57. https://doi.org/10.3897/pharmacia.66.e35162
Figure 4 Tornado diagram for LYS
Figure 2 from: Tachkov K, Mitov K, Savova A (2019) Predicting the outcomes and costs for a cohort of 426 patients with Chronic Obstructive Pulmonary Disease (COPD) in Bulgaria through a Markov model. Pharmacia 66(2): 53-57. https://doi.org/10.3897/pharmacia.66.e35162
Figure 2 CEAC of all data points
Figure 1 from: Tachkov K, Mitov K, Savova A (2019) Predicting the outcomes and costs for a cohort of 426 patients with Chronic Obstructive Pulmonary Disease (COPD) in Bulgaria through a Markov model. Pharmacia 66(2): 53-57. https://doi.org/10.3897/pharmacia.66.e35162
Figure 1 ICER points and dispersion cloud of Monte-Carlo simulation
Figure 3 from: Tachkov K, Mitov K, Savova A (2019) Predicting the outcomes and costs for a cohort of 426 patients with Chronic Obstructive Pulmonary Disease (COPD) in Bulgaria through a Markov model. Pharmacia 66(2): 53-57. https://doi.org/10.3897/pharmacia.66.e35162
Figure 3 Tornado diagram for QALYs
Datasets used in the article "Impact of the Initial Stratospheric Polar Vortex State on East Asian Spring Rainfall Prediction in Seasonal Forecast Models"
<p>The seasonal forecast model, National (Beijing) Climate Center (BCC) of China Meteorological Administration (CMA) (<a href="http://s2s.cma.cn/centers?mo=babj_CMA_37">http://s2s.cma.cn/centers?mo=babj_CMA_37</a>) has performed large ensemble forecasts from 2001–2021. The ensemble means of the seasonal forecasts in the first three months are publicly available here. The forecasts begin from late February, and the forecasts in the following three months are analyzed in our submitted paper (22 years, 3 months for per year).</p>
JAMES technical report, scripts and data for "'Comparison of C3 photosynthetic responses to light and CO2 predicted by the leaf photosynthesis models of Farquhar et al. (1980) and Goudriaan et al. (1985)"
<p>Dear reader,</p> <p>In this repository you will find 7 MATLAB scripts and 3 Excel datasets. The script "A_curves_Figure2.m" calls the function scripts "FvCB_model_Figure2.m", "FvCB_model_Figure2_noTPU" and "G85_model_Figure2" to create Figure 2 of the JAMES publication. The script "G85_Rd_FigureS1" calls the function script "G85_model_FigureS1" to create Figure S1 of the Supporting Information. The script "r2_RMSE_Table2" calculates statistics displayed in Table 2 of the JAMES publication. The Excel worksheet "Table2.xlsx" contains the values in Table 2 of the JAMES publication for quick data copying. The Excel worksheets "FvCBparameters_fitted_withTPU.xlsx" and "FvCBparameters_fitted_noTPU.xlsx" contains fitted FvCB model parameter values that were obtained using the "fitaci" function from the plantecophys R package (Duursma, 2015). <br> <br> Kind regards,</p> <p>Kevin van Diepen </p>
Effective Molecular Dynamics from Neural-Network Based Structure Prediction Models
<p>Molecular dynamics simulation (61.5 us) data of 28 one- and two-domain proteins from Jussupow & Kaila: Effective Molecular Dynamics from Neural-Network Based Structure Prediction Models</p>
The dataset of predicting the potential habitat suitability of Saussurea species in China under future climate change using the optimized Maximum Entropy (MaxEnt) model
<p><strong>Description:</strong></p> <p>This dataset accompanies the study on the Saussurea species, renowned for its biodiversity and medicinal significance in high-elevation regions, which faces endangerment due to climate change and human activities. Despite its importance, conservation research on Saussurea has been limited. To address this gap, the study employed the optimized MaxEnt model to simulate Saussurea's habitat suitability and analyze key environmental factors influencing its distribution.</p> <p>The dataset includes:</p> <ol> <li><strong>Model and Parameter Optimization Code</strong>: The code used for optimizing the MaxEnt model parameters, ensuring reproducibility of the habitat suitability models.</li> <li><strong>Saussurea Distribution Points</strong>: Georeferenced points indicating the observed locations of Saussurea species.</li> <li><strong>Current Environmental Variables</strong>: Data on key environmental factors influencing Saussurea distribution, such as Elevation, Isothermality (Bio3), and Temperature Annual Range (Bio7).</li> <li><strong>Future Environmental Variables: </strong>Data on key environmental factors influencing Saussurea distribution under SSP126, SSP245, SSP370 and SSP585 in 2020-2100s.</li> </ol>
Datasets and Model Predictions for "Scalable Spatiotemporal Prediction with Bayesian Neural Fields"
<p>This repository contains the evaluation datasets and model predictions that appear in the research article <em>Scalable Spatiotemporal Prediction with Bayesian Neural Fields</em> (Saad et al., 2024). Refer to the README files in each zip file for additional information.</p>
Data set for training a ML model to predict duration of MPI application phases (HPC system) - with previous phase info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 10 different data sets corresponding to different HPC applications.</p> <p>These data sets contain information regarding the previous MPI call with same ID and type.</p>
Data set for training a ML model to predict duration of MPI application phases (HPC system) - without previous phases info
<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 11 different data sets corresponding to different HPC applications.</p> <p>These data sets do not contain information regarding previous MPI calls</p>
Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk (sequence model release)
<p>(This is the updated version that has been converted a standard pytorch model format)</p> <p>This is the deep learning sequence model used in </p> <p>Jian Zhou, Chandra L. Theesfeld, Kevin Yao, Kathleen M. Chen, Aaron K. Wong, and Olga G. Troyanskaya, Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk, Nature Genetics, 2018.</p> <p>Note the full software is available from https://github.com/FunctionLab/ExPecto and this release is created for the convenience of use and under the same non-commercial license. The model weights can be loaded with pytorch load_state_dict function (for an example please find <a href="https://github.com/FunctionLab/ExPecto/blob/master/chromatin.py">https://github.com/FunctionLab/ExPecto/blob/master/chromatin.py</a>). We also provide a web server for browsing mutations with strong predicted effects at https://hb.flatironinstitute.org/expecto/, which are currently limited to mutations within 1kb to TSS or are 1000 Genomes variants.</p> <p>Trivia: we code-named our models with whale names. This model has an unofficial codename DeepSEA "Beluga".</p>
Figure 7 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 7. Location of the protected area containing goitered gazelle in central Iran.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.