Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Modeling management strategies for chronic disease in wildlife: predictions for the control of respiratory disease in bighorn sheep
Open the record for dataset details and reuse information.
Data from: Predicting species occurrences with habitat network models
Open the record for dataset details and reuse information.
Data from: Ocean circulation model predicts high genetic structure in a long-lived pelagic developer
Open the record for dataset details and reuse information.
Experimental evidence of warming-induced disease emergence and its prediction by a trait-based mechanistic model
Open the record for dataset details and reuse information.
Data from: Integrated modeling predicts shifts in waterbird population dynamics under climate change
Open the record for dataset details and reuse information.
Data from: Genome-wide prediction models that incorporate de novo GWAS are a powerful new tool for tropical rice improvement
Open the record for dataset details and reuse information.
A new model of forelimb ecomorphology for predicting the ancient habitats of fossil turtles
Open the record for dataset details and reuse information.
Data for: The biomechanics of tooth strength: testing the utility of simple models for predicting fracture in geometrically complex teeth
Open the record for dataset details and reuse information.
Benchmarking parametric and machine learning models for genomic prediction of complex traits
Open the record for dataset details and reuse information.
Calibration of probability predictions from machine-learning and statistical models
Open the record for dataset details and reuse information.
Lineage-level distribution models lead to more realistic climate change predictions for a threatened crayfish
Open the record for dataset details and reuse information.
Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"
<p><strong>Models and Predictions</strong></p> <p>This dataset contains the trained XGBoost and EA-LSTM models and the models' predictions for the paper <a href="https://github.com/gauchm/ealstm_regional_modeling"><em>The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction</em></a>.</p> <p>For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there's only one folder, since it's all basins).<br> In each folder, there are three files:</p> <ul> <li><span class="math-tex">\(\texttt{model.pkl}\)</span> (XGBoost) or <em><span class="math-tex">\(\texttt{model_epoch30.pt}\)</span></em> (EA-LSTM), which stores the pickled trained model</li> <li><em><span class="math-tex">\(\texttt{xgboost_seedNNN.p}\)</span></em> or <em><span class="math-tex">\(\texttt{ealstm_seedNNN.p}\)</span></em>, which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow.</li> <li><span class="math-tex">\(\texttt{attributes.db}\)</span>, which stores static catchment attributes needed for inference.</li> </ul> <p>In addition to each folder, there is a SLURM submission script called <em><span class="math-tex">\(\texttt{<foldername>.sbatch}\)</span></em> that was used to create and evaluate the model in the folder.</p> <p> </p> <p>* sequence lengths 270 and 365 only contain data for EA-LSTM.</p>
Data and code for "Predicting evaporation in stream temperature models – Penman, Dalton or something else?"
<p>The uploaded files contain the data set and code used in an empirical evaluation of the application of the Penman equation for predicting evaporation from streams.</p> <ul> <li>Fishtrap_for_stream_evap_analysis.csv - data set used in the analysis</li> <li>streamEvapAnalysis_final.r - code used to analyse the data</li> </ul>
Sex prediction based on teeth mesiodistal width: model development in a Portuguese population
<p>168 pretreatment dental casts of orthodontics Portuguese subjects (59 males and 109 females) were included. Mesiodistal widths from right first molar to left first molar were measured on each pretreatment cast to the nearest 0.01 mm using digital caliper.</p>
Development and validation of a machine learning model for use as an automated artificial intelligence tool to predict mortality risk in patients with COVID-19
<p><strong>Background</strong></p> <p>New York City quickly became an epicenter of the COVID-19 pandemic. Due to a sudden and massive increase in patients during COVID-19 pandemic, healthcare providers incurred an exponential increase in workload which created a strain on the staff and limited resources. As this is a new infection, predictors of morbidity and mortality are not well characterized.</p> <p><strong>Methods</strong></p> <p>We developed a prediction model to predict patients at risk for mortality using only laboratory, vital and demographic information readily available in the electronic health record on more than 3000 hospital admissions with COVID-19. A variable importance algorithm was used for interpretability and understanding of performance and predictors.</p> <p><strong>Findings</strong></p> <p>We built a model with 84-97% accuracy to identify predictors and patients with high risk of mortality, and developed an automated artificial intelligence (AI) notification tool that does not require manual calculation by the busy clinician. Oximetry, respirations, blood urea nitrogen, lymphocyte percent, calcium, troponin and neutrophil percentage were important features and key ranges were identified that contributed to a 50% increase in patients’ mortality prediction score. With an increasing negative predictive value (NPV) starting 0.90 after the second day of admission, we are able more confidently able identify likely survivors. This study serves as a use case of a model with visualizations to aide clinicians with a better understanding of the model and predictors of mortality. Additionally, an example of the operationalization of the model via an AI notification tool is illustrated.</p>
Supplementary Structural Models (SARS-CoV-2 Spike-RBD:ACE2 complex and TMPRSS2) - SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals
<p>Structural Models (PDB) of SARS-CoV-2 Spike RBD bound to ACE2 receptors of 215 animals.</p> <p>Structural model of Human TMPRSS2.</p> <p>Modelled using the FunMod pipeline and referenced in the preprint</p> <p><a href="https://www.biorxiv.org/content/10.1101/2020.05.01.072371v5">SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals</a></p> <p> </p>
Dataset and codes for "BaHSYM: parsimonious Bayesian Hierarchical Model to predict river Sediment Yield"
<p>This folder contains:</p> <ul> <li>R project file</li> <li>R code for Best Fit model</li> <li>R code for temporal cross-validation</li> <li>R code for spatial cross-validation</li> <li>R code for cluster analysis</li> <li>dataset containing all input variables for the river gauges (and catchments) used for the development and testing of the BaHSYM model in Austria</li> </ul> <p>It also contains the same codes and datasets adapted to reproduce the model by de Vente et al. (2011), i.e. with the same structure but with the variables used in such model.</p>
Regression committee machine and petrophysical model jointly driven parameter reservoirs prediction from wireline logs for tight sandstone
<pre>This data comes from this study: "Regression committee machine and petrophysical model jointly driven parameter reservoirs prediction from wireline logs for tight sandstone". It is the intelligent prediction result of porosity, permeability and water saturation of two wells in the Ordos Basin, China</pre>
Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models
The location and modular structure of eukaryotic protein-coding genes in genomic sequences can be automatically predicted by gene annotation algorithms. These predictions are often used for comparative studies on gene structure, gene repertoires, and genome evolution. However, automatic annotation algorithms do not yet correctly identify all genes within a genome, and manual annotation is often necessary to obtain accurate gene models and gene sets. As manual annotation is time-consuming, only a fraction of the gene models in a genome is typically manually annotated, and this fraction often differs between species. To assess the impact of manual annotation efforts on genome-wide analyses of gene structural properties, we compared the structural properties of protein-coding genes in seven diverse insect species sequenced by the i5k initiative. Our results show that the subset of genes chosen for manual annotation by a research community (3.5-7% of gene models) may have structural properties (e.g., lengths and exon counts) that are not necessarily representative for a species' gene set as a whole. Nonetheless, the structural properties of automatically generated gene models are only altered marginally (if at all) through manual annotation. Major correlative trends, for example a negative correlation between genome size and exonic proportion, can be inferred from either the automatically predicted or manually annotated gene models alike. Vice versa, some previously reported trends did not appear in either the automatic or manually annotated gene sets, pointing towards insect-specific gene structural peculiarities. In our analysis of gene structural properties, automatically predicted gene models proved to be sufficiently reliable to recover the same gene-repertoire-wide correlative trends that we found when focusing on manually annotated gene models only. We acknowledge that analyses on the individual gene level clearly benefit from manual curation. However, as genome sequencing and annotation projects often differ in the extent of their manual annotation and curation efforts, our results indicate that comparative studies analyzing gene structural properties in these genomes can nonetheless be justifiable and informative.
Depressurization of CO2 in a pipe: High-resolution pressure and temperature data and comparison with model predictions – dataset
<p>This dataset contains data from depressurization of pure CO<sub>2</sub> and nitrogen in a tube from a gaseous and a dense-liquid state. The data are described in the accompanying paper (DOI: <a href="https://doi.org/10.1016/j.energy.2020.118560">10.1016/j.energy.2020.118560</a>).</p> <p>Test number; fluid; pressure (MPa); temperature (deg C):<br> 3; CO2; 4.04; 10.2<br> 4; CO2; 12.54; 21.1<br> 6; CO2; 10.40; 40.0<br> 8; CO2; 12.22; 24.6<br> 11; N2; 5.13; 10.0</p> <p><br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.