Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
dryad32/100

Modeling management strategies for chronic disease in wildlife: predictions for the control of respiratory disease in bighorn sheep

Open the record for dataset details and reuse information.

publicFeb 2022View details →
dryad32/100

Data from: Predicting species occurrences with habitat network models

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad32/100

Data from: Ocean circulation model predicts high genetic structure in a long-lived pelagic developer

Open the record for dataset details and reuse information.

publicSep 2014View details →
dryad32/100

Experimental evidence of warming-induced disease emergence and its prediction by a trait-based mechanistic model

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad32/100

Data from: Integrated modeling predicts shifts in waterbird population dynamics under climate change

Open the record for dataset details and reuse information.

publicMay 2019View details →
dryad32/100

Data from: Genome-wide prediction models that incorporate de novo GWAS are a powerful new tool for tropical rice improvement

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad32/100

A new model of forelimb ecomorphology for predicting the ancient habitats of fossil turtles

Open the record for dataset details and reuse information.

publicJan 2022View details →
dryad32/100

Data for: The biomechanics of tooth strength: testing the utility of simple models for predicting fracture in geometrically complex teeth

Open the record for dataset details and reuse information.

publicJul 2023View details →
dryad32/100

Benchmarking parametric and machine learning models for genomic prediction of complex traits

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

Calibration of probability predictions from machine-learning and statistical models

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad32/100

Lineage-level distribution models lead to more realistic climate change predictions for a threatened crayfish

Open the record for dataset details and reuse information.

publicSep 2021View details →
zenodo28/100

Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"

<p><strong>Models and Predictions</strong></p> <p>This dataset contains the trained XGBoost and EA-LSTM models and the models&#39; predictions for the paper <a href="https://github.com/gauchm/ealstm_regional_modeling"><em>The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction</em></a>.</p> <p>For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there&#39;s only one folder, since it&#39;s all basins).<br> In each folder, there are three files:</p> <ul> <li><span class="math-tex">\(\texttt{model.pkl}\)</span> (XGBoost) or <em><span class="math-tex">\(\texttt{model_epoch30.pt}\)</span></em> (EA-LSTM), which stores the pickled trained model</li> <li><em><span class="math-tex">\(\texttt{xgboost_seedNNN.p}\)</span></em> or <em><span class="math-tex">\(\texttt{ealstm_seedNNN.p}\)</span></em>, which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow.</li> <li><span class="math-tex">\(\texttt{attributes.db}\)</span>, which stores static catchment attributes needed for inference.</li> </ul> <p>In addition to each folder, there is a SLURM submission script called <em><span class="math-tex">\(\texttt{&lt;foldername&gt;.sbatch}\)</span></em> that was used to create and evaluate the model in the folder.</p> <p>&nbsp;</p> <p>* sequence lengths 270 and 365 only contain data for EA-LSTM.</p>

opencc-by-4.0Nov 2019View details →
zenodo28/100

Data and code for "Predicting evaporation in stream temperature models – Penman, Dalton or something else?"

<p>The uploaded files contain the data set and code used in an empirical evaluation of the application of the Penman equation for predicting evaporation from streams.</p> <ul> <li>Fishtrap_for_stream_evap_analysis.csv - data set used in the analysis</li> <li>streamEvapAnalysis_final.r - code used to analyse the data</li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo28/100

Sex prediction based on teeth mesiodistal width: model development in a Portuguese population

<p>168 pretreatment dental casts of orthodontics Portuguese subjects (59 males and 109 females) were included. Mesiodistal widths from right first molar to left first molar were measured on each pretreatment cast to the nearest 0.01 mm using digital caliper.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Development and validation of a machine learning model for use as an automated artificial intelligence tool to predict mortality risk in patients with COVID-19

<p><strong>Background</strong></p> <p>New York City quickly became an epicenter of the COVID-19 pandemic. Due to a sudden and massive increase in patients during COVID-19 pandemic, healthcare providers incurred an exponential increase in workload which created a strain on the staff and limited resources. As this is a new infection, predictors of morbidity and mortality are not well characterized.</p> <p><strong>Methods</strong></p> <p>We developed a prediction model to predict patients at risk for mortality using only laboratory, vital and demographic information readily available in the electronic health record on more than 3000 hospital admissions with COVID-19. A variable importance algorithm was used for interpretability and understanding of performance and predictors.</p> <p><strong>Findings</strong></p> <p>We built a model with 84-97% accuracy to identify predictors and patients with high risk of mortality, and developed an automated artificial intelligence (AI) notification tool that does not require manual calculation by the busy clinician. Oximetry, respirations, blood urea nitrogen, lymphocyte percent, calcium, troponin and neutrophil percentage were important features and key ranges were identified that contributed to a 50% increase in patients&rsquo; mortality prediction score. With an increasing negative predictive value (NPV) starting 0.90 after the second day of admission, we are able more confidently able identify likely survivors. This study serves as a use case of a model with visualizations to aide clinicians with a better understanding of the model and predictors of mortality. Additionally, an example of the operationalization of the model via an AI notification tool is illustrated.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Supplementary Structural Models (SARS-CoV-2 Spike-RBD:ACE2 complex and TMPRSS2) - SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals

<p>Structural Models (PDB) of SARS-CoV-2 Spike RBD bound to ACE2 receptors of 215 animals.</p> <p>Structural model of Human TMPRSS2.</p> <p>Modelled using the FunMod pipeline and referenced in the preprint</p> <p><a href="https://www.biorxiv.org/content/10.1101/2020.05.01.072371v5">SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals</a></p> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Dataset and codes for "BaHSYM: parsimonious Bayesian Hierarchical Model to predict river Sediment Yield"

<p>This folder contains:</p> <ul> <li>R project file</li> <li>R code for Best Fit model</li> <li>R code for temporal cross-validation</li> <li>R code for spatial cross-validation</li> <li>R code for cluster analysis</li> <li>dataset containing all input variables for the river gauges (and catchments) used for the development and testing of the BaHSYM model in Austria</li> </ul> <p>It also contains the same codes and datasets adapted to reproduce the model by de Vente et al. (2011), i.e. with the same structure but with the variables used in such model.</p>

opencc-by-4.0Mar 2020View details →
zenodo28/100

Regression committee machine and petrophysical model jointly driven parameter reservoirs prediction from wireline logs for tight sandstone

<pre>This data comes from this study: &quot;Regression committee machine and petrophysical model jointly driven parameter reservoirs prediction from wireline logs for tight sandstone&quot;. It is the intelligent prediction result of porosity, permeability and water saturation of two wells in the Ordos Basin, China</pre>

opencc-by-4.0Aug 2020View details →
dryad28/100

Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models

The location and modular structure of eukaryotic protein-coding genes in genomic sequences can be automatically predicted by gene annotation algorithms. These predictions are often used for comparative studies on gene structure, gene repertoires, and genome evolution. However, automatic annotation algorithms do not yet correctly identify all genes within a genome, and manual annotation is often necessary to obtain accurate gene models and gene sets. As manual annotation is time-consuming, only a fraction of the gene models in a genome is typically manually annotated, and this fraction often differs between species. To assess the impact of manual annotation efforts on genome-wide analyses of gene structural properties, we compared the structural properties of protein-coding genes in seven diverse insect species sequenced by the i5k initiative. Our results show that the subset of genes chosen for manual annotation by a research community (3.5-7% of gene models) may have structural properties (e.g., lengths and exon counts) that are not necessarily representative for a species' gene set as a whole. Nonetheless, the structural properties of automatically generated gene models are only altered marginally (if at all) through manual annotation. Major correlative trends, for example a negative correlation between genome size and exonic proportion, can be inferred from either the automatically predicted or manually annotated gene models alike. Vice versa, some previously reported trends did not appear in either the automatic or manually annotated gene sets, pointing towards insect-specific gene structural peculiarities. In our analysis of gene structural properties, automatically predicted gene models proved to be sufficiently reliable to recover the same gene-repertoire-wide correlative trends that we found when focusing on manually annotated gene models only. We acknowledge that analyses on the individual gene level clearly benefit from manual curation. However, as genome sequencing and annotation projects often differ in the extent of their manual annotation and curation efforts, our results indicate that comparative studies analyzing gene structural properties in these genomes can nonetheless be justifiable and informative.

opencc-zeroAug 2020View details →
zenodo28/100

Depressurization of CO2 in a pipe: High-resolution pressure and temperature data and comparison with model predictions – dataset

<p>This dataset contains data from depressurization of pure CO<sub>2</sub> and nitrogen in a tube from a gaseous and a dense-liquid state. The data are described in the accompanying paper (DOI: <a href="https://doi.org/10.1016/j.energy.2020.118560">10.1016/j.energy.2020.118560</a>).</p> <p>Test number; fluid; pressure (MPa); temperature (deg C):<br> 3; CO2; 4.04; 10.2<br> 4; CO2; 12.54; 21.1<br> 6; CO2; 10.40; 40.0<br> 8; CO2; 12.22; 24.6<br> 11; N2; 5.13; 10.0</p> <p><br> &nbsp;</p>

opencc-by-4.0Aug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record