Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
dryad32/100

Predictive modeling for clinical features associated with Neurofibromatosis Type 1

<p>Objective: Perform a longitudinal analysis of clinical features associated with Neurofibromatosis Type 1 (NF1) based on demographic and clinical characteristics, and to apply a machine learning strategy to determine feasibility of developing exploratory predictive models of optic pathway glioma (OPG) and attention-deficit/hyperactivity disorder (ADHD) in a pediatric NF1 cohort.</p> <p>Methods: Using NF1 as a model system, we perform retrospective data analyses utilizing a manually-curated NF1 clinical registry and electronic health record (EHR) information, and develop machine-learning models. Data for 798 individuals were available, with 578 comprising the pediatric cohort used for analysis.</p> <p>Results: Males and females were evenly represented in the cohort. White children were more likely to develop OPG (OR: 2.11, 95%CI: 1.11-4.00, p=0.02) relative to their non-white peers. Median age at diagnosis of OPG was 6.5 years (1.7-17.0), irrespective of sex. Males were more likely than females to have a diagnosis of ADHD (OR: 1.90, 95%CI: 1.33-2.70, p&lt;0.001), and earlier diagnosis in males relative to females was observed. The gradient boosting classification model predicted diagnosis of ADHD with an AUROC of 0.74, and predicted diagnosis of OPG with an AUROC of 0.82.</p> <p>Conclusions: Using readily available clinical and EHR data, we successfully recapitulated several important and clinically-relevant patterns in NF1 semiology specifically based on demographic and clinical characteristics. Naïve machine learning techniques can be potentially used to develop and validate predictive phenotype complexes applicable to risk stratification and disease management in NF1.</p>

opencc-zeroMar 2022View details →
dryad32/100

Worldclim 2.1 versus Worldclim 1.4: climatic niche and grid resolution affect between-version mismatches in habitat suitability models predictions across Europe

<p>The influence of climate on the distribution of taxa has been extensively investigated in the last two decades through Habitat Suitability Models (HSMs). In this context, the Worldclim database represents an invaluable data source as it provides worldwide climate surfaces for both historical and future time horizons. Thousands of HSMs-based papers have been published taking advantage of Worldclim 1.4, the first online version of this repository. In 2017, Worldclim 2.1 was released. Here, we evaluated spatially explicit prediction mismatch at continental scale, focusing on Europe, between HSMs fitted using climate surfaces from the two Worldclim versions (between-version differences). To this aim, we simulated occurrence probability and presence-absence across Europe of four virtual species (VS) with differing climate-occurrence relationships. For each VS, we fitted HSMs upon uncorrelated bioclimatic variables derived from each Worldclim version at three grid resolutions. For each factor combination, HSMs attaining sufficient discrimination performance on spatially independent test data were projected across Europe under current conditions and various future scenarios, and importance scores of the single variables were computed. HSMs failed in accurately retrieving the simulated climate-occurrence relationships for the climate-tolerant VS and the one occurring under a narrow combination of climatic conditions. Under current climate, noticeable between-version prediction mismatch emerged across most of Europe for these two VSs, whose simulated suitability mainly depended upon diurnal or yearly variability in temperature; differently, between-version differences were more clustered toward areas showing extreme values, like mountainous massifs or southern regions, for VSs responding to average temperature and precipitation trends. Under future climate, the chosen emission scenarios and Global Climate Models did not evidently influence between-version prediction discrepancies, while grid resolution synergistically interacted with VSs' niche characteristics in determining extent of such differences. Our findings could help in re-evaluating previous biodiversity-related works relying on geographical predictions from Worldclim-based HSMs.</p>

opencc-zeroDec 2022View details →
zenodo32/100

A calcium-based plasticity model for predicting long-term potentiation and depression in the neocortex

<p>This dataset contains all (&gt;1000) cell pairs as the Blue Brain Projects <a href="https://github.com/BlueBrain/EModelRunner">EModelRunner</a> packages as well as analysis code and analysed data used for the figures of our <a href="https://www.biorxiv.org/content/biorxiv/early/2020/04/20/2020.04.19.043117.full.pdf">preprint</a>: <em><strong>&quot;A calcium-based plasticity model predicts long-term potentiation and depression in the neocortex&quot;</strong></em>.</p> <p>More documentation will follow in the upcoming days.</p> <p><strong>Updates:</strong><br> v1.1 (31/12/2021): updated READMEs within cell_packages, added 2 extra authors for their contribution in EModelRunner.<br> v1.2 (31/12/2021): same as v1.1 but w/o MacOS junk<br> v1.3 (11/01/2022): added analysis notebooks<br> v2.0 (11/01/2022): same as v1.3 but w/o MacOS junk and proper version number<br> v2.1 (14/03/2022): fetching data from websites when possible instead of providing the downloaded csv files.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Plots for the publication "Lidar-assisted model predictive control of wind turbine fatigue via online rainflow-counting considering stress history"

<p>These are the raw plot files from the publication &quot;Lidar-assisted model predictive control of wind turbine fatigue via online rainflow-counting considering stress history&quot;.</p> <p>The files have been created with MATLAB 2019, and labeled according to their corresponding figure number(s) in the publication.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Model weights from "The Effects of Nonlinear Signal on Expression-Based Prediction Performance"

<p>This file stores a representative set of saved models from the manuscript &quot;The Effects of Nonlinear Signal on Expression-Based Prediction Performance&quot;.&nbsp;</p> <p>To simplify the uploading process, the file was split into chunks using the `split` utility in Linux. They can be joined back together with the command `cat model_weights* &gt; model_weights.tar.gz`</p> <p>These saved files correspond to the weights and optimizer state of models trained on various biological tasks. They can be &quot;rehydrated&quot; using the `load_model` function of the associated model class from this repo: https://github.com/greenelab/linear_signal</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Data Set of Publication on Accurate Performance Predictions with Component-based Models of Data Streaming Applications

<p>This is the data set for the article &quot;Accurate Performance Predictions with Component-based Models of Data Streaming Applications&quot; by Dominik Werle, Stephan Seifermann and Anne Koziolek which appears in the proceedings of the 16th European Conference on Software Architecture (ECSA).</p> <p>The data set contains measurements of the evaluation system, models of the system, simulation results, derived analysis results and code for running the simulation.</p> <p>This work was supported by KASTEL Security Research Labs and by the German Research Foundation (DFG) under project number 432576552, HE8596/1-1 (FluidTrust).</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Data for figures in Model predictions of wave overwash extent into the marginal ice zone

<p>The data required to reproduce the figures in &#39;Model predictions of wave overwash extent into the marginal ice zone&#39;.</p> <p>The data includes:</p> <ul> <li>Coefficient values produced by coupled floe-wave motions, required to determine overwash of a single floe and thus the overwash extent model.</li> <li>The data for the transects used to predict overwash in Figure 11 (Agulhas II), includes wave conditions and the ice concentration along the transect (other floe field properties remains constant).</li> </ul> <p>The code is uploaded at -&nbsp;</p> <pre>https://doi.org/10.5281/zenodo.7059554</pre>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Climate data for Machine Learning based 100-year flood flow prediction model

<p>This study evaluates the application of ML technique over northeast United States regions and compares its performance to the U.S. Geological Survey (USGS) Streamflow Statistics (StreamStats)</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Uplink-based Live Session Model for Stalling Prediction in Video Streaming

<p>Dataset to model streaming session only by uplink requests with the goal of quality impairment estimation like quality changes or stallings.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Dataset for Uplink-based Live Session Model for Stalling Prediction in Video Streaming

<p>This dataset presents aggregated YouTube streaming data used for uplink based quality impairment estimation.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models

<p>Training and test datasets and scripts for training models.</p>

openmit-licenseSep 2022View details →
zenodo32/100

Data for Streamflow Prediction: Comparison of SWAT vs. Random Forest Models in Diverse Catchments

<p>This study introduces a time-lag-informed Random Forest (RF) framework for streamflow time series prediction across diverse catchments, and compares its results against SWAT predictions. We found strong evidence of RF's better performance by adding historical flows and time-lags for meteorological values over using only actual meteorological values. On a daily scale, RF demonstrated robust performance (Nash&ndash;Sutcliffe efficiency [<em>NSE</em>] &gt; 0.5), whereas SWAT generally yielded unsatisfactory results (<em>NSE</em> &lt; 0.5) and tended to overestimate daily streamflow by up to 27% (<em>PBIAS</em>). However, SWAT provided better monthly predictions, particularly in catchments with irregular flow patterns. Although both models faced challenges in predicting peak flows in snow-influenced catchments, RF outperformed SWAT in an arid catchment. RF also exhibited a notable advantage over SWAT in terms of computational efficiency. Overall, RF is a good choice for daily predictions with limited data, whereas SWAT is preferable for monthly predictions and understanding hydrological processes in depth.</p> <p>This repository contains the input data used for building the RF and SWAT models and the files describing the modeling results.</p> <p>The corresponding Zenodo code repository is available at <a href="../doi/10.5281/zenodo.11064973" target="_blank" rel="noopener">https://zenodo.org/doi/10.5281/zenodo.11064973</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Universal Rapid Weather Prediction Model (Sonagi Model) Korea Peninsula 5 Days Forecast Result (Pressure Isobar Map)

<p>Universal Rapid Weather Prediction Model (Sonagi Model) Korea Peninsula 5 Days Forecast Result (Pressure Isobar Map)</p> <p>Each file has its altitude in front of the name of the file, and by each isobar in map directs the air current heading higher or lower altitude. And rest of the name follows the target ed UTC time. Generally, iso-temperature lines are used to indicate where air flows, but it was simulatable in Sonagi model to where air is heading by pressure, so that pressure isobar is used to indicate where the air flows.</p> <p>Input data for the prediction in Universal Rapid Weather Prediction Model (Sonagi Model) is originated from GK2A satelite of Korea Meteorological Administration. (https://apihub.kma.go.kr/) For sharing the original prediction data, contact me at somehowme@gmail.com or flyingtext@nate.com (Prediction netCDF4 files are almost 16GB in sum total.)</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Universal Rapid Weather Prediction Model (Sonagi Model) Pacific Ocean 5 Days Forecast Result (Pressure Isobar Map)

<p>Universal Rapid Weather Prediction Model (Sonagi Model) Pacific Ocean 5 Days Forecast Result (Pressure Isobar Map)</p> <p>Each file has its altitude in front of the name of the file, and by each isobar in map directs the air current heading higher or lower altitude. And rest of the name follows the target ed UTC time. Generally, iso-temperature lines are used to indicate where air flows, but it was simulatable in Sonagi model to where air is heading by pressure, so that pressure isobar is used to indicate where the air flows.</p> <p>Input data for the prediction in Universal Rapid Weather Prediction Model (Sonagi Model) is originated from GK2A satelite of Korea Meteorological Administration. (https://apihub.kma.go.kr/) For sharing the original prediction data, contact me at somehowme@gmail.com or flyingtext@nate.com (Prediction netCDF4 files are almost 16GB in sum total.)</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Predictive Modeling of Bearing Degradation: LSTM Neural Networks for Uncertainty Quantification

<p>These MATLAB codes are part of a research project focused on predicting bearing degradation through vibration measurements. The codes implement LSTM (Long Short-Term Memory) neural network models trained under different objectives, including uncertainty quantification and RMSE (Root Mean Square Error) minimization. The objective of the research is to compare the performance of these models in predicting bearing health and assessing the associated uncertainty.</p> <p><strong>Note:</strong> The current codes are under embargo access as the corresponding paper has been submitted to the ESCA 11 conference. The codes will be made openly accessible upon acceptance of the paper and during the presentation dates. Please cite our paper when using these codes.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

ML models for dengue prediction using Orange

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

PhysicsGen - Can Generative Models Learn from Images to Predict Complex Physical Relations?

<p>This dataset comprises 300,000 pairs of images designed for the advancement of generative model applications in physical simulations. Each pair consists of an input image and its corresponding output image that represents a physical simulation. The dataset aims to facilitate research into whether generative models can effectively learn and reproduce complex physical dynamics from visual data, potentially replacing traditional differential equation-based methods with significant computational speedups.</p> <p>Data, baseline models and evaluation code: <a href="https://www.physics-gen.org">https://www.physics-gen.org</a></p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems

<p>F-DATA is a novel workload dataset containing the data of around 24 million jobs executed on <a href="https://www.r-ccs.riken.jp/en/fugaku/">Supercomputer Fugaku</a>, over the three years of public system usage (March 2021-April 2024). Each job data contains an extensive set of features, such as exit code, duration, power consumption and performance metrics (e.g. #flops, memory bandwidth, operational intensity and memory/compute bound label), which allows for a multitude of job characteristics prediction. The full list of features can be found in the file&nbsp;<code>feature_list.csv</code>.</p> <p>The sensitive data appears both in anonymized and encoded versions. The encoding is based on a Natural Language Processing model and retains sensitive but useful job information for prediction purposes, without violating data privacy. The scripts used to generate the dataset are available in the<a href="https://github.com/francescoantici/F-DATA"> F-DATA GitHub repository</a>, along with a series of plots and instruction on how to load the data.</p> <p>F-DATA is composed of 38 files, with each&nbsp;<code>YY_MM.parquet</code> file containing the data of the jobs submitted in the month MM of the year YY.&nbsp;</p> <div> <div>The files of F-DATA are saved as <code>.parquet</code> files. It is possible to load such files as dataframes by leveraging the <code>pandas</code> APIs, after installing <code>pyarrow</code> (<code>pip install pyarrow</code>). A single file can be read with the following <code>Python</code> instrcutions:</div> <br> <blockquote> <div><code># Importing pandas library</code></div> <div><code>import pandas as pd</code></div> <div>&nbsp;</div> <div><code># Read the 21_01.parquet file in a dataframe format</code></div> <div><code>df = pd.read_parquet("21_01.parquet")</code></div> <div><code>df.head()</code></div> </blockquote> <div>&nbsp;</div> <div>Please cite this work as:<br><br> <div> <div>@article{antici2025fdata,</div> <div>title={F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems},</div> <div>author={Antici, Francesco and Bartolini, Andrea and Domke, Jens and Kiziltan, Zeynep and Yamamoto, Keiji},</div> <div>journal = {Scientific Data},</div> <div>volume={12},</div> <div>pages={1321},</div> <div>year={2025},</div> <div>publisher={Nature Publishing Group},</div> <div>doi={https://doi.org/10.1038/s41597-025-05633-1}</div> <div>}</div> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo32/100

A Riskscore Model for Predicting Survival, Tumor Microenvironment, Immunotherapy and drug sensitivity of Lung Squamous Cell Carcinoma Based on PI3K/AKT/MTOR Pathway-Related Genes

<p>Firstly, the data we provide is the raw data downloaded from the TCGA database. Secondly, we provide the following explanations for the raw data, taking Figure 1 as an example:</p> <p>First, we selected a LUSC dataset from the TCGA database that includes RNA-seq data (FPKM values) and clear clinical information, consisting of 51 normal samples and 501 LUSC samples. The raw data <span>can be found in</span>&nbsp;the "mRNA" document in the "Fig.1" file<span>, and t</span>he data includes TCGA <span>id</span>&nbsp;information.</p> <p>Next, the data was imported into "mRNA_edgeR" to analyze whether the RNA-seq data from these 552 samples meet the criteria of |logFC| &gt; 0.585 and FDR &lt; 0.05, <span>and the results are shown in</span>&nbsp;the "diffSig" table. Subsequently, the genes in the "diffSig" table were intersected with 105 PAGs,&nbsp;<span>and</span>&nbsp;44 key PAGs <span>were identified, </span>as shown in Figure 1.</p> <p>The <span>other</span>&nbsp;figures can <span>also </span>be reproduced sequentially based on the methods described <span>above</span>.</p> <p>Furthermore, due to the limited number of LUSC patients, the clinical information in the database is relatively incomplete. Hence, the clinical information that we provide constitutes the entire content available in the database.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Identification of Potential Biomarkers for Cerebral Palsy and the Development of Prediction Models

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record