Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
geo24/100

Improved Prediction of Smoking Status via Isoform-Aware RNA-seq Deep Learning Models

GEO Series GSE158699. Homo sapiens. 2653 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2020View details →
geo24/100

Integrative analysis of differentially expressed genes and miRNAs predicts complex T3-mediated protective circuits in a rat model of cardiac ischemia reperfusion

GEO Series GSE116046. Rattus norvegicus. 5 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenJul 2018View details →
geo24/100

Identification of Potential Models for Predicting Progestin Insensitivity in Patients with Endometrial Atypical Hyperplasia and Endometrial Cancer Based on Integration of ATAC-Seq and RNA-Seq Analysis

GEO Series GSE201927. Homo sapiens. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2022View details →
geo24/100

Mouse models of genetically heterogeneous multiple myeloma reveal mechanisms of immune evasion that predict clinical immunotherapy responses

GEO Series GSE205447. Homo sapiens; Mus musculus. 133 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2023View details →
geo24/100

GATA6 is predicted to regulate DNA methylation in an in vitro model of human hepatocyte differentiation (ChIPmentation)

GEO Series GSE163330. Homo sapiens. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo24/100

GATA6 is predicted to regulate DNA methylation in an in vitro model of human hepatocyte differentiation

GEO Series GSE163331. Homo sapiens. 68 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other; Methylation profiling by array.

openGEO-OpenMay 2022View details →
geo24/100

Massively parallel profiling and predictive modeling of the outcomes of CRISPR/Cas9-mediated double-strand break repair

GEO Series GSE131421. Homo sapiens. 4 samples. Type: Other.

openGEO-OpenJun 2019View details →
geo24/100

Establishment of interpretable cytotoxicity prediction models using machine learning analysis of transcriptome features

GEO Series GSE252529. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2025View details →
geo24/100

Vascularized immunocompetent patient-derived tumor model to predict treatment response for lung cancer patients

GEO Series GSE240662. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2023View details →
geo24/100

A functional genomics predictive network model identifies regulators of inflammatory bowel disease: Ribo-zero RNAseq mouse distal colon DSS Colitis model of genetically engineered mice

GEO Series GSE83550. Mus musculus. 194 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2017View details →
geo24/100

Creation and validation of models to predict response to chemotherapy before treatment in serous ovarian cancer

GEO Series GSE156699. Homo sapiens. 88 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
zenodo24/100

Climate model variability leads to uncertain predictions of the future abundance of stream macroinvertebrates

<p>All these files belong to the figures of the paper entitled &quot;Climate model variability leads to uncertain predictions of the future abundance of stream macroinvertebrates&quot; which is about to be published in the journal &quot;<em>Scientific Reports</em>&quot;.</p>

opencc-by-4.0Feb 2020View details →
zenodo24/100

A hybrid model for fast and probabilistic urban pluvial flood prediction - dataset

<p>A hybrid model for fast and probabilistic urban pluvial flood prediction - dataset&nbsp;</p> <p>This dataset presents the data used in the paper &quot;A hybrid model for fast and probabilistic urban pluvial flood prediction&quot;.&nbsp;Data used to build, calibrate and evaluate the model and data that support the findings of the study are given for two cases, namely Gent and Antwerp.&nbsp;</p> <p>Under the folder of each case, one could find the following content:</p> <ul> <li>Model building: Data used to build the model topology. It includes geo-data of the 1D sewer network, boundary conditions of the catchment and spatial rainfall configurations&nbsp;&nbsp;</li> <li>Calibration: Rainfall data and flood records data used for calibration of the model.</li> <li>Results and evaluation: Rainfall events used for evaluation, flood probability predicted by the hybrid models and flood records used for comparing with the hybrid models.&nbsp;</li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo24/100

Raw Data of SLR for Parallelization, Modeling, and Performance Prediction in the Multi-/Many Core Area

<p>Contains the raw data and paper list of the SLR performed for the following publication:</p> <p>Parallelization, Modeling, and Performance Prediction in the Multi-/Many Core Area: A Systematic Literature Review</p>

opencc-by-4.0Aug 2020View details →
zenodo24/100

MLP link prediction models in H5 format

<p>The KG-COVID-19&nbsp;graph from this Zenodo URL was used:</p> <p>https://zenodo.org/record/4011267/files/kg-covid-19-skipgram-aug-2020.tar.gz</p> <p>To produce these embeddings (Skipgram, 80/20 training/test split, seed=42, 500 epochs max, delta 0.0001)</p> <p>https://zenodo.org/record/4019808/files/SkipGram_80_20_training_test_epoch_500_delta_0.0001_embedding.npy</p> <p>These embeddings were used to train link prediction classifiers using this Jupyter notebook:</p> <p>https://github.com/justaddcoffee/kg_covid_19_drug_analyses/blob/master/Link%20prediction.ipynb</p> <p>SkipGram_weightedL2_finalized_model.h5</p> <p>SkipGram_weightedL2.csv</p> <p>SkipGram_weightedL1_finalized_model.h5</p> <p>SkipGram_weightedL1.csv</p> <p>SkipGram_hadamard_finalized_model.h5</p> <p>SkipGram_hadamard.csv</p> <p>SkipGram_average_finalized_model.h5</p> <p>SkipGram_average.csv all_reports.csv</p>

opencc-by-4.0Sep 2020View details →
zenodo24/100

Model data for "A hybrid dynamical-statistical model for advancing subseasonal tropical cyclone prediction over the western North Pacific"

<p>The dataset is the potential predictors of the statistical forecast model based on the&nbsp;4 methods.</p> <p>The files in the document&nbsp;named &quot;Train&quot; are the potential predictors and TC anomalous counts for C1-C7 and TCall in the training period of 1979-2002 with 480-time points. For example, &quot;./data/Train/M1/prepar/Obs_C1_pre-data.txt&quot; contains 7 potential predictors of OLR, SSTA, specific humidity at 700 hPa, omega at 500 hPa, divergence and vorticity at 850hPa defined with method 1, and&nbsp;TC anomalous for TC of C1 prediction.</p> <p>The files named &quot;Model_Lead*_C*_pre-data_*.txt&quot; in &quot;Frcst&quot; the document&nbsp;are the potential predictors in the forecast period of 2003-2013 at lead times of 10, 15, 20, 25, 30,&nbsp;and 35 days with 220-time points.&nbsp;For example, &quot;./data/Frcst/M1/prepar/Model_Lead10_C1_pre-data_00.txt&quot; contains 7 potential predictors from the output of the FLOR model at lead 10 days initialized at Z00 time defined with method 1. Besides, &quot;./data/Frcst/M1/prepar/Obs_C1_pre-data.txt&quot;&nbsp;contains 7 potential predictors from the observation, which is the result of lead 0 days.</p>

opencc-by-4.0Oct 2020View details →
dryad24/100

Data from: Plant water potential improves prediction of empirical stomatal models

Climate change is expected to lead to increases in drought frequency and severity, with deleterious effects on many ecosystems. Stomatal responses to changing environmental conditions form the backbone of all ecosystem models, but are based on empirical relationships and are not well-tested during drought conditions. Here, we use a dataset of 34 woody plant species spanning global forest biomes to examine the effect of leaf water potential on stomatal conductance and test the predictive accuracy of three major stomatal models and a recently proposed model. We find that current leaf-level empirical models have consistent biases of over-prediction of stomatal conductance during dry conditions, particularly at low soil water potentials. Furthermore, the recently proposed stomatal conductance model yields increases in predictive capability compared to current models, and with particular improvement during drought conditions. Our results reveal that including stomatal sensitivity to declining water potential and consequent impairment of plant water transport will improve predictions during drought conditions and show that many biomes contain a diversity of plant stomatal strategies that range from risky to conservative stomatal regulation during water stress. Such improvements in stomatal simulation are greatly needed to help unravel and predict the response of ecosystems to future climate extremes.

opencc-zeroDec 2016View details →
dryad24/100

Data from: Assimilating MODIS data-derived minimum input data set and water stress factors into CERES-Maize model improves regional corn yield predictions

Crop growth models and remote sensing are useful tools for predicting crop growth and yield, but each tool has inherent drawbacks when predicting crop growth and yield at a regional scale. To improve the accuracy and precision of regional corn yield predictions, a simple approach for assimilating Moderate Resolution Imaging Spectroradiometer (MODIS) products into a crop growth model was developed, and regional yield prediction performance was evaluated in a major corn-producing state, Illinois, USA. Corn growth and yield were simulated for each grid using the Crop Environment Resource Synthesis (CERES)-Maize model with minimum inputs comprising planting date, fertilizer amount, genetic coefficients, soil, and weather data. Planting date was estimated using a phenology model with a leaf area duration (LAD)-logistic function that describes the seasonal evolution of MODIS-derived leaf area index (LAI). Genetic coefficients of the corn cultivar were determined to be the genetic coefficients of the maturity group [included in Decision Support System for Agrotechnology Transfer (DSSAT) 4.6], which shows the minimum difference between the maximum LAI derived from the LAD-logistic function and that simulated by the CERES-Maize model. In addition, the daily water stress factors were estimated from the ratio between daily leaf area/weight growth rates estimated from the LAD-logistic function and that simulated by the CERES-Maize model under the rain-fed and auto-irrigation conditions. The additional assimilation of MODIS data-derived water stress factors and LAI under the auto-irrigation condition showed the highest prediction accuracy and precision for the yearly corn yield prediction (R2 is 0.78 and the root mean square error is 0.75 t ha-1). The present strategy for assimilating MODIS data into a crop growth model using minimum inputs was successful for predicting regional yields, and it should be examined for spatial portability to diverse agro-climatic and agro-technology regions.

opencc-zeroDec 2018View details →
dryad24/100

Data from: Improving accuracies of genomic predictions for drought tolerance in maize by joint modeling of additive and dominance effects in multi-environment trials

Breeding for drought tolerance is a challenging task that requires costly, extensive and precise phenotyping. Genomic selection (GS) can be used to maximize selection efficiency and the genetic gains in maize (Zea mays L.) breeding programs for drought tolerance. Here we evaluated the accuracy of genomic selection of additive (A) against additive+dominance (AD) models to predict the performance of untested maize single-cross hybrids for drought tolerance in multi-environment trials. Phenotypic data of five drought-tolerance traits were measured in 308 hybrids in eight trials under water-stressed (WS) and well-watered (WW) conditions over two years and two locations in Brazil. Hybrids' genotypes were inferred based on their parents' genotypes (inbred lines) using single nucleotide polymorphism data obtained via genotyping-by-sequencing. GS analyses were performed using genomic best linear unbiased prediction by fitting a factor analytic (FA) multiplicative mixed model. Results showed differences in the predictive accuracy between A and AD models for the five traits under consideration in both water conditions. For grain yield (GY), the AD model doubled the predictive accuracy in comparison to the A model. FA framework allowed for investigating the stability of additive and dominance effects across environments, as well as the additive- and dominance-by-environment interactions, with interesting applications for parental and hybrid selection. Prediction performance of untested hybrids using GS that benefit from borrowing information from correlated trials increased 40% and 9% for A and AD models, respectively. These results highlighted the importance of multi-environment trial analysis with GS that incorporate dominance effects into genomic predictions of GY in maize single-cross hybrids.

opencc-zeroDec 2016View details →
zenodo24/100

Dynamic Prediction of HIV-Related Incomplete Immune Reconstitution: A Multi-Center, Large Cohort Study Using Advanced Joint Modeling

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record