Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Improved Prediction of Smoking Status via Isoform-Aware RNA-seq Deep Learning Models
GEO Series GSE158699. Homo sapiens. 2653 samples. Type: Expression profiling by high throughput sequencing.
Integrative analysis of differentially expressed genes and miRNAs predicts complex T3-mediated protective circuits in a rat model of cardiac ischemia reperfusion
GEO Series GSE116046. Rattus norvegicus. 5 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Identification of Potential Models for Predicting Progestin Insensitivity in Patients with Endometrial Atypical Hyperplasia and Endometrial Cancer Based on Integration of ATAC-Seq and RNA-Seq Analysis
GEO Series GSE201927. Homo sapiens. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Mouse models of genetically heterogeneous multiple myeloma reveal mechanisms of immune evasion that predict clinical immunotherapy responses
GEO Series GSE205447. Homo sapiens; Mus musculus. 133 samples. Type: Expression profiling by high throughput sequencing.
GATA6 is predicted to regulate DNA methylation in an in vitro model of human hepatocyte differentiation (ChIPmentation)
GEO Series GSE163330. Homo sapiens. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
GATA6 is predicted to regulate DNA methylation in an in vitro model of human hepatocyte differentiation
GEO Series GSE163331. Homo sapiens. 68 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other; Methylation profiling by array.
Massively parallel profiling and predictive modeling of the outcomes of CRISPR/Cas9-mediated double-strand break repair
GEO Series GSE131421. Homo sapiens. 4 samples. Type: Other.
Establishment of interpretable cytotoxicity prediction models using machine learning analysis of transcriptome features
GEO Series GSE252529. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.
Vascularized immunocompetent patient-derived tumor model to predict treatment response for lung cancer patients
GEO Series GSE240662. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing.
A functional genomics predictive network model identifies regulators of inflammatory bowel disease: Ribo-zero RNAseq mouse distal colon DSS Colitis model of genetically engineered mice
GEO Series GSE83550. Mus musculus. 194 samples. Type: Expression profiling by high throughput sequencing.
Creation and validation of models to predict response to chemotherapy before treatment in serous ovarian cancer
GEO Series GSE156699. Homo sapiens. 88 samples. Type: Expression profiling by high throughput sequencing.
Climate model variability leads to uncertain predictions of the future abundance of stream macroinvertebrates
<p>All these files belong to the figures of the paper entitled "Climate model variability leads to uncertain predictions of the future abundance of stream macroinvertebrates" which is about to be published in the journal "<em>Scientific Reports</em>".</p>
A hybrid model for fast and probabilistic urban pluvial flood prediction - dataset
<p>A hybrid model for fast and probabilistic urban pluvial flood prediction - dataset </p> <p>This dataset presents the data used in the paper "A hybrid model for fast and probabilistic urban pluvial flood prediction". Data used to build, calibrate and evaluate the model and data that support the findings of the study are given for two cases, namely Gent and Antwerp. </p> <p>Under the folder of each case, one could find the following content:</p> <ul> <li>Model building: Data used to build the model topology. It includes geo-data of the 1D sewer network, boundary conditions of the catchment and spatial rainfall configurations </li> <li>Calibration: Rainfall data and flood records data used for calibration of the model.</li> <li>Results and evaluation: Rainfall events used for evaluation, flood probability predicted by the hybrid models and flood records used for comparing with the hybrid models. </li> </ul>
Raw Data of SLR for Parallelization, Modeling, and Performance Prediction in the Multi-/Many Core Area
<p>Contains the raw data and paper list of the SLR performed for the following publication:</p> <p>Parallelization, Modeling, and Performance Prediction in the Multi-/Many Core Area: A Systematic Literature Review</p>
MLP link prediction models in H5 format
<p>The KG-COVID-19 graph from this Zenodo URL was used:</p> <p>https://zenodo.org/record/4011267/files/kg-covid-19-skipgram-aug-2020.tar.gz</p> <p>To produce these embeddings (Skipgram, 80/20 training/test split, seed=42, 500 epochs max, delta 0.0001)</p> <p>https://zenodo.org/record/4019808/files/SkipGram_80_20_training_test_epoch_500_delta_0.0001_embedding.npy</p> <p>These embeddings were used to train link prediction classifiers using this Jupyter notebook:</p> <p>https://github.com/justaddcoffee/kg_covid_19_drug_analyses/blob/master/Link%20prediction.ipynb</p> <p>SkipGram_weightedL2_finalized_model.h5</p> <p>SkipGram_weightedL2.csv</p> <p>SkipGram_weightedL1_finalized_model.h5</p> <p>SkipGram_weightedL1.csv</p> <p>SkipGram_hadamard_finalized_model.h5</p> <p>SkipGram_hadamard.csv</p> <p>SkipGram_average_finalized_model.h5</p> <p>SkipGram_average.csv all_reports.csv</p>
Model data for "A hybrid dynamical-statistical model for advancing subseasonal tropical cyclone prediction over the western North Pacific"
<p>The dataset is the potential predictors of the statistical forecast model based on the 4 methods.</p> <p>The files in the document named "Train" are the potential predictors and TC anomalous counts for C1-C7 and TCall in the training period of 1979-2002 with 480-time points. For example, "./data/Train/M1/prepar/Obs_C1_pre-data.txt" contains 7 potential predictors of OLR, SSTA, specific humidity at 700 hPa, omega at 500 hPa, divergence and vorticity at 850hPa defined with method 1, and TC anomalous for TC of C1 prediction.</p> <p>The files named "Model_Lead*_C*_pre-data_*.txt" in "Frcst" the document are the potential predictors in the forecast period of 2003-2013 at lead times of 10, 15, 20, 25, 30, and 35 days with 220-time points. For example, "./data/Frcst/M1/prepar/Model_Lead10_C1_pre-data_00.txt" contains 7 potential predictors from the output of the FLOR model at lead 10 days initialized at Z00 time defined with method 1. Besides, "./data/Frcst/M1/prepar/Obs_C1_pre-data.txt" contains 7 potential predictors from the observation, which is the result of lead 0 days.</p>
Data from: Plant water potential improves prediction of empirical stomatal models
Climate change is expected to lead to increases in drought frequency and severity, with deleterious effects on many ecosystems. Stomatal responses to changing environmental conditions form the backbone of all ecosystem models, but are based on empirical relationships and are not well-tested during drought conditions. Here, we use a dataset of 34 woody plant species spanning global forest biomes to examine the effect of leaf water potential on stomatal conductance and test the predictive accuracy of three major stomatal models and a recently proposed model. We find that current leaf-level empirical models have consistent biases of over-prediction of stomatal conductance during dry conditions, particularly at low soil water potentials. Furthermore, the recently proposed stomatal conductance model yields increases in predictive capability compared to current models, and with particular improvement during drought conditions. Our results reveal that including stomatal sensitivity to declining water potential and consequent impairment of plant water transport will improve predictions during drought conditions and show that many biomes contain a diversity of plant stomatal strategies that range from risky to conservative stomatal regulation during water stress. Such improvements in stomatal simulation are greatly needed to help unravel and predict the response of ecosystems to future climate extremes.
Data from: Assimilating MODIS data-derived minimum input data set and water stress factors into CERES-Maize model improves regional corn yield predictions
Crop growth models and remote sensing are useful tools for predicting crop growth and yield, but each tool has inherent drawbacks when predicting crop growth and yield at a regional scale. To improve the accuracy and precision of regional corn yield predictions, a simple approach for assimilating Moderate Resolution Imaging Spectroradiometer (MODIS) products into a crop growth model was developed, and regional yield prediction performance was evaluated in a major corn-producing state, Illinois, USA. Corn growth and yield were simulated for each grid using the Crop Environment Resource Synthesis (CERES)-Maize model with minimum inputs comprising planting date, fertilizer amount, genetic coefficients, soil, and weather data. Planting date was estimated using a phenology model with a leaf area duration (LAD)-logistic function that describes the seasonal evolution of MODIS-derived leaf area index (LAI). Genetic coefficients of the corn cultivar were determined to be the genetic coefficients of the maturity group [included in Decision Support System for Agrotechnology Transfer (DSSAT) 4.6], which shows the minimum difference between the maximum LAI derived from the LAD-logistic function and that simulated by the CERES-Maize model. In addition, the daily water stress factors were estimated from the ratio between daily leaf area/weight growth rates estimated from the LAD-logistic function and that simulated by the CERES-Maize model under the rain-fed and auto-irrigation conditions. The additional assimilation of MODIS data-derived water stress factors and LAI under the auto-irrigation condition showed the highest prediction accuracy and precision for the yearly corn yield prediction (R2 is 0.78 and the root mean square error is 0.75 t ha-1). The present strategy for assimilating MODIS data into a crop growth model using minimum inputs was successful for predicting regional yields, and it should be examined for spatial portability to diverse agro-climatic and agro-technology regions.
Data from: Improving accuracies of genomic predictions for drought tolerance in maize by joint modeling of additive and dominance effects in multi-environment trials
Breeding for drought tolerance is a challenging task that requires costly, extensive and precise phenotyping. Genomic selection (GS) can be used to maximize selection efficiency and the genetic gains in maize (Zea mays L.) breeding programs for drought tolerance. Here we evaluated the accuracy of genomic selection of additive (A) against additive+dominance (AD) models to predict the performance of untested maize single-cross hybrids for drought tolerance in multi-environment trials. Phenotypic data of five drought-tolerance traits were measured in 308 hybrids in eight trials under water-stressed (WS) and well-watered (WW) conditions over two years and two locations in Brazil. Hybrids' genotypes were inferred based on their parents' genotypes (inbred lines) using single nucleotide polymorphism data obtained via genotyping-by-sequencing. GS analyses were performed using genomic best linear unbiased prediction by fitting a factor analytic (FA) multiplicative mixed model. Results showed differences in the predictive accuracy between A and AD models for the five traits under consideration in both water conditions. For grain yield (GY), the AD model doubled the predictive accuracy in comparison to the A model. FA framework allowed for investigating the stability of additive and dominance effects across environments, as well as the additive- and dominance-by-environment interactions, with interesting applications for parental and hybrid selection. Prediction performance of untested hybrids using GS that benefit from borrowing information from correlated trials increased 40% and 9% for A and AD models, respectively. These results highlighted the importance of multi-environment trial analysis with GS that incorporate dominance effects into genomic predictions of GY in maize single-cross hybrids.
Dynamic Prediction of HIV-Related Incomplete Immune Reconstitution: A Multi-Center, Large Cohort Study Using Advanced Joint Modeling
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.