Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
zenodo32/100

Prediction of the ground state for indenofluorene-type systems with Clar's π-sextet model

<p>This dataset contains the computational data associated with "Prediction of the ground state for indenofluorene-type systems with Clar's &pi;-sextet model".&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Abundance Trend Indicator - Models, Prediction, Stacked Environmental Data and Training Set Similarity

<p># Readme</p> <p>These trained models can be used to predict the abundance trends of New Zealand's forest species and can be used together with the code in https://github.com/lnilya/abundance-trend-indicator</p> <p>Since the process of using the models requires coding expertise and some setting up, please make sure to reach out to ilya.shabanov@vuw.ac.nz for any questions. All files will require the code in the repository to be read and used.&nbsp;</p> <p>If you want to explore the results generated with these models, please visit https://ati-nz-predictions-7e6f3d514735.herokuapp.com/ for a user-friendly, interactive UI.</p> <p>## Contents</p> <p>_models: Contains the trained models (Artificial Neural Network (ANN), Random Forest (RF), SVMW (Support vector machine) and GLM (logistic regression)) at different degrees of noise filtering, different datasets and variable sets. The model files also contain test and training scores. To load the files please refer to the readme in the code repository: ttps://github.com/lnilya/abundance-trend-indicator</p> <p><br>_predictions/_environment: Contains the predictor variables for the study area (New Zealand, 1950-2019) that are needed by the models to make predictions.&nbsp;</p> <p>_predictions/_similarity: Contains the masks of areas that can be predicted by models and are similar to the training set.</p> <p>_predictions/_ati: Contain the predicted results for the abundance trend. These can be explored on https://ati-nz-predictions-7e6f3d514735.herokuapp.com/&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

The features of Tissues and Patches for "Predicting microsatellite instabilitiy from histology images with a three-level hierarchical graph fusion model"

<p>This repository contains features and corresponding coordinates of patches and tissues extracted from 430 and 326 histologic images from patients with colorectal and gastric cancers from the TCGA cohort (original whole section SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this library are from formalin-fixed paraffin-embedded (FFPE) diagnostic sections (&ldquo;DX&rdquo; on the GDC Data Portal). This blog explains this in detail: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing.</strong></p> <p>All SVS slices were pre-processed as follows.</p> <p>According to &ldquo;Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer&rdquo; these histology images were categorized into The histology images were classified as &ldquo;MSS&rdquo; (microsatellite stable) or &ldquo;MSIMUT&rdquo; (microsatellite unstable or highly mutated) according to &ldquo;Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer&rdquo;, which corresponds to the division of the training and test sets in the article.<br><br></p> <p>Patches were extracted at 40x objective magnification and 20x objective magnification, respectively, and the corresponding features were extracted by pre-training resnet48, respectively</p> <p>The features of Tissues are thumbnails obtained at 2.5x objective magnification and further extracted by MedSAM after extracting the masks of the tissues.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

An adaptive and interpretable modeling architecture assisted rapid and reliable consensus prediction for hazardous properties of chemicals

<p>*the computational results of interpretable cases are available in the supporting information for interpretable case.xlsx ;</p> <p>*the training dataset is utilized for model training while the validation dataset is utilized for evaluating.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Cell features for "Datasets for "Predicting microsatellite instabilitiy from histology images with a three-level hierarchical graph fusion model""

<p>This repository contains features and corresponding coordinates of cells extracted from 430 and 326 histologic images from patients with colorectal and gastric cancers from the TCGA cohort (original whole section SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this library are from formalin-fixed paraffin-embedded (FFPE) diagnostic sections (&ldquo;DX&rdquo; on the GDC Data Portal). This blog explains this in detail: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing.</strong></p> <p>All SVS slices were pre-processed as follows.</p> <p>According to &ldquo;Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer&rdquo; these histology images were categorized into &ldquo;MSS&rdquo; (microsatellite stable) or &ldquo;MSIMUT&rdquo; (microsatellite unstable or highly mutated) and corresponded to the article dividing the training and test sets.<br><br></p> <p>The features of all cells were extracted by Hovernet and Transnuseg at 40x objective magnification for extraction masking and further feature extraction</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

A Model Predicting The Relationships Between Seed Size, Yield, and Actual Yield in Dry Field Peas

<p>This is a model that predicts yield and actual yield in dry pea based on seed size. Seeding rate, seed cost, pod length, and expected grain yield can be varied in the model. The ideal seed size is displayed at the vertex of the curve in three different graphs.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo32/100

The replication page for paper: Revisiting Supervised and Unsupervised Models for Effort-Aware Just-in-Time Defect Prediction

<p>The replication page for the paper submitted to EMSE: Revisiting Supervised and Unsupervised Models for Effort-Aware Just-in-Time Defect Prediction.</p> <p>There are four java files in the package named &quot;model&quot;. Each file corresponds to one specific model (CBS and CBS+ are in the same file). Each file has a main function and you will get the experiment results when running the main function.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2018View details →
zenodo32/100

PATTERNS OF RICHNESS OF FRESHWATER MOLLUSCS FROM CHILE: PREDICTIONS OF ITS DISTRIBUTION BASED ON NULL MODELS

<p>Script and data for GLM analysis and co-occurrence analysis.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo32/100

Modelling and prediction of the thermophysical properties of aqueous mixtures of choline geranate and geranic acid (CAGE) using SAFT-g Mie

<p>Calculated and experimental data for all the figures in the publication&nbsp;</p>

opencc-byNov 2019View details →
zenodo32/100

MvGraphDTA: Multi-view-based graph deep model for drug-target affinity prediction by introducing the graphs and line graphs

<h1>MvGraphDTA</h1> <p>MvGraphDTA:通过引入图形和折线图,基于多视图的图形深度模型用于药物-靶点亲和力预测</p> <h2>要求</h2> <p><span><span>numpy</span></span><span>==1.23.5</span></p> <p><span><span>pandas</span></span><span>==1.5.2</span></p> <p><span><span>biopython</span></span><span>==1.79</span></p> <p><span><span>scipy</span></span><span>==1.9.3</span></p> <p><span><span>torch</span></span><span>==2.0.1</span></p> <p><span><span>torch_geometric</span></span><span>==2.3.1</span></p> <h2>示例用法</h2> <h3>1. 使用我们的预训练模型</h3> <p>在本节中,我们提供了 pdbbindv2016 的核心集数据和 Li 的数据(过滤后的 casf2013 和 casf2016),您可以直接执行以下命令来运行我们的预训练模型并在核心集上获取结果。</p> <pre># Run the following command.<br>python test_pretrain.py</pre> <h3>2. 在数据集上运行</h3> <p>在本节中,您必须提供药物的 .sdf 文件以及靶标的 .pdb 文件。</p> <div> <p># 您可以通过运行以下命令获取药物和靶点的图形和折线图。<br>Python data_process.py</p> <p># 当所有数据都准备好后,您可以通过运行以下命令来训练自己的模型。<br>Python training.py</p> </div>

opencc-by-4.0Aug 2024View details →
zenodo32/100

The source code for a new capillary and adsorption‒force model predicting hydraulic conductivity of soil during freeze‒thaw processes

<p>The source code is related to "A New Capillary and Adsorption‒Force Model Predicting Hydraulic Conductivity of Soil during Freeze‒thaw Processes" (Shufeng Qiao, Rui Ma, Yunquan Wang, Ziyong Sun, Helen Kristine French, Yanxin Wang)</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Pretrained Model of Channel Mixer Layer for Spatiotemporal Predictive Learning

<p>Model supplementary material for <em><strong>Space Weather</strong></em> journal paper with the title: Channel Mixer Layer: Multimodal Fusion Towards Machine Reasoning for Spatiotemporal Predictive Learning of Ionospheric Total Electron Content</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Predicting the pro-longevity or anti-longevity effect of model organism genes with enhanced Gaussian noise augmentation-based contrastive learning on protein-protein interaction networks

<p>The datasets used to evaluate Enhanced Gaussian noise augmentation-based contrastive learning (EGsCL) against predicting the pro-longevity or anti-longevity effect of model organism gene. This repo also includes the pretrained encoders that obtained the best predictive performance for each organism (see Table 2).</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

First global machine learning model to predict solar flare impact on Earth's ionosphere

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Trained Surface Layer Models and Metrics for "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications"

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Towards Dynamical Annual To Decadal Climate Prediction Using the IAP-CAS Model

<p>We execute a set of Annual to Decadal (A2D) hindcast experiments using the IAP-CAS model (FGOALS-f2), which integrates 129 months of each prediction initialized from 1981 to 2015 annually, and evaluate the results of this experiment. The IAP-CAS model output associated with this work is stored here.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Data of Spatio-Temporal deep learning model for regional EPB irregularities short-term Prediction

<p>Using the dense ground-based GNSS receiver network and ionosonde data from East and Southeast Asia during 2010-2021, a novel Spatio-Temporal deep learning model for regional EPB irregularities short-term Prediction (STEP) was developed. The model integrates the convolutional neural network (CNN) and long short-term memory (LSTM) network, together with attention mechanisms, to capture both spatial and temporal features of regional ionospheric irregularities.<br>This dataset includes both the model and the results generated by STEP. The parameters provided are: UT (hours), Latitude (&deg;), Longitude (&deg;), Date, Y_pred (TECU/min), and Y_true (TECU/min). The dimensions of Y_pred and Y_true are 10812 x 610, where 10812 represents the product of the number of date and the number of UT (minus 18), and 610 corresponds to the product of the number of Latitude and Longitude. The model with a .pth extension can be loaded using PyTorch.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Predicting the solubility of amino acids and peptides with the SAFT-γ Mie approach: Neutral and charged models

<p>Calculated data accompanying the IECR 2024 publication "Predicting the Solubility of Amino Acids and Peptides with the&nbsp;SAFT‑&gamma; Mie Approach: Neutral and Charged Models", by Ahmed Alyazidi, Shubhani Paliwal, Felipe A. Perdomo, Amy Mead, Mingxia Guo, Jerry Y. Y. Heng,&nbsp;Thomas Bernet, Andrew Haslam, Claire S. Adjiman, George Jackson, and Amparo Galindo.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
dryad32/100

Prediction model of in-hospital mortality in intensive care unit patients with heart failure: machine learning-based, retrospective analysis of the MIMIC-III database

<p><b>Objective:</b> The predictors of in-hospital mortality for intensive care units (ICU)-admitted HF patients remain poorly characterized.We aimed to develop and validate a prediction model for all-cause in-hospital mortality among ICU-admitted HF patients.</p> <p><b>Design: </b>A retrospective cohort study.</p> <p><b>Setting and Participants: </b>Data were extracted from the MIMIC-III database. Data on 1,177 heart failure patients were analysed.</p> <p><strong>Methods</strong>: Patients meeting the inclusion criteria were identified from the MIMIC-III database and randomly divided into derivation and validation groups. Independent risk factors for in-hospital mortality were screened using XGBoost and LASSO regression models in the derivation sample. Multivariable logistic regression analysis was used to build prediction models. Discrimination, calibration, and clinical usefulness of the predicting model were assessed using the C-index, calibration plot, and decision curve analysis. After pairwise comparison, the best performing model was chosen to build a nomogram according to the regression coefficients.</p> <p><b>Results:</b> Among the 1,177 admissions, in-hospital mortality was 13.52%. In both groups, the XGBoost, LASSO regression, and GWTG-HF risk score models showed acceptable discrimination. The XGBoost and LASSO regression models also showed good calibration. In pairwise comparison, the prediction effectiveness was higher with the XGBoost and LASSO regression models than with the GWTG-HF risk score model (P&lt;0.05). The XGBoost model was chosen as our final model for its more concise and wider net benefit threshold probability range and was presented as the nomogram.</p> <p><b>Conclusions</b><b>:</b> Our nomogram enabled good prediction of in-hospital mortality in ICU-admitted HF patients, which may help clinical decision-making for such patients.</p>

opencc-zeroJun 2021View details →
dryad32/100

Do the predicted suitability scores from species distribution models correlate with species performance on-ground?

<p>Species distribution models are a very popular statistical tool for inferring potential distribution range of species across space and time and are thought to be a good predictor for habitat suitability. Some studies have suggested that if these models are reliable, predicted habitat suitability (PHS) should relate to species traits visualization, growth potential, body size, abundance. We validated this hypothesis by estimating association between the PHS and species abundance for 17 avian species endemic to the Western Ghats - Sri Lanka biodiversity hotspot. Additionally, we compared the PHS of sites where species were detected in both seasons (wet and dry) against sites where they were detected in the dry season alone. As a proxy for abundance, we estimated single-season occupancy estimates (ψ) using detection/non-detection data from multiple visits to the survey sites. We report significant and positive PHS-ψ correlation, though the strength of this association varied across species and models. Half of the species showed higher suitability scores for the sites where they were detected year round. The results presented here suggest that the predictive models can be used as a proxy for habitat quality, in addition to inferring the potential distribution.</p>

opencc-zeroJul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record