Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
zenodo36/100

molxspec: Deep learning models for predicting MS2 spectra from molecular structures

<p>This repository contains a pre-processed dataset derived from the <a href="https://gnps.ucsd.edu">GNPS public repository</a> of natural product mass spectra as well as pretrained model weights for four different types of model architectures using pytorch (version 1.9.0). The contents are as follows:</p> <ul> <li>gnps_processed_data.tgz: Contains tab separated files of molecule/MS2 spectra pairs derived from GNPS after filtering for invalid structures, too large molecules (bigger than 2000 M/Z spectra), and structures that yielded valid 3D geometry optimization. The processing steps were done for positive ionization mode (pos_* files), though negative ionization data is also included (neg_* files)</li> <li>models.tgz: Contains pytorch format pretrained models for four different architecutres: MLP (a residual block multilayer perceptron trained on ECFP molecular fingerprints), BERT (the same MLP but trained on pretrained representations from the Zinc V1 pretrained ChemBERTa models on SMILES), GCN (a graph convolution architecture), and EGNN (an equivariant graph neural network). Models were trained on&nbsp;pos_processed_gnps_shuffled_with_3d_train.tsv found in the&nbsp;gnps_processed_data.tgz file described previously.</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Whole-cell modeling in yeast predicts compartment-specific proteome constraints that drive metabolic strategies

<p>The pcYeast7.6 model and files, required to reproduce the figures, provided in the publication &quot;Whole-cell modeling in yeast predicts compartment-specific proteome constraints that drive metabolic strategies&quot;, accepted in <em>Nat Commun</em>. The Zenodo upload was created by Pranas Grigaitis, p.grigaitis [at] vu.nl.</p> <p>Abstract</p> <p>When conditions change, unicellular organisms rewire their metabolism to sustain cell maintenance and cellular growth. Such rewiring may be understood as resource re-allocation under cellular constraints. Eukaryal cells contain metabolically active organelles such as mitochondria, competing for cytosolic space and resources, and the nature of the relevant cellular constraints remain to be determined for such cells. Here we present a comprehensive metabolic model of the yeast cell, based on its full metabolic reaction network extended with protein synthesis and degradation reactions. The model predicts metabolic fluxes and corresponding protein expression by constraining compartment-specific protein pools and maximising growth rate. Comparing model predictions with quantitative experimental data suggests that under glucose limitation, a mitochondrial constraint limits growth at the onset of ethanol formation - known as the Crabtree effect. Under sugar excess, however, a constraint on total cytosolic volume dictates overflow metabolism. Our comprehensive model thus identifies condition-dependent and compartment-specific constraints that can explain metabolic strategies and protein expression profiles from growth rate optimization, providing a framework to understand metabolic adaptation in eukaryal cells.</p>

opencc-by-4.0Nov 2021View details →
dryad36/100

Effects of enhanced productivity of resources shared by predators in a food-web module: Comparing results of a field experiment to predictions of mathematical models of intra-guild predation

<p>This dataset contains data from a field experiment described in the publication "Wise, D. H. &amp; Farfan, M.A. (2021) Effects of enhanced productivity of resources shared by predators in a food-web module:  Comparing results of a field experiment to predictions of mathematical models of intra-guild predation. Ecology and Evolution, 00: 1-11. <a href="https://doi.org/10.1002/ece3.8375">https://doi.org/10.1002/ece3.8375</a>".</p> <p>The field experiment compared the response to increased input of nutrients and energy (artificial detritus) to an empirical model of intra-guild predation (IGP) to the predictions of published, simple mathematical models of asymmetric IGP (a generalist IG Predator that feeds both on a specialist IG Prey and a Resource that it shares with the IG Prey). The empirical model was a food-web module created by pooling species abundances across many families in a community of soil micro-arthropods into three response variables: IG Predator (large predatory mites), IG Prey (small predatory mites) and a shared Resource (fungivorous mites and springtails). The pattern of change over time in densities of the three response variables (IG Predator, IG Prey and Resource) was compared to the predictions of mathematical models of IGP to determine if the feeding relationships in this community of soil micro-arthropods could be abstracted into a simple IGP module. Thus, we were testing the hypothesis that IGP is a dominant organizing principle in this community.</p> <p>Simple mathematical models predict that increased input of nutrients and energy to the shared Resource will increase the equilibrium density of Resource and IG Predator, but will decrease that of IG Prey. By the experiment's end, densities of fungivores (Resource) had increased ~1.5x (ratio of pooled fungivore densities in the High treatment to plots with no addition of detritus (None treatment); and IG Predator densities had increased ~4x. Contrary to the prediction of mathematical models, IG Prey had not decreased, but instead had increased ~1.5x. We discuss possible reasons for the failure of the empirical model to agree with IGP theory.</p>

opencc-zeroDec 2021View details →
zenodo36/100

The prediction data analyzed in the article: "An improved regional coupled modeling system for Arctic sea ice simulation and prediction: a case study for 2018"

<p>The outputs of seasonal predictions with the Coupled Arctic Prediction System version 1 analyzed in the article, &quot;An improved regional coupled modeling system for Arctic sea ice simulation and prediction: a case study for 2018&quot;,&nbsp;including:</p> <p>Sea ice concentration (SIC)</p> <p>Sea ice thickness (SIT)</p> <p>Sea surface temperature (SST)</p> <p>Ice mass budget diagnostics</p> <p>Accumulated downward shortwave radiation at the surface (ASWDN)</p> <p>Accumulated downward longwave radiation at the surface (ALWDN)</p> <p>Near surface air temperature (T2)&nbsp;</p> <p>Temperature and salinity profile of the upper ocean under sea ice &nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

DATASET RELATED TO ARTICLE "Brain Tumor Resection in Elderly Patients Potential Factors of Postoperative Worsening in a Predictive Outcome Model"

<p><span>clinical database including information on patients included in the study at title</span></p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

A Survival Analysis based Volatility and Sparsity Modeling Network for Student Dropout Prediction

<p>KDDCup15.rar and&nbsp;XuetangX.rar are the metadata&nbsp;used in the paper&nbsp;&quot;A Survival Analysis based Volatility and Sparsity ModelingNetwork for Student Dropout Prediction&quot;. Both of them have be drawn from the largest MOOC platform in China, XuetangX (see https://www.xuetangx.com/). If any&nbsp;interested parties&nbsp;want to fetch the original dataset, they may follow the URLs bellow:</p> <p>KDDCup 2015 dataset is available at <a href="https://www.biendata.xyz/competition/kddcup2015/data/">https://www.biendata.xyz/competition/kddcup2015/data/</a>.</p> <p>XuetangX dataset is available at <a href="http://moocdata.cn/data/user-activity">http://moocdata.cn/data/user-activity</a>.</p>

opencc-by-4.0Jan 2022View details →
dryad36/100

A predictive flight-altitude model for avoiding future conflicts between an emblematic raptor and wind energy development in the Swiss Alps

<p>Deployment of wind energy is proposed as a mechanism to reduce greenhouse gas emissions. Yet, wind energy and large birds, notably soaring raptors, both depend on suitable wind conditions. Conflicts in airspace use may thus arise between wind energy development and wildlife protection due to the risks of collisions of birds with the blades of wind turbines. Using locations of GPS-tagged bearded vultures, a rare scavenging raptor reintroduced into the Alps, we built a spatially-explicit model to predict potential areas of conflict with future wind turbines deployments in the Swiss Alps. We modelled the probability of bearded vultures flying within or below the rotor-swept zone of wind turbines as a function of wind and environmental conditions, including food supply (presence of wild ungulates). Flight activity at potential risk of collision was generally high, concentrating on south-exposed mountainsides, especially in areas where ibex carcasses have a high occurrence probability, with critical areas covering vast expanses throughout the Swiss Alps. Our model provides a spatially-explicit decision tool that will guide authorities and energy companies for planning the deployment of wind farms in a proactive manner to reduce risk to emblematic Alpine wildlife.</p>

opencc-zeroJan 2022View details →
zenodo36/100

PIGNet: A physics-informed deep learning model toward generalized drug-target interaction predictions

<p>Training, test datasets of the paper &quot;PIGNet: A physics-informed deep learning model toward generalized drug-target interaction predictions&quot;.</p>

opencc-by-4.0Feb 2022View details →
dryad36/100

Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges

<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable;  ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>

opencc-zeroFeb 2022View details →
zenodo36/100

MeDeMo - a dependency model for DNA methylation-aware transcription factor binding predictions

<p>The uploaded <em>fasta </em>files contain extended reference genomes for three cell lines HepG2, GM12878, K562 (ENCODE) and two primary liver hepatocyte samples from the german epigenomics consortium (DEEP). The extended reference genomes&nbsp;contain information on DNA methylation in a CpG context. They can be used as input for <em>MeDeMo</em>, a tool to infer transcription factor binding sites incorporating not only sequence specificity but also DNA methylation. <em>MeDeMo&nbsp;</em>is available online at:&nbsp;<a href="http://www.jstacs.de/index.php/MeDeMo">http://www.jstacs.de/index.php/MeDeMo</a>.</p> <p>We considered the files ENCFF279HCL and ENCFF835NTC for GM12878, ENCFF867JRG and&nbsp;ENCFF721JMB for K562 as well as ENCFF064GJQ and ENCFF369YQW for HepG2. From DEEP, we considered samples&nbsp;41_Hf01 and&nbsp;41_Hf03 which are available through the International Human Epigenomics Consortium (IHEC).</p> <p>In addition, we provide all models trained using the mentioned data sets as well models for and motifs from genome wide predictions.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Data for Floor heating pre-on/off parameters based on Model Predictive Control feature extrapolation Paper in CLIMA2022 conference proceedings

<p>This is a collection of time series results used to obtain all the results shown in the paper. The tags of the .csv or .mat files are self explanatory and easy to use.</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Data from: How well do embryo development rate models derived from laboratory data predict embryo development in sea turtle nests?

<p>Development rate of ectothermic animals varies with temperature. Here we use data derived from laboratory constant temperature incubation experiments to formulate development rate models that can be used to model embryonic development rate in sea turtle nests. We then use a novel method for detecting the time of hatching to measure the in situ incubation period of sea turtle clutches to test the accuracy of our models in predicting the incubation period from nest temperature traces. We found that all our models overestimated the incubation period. We hypothesize three possible explanations which are not mutually exclusive for the mismatch between our modeling and empirically measured in situ incubation period: (1) a difference in the way the incubation period is calculated in laboratory data and in our field nests, (2) inaccuracies in the assumptions made by our models at high incubation temperatures where there is no empirical laboratory data, and (3) a tendency for development rate in laboratory experiments to be progressively slower as temperature decreases compared with in situ incubation.</p>

opencc-zeroApr 2022View details →
dryad36/100

Publication release: How well do species distribution models predict occurrences in exotic ranges?

<div class="record-description"> <p>Species distribution models (SDMs) are widely used predictive tools to forecast potential biological invasions. However, the reliability of SDMs extrapolated to exotic ranges remains understudied, with most analyses restricted to few species and equivocal results. We examined the spatial transferability of SDMs for 647 non-indigenous species extrapolated across 1,867 invaded ranges, and identify what factors may help differentiate predictive success from failure. We performed a large-scale assessment of the transferability of SDMs using two modelling approaches: generalized additive models (GAMs) and MaxEnt. We fitted SDMs on the native ranges of species and extrapolated them to exotic ranges. We examined the influence of general factors and factors related to biological invasions on spatial transferability.</p> <p>Here, we provide the code and data for publication in Global Ecology and Biogeography as part of Nguyen and Leung 2022 "How well do species distribution models predict occurrences in exotic ranges?". Provided are the files and scripts necessary to fit and validate the SDMs using distirbutional data from their native and exotic ranges, respectively, formulated as generalized additive models (GAMs) or MaxEnt models. Additionally, provided is a script to validate the SDMs on their native fitting range using 10-fold cross-validation, and to fit the transferability model, as a linear mixed model (LMM), with a provided cleaned data.frame. The dataset provided includes a full species list with GBIF occurrence records, target-group background (TGB) records to use with model fitting and validation, as well as environmental data associated with the sightings.</p> </div>

opencc-zeroApr 2022View details →
dryad36/100

Data from: a physics-based digital twin for model predictive control of autonomous unmanned aerial vehicle landing

<p>This paper proposes a two-level, data-driven, digital twin concept for the autonomous landing of aircraft, under some assumptions. It features a digital twin instance for model predictive control; and an innovative, real-time, digital twin prototype for fluid-structure interaction and flight dynamics to inform it. The latter digital twin is based on the linearization about a pre-designed glideslope trajectory of a high-fidelity, viscous, nonlinear computational model for flight dynamics; and its projection onto a low-dimensional approximation subspace to achieve real-time performance, while maintaining accuracy. Its main purpose is to predict in real-time, during flight, the state of an aircraft and the aerodynamic forces and moments acting on it. Unlike static lookup tables or regression-based surrogate models based on steady-state wind tunnel data, the aforementioned real-time digital twin prototype allows the digital twin instance for model predictive control to be informed by a truly dynamic flight model, rather than a less accurate set of steady-state aerodynamic force and moment data points. The paper describes in detail the construction of the proposed two-level digital twin concept and its verification by numerical simulation. It also reports on its preliminary flight validation in autonomous mode for an off-the-shelf unmanned aerial vehicle instrumented at Stanford University.</p>

opencc-zeroMay 2022View details →
zenodo36/100

Predicting Software Refactoring - joblib models

<p>Zip file contains Machine Learning models trained as a partial result of our work described in our paper</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Towards Developing and Analysing The Metric-Based Software Defect Severity Prediction Model

<p>This is a metric based approach to solve software defect severity prediction problem. In addition to that, this work proposes a new evaluation scheme that comprised of five metrics to analyze the performances.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Prediction Result of Next Monochrome Ionogram frame using LSTM-CNN Model

<p>Video of next frame ionogram prediction result using simple LSTM-CNN model</p> <p>Paper published in The 9Th&nbsp;International Seminar on Aerospace Science and Technology ISAST 2022</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Input files for "Predicting the orientation of adsorbed proteins steered with electric fields using a simple electrostatic model"

<p>In files to run PyGBe (https://github.com/pygbe/pygbe), and reproduce the results from the J. Phys. Chem. B&nbsp;article&nbsp;&quot;Predicting the orientation of adsorbed proteins steered with electric fields using a simple electrostatic model&quot;.&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Predicting past and future SARS-CoV-2-related sick leave using discrete time Markov modelling

<p><strong>Background: </strong>Prediction of SARS-CoV-2-induced sick leave among healthcare workers (HCWs) is essential for being able to plan the healthcare response to the epidemic.</p> <p><strong>Methods: </strong>During first wave of the SARS-Cov-2 epidemic (April 23<sup>rd </sup>to June 24<sup>th</sup>, 2020), the HCWs in the greater Stockholm region in Sweden were invited to a study of past or present SARS-CoV-2 infection. We develop a discrete time Markov model using a cohort of 9449 healthcare workers (HCWs) who had complete data on SARS-CoV-2 RNA and antibodies as well as sick leave data for the calendar year 2020. The one-week and standardized longer term transition probabilities of sick leave and the ratios of the standardized probabilities for the baseline covariate distribution were compared with the referent period (an independent period when there were no SARS-CoV-2 infections) in relation to PCR results, serology results and gender.</p> <p><strong>Results:</strong> The one-week probabilities of transitioning from healthy to partial sick leave or full sick leave during the outbreak as compared to after the outbreak were highest for healthy HCWs testing positive for large amounts of virus (ratio: 3.69, (95% confidence interval, CI: 2.44-5.59) and 6.67 (95% CI: 1.58-28.13), respectively). The proportion of all sick leaves attributed to COVID-19 during outbreak was at most 55% (95% CI: 50%-59%).</p> <p><strong>Conclusions: </strong>A robust Markov model enabled use of simple SARS-CoV-2 testing data for quantifying past and future COVID-related sick leave among HCWs, which can serve as a basis for planning of healthcare during outbreaks.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Prediction of the broadcasting model and various baselines

<p>The ground truth values, and the predictions from the broadcasting model and the baselines (SC and CMAQ) in the submission&nbsp;<em>Development of an LSTM-Broadcasting deep-learning framework for regional air pollution forecast improvement</em>. The namelists of the WRF and CMAQ models are also included.</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record