Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,773 results for “Predictive model”

Learn how ShareScore rates datasets ↗
zenodo36/100

Dataset of paper "Predicting the size of silver nanoparticles synthesised in flow reactors: Coupling population balance models with fluid dynamic simulations"

<p>Dataset of paper "Predicting the size of silver nanoparticles synthesised in flow reactors: Coupling population balance models with fluid dynamic simulations"</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Prediction of mechanistic subtypes of Parkinson's using patient-derived stem cell models

<p>Data and pipelines used to&nbsp;predict mechanistic subtypes of Parkinson's disease using patient-derived stem cell model.</p> <p><strong>Lists of files included;</strong></p> <ul> <li>chemPredPD2022_Imaging process pipelines: pipelines to extract tabular data in&nbsp;Columbus Image Data Storage and Analysis System</li> <li>demo_data_images: a set of data for the Demo</li> <li>ImageData</li> <li>TabularData</li> <li>New test data_PINK1_isoCTRL</li> <li>New test data_SNCA_isoCTRL</li> <li>Tabular_demo_data</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Prediction model of the temporal dynamics of severe pest cashew Anacampsis phytomiella using artificial neural networks

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo36/100

Assessing Heavy Metal Contamination in Agricultural Soils: A Predictive Model Integrating GIS Tools and Probability-Risk Matrix – Case Study: Guarda Region, Portugal

<p>In these files we can find the final risk map of heavy metal contamination for the guarding area in Portugal obtained according to the methodology explained in the paper "Assessing Heavy Metal Contamination in Agricultural Soils: A Predictive Model Instegrating GIS Tools and Probability-Risk Matrix - Case Study: Guarda Region (Portugal)</p> <p>Final Risk Equal.tiff:&nbsp; GeoTiff with a pixel size of 30m. EPSG:3763 - ETRS89 / Portugal TM06</p> <p>Also attached is the symbolisation for the image in .qml (Quantum GIS Layer Style File) format.</p> <p>A file called RISK RECLASS is also available, where you can find the risk classification maps for each of the studied factors:&nbsp;</p> <ul> <li>Proximity to roads</li> <li>Proximity to industrial areas</li> <li>Ph</li> <li>Soil organic content</li> <li>Slope</li> <li>Soil texture</li> <li>Mining extraction areas&nbsp;</li> <li>Drainage</li> </ul> <p>finally a DATABASE file where the data of the 360 points for the calculation of the risk maps can be found.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

An ulvophycean marine green alga produces large parthenogenetic isogametes as predicted by the gamete dynamics model for the evolution of anisogamy

<p>In eukaryotes, the gamete size difference between the two sexes (anisogamy) evolved from gametes of equal size in both mating types (isogamy) and is plausibly claimed to generate sexual selection in morphology and behaviour. The gamete dynamics (GD) model for anisogamy evolution combines gamete limitation and competition and predicts that, if gametes of both mating types can develop parthenogenetically (i.e. without fusing with the opposite mating type), large isogamy can evolve under gamete-limited conditions. Ulvophycean marine green algae that exhibit various gametic systems from isogamy to anisogamy are important models for testing such theories. However, in most previous papers, whether a species is isogamous or anisogamous has not been examined statistically, which leaves the above theoretical prediction untested. We reveal (i) that the gametic system of <em>Struvea okamurae</em> is large isogamy using a generalized linear mixed model (GLMM), which accounted for the variation of gamete size among individual gametophytes, and (ii) that gametes of this alga can actually develop parthenogenetically, contrary to a previous report. Habitat environments and gametic behaviour suggest that this alga might experience gamete-limited conditions. <em>S. okamurae</em> seems to produce large parthenogenetic isogametes following GD model predictions, as an adaptation to deep waters.</p>

opencc-zeroMar 2024View details →
zenodo36/100

The energy bands predicted by the universal HamGNN model

<p>universal_Hamiltonian.ckpt is the network weights for the universal HamGNN model. flat_systems.zip is the file contains the structures and energy bands for crystals with flat bands in GeNOME dataset. Energy_bands_prediction.zip is the file contains the structures and predicted energy bands for crystals in the test dataset of Materials Project. Energy_bands_DFT.zip is the file contains the structures and DFT calculated energy bands for crystals in the test dataset of Materials Project.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

NOAA PSL thermodynamic profiles retrieved from a combination of active and passive remote sensors and numerical weather prediction models with the optimal estimation physical retrieval TROPoe at Platteville, CO, USA

<p>This dataset contains retrieved profiles of thermodynamic variables obtained using the Tropospheric Remotely Observed Profiling via Optimal Estimation (TROPoe) physical retrieval from various combinations of input data collected by passive and active remote sensing instruments, in-situ surface platforms, and numerical weather prediction models deployed at the Platteville, CO, USA, site in fall 20221-winter 2022. Among the employed instruments are Microwave Radiometers (MWRs), Infrared Spectrometers (IRS), Radio Acoustic Sounding Systems (RASS), ceilometers, surface sensors, and information from the operational Rapid Refresh numerical weather prediction model.</p> <p>The dataset also includes 15 radiosounding launched for assessing the retrievals.</p> <p>For further information, please see:</p> <p>Bianco, L., Adler, B., Bariteau, L., Djalalova, I. V., Myers, T., Pezoa, S., Turner, D. D., and Wilczak, J. M.: Sensitivity of thermodynamic profiles retrieved from ground-based microwave and infrared observations to additional input data from active remote sensing instruments and numerical weather prediction models, Atmos. Meas. Tech. Discuss. [preprint], https://doi.org/10.5194/amt-2023-263, in review, 2024.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Predicting zeta potential of liposomes from their structure: A nano-QSAR model for DOPE, DC-Chol, DOTAP, and EPC formulations

<p>Data set for publication:&nbsp; "<em>Predicting zeta potential of liposomes from their structure: A nano-QSPR model for DOPE, DC-Chol, DOTAP, and EPC formulations.</em>"; Computational and Structural Biotechnology Journal 25 (2024) 3&ndash;8; https://doi.org/10.1016/j.csbj.2024.01.012&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Dataset, splits, models, and scripts for the QM descriptors prediction

<p>Dataset, splits, models, and scripts from the manuscript "When Do Quantum Mechanical Descriptors Help Graph Neural Networks Predict Chemical Properties?" are provided. The curated dataset includes 37 QM descriptors for 64,921 unique molecules across six levels of theory: wB97XD, B3LYP, M06-2X, PBE0, TPSS, and BP86. This dataset is stored in the data.tar.gz file, which also contains a file for multitask constraints applied to various atomic and bond properties. The data splits (training, validation, and test splits) for both random and scaffold-based divisions are saved as separate index files in splits.tar.gz. The trained D-MPNN models for predicting QM descriptors are saved in the models.tar.gz file. The scripts.tar.gz file contains ready-to-use scripts for training machine learning models to predict QM descriptors, as well as scripts for predicting QM descriptors using our trained models on unseen molecules and for applying radial basis function (RBF) expansion to QM atom and bond features.</p> <p>Below are descriptions of the available scripts:</p> <ol> <li><code>atom_bond_descriptors.sh</code>: Trains atom/bond targets.</li> <li><code>atom_bond_descriptors_predict.sh</code>: Predicts atom/bond targets from pre-trained model.</li> <li><code>dipole_quadrupole_moments.sh</code>: Trains dipole and quadrupole moments.</li> <li><code>dipole_quadrupole_moments_predict.sh</code>: Predicts dipole and quadrupole moments from pre-trained model.</li> <li><code>energy_gaps_IP_EA.sh</code>: Trains energy gaps, ionization potential (IP), and electron affinity (EA).</li> <li><code>energy_gaps_IP_EA_predict.sh</code>: Predicts energy gaps, IP, and EA from pre-trained model.</li> <li><code>get_constraints.py</code>: Generates constraints file for testing dataset. This generated file needs to be provided before using our trained models to predict the atom/bond QM descriptors of your testing data.</li> <li><code>csv2pkl.py</code>: Converts QM atom and bond features to .pkl files using RBF expansion for use with Chemprop software.</li> </ol> <p>Below is the procedure for running the ml-QM-GNN on your own dataset:</p> <ol> <li>Use <code>get_constraints.py</code> to generate a constraint file required for predicting atom/bond QM descriptors with the trained ML models.</li> <li>Execute <code>atom_bond_descriptors_predict.sh</code> to predict atom and bond properties. Run <code>dipole_quadrupole_moments_predict.sh</code> and <code>energy_gaps_IP_EA_predict.sh</code> to calculate molecular QM descriptors.</li> <li>Utilize <code>csv2pkl.py</code> to convert the data from predicted atom/bond descriptors .csv file into separate atom and bond feature files (which are saved as .pkl files here).</li> <li>Run <a href="https://github.com/chemprop/chemprop">Chemprop</a> to train your models using the additional predicted features supported here.</li> </ol>

opencc-by-4.0Feb 2024View details →
zenodo36/100

On the prediction of the time-varying behaviour of dynamic systems by interpolating state-space models

<p>In this article, a local Linear Parameter Varying (LPV) model identification approach is exploited to analyze the dynamic behaviour of a structure whose dynamics varies over time. This structure is composed by two aluminum crosses connected by a rubber mount. To observe time-dependent variations on the dynamics of this assembly, it is placed in a climate chamber and submitted to a six minute temperature run-up. During this run-up the structure is continuously excited by a shaker. The load provided by this device is measured by a load cell, while six accelerometers are measuring the responses of the system. The temperatures of the air inside the climate chamber and at the surface of the mount are also continuously measured. It is found that during the performed temperature run-up, the rubber mount temperature increased from, roughly, 14℃ to, approximately, 35.2℃. By using the measured load provided by the shaker and the measured accelerations, Frequency Response Functions (FRFs) at five different rubber mount temperatures are computed. From each of these sets of FRFs, state-space models are estimated. Afterwards, these models are used to define an interpolating LPV model, which enables the computation of interpolated state-space models representative of the dynamics of the system at each time sample. It is found that by feeding the interpolated state-space models with the measured load, an accurate simulation of the measured accelerations is obtained. Moreover, by exploiting a joint input state estimation algorithm with the interpolated state-space models and with the measured accelerations, a very good prediction of the applied load can be obtained. It is also shown that if the time dependency of the dynamics of the system is ignored, the results are less accurate.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

The Role of Genomic Data in Stratifying Patients within Predictive Models for Breast Cancer Survival Outcome

<p>Data associated with my PhD thesis titled "The Role of Genomic Data in Stratifying Patients within Predictive Models for Breast Cancer Survival Outcome".</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Lightning Prediction in the Tehran Region Using the WRF Model with Multiple Physical Parameterizations and an Ensemble Approach

<p><span>The Grid Analysis and Display System (</span>GrADS)<span>&nbsp;</span><span>and</span><span> </span><span>Python</span><span>&nbsp;</span><span>scripts and the output data from simulations that we used in this study.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Aerial Images_Part 2_Integrating Remote Sensing and Machine Learning for Developing Spatio-Temporal Model to Predict Aquatic Larval Habitats of Malaria

<p>Aerial Images_Part 2_Integrating Remote Sensing and Machine Learning for Developing Spatio-Temporal Model to Predict Aquatic Larval Habitats of Malaria</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Aerial Images_Part 1_Integrating Remote Sensing and Machine Learning for Developing Spatio-Temporal Model to Predict Aquatic Larval Habitats of Malaria

<p>Aerial Images_Part 1_Integrating Remote Sensing and Machine Learning for Developing Spatio-Temporal Model to Predict Aquatic Larval Habitats of Malaria</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Prediction of COVID-19 case numbers using state-space modeling and wastewater virus datasets

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo36/100

Processed Sentinel 1, Sentinel 2 and Copernicus Emergency Management Service data for fine tuning and predicting flood extent with IBM's granite-geospatial-uki-flood-detection model

<p>This dataset contains processed Sentinel 1 Sentinel 2 imagery together with flood event labels extracted from the Copernicus Emergency Management Service. It has been assembled to demonstrate fine tuning and inference of flood event segmentation using granite geospatial foundation models developed by IBM Research. Please see <a href="https://huggingface.co/ibm-granite/granite-geospatial-uki-flooddetection">https://huggingface.co/ibm-granite/granite-geospatial-uki-flooddetection</a> for more information on models and use.</p> <p>Sentinel-1</p> <p>The European Space Agency. 2014. Sentinel-1 Mission. <a href="https://sentinel.esa.int/web/sentinel/copernicus/sentinel-1">https://sentinel.esa.int/web/sentinel/missions/sentinel1</a>. Accessed: 2024-11-25.</p> <p>Sentinel-2</p> <p>The European Space Agency. 2015. Sentinel-2 Mission. <a href="https://sentinel.esa.int/web/sentinel/copernicus/sentinel-2">https://sentinel.esa.int/web/sentinel/missions/sentinel2</a>. Accessed: 2024-11-25.</p> <p>Copernicus Emergency Management Service</p> <p><a href="https://emergency.copernicus.eu/mapping/list-of-activations-rapid">https://emergency.copernicus.eu/mapping/list-of-activations-rapid</a>. Accessed: 2024-11-25.&nbsp;</p> <p><strong>Attribution</strong></p> <p>Contains modified Copernicus Sentinel data [2019-2024]</p> <p>Contains modified Copernicus Service information [2019-2023]</p>

openNov 2024View details →
zenodo36/100

Predicting transcriptional responses to novel chemical perturbations using deep generative model for drug discovery

<p>Understanding transcriptional responses to chemical perturbations is central to drug discovery, but exhaustive experimental screening of diseasecompound combinations is unfeasible. To overcome this limitation, here we introduce PRnet, a perturbation-conditioned deep generative model that predicts transcriptional responses to novel chemical perturbations that have never experimentally perturbed at bulk and single-cell levels. Evaluations indicate that PRnet outperforms alternative methods in predicting responses across novel compounds, pathways, and cell lines. PRnet enables gene-level response interpretation and in-silico drug screening for diseases based on gene signatures. PRnet further identifies and experimentally validates novel compound candidates against small cell lung cancer and colorectal cancer. Lastly, PRnet generates a large-scale integration atlas of perturbation profiles, covering 88 cell lines, 52 tissues, and various compound libraries. PRnet provides a robust and scalable candidate recommendation workflow and successfully recommends drug candidates for 233 diseases. Overall, PRnet is an effective and valuable tool for gene-based therapeutics screening.</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

Bushmeat yields, species extinction rates and ecosystem-level impacts of bushmeat harvesting as predicted by the Madingley General Ecosystem Model

<p>The datasets contain data generated using the Madingley General Ecosystem Model for experiments decribed in the paper: "T. Barychka, G.M.Mace and D.W.Purves (2021) The Madingley General Ecosystem Model predicts bushmeat yields, species extinction rates and ecosystem-level impacts of bushmeat harvesting. Oikos."</p> <p>The Madingley General Ecosystem Model was used to generate predictions of bushmeat yields, extinction rates and broader ecosystem impacts for a range of harvesting intensities of duiker-sized endothermic herbivores. Duiker antelope (such as <i>Cephalophus callipygus</i> and <i>Cephalophus dorsalis</i>) are the most heavily hunted species in sub-Saharan Africa, contributing 34%-95% of all bushmeat in the Congo Basin. In the Madingley, the harvested group was described as "Heterotroph – Herbivore – Terrestrial – Mobile – Iteroparous– Endotherm", with adult bodymasses of 13-21 kg and juvenile bodymasses of &gt;100 g. Harvesting period was set at 30 years (<em>n </em>=30).</p> <p>In the first experiment, we used the Madingley model to predict bushmeat yields ("Harvested Biomasses") and extinction rates ("Density") from harvesting duiker-sized herbivores using proportional harvesting strategy, with harvest rates ranging from 0 to 0.90. These were compared to the estimates of bushmeat yields and survival probabilities for two duiker antelope species (<i>Cephalophus callipygus</i> and <i>Cephalophus dorsalis) </i>from conventional single-species Beverton-Holt model. </p> <p>In the second experiment, we used the Madingley model to generate data on the state ("State") of the harvested ecosystem. The datasets contain estimates of biomasses, abundances, adult and juvenile bodymasses, etc. of the harvested duiker-sized herbivores as well as unharvested herbivores, omnivores and carnivores present in the simulated ecosystem. In the paper we focused on abundances; however, other estimates e.g., adult bodymasses can be used to conduct further studies on the effects of harvesting in tropical ecosystems.</p> <p>Main results of the experiments are that: 1) the Madingley model gave estimates for optimal harvesting rate, and extinction rate, that were qualitatively and quantitatively similar to the estimates from conventional single-species Beverton-Holt model; 2) the Madingley model predicted a background local extinction probability for the target species of at least 10%; 3) at medium and high levels of harvesting of duiker-sized herbivores, the Madingley model predicted  statistically significant, but moderate, reductions in the densities of the targeted functional group; increases in small-bodied herbivores; decreases in large-bodied carnivores; and minimal ecosystem-level impacts overall.</p>

opencc-zeroOct 2021View details →
zenodo36/100

Fine-tuning of predictive microbiology models through microlocal characterization of foods by Nuclear Magnetic Resonance (NMR)

<p>Fine-tuning of predictive microbiology models through microlocal characterization of foods by Nuclear Magnetic Resonance (NMR)</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Random Forest model for PPI predictions

<p>The uploaded model contains the trained random forest model to predict protein-protein Interactions based on their amino acid sequence. The model is described in our preprint &quot;ProteinPrompt: a webserver for predicting protein-protein interactions&quot; which can be found on bioRxiv.org: https://doi.org/10.1101/2021.09.03.458859</p> <p>The model was trained with the scikit-learn package in version 0.20.3 under Python 3.7.12</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record