Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,773 results for “predictive modeling”

Learn how ShareScore rates datasets ↗
zenodo40/100

Supplements for "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"

<p><strong>Supplements for "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"</strong></p> <p>This supplement includes the following files:</p> <ul> <li>Main_v02.R --- Script in R language to process files (in "PDBs.zip") and produce the dataset&nbsp;("Dataset_v02.csv")</li> <li>PDBs.zip --- PDB files include filtered structures processed by three programs</li> <li>Dataset_v02.csv --- final filtered version&nbsp;of dataset produced in R language (by "Main_v02.R").&nbsp; &nbsp;</li> <li>Dataset key.txt --- key to column names in Dataset_v02.csv&nbsp; &nbsp; &nbsp;</li> </ul> <p>The article featuring this dataset is published in:</p> <p>Journal: &nbsp;Journal of Chemical Information and Modeling<br>Title: "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"<br>Author(s): Hameduh, Tareq; Mokry, Michal ; Miller, Andrew &nbsp;; Heger, Zbynek; Haddad, Yazan</p> <p><a href="https://emea01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fpubs.acs.org%2Fdoi%2F10.1021%2Facs.jcim.3c00134&amp;data=05%7C01%7C%7C4876a5c33e5f4bd8daf008db7e55a28d%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C638242678496577396%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=NE%2F0dYvc4m5%2B8T0YQtqxY21yk1pA9ZpiADqzx3vEJ4Q%3D&amp;reserved=0">https://pubs.acs.org/doi/10.1021/acs.jcim.3c00134</a></p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

BRCA1-specific machine learning model predicts variant pathogenicity with high accuracy - Supplementary material

<p>Figure S1: Distribution of the reviewed 141&nbsp;<em>BRCA1</em>&nbsp;missense variants; Figure S2:&nbsp;The Shapely values for the&nbsp;<em>BRCA1</em>&nbsp;XGBoost models; Figure S3: The Shapely values of the&nbsp;<em>BRCA1</em>&nbsp;XGBoost model used to predict the functional assays&rsquo; results for variants of uncertain significance;&nbsp;Table S1: The receiver operating characteristic (ROC) curve analysis for the different in silico predictions; Table S2: Cross validation of the BRCA1 model in 5 different random training and test samples; Table S3: Pathogenicity prediction and prioritization of the 31,058 unreviewed BRCA1 variants from the BRCA Exchange database.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

The Prairie State: Using Ecological Niche Modeling to Predict Distributions of Early Land Plants

<p>This data includes raw data of over 12,000 occurrences were downloaded from the<strong>&nbsp;Consortium of Bryophyte Herbaria (<a href="http://www.bryophyteportal.org/portal">www.bryophyteportal.org/portal</a>),&nbsp;</strong>that were listed to be in Illinois and included longitude and latitude data. This data set was screened and cleaned to investigate species distribution models as well as generate&nbsp;models of selected bryophytes investigating future changes in distribution across climate change scenarios.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

In silico design, docking simulation, and ANN-QSAR model for predicting the anticoagulant activity of thiourea isosteviol compounds as FXa inhibitors

<p>The present work combined molecular modeling and docking approach for searching and designing novel thiourea isosteviol-based compounds as potential FXa inhibitors. Elaborated regression model establishes the relationships between experimentally determined anticoagulant activity and molecular descriptors and enables the prediction of FXa inhibitory activity for novel compounds. The obtained results proved that the Artificial Neural Network algorithm facilitates the search for the most promising isosteviol derivatives incorporating thiourea fragments as FXa inhibitors. Moreover, docking simulation confirms the prominent binding of the newly in silico designed molecules with the active sites of the protein, which may be the lead molecules and can be further optimized for the efficient pharmacodynamic and pharmacokinetic profiles.&nbsp;The enclosed files are representations of&nbsp;molecular structures of thiourea isosteviol compounds with experimentally tested FXa inhibitory activity (i20-i39) geometrically optimized in hyperchem, newly in silico designed thiourea isosteviol compounds geometrically optimized in hyperchem (e1-e11), one file contains molecular descriptors for optimized structures calculated in Dragon and there is also a code for ANN QSAR model for predicting activity of novel thiourea isosteviol compounds.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Predicted models of S receptor kinase ectodomain (eSRK), S-locus protein 11 (SP11), and eSRK-SP11 complexes

<p>Predicted models of <em>S</em> receptor kinase ectodomain (eSRK), <em>S</em>-locus protein 11 (SP11), and their complexes using ColabFold and curated multiple sequence alignments (MSAs).</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

CPT-1 pre-computed whole-proteome variant effect predictions and model source code

<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> &amp; <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository&nbsp;songlab-cal/CPT version as of Jul 12, 2023)</p><p>&nbsp;</p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed:&nbsp;<a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI:&nbsp;<a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Model output from Snow Ensemble Uncertainty Project (SEUP) as used in Seasonal Snow Predictability Derived from Early-Season Snow in North America

<p>The files provided here are the output from the median peak snow water equivalent (peak_SWE.mat), 1 December snow water equivalent (Dec1_SWE.mat), and 1 January snow water equivalent (Jan1_SWE.mat)&nbsp;model simulations for the Noah-MP run with MERRA-2 forcing, as used in Lundquist et al. (2023) and&nbsp;described in Kim et al. (2021).&nbsp; We also include the model grid&#39;s latitude, longitude and elevation data (SEUPlatlon.mat), and example code for plotting the data (Plotmodeldata.m) as in the Lundquist et al. (2023) paper.&nbsp;</p> <p>Kim, R. S., Kumar, S., Vuyovich, C., Houser, P., Lundquist, J., Mudryk, L., et al. (2021). Snow Ensemble Uncertainty Project (SEUP): Quantification of snow water equivalent uncertainty across North America via ensemble land surface modeling. <em>The Cryosphere, 15</em>(2), 771-791.</p> <p>Lundquist, J. D., R. S. Kim, M. Durand, and L. R. Prugh, 2023, Seasonal Peak Snow Predictability Derived from Early-Season Snow in North America, Geophysical Research Letters, (submitted 2023)</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Multi-Omics Visible Drug Activity Prediction with a Biologically Informed Neural Network Model

<p>Drug discovery is a challenging task, it takes several years for a drug to be introduced on the market, with most of<br> the studied drugs not even passing the first phase. The understanding of the mechanisms influencing response to drugs<br> can reduce failures and accelerate drug development. Virtual drug screening, based on Machine Learning models, is a<br> promising field for the prediction of the outcome of a treatment. However, the complex relationships between the features<br> learned by these models are still poorly understood and not easy to interpret.<br> We have designed a Neural Network model for drug sensitivity prediction that leverages a Visible Neural Network, an<br> easily interpretable model, due to its biologically informed nature. The trained model can be inspected to study which<br> biological processes were fundamental for the prediction and to identify the drug properties that affect sensitivity. It<br> combines multi-omics data from various types of tumor tissues and drug representations based on molecular descriptors.<br> The mechanisms learned from the network can also be exploited to find candidate drugs for synergy to predict the effect<br> of combined therapies. We consider the unbalanced nature of public drug screening datasets and show that our model<br> outperforms state-of-the-art visible machine learning models.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Advancements in QSAR modelling: Decision trees and rotation forest for prediction of Aspergillus anti-inflammatory metabolites

<p>This study presents applications of advancements in QSAR modelling for predicting nitric oxide (NO) inhibitors and anti-inflammatory metabolites from the <em>Aspergillus</em> genus. Inflammation-related diseases remain a pressing concern, necessitating the identification of effective anti-inflammatory compounds. Using decision trees, the Ranker method, and CorrelationAtrributeEval as a base classifier for attribute selection together with Rotation Forest and Adaboost as enhancers, we explored their potential with different classifiers including Artificial Neural Networks and J48 Trees. The proposed QSAR models employed an ensemble approach with Rotation Forest and Adaboost.M1, applying an automated KNIME workflow. Seven molecular descriptors were selected and trained on a comprehensive dataset of diverse anti-inflammatory <em>Aspergillus</em> specialised metabolites. Results showed that the Rotation Forest-enhanced version outperformed other models, capturing complex structure-activity relationships and improving predictive performance. Chemical characteristics of electrotopological state, topological distances, and functional groups including secondary amides and alcohols contribute to important anti-inflammatory effects. The developed QSAR model showed good predictive performance for anti-inflammatory <em>Aspergillus</em> metabolites, focusing on their NO inhibitory activity. These results can contribute to the discovery of novel anti-inflammatory drugs based on computational techniques.</p> <p>&nbsp;</p>

opencc-bySep 2023View details →
zenodo40/100

Soil chemistry dataset from the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica"

<p>Bases sum, H+Al (potential acidity), pH, phosphorous, remaining P (P-rem), sodium and total organic carbon&nbsp;distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The quantile and prediction interval data represent the spatial uncertainty of the predictions.</p> <p>As soon as the work&nbsp;&quot;Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica&quot; is published, the paper will be cited here.&nbsp;</p> <p>The .zip file contains the following folders:</p> <p>1) soil_chemistry_antarctica: data containing the soil chemical attributes distribution</p> <p>2) soil_chemistry_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil attributes prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil attributes prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil attributes&nbsp;prediction</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Prediction of the potentially suitable areas of Leonurus japonicus with the optimized MaxEnt model

<p><em><span>Leonurus japonicus </span></em><span>Houtt.</span><span> is a traditional Chinese medicinal plant with high medicinal and edible value.</span><span> Wild <em>L. japonicus</em> resources have been reduced dramatically in recent years. This study predicted the response of distribution range of <em>L. japonicus</em> to climate change in China, which provided the scientific basis for the conservation and utilization. In this study, 489 occurrence poin</span><span>ts</span><span> of <em>L. japonicus</em> were selected based on GIS technology and spThin package. The default parameters of the Maxent model were adjusted by using ENMeva1 package of the R environment, and the optimized Maxent model was used to analyze the distribution of <em>L. japonicus</em>. When the feature combination in the model parameters is hing and the regularization multiplier is 1.5, the Maxent model has a higher degree of optimization. With the AUC of 0.830 our model showed a good predictive performance The results showed that <em>L. japonicus</em> was widely distributed in the current period. The maximum temperature of the warmest month, the minimum temperature of the coldest </span><span>month</span><span>, the precipitation of the wettest month, the precipitation of the driest month and altitude were the main environmental factors affecting the distribution of <em>L. japonicus</em>. Under the three climate change scenarios, the suitable distribution area of <em>L. japonicus</em> will range-shift to high latitudes, indicating that the distribution of <em>L. japonicus</em> has a strong response to climate change. The regional change rate is the lowest under the SSP126-2090s scenario and the highest under the SSP585-2090s scenario.</span></p>

opencc-zeroSep 2023View details →
zenodo40/100

Data --- "Optimization of Convolutional Neural Network models for spatially coherent multi-site fire danger predictions"

<p>Data to reproduce the results of the manuscript entitled "Optimization of Convolutional Neural Network models for spatially coherent multi-site fire danger predictions" submitted to Geophysical Research Letters. The companion jupyter notebook can be found in DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.8387558">10.5281/zenodo.8387558</a></p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Predicting the distribution of serotonergic axons: A supercomputing simulation of reflected fractional Brownian motion in a 3D-mouse brain model

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

A mathematical model to predict network growth in physarum polycephalum as a function of extracellular matrix viscosity, measured by a novel viscometer

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad40/100

Modelling heterogeneity in the classification process in multi-species distribution models can improve predictive performance

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Comparative ecological analysis and predictive modeling of tick-borne pathogens

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad40/100

Data from: Frugivore traits predict plant-frugivore interactions using generalized joint attribute modeling

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad40/100

Development of whole-genome prediction models to increase the rate of genetic gain in intermediate wheatgrass (Thinopyrum intermedium) breeding

Open the record for dataset details and reuse information.

publicApr 2021View details →
dryad40/100

Data from: Exploiting nozzle geometry to predict resolution in extrusion-based bioprinting: mathematical modelling of a power-law fluid

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

Data for: Modeling the transition of death assemblages through the mixed layer predicts a downcore increase in time averaging

Open the record for dataset details and reuse information.

publicNov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record