Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.7.1
Dataset results
1,773 results for “predictive modeling”
Supplements for "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"
<p><strong>Supplements for "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"</strong></p> <p>This supplement includes the following files:</p> <ul> <li>Main_v02.R --- Script in R language to process files (in "PDBs.zip") and produce the dataset ("Dataset_v02.csv")</li> <li>PDBs.zip --- PDB files include filtered structures processed by three programs</li> <li>Dataset_v02.csv --- final filtered version of dataset produced in R language (by "Main_v02.R"). </li> <li>Dataset key.txt --- key to column names in Dataset_v02.csv </li> </ul> <p>The article featuring this dataset is published in:</p> <p>Journal: Journal of Chemical Information and Modeling<br>Title: "Solvent Accessibility Promotes Rotamer Errors During Protein Modelling with Major Side-Chain Prediction Programs"<br>Author(s): Hameduh, Tareq; Mokry, Michal ; Miller, Andrew ; Heger, Zbynek; Haddad, Yazan</p> <p><a href="https://emea01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fpubs.acs.org%2Fdoi%2F10.1021%2Facs.jcim.3c00134&data=05%7C01%7C%7C4876a5c33e5f4bd8daf008db7e55a28d%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C638242678496577396%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=NE%2F0dYvc4m5%2B8T0YQtqxY21yk1pA9ZpiADqzx3vEJ4Q%3D&reserved=0">https://pubs.acs.org/doi/10.1021/acs.jcim.3c00134</a></p> <p> </p>
BRCA1-specific machine learning model predicts variant pathogenicity with high accuracy - Supplementary material
<p>Figure S1: Distribution of the reviewed 141 <em>BRCA1</em> missense variants; Figure S2: The Shapely values for the <em>BRCA1</em> XGBoost models; Figure S3: The Shapely values of the <em>BRCA1</em> XGBoost model used to predict the functional assays’ results for variants of uncertain significance; Table S1: The receiver operating characteristic (ROC) curve analysis for the different in silico predictions; Table S2: Cross validation of the BRCA1 model in 5 different random training and test samples; Table S3: Pathogenicity prediction and prioritization of the 31,058 unreviewed BRCA1 variants from the BRCA Exchange database.</p>
The Prairie State: Using Ecological Niche Modeling to Predict Distributions of Early Land Plants
<p>This data includes raw data of over 12,000 occurrences were downloaded from the<strong> Consortium of Bryophyte Herbaria (<a href="http://www.bryophyteportal.org/portal">www.bryophyteportal.org/portal</a>), </strong>that were listed to be in Illinois and included longitude and latitude data. This data set was screened and cleaned to investigate species distribution models as well as generate models of selected bryophytes investigating future changes in distribution across climate change scenarios.</p>
In silico design, docking simulation, and ANN-QSAR model for predicting the anticoagulant activity of thiourea isosteviol compounds as FXa inhibitors
<p>The present work combined molecular modeling and docking approach for searching and designing novel thiourea isosteviol-based compounds as potential FXa inhibitors. Elaborated regression model establishes the relationships between experimentally determined anticoagulant activity and molecular descriptors and enables the prediction of FXa inhibitory activity for novel compounds. The obtained results proved that the Artificial Neural Network algorithm facilitates the search for the most promising isosteviol derivatives incorporating thiourea fragments as FXa inhibitors. Moreover, docking simulation confirms the prominent binding of the newly in silico designed molecules with the active sites of the protein, which may be the lead molecules and can be further optimized for the efficient pharmacodynamic and pharmacokinetic profiles. The enclosed files are representations of molecular structures of thiourea isosteviol compounds with experimentally tested FXa inhibitory activity (i20-i39) geometrically optimized in hyperchem, newly in silico designed thiourea isosteviol compounds geometrically optimized in hyperchem (e1-e11), one file contains molecular descriptors for optimized structures calculated in Dragon and there is also a code for ANN QSAR model for predicting activity of novel thiourea isosteviol compounds. </p>
Predicted models of S receptor kinase ectodomain (eSRK), S-locus protein 11 (SP11), and eSRK-SP11 complexes
<p>Predicted models of <em>S</em> receptor kinase ectodomain (eSRK), <em>S</em>-locus protein 11 (SP11), and their complexes using ColabFold and curated multiple sequence alignments (MSAs).</p> <p> </p>
CPT-1 pre-computed whole-proteome variant effect predictions and model source code
<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> & <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository songlab-cal/CPT version as of Jul 12, 2023)</p><p> </p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed: <a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI: <a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p> </p>
Model output from Snow Ensemble Uncertainty Project (SEUP) as used in Seasonal Snow Predictability Derived from Early-Season Snow in North America
<p>The files provided here are the output from the median peak snow water equivalent (peak_SWE.mat), 1 December snow water equivalent (Dec1_SWE.mat), and 1 January snow water equivalent (Jan1_SWE.mat) model simulations for the Noah-MP run with MERRA-2 forcing, as used in Lundquist et al. (2023) and described in Kim et al. (2021). We also include the model grid's latitude, longitude and elevation data (SEUPlatlon.mat), and example code for plotting the data (Plotmodeldata.m) as in the Lundquist et al. (2023) paper. </p> <p>Kim, R. S., Kumar, S., Vuyovich, C., Houser, P., Lundquist, J., Mudryk, L., et al. (2021). Snow Ensemble Uncertainty Project (SEUP): Quantification of snow water equivalent uncertainty across North America via ensemble land surface modeling. <em>The Cryosphere, 15</em>(2), 771-791.</p> <p>Lundquist, J. D., R. S. Kim, M. Durand, and L. R. Prugh, 2023, Seasonal Peak Snow Predictability Derived from Early-Season Snow in North America, Geophysical Research Letters, (submitted 2023)</p>
Multi-Omics Visible Drug Activity Prediction with a Biologically Informed Neural Network Model
<p>Drug discovery is a challenging task, it takes several years for a drug to be introduced on the market, with most of<br> the studied drugs not even passing the first phase. The understanding of the mechanisms influencing response to drugs<br> can reduce failures and accelerate drug development. Virtual drug screening, based on Machine Learning models, is a<br> promising field for the prediction of the outcome of a treatment. However, the complex relationships between the features<br> learned by these models are still poorly understood and not easy to interpret.<br> We have designed a Neural Network model for drug sensitivity prediction that leverages a Visible Neural Network, an<br> easily interpretable model, due to its biologically informed nature. The trained model can be inspected to study which<br> biological processes were fundamental for the prediction and to identify the drug properties that affect sensitivity. It<br> combines multi-omics data from various types of tumor tissues and drug representations based on molecular descriptors.<br> The mechanisms learned from the network can also be exploited to find candidate drugs for synergy to predict the effect<br> of combined therapies. We consider the unbalanced nature of public drug screening datasets and show that our model<br> outperforms state-of-the-art visible machine learning models.</p>
Advancements in QSAR modelling: Decision trees and rotation forest for prediction of Aspergillus anti-inflammatory metabolites
<p>This study presents applications of advancements in QSAR modelling for predicting nitric oxide (NO) inhibitors and anti-inflammatory metabolites from the <em>Aspergillus</em> genus. Inflammation-related diseases remain a pressing concern, necessitating the identification of effective anti-inflammatory compounds. Using decision trees, the Ranker method, and CorrelationAtrributeEval as a base classifier for attribute selection together with Rotation Forest and Adaboost as enhancers, we explored their potential with different classifiers including Artificial Neural Networks and J48 Trees. The proposed QSAR models employed an ensemble approach with Rotation Forest and Adaboost.M1, applying an automated KNIME workflow. Seven molecular descriptors were selected and trained on a comprehensive dataset of diverse anti-inflammatory <em>Aspergillus</em> specialised metabolites. Results showed that the Rotation Forest-enhanced version outperformed other models, capturing complex structure-activity relationships and improving predictive performance. Chemical characteristics of electrotopological state, topological distances, and functional groups including secondary amides and alcohols contribute to important anti-inflammatory effects. The developed QSAR model showed good predictive performance for anti-inflammatory <em>Aspergillus</em> metabolites, focusing on their NO inhibitory activity. These results can contribute to the discovery of novel anti-inflammatory drugs based on computational techniques.</p> <p> </p>
Soil chemistry dataset from the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica"
<p>Bases sum, H+Al (potential acidity), pH, phosphorous, remaining P (P-rem), sodium and total organic carbon distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The quantile and prediction interval data represent the spatial uncertainty of the predictions.</p> <p>As soon as the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica" is published, the paper will be cited here. </p> <p>The .zip file contains the following folders:</p> <p>1) soil_chemistry_antarctica: data containing the soil chemical attributes distribution</p> <p>2) soil_chemistry_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil attributes prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil attributes prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil attributes prediction</p>
Prediction of the potentially suitable areas of Leonurus japonicus with the optimized MaxEnt model
<p><em><span>Leonurus japonicus </span></em><span>Houtt.</span><span> is a traditional Chinese medicinal plant with high medicinal and edible value.</span><span> Wild <em>L. japonicus</em> resources have been reduced dramatically in recent years. This study predicted the response of distribution range of <em>L. japonicus</em> to climate change in China, which provided the scientific basis for the conservation and utilization. In this study, 489 occurrence poin</span><span>ts</span><span> of <em>L. japonicus</em> were selected based on GIS technology and spThin package. The default parameters of the Maxent model were adjusted by using ENMeva1 package of the R environment, and the optimized Maxent model was used to analyze the distribution of <em>L. japonicus</em>. When the feature combination in the model parameters is hing and the regularization multiplier is 1.5, the Maxent model has a higher degree of optimization. With the AUC of 0.830 our model showed a good predictive performance The results showed that <em>L. japonicus</em> was widely distributed in the current period. The maximum temperature of the warmest month, the minimum temperature of the coldest </span><span>month</span><span>, the precipitation of the wettest month, the precipitation of the driest month and altitude were the main environmental factors affecting the distribution of <em>L. japonicus</em>. Under the three climate change scenarios, the suitable distribution area of <em>L. japonicus</em> will range-shift to high latitudes, indicating that the distribution of <em>L. japonicus</em> has a strong response to climate change. The regional change rate is the lowest under the SSP126-2090s scenario and the highest under the SSP585-2090s scenario.</span></p>
Data --- "Optimization of Convolutional Neural Network models for spatially coherent multi-site fire danger predictions"
<p>Data to reproduce the results of the manuscript entitled "Optimization of Convolutional Neural Network models for spatially coherent multi-site fire danger predictions" submitted to Geophysical Research Letters. The companion jupyter notebook can be found in DOI: <a href="https://doi.org/10.5281/zenodo.8387558">10.5281/zenodo.8387558</a></p>
Predicting the distribution of serotonergic axons: A supercomputing simulation of reflected fractional Brownian motion in a 3D-mouse brain model
Open the record for dataset details and reuse information.
A mathematical model to predict network growth in physarum polycephalum as a function of extracellular matrix viscosity, measured by a novel viscometer
Open the record for dataset details and reuse information.
Modelling heterogeneity in the classification process in multi-species distribution models can improve predictive performance
Open the record for dataset details and reuse information.
Comparative ecological analysis and predictive modeling of tick-borne pathogens
Open the record for dataset details and reuse information.
Data from: Frugivore traits predict plant-frugivore interactions using generalized joint attribute modeling
Open the record for dataset details and reuse information.
Development of whole-genome prediction models to increase the rate of genetic gain in intermediate wheatgrass (Thinopyrum intermedium) breeding
Open the record for dataset details and reuse information.
Data from: Exploiting nozzle geometry to predict resolution in extrusion-based bioprinting: mathematical modelling of a power-law fluid
Open the record for dataset details and reuse information.
Data for: Modeling the transition of death assemblages through the mixed layer predicts a downcore increase in time averaging
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.