Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,773 results for “predictive modeling”

Learn how ShareScore rates datasets ↗
zenodo40/100

Potential distribution of invasive boxwood blight pathogen (Calonectria pseudonaviculata) as predicted by process-based and correlative models

<p>R project, R scripts, and data files for reproducing most of the analyses presented in a climatic suitability study for boxwood blight. The README. md file describes how to run the scripts and provides details on data inputs.</p> <p><strong>Abstract: </strong>Boxwood blight caused by <em>Cps</em> is an emerging disease that has had devastating impacts on <em>Buxus</em> spp. in the horticultural sector, landscapes, and native ecosystems. In this study, we produced a process-based climatic suitability model in the CLIMEX program and combined outputs of four different correlative modeling algorithms to generate an ensemble correlative model. All models were fit and validated using a presence record dataset comprised of <em>Cps</em> detections across its entire known invaded range. Evaluations of model performance provided validation of good model fit for all models. A consensus map of CLIMEX and ensemble correlative model predictions indicated that not-yet-invaded areas in eastern and southern Europe and in the southeastern, midwestern, and Pacific coast regions of North America are climatically suitable for <em>Cps</em> establishment. Most regions of the world where<em> Buxus</em> and its congeners are native are also at risk of establishment. These findings provide the first insights into <em>Cps</em> global invasion threat, suggesting that this invasive pathogen has the potential to significantly expand its range.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Prediction and Visualization of Human Transmembrane Proteins using AlphaFold and Protein Language Models

<p><strong>Description:</strong> <strong>TMvis</strong> (&quot;TMvis496.tar.gz&quot;) is a dataset containing 496 3D-structures of predicted human transmembrane proteins (TMP) and their predicted membrane embedding. The method TMbed [1], based on the protein language model ProtT5 [2] predicted 4.967 TMP for the human proteome (20,375 proteins, UniProt [3] version April 2022; excluding TITIN_HUMAN due to length). For these proteins, we obtained AlphaFold [4] structures from AlphaFoldDB [5] with an average per-residue confidence score (pLDDT) of more than 90%. This resulted in the 496 proteins of TMvis, as can be found in &quot;TMvis496.fasta&quot;. The membrane embedding was predicted using the methods ANVIL [6], PPM3 [7], and per-residue TMbed predictions. As the three methods are based on different approaches, we decided to publish results for all. The figure &ldquo;TMvis_project_overview.png&rdquo; provides a graphical overview for each step described above.</p> <p><strong>TMvis Folder Structure:</strong> TMvis is separated into &ldquo;alpha&rdquo; containing predicted alpha-helical TMPs, and &ldquo;beta&rdquo; containing predicted beta-barrel TMPs. Within these folders, each protein is assigned one folder, identifiable by the respective unique UniProt ID. Each protein folder consists of:<br> - &ldquo;UniprotID.fasta&rdquo; with UniProt ID, sequence, TMbed per-residue prediction<br> - &ldquo;AF-UniprotID-F1-model_v2.pdb&rdquo; with the AlphaFold structure<br> - &ldquo;AF-UniprotID-F1-model_v2.cif&rdquo; with the AlphaFold structure<br> - &ldquo;AF-UniprotID-F1-model_v2_ANVIL.pdb&rdquo; with predicted ANVIL membrane embedding<br> - &ldquo;AF-UniprotID-F1-model_v2_ppm.pdb&rdquo; predicted PPM3 membrane embedding</p> <p>TMvis&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> ├── alpha&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; ├── A0A087X1C5&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp; ├── A0A087X1C5.fasta&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp; ├── AF-A0A087X1C5-F1-model_v2.pdb&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp; ├── AF-A0A087X1C5-F1-model_v2.cif&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp; ├── AF-A0A087X1C5-F1-model_v2_ANVIL.pdb&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; │&nbsp;&nbsp; └── AF-A0A087X1C5-F1-model_v2_ppm.PDB&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> │&nbsp;&nbsp; └── ...&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> └── beta&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> &nbsp;&nbsp;&nbsp; └── P45880</p> <p><strong>TMvis visualization:</strong> The 3D-visualization of every protein in the dataset TMvis can be easily accessed using the Jupyter Notebook &ldquo;TMvis.ipynb&rdquo;. It contains detailed descriptions the different membrane prediction tools ANVIL, PPM3, and TMbed as well as the respective code. Additionally, it allows to visualize the per-residue confidence scores (pLDDT) of AlphaFold.</p> <p>&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;</p> <p><strong>References:</strong></p> <p>[1] TMbed - TMbed Bernhofer, Michael, and Burkhard Rost. 2022. &ldquo;TMbed &ndash; Transmembrane Proteins Predicted through Language Model Embeddings.&rdquo; bioRxiv.</p> <p>[2] ProtT5 - A. Elnaggar et al., &quot;ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing,&quot; in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: 10.1109/TPAMI.2021.3095381.</p> <p>[3] UniProt - UniProt Consortium (2021). UniProt: the universal protein knowledgebase in 2021. Nucleic acids research, 49(D1), D480&ndash;D489.</p> <p>[4] AlphaFold - AlphaFold Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, et al. 2021. &ldquo;Highly Accurate Protein Structure Prediction with AlphaFold.&rdquo; Nature 596 (7873): 583&ndash;89.</p> <p>[5] Alphafold DB - Varadi, Mihaly, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, et al. 2022. &ldquo;AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models.&rdquo; Nucleic Acids Research 50 (D1): D439&ndash;44.</p> <p>[6] ANVIL - ANVIL Postic, Guillaume, Yassine Ghouzam, Vincent Guiraud, and Jean-Christophe Gelly. 2016. &ldquo;Membrane Positioning for High- and Low-Resolution Protein Structures through a Binary Classification Approach.&rdquo; Protein Engineering, Design &amp; Selection: PEDS 29 (3): 87&ndash;91.</p> <p>[7] PPM3 - PPM3 Lomize, Mikhail A., Irina D. Pogozheva, Hyeon Joo, Henry I. Mosberg, and Andrei L. Lomize. 2012. &ldquo;OPM Database and PPM Web Server: Resources for Positioning of Proteins in Membranes.&rdquo; Nucleic Acids Research 40 (Database issue): D370&ndash;76.</p> <p>&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;</p> <p><strong>License:</strong></p> <p>This work is licensed under a Creative Commons Attribution 4.0 International License (CC-BY 4.0).</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data used in a manuscript entitled "Large ensemble simulation for investigating predictability of precursor vortices of Typhoon Faxai in 2019 with a 14-km mesh global nonhydrostatic atmospheric model" submitted to Geophysical Research Letters

<p>This include a dataset used in a manuscript entitled &ldquo;Large ensemble simulation for investigating predictability of precursor vortices of Typhoon Faxai in 2019 with a 14-km mesh global nonhydrostatic atmospheric model&rdquo; by Yamada and co-authors, which is submitted to Geophysical Research Letters.</p> <p>Contact: Yohei Yamada (yoheiy@jamstec.go.jp)</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models

<p>This data set contains the data, JMP scripts, and figures of the article titled &quot;Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models&quot; to be published in the journal Animal - Open Space.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Resource-Rational Lossy-Context Surprisal (Model Predictions)

<p>Resource-Rational Lossy-Context Surprisal is a computationally implemented model of how humans process language, predicting at what points in complex sentences they experience comprehension difficulty. It unifies the memory-based and expectation-based paradigms in psycholinguistics, and provides a more refined account of when hierarchical structure is difficult to comprehend for humans.</p> <p>This repository contains output of the model on a battery of test sentences exhibiting iterated recursive structure, described in associated publications on Resource-Rational Lossy-Context Surprisal. The filenames are referred to in the source code, to be published together with a forthcoming journal publication on the model.</p> <p>The model was first described in the following publication:</p> <p><em>Lexical Effects in Structural Forgetting: Evidence for Experience-Based Accounts and a Neural Network Model</em></p> <p>(Michael Hahn, Richard Futrell, Edward Gibson), 33rd Annual CUNY Human Sentence Processing Conference, 2020</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Codes for Purgar et al. 2022: Investigating the ability of growth models to predict in situ Vibrio spp. abundances

<p>Model simulations and analysis of the Vibrio spp. growth models.&nbsp;<br> Prepared to accompany the publication, Purgar et al. 2022 &quot;Investigating the ability of growth models to predict in situ<br> Vibrio spp. abundances&quot; in Microorganisms, Special Issue &bdquo;Microbial Communities in Changing Aquatic Environments&ldquo;.&nbsp;</p> <p>Description of the files&nbsp;can be found in the Readme file.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

A cooperative deep learning model for stock market prediction using deep autoencoder and sentiment analysis

<p>This data is used for Stock Market Prediction.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

On the Application of Machine Learning Models to Assess and Predict Software Reusability

<p>This is the dataset, results and notebook for the submission into Maltesque 2022 conference.</p>

opencc-by-4.0Nov 2022View details →
dryad40/100

Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model

<p>The diamond is 58 times harder than any other mineral in the world, and its elegance as a jewel has long been appreciated. Forecasting diamond prices is challenging due to nonlinearity in important features such as carat, cut, clarity, table, and depth. Against this backdrop, the study conducted a comparative analysis of the performance of multiple supervised machine learning models (regressors and classifiers) in predicting diamond prices. Eight supervised machine learning algorithms were evaluated in this work including Multiple Linear Regression, Linear Discriminant Analysis, eXtreme Gradient Boosting, Random Forest, k-Nearest Neighbors, Support Vector Machines, Boosted Regression and Classification Trees, and Multi-Layer Perceptron. The analysis is based on data preprocessing, exploratory data analysis (EDA), training the aforementioned models, assessing their accuracy, and interpreting their results. Based on the performance metrics values and analysis, it was discovered that eXtreme Gradient Boosting was the most optimal algorithm in both classification and regression, with a R<sup>2</sup> score of 97.45% and an Accuracy value of 74.28%. As a result, eXtreme Gradient Boosting was recommended as the optimal regressor and classifier for forecasting the price of a diamond specimen.</p>

opencc-zeroOct 2022View details →
zenodo40/100

Supplementary material 1 from: Motloung R, Robertson M, Rouget M, Wilson J (2014) Forestry trial data can be used to evaluate climate-based species distribution models in predicting tree invasions. NeoBiota 20: 31-48. https://doi.org/10.3897/neobiota.20.5778

Current and potential distributions of sixteen species that are not widespread in southern Africa arranged on the basis of their suitable range size : a) Acacia paradoxa, b) A. cultriformis, c) A. falciformis, d) A. pendula, e) A. rubida, f) A. stricta, g) A. retinodes, h) A. fimbriata, i) A. aneura, j) A. viscidula, k) A. acuminata, l) A. adunca, m) A. binervata, n) A. schinoides, o) A. prominens, p) A. mangium. The grey shading indicates areas that SDMs have identified as suitable by SDMs while the white ones are unsuitable.

opencc-by-4.0Jan 2014View details →
dryad40/100

Comparative ecological analysis and predictive modeling of tick-borne pathogens

<p>Tick-borne diseases constitute the predominant vector-borne health threat in North America. Recent observations have noted a significant expansion in the range of the black-legged tick (<em>Ixodes scapularis</em> Say, Acari: Ixodidae), alongside a rise in the incidence of diseases caused by its vectored pathogens: <em>Borrelia burgdorferi</em> (Spirochaetales: Spirochaetaceae), <em>Babesia microti</em> (Piroplasmida: Babesiidae), and <em>Anaplasma phagocytophilium</em> (Rickettsiales: Anaplasmataceae), the causative agents of Lyme disease, babesiosis, and anaplasmosis, respectively. Prior research identified environmental features that influence the ecological dynamics of <em>I. scapularis</em> and <em>B. burgdorferi</em> that can be used to predict the distribution and abundance of these organisms, and thus Lyme disease risk. In contrast, there is a paucity of research into the environmental determinants of <em>B. microti</em> and <em>A. phagocytophilium</em>. Here we use over a decade of surveillance data to model the impact of environmental features on the infection prevalence of these increasingly common human pathogens in ticks across New York State (NYS). Our findings reveal a consistent northward and westward expansion of <em>B. microti</em> in NYS from 2009 to 2019, while the range of <em>A. phagocytophilum</em> varied at fine spatial scales. We constructed biogeographic models using data from over 1000 site-year visits and encompassing more than 250 environmental variables to accurately forecast infection prevalence for each pathogen to future years that were not included in model training. Several environmental features were identified to have divergent effects on the pathogens, revealing potential ecological differences governing their distribution and abundance. These validated biogeographic models are immediately useful for disease prevention efforts.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Data for 'VespaG: Expert-guided protein language models enable accurate and blazingly fast fitness prediction'

<div>Datasets used for development of VespaG and VespaG predictions generated with <a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.&nbsp;</div> <div>&nbsp;</div> <div>Uploads contain:</div> <div> <ol> <li><strong>Performance</strong> summaries for ProteinGym [1]:<br>- Spearman and Pearson correlation for VespaG:&nbsp;<em>proteingym_performance_vespag.csv&nbsp;</em>(columns: 'DMS_id', 'Spearman', 'Pearson')<br>- Spearman correlation for evaluated methods VespaG, GEMME [2], VESPA [3], TranceptEVE [4], AlphaMissense [5], PoET [6]: <em>proteingym_spearman_allmethods.csv&nbsp;</em>(columns: 'DMS_id', 'Trancept EVE-L', 'VESPA', 'VespaG', 'GEMME', 'AlphaMissense', 'PoET', 'UniProt_ID', 'coarse_selection_type' (function), 'taxon')</li> <li><strong>Fasta</strong> files with sequences for all train sets (<em>vespag_fasta_training_datasets.zip</em> with seq_all9k.fasta, seq_human5k.fasta, seq_droso4k.fasta, seq_ecoli2k.fasta, seq_virus1k.fasta) and test set (<em>proteingym_217.fasta</em>)</li> <li><strong>VespaG</strong> <strong>Predictions</strong> for test set:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip</em> with raw_preds_ecoli.csv, raw_preds_human.csv, raw_preds_virus.csv, raw_preds_all.csv, raw_preds_droso.csv (columns: 'DMS_id', 'mutation', 'DMS_score', 'VespaG'). Predictions are based on different training data, the final model VespaG was trained on a subset of the human proteome and <strong>raw VespaG predictions</strong> <strong>for</strong> <strong>the</strong> <strong>ProteinGym benchmark are in&nbsp;raw_preds_human.csv </strong>(used to calculate the performances above).</li> <li><strong>GEMME predictions</strong> for train sets:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip&nbsp;</em>with folders 'human', 'droso', 'ecoli', 'virus', 'all' for respective fasta file (each containing GEMME mutational landscape output files named '<em>ID' + '</em>_normPred_evolCombi.txt')</li> <li><strong>ESM-2</strong> <strong>embeddings</strong> [7] for test set (<em>proteingym_217_esm2.h5</em>)</li> </ol> </div> <div>For details on VespaG see:</div> <div> <div> <div>VespaG: Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction</div> </div> <div>Celine Marquet, Julius Schlensok, Marina Abakarova, Burkhard Rost, Elodie Laine</div> <div>bioRxiv 2024.04.24.590982; doi: https://doi.org/10.1101/2024.04.24.590982</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see&nbsp;<a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <div>Exhaustive experimental annotation of the effect of all known protein variants remains daunting and expensive, stressing the need for scalable effect predictions. We introduce VespaG, a blazingly fast single amino acid variant effect predictor, leveraging embeddings of protein Language Models as input to a minimal deep learning model. To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME as a pseudo standard-of-truth. Assessed against the ProteinGym Substitution Benchmark (217 multiplex assays of variant effect with 2.5 million variants), VespaG achieved a mean Spearman correlation of 0.48 +/- 0.01, matching state-of-the-art methods such as GEMME, TranceptEVE, PoET, AlphaMissense, and VESPA. VespaG reached its top-level performance several orders of magnitude faster, predicting all mutational landscapes of the human proteome in 30 minutes on a consumer laptop (12-core CPU, 16 GB RAM).</div> <div>&nbsp;</div> <div>[1] Notin, Pascal, et al. "ProteinGym: large-scale benchmarks for protein fitness prediction and design." <em>Advances in Neural Information Processing Systems</em> 36 (2024).<br>[2] Laine, Elodie, Yasaman Karami, and Alessandra Carbone. "GEMME: a simple and fast global epistatic model predicting mutational effects." <em>Molecular biology and evolution</em> 36.11 (2019): 2604-2619.</div> <div>[3] Marquet, C&eacute;line, et al. "Embeddings from protein language models predict conservation and variant effects." <em>Human genetics</em> 141.10 (2022): 1629-1647.</div> <div>[4] Notin, Pascal, et al. "TranceptEVE: Combining family-specific and family-agnostic models of protein sequences for improved fitness prediction." <em>bioRxiv</em> (2022): 2022-12.</div> <div>[5] Cheng, Jun, et al. "Accurate proteome-wide missense variant effect prediction with AlphaMissense." <em>Science</em> 381.6664 (2023): eadg7492.</div> <div>[6] Truong Jr, Timothy, and Tristan Bepler. "PoET: A generative model of protein families as sequences-of-sequences." <em>Advances in Neural Information Processing Systems</em> 36 (2024).</div> <div>[7] Lin, Zeming, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." <em>Science</em>379.6637 (2023): 1123-1130.</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Figure 6 in Forest yield prediction under different climate change scenarios using data intelligent models in Pakistan

Figure 6. Empirical cumulative distribution function (ECDF) of the Predicted error |PE| (cft) in testing period for the RF and KRR models between the predicted and observed yields of Blue pine and Silver fir species.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 4 in Forest yield prediction under different climate change scenarios using data intelligent models in Pakistan

Figure 4. Box-plots of the Predicted error | PE| (cft) in testing period (1996-2016) for the RF and KRR models between the predicted and observed yields of Blue pine and Silver fir species.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 7 in Forest yield prediction under different climate change scenarios using data intelligent models in Pakistan

Figure 7. Taylor diagram showing the correlation coefficient between the predicted and observed yields (Blue pine and Silver fir) (cft) and standard deviation for the RF and KRR models.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 5 in Forest yield prediction under different climate change scenarios using data intelligent models in Pakistan

Figure 5. Polar plots show the Predicted error |PE|(cft) in testing period (1996-2016) for the RF and KRR models between the predicted and observed yields of Blue pine and Silver fir species.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 5 in Establishment of an expansion-predicting model for invasive alien cerambycid beetle Aromia bungii based on a virtual ecology approach

Figure 5. Map of predicted occurrence units for the whole of Saitama Prefecture using both the river density model and river single model. The degree of shading reflects the theoretical invasion number predicted by each simulation model.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 4 in Establishment of an expansion-predicting model for invasive alien cerambycid beetle Aromia bungii based on a virtual ecology approach

Figure 4. (a) Map of occurrence records for A. bungii through 2019. (b–g) Predicted occurrence units based on our models for each habitat variable. The degree of shading reflects the theoretical invasion number predicted by each model.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 2 in Establishment of an expansion-predicting model for invasive alien cerambycid beetle Aromia bungii based on a virtual ecology approach

Figure 2. Basic structure of the cellular automata model. (A) Two values are associated with each cell: 1) the cell ID "x," a unique ID for each cell, and 2) the expansion probability "ex" indicating four directional vectors into adjacent cells (described below). (B) Values e1, e2, e3, and e4 indicate the probability of dispersion using the path to the top, left, bottom, and right cells, respectively. If the dispersion path value is 1, the insect population in this cell can expand to the adjacent cell.

opencc-by-4.0Dec 2021View details →
zenodo40/100

The state-of-the-art machine learning model for Plasma Protein Binding Prediction: computational modeling with OCHEM and experimental validation

<p><span>Institute of Materia Medica,&nbsp;Chinese Academy of Medical Sciences purchased 10,000 ChemDiv databases.</span></p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record