Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “protein structure prediction”
Structure prediction from SARS-CoV-2 accessory proteins ORF-7B
<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-7B.</p> <p>The archive contains both the structure and the logs from the prediction.</p>
Structure-based self-supervised learning enables ultrafast prediction of stability changes upon mutation at the protein universe scale
<p>Pythia computed all single mutations of <em>E.coli</em> proteome, high quality high quality of Swiss-Prot structures and thermophilic proteins used in analysis.</p>
A3D database: structure-based predictions of protein aggregation for the human proteome
<p>A3D database: structure-based predictions of protein aggregation for the human proteome</p>
Supplementary Structural Models (SARS-CoV-2 Spike-RBD:ACE2 complex and TMPRSS2) - SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals
<p>Structural Models (PDB) of SARS-CoV-2 Spike RBD bound to ACE2 receptors of 215 animals.</p> <p>Structural model of Human TMPRSS2.</p> <p>Modelled using the FunMod pipeline and referenced in the preprint</p> <p><a href="https://www.biorxiv.org/content/10.1101/2020.05.01.072371v5">SARS-CoV-2 spike protein predicted to form complexes with host receptor protein orthologues from a broad range of mammals</a></p> <p> </p>
Data from: Knowledge-based prediction of protein backbone conformation using a structural alphabet
Libraries of structural prototypes that abstract protein local structures are known as structural alphabets and have proven to be very useful in various aspects of protein structure analyses and predictions. One such library, Protein Blocks, is composed of 16 standard 5-residues long structural prototypes. This form of analyzing proteins involves drafting its structure as a string of Protein Blocks. Predicting the local structure of a protein in terms of protein blocks is the general objective of this work. A new approach, PB-kPRED is proposed towards this aim. It involves (i) organizing the structural knowledge in the form of a database of pentapeptide fragments extracted from all protein structures in the PDB and (ii) applying a knowledge-based algorithm that does not rely on any secondary structure predictions and/or sequence alignment profiles, to scan this database and predict most probable backbone conformations for the protein local structures. Though PB-kPRED uses the structural information from homologues in preference, if available. The predictions were evaluated rigorously on 15,544 query proteins representing a non-redundant subset of the PDB filtered at 30% sequence identity cut-off. We have shown that the kPRED method was able to achieve mean accuracies ranging from 40.8% to 66.3% depending on the availability of homologues. The impact of the different strategies for scanning the database on the prediction was evaluated and is discussed. Our results highlights the usefulness of the method in the context of proteins without any known structural homologues. A scoring function that gives a good estimate of the accuracy of prediction was further developed. This score estimates very well the accuracy of the algorithm (R2 of 0.82). An online version of the tool is provided freely for non-commercial usage at http://www.bo-protscience.fr/kpred/.
A joint embedding of protein sequence and structure enables robust variant effect predictions
Open the record for dataset details and reuse information.
Sequence data and structural data utilized in the study and analysis of grain protein function prediction.
Open the record for dataset details and reuse information.
Conformation Database for Publication: Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction
<p><strong>Conformation database</strong> for 2022 Publication "Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction"</p> <ul> <li>DOI of Physica A publication: <a href="https://doi.org/10.1016/j.physa.2022.128395">https://doi.org/10.1016/j.physa.2022.128395</a></li> <li>GitHub source code: <a href="https://github.com/CompSoftMatterBiophysics-CityU-HK/Applying-DRL-to-HP-Model-for-Protein-Structure-Prediction">https://github.com/CompSoftMatterBiophysics-CityU-HK/Applying-DRL-to-HP-Model-for-Protein-Structure-Prediction</a></li> </ul> <p>This conformation database shows the distinct conformations of best-known and next best energies:</p> <p>├── <strong>20merA</strong><br> │ ├── <strong>20merA_E8_set</strong><br> │ ├── <strong>20merA_E9_set</strong><br> │ ├── confs_20merA_E8.txt<br> │ └── confs_20merA_E9.txt<br> ├── <strong>20merB</strong><br> │ ├── <strong>20merB_E10_set</strong><br> │ ├── <strong>20merB_E9_set</strong><br> │ ├── confs_20merB_E10.txt<br> │ └── confs_20merB_E9.txt<br> ├── <strong>24mer</strong><br> │ ├── <strong>24mer_E8_set</strong><br> │ ├── <strong>24mer_E9_set</strong><br> │ ├── confs_24mer_E8.txt<br> │ └── confs_24mer_E9.txt<br> ├── <strong>25mer</strong><br> │ ├── <strong>25mer_E7_set</strong><br> │ ├── <strong>25mer_E8_set</strong><br> │ ├── confs_25mer_E7.txt<br> │ └── confs_25mer_E8.txt<br> ├── <strong>36mer</strong><br> │ ├── <strong>36mer_E13_set</strong><br> │ ├── <strong>36mer_E14_set</strong><br> │ ├── confs_36mer_E13.txt<br> │ └── confs_36mer_E14.txt<br> ├── <strong>48mer</strong><br> │ ├── <strong>48mer_E22_set</strong><br> │ ├── <strong>48mer_E23_set</strong><br> │ ├── confs_48mer_E22.txt<br> │ └── confs_48mer_E23.txt<br> └── <strong>50mer</strong><br> ├── <strong>50mer_E20_set</strong><br> ├── <strong>50mer_E21_set</strong><br> ├── confs_50mer_E20.txt<br> └── confs_50mer_E21.txt</p>
Mapping Synthetic Binding Proteins Epitopes on Diverse Protein Targets by Protein Structure Prediction and Protein-Protein Docking
<p>The predicted 3D structures of 145 SBPs and the 96 models of SBPs in complex with protein targets.</p>
Data for CASP15 performance benchmarking of the state-of-the-art protein structure prediction methods
<p>CASP15 performance benchmarking of the state-of-the-art protein structure prediction methods</p>
Automated protein-protein structure prediction of the T cell receptor-peptide major histocompatibility complex
Open the record for dataset details and reuse information.
Data from: Knowledge-based prediction of protein backbone conformation using a structural alphabet
Open the record for dataset details and reuse information.
Structure prediction of protein-ligand complexes from sequence information with Umol
<p>posebusters_benchmark_set.tar.zst - files for the prediction (features to Umol) and scoring of the pose busters benchmark </p><p>posebusters_pred_native.tar.zst - pdb and sdf files of proteins and ligands. Includes native structures, predicted structures and relaxed predicted structures with plDDT in the B factor column.</p><p>posebusters_scores.csv - contains ligand RMSD and other metrics for the unrelaxed structures predicted with Umol.</p><p>PDBBind_processed.tar.zst - files for the training (features to Umol) using PDBbind version 2020</p><p> </p><p> </p>
Dataset from: Structure-based prediction of protein-nucleic acid binding using graph neural networks
<p>Datasets used for training/evaluating the neural network models described in our article. Additional documentation related to how these datasets were constructed can be found in the github repository https://github.com/jaredsagendorf/pnabind/tree/master/datasets</p>
A joint embedding of protein sequence and structure enables robust variant effect predictions
<p>Data related to the GitHub repository KULL-Centre/_2023_Blaabjerg_SSEmb, which is also stored on Zenodo here: <span><span><a href="../doi/10.5281/zenodo.13765792" target="_blank" rel="noopener noreferrer">https://zenodo.org/doi/10.5281/zenodo.13765792</a>.</span></span></p>
Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction, 2014
<p>This contains the protein sequence and secondary structure dataset from <a href="https://proceedings.mlr.press/v32/zhou14.html"><strong>Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction</strong></a><strong>, ICML, 2014</strong></p> <p>This dataset was originally hosted at http://www.princeton.edu/~jzthree/datasets/ICML2014/. Since the original URL is no longer available and the dataset is still used by many, I moved the dataset here.</p>
RBP Footprint Grand Challenge: An evaluation of novel computational approaches to RNA-binding protein target prediction from structural data
GEO Series GSE227455. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing; Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.