Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “protein structure prediction”
Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15
<p>Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15</p>
Automated benchmarking of combined protein structure and ligand conformation prediction
<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB. </p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>
Host-pathogen protein interactions predicted using structure
<p>This dataset accompanies a manuscript describing a method to predict host-pathogen protein interactions using structure:</p> <p>Host-pathogen protein interactions predicted by comparative modeling.<br /> Davis FP, Barkan DT, Eswar N, McKerrow JH, Sali A. Protein Sci (2007) 16:2585-2596.<br /> http://www.proteinscience.org/cgi/doi/10.1110/ps.073228407</p> <p>The files contain predictions made for 10 human pathogens including species of Mycobacterium, Apicomplexa, and Kinetoplastida. The species.all.zip files contain all interactions predictions for each species along with the filter criteria that each interaction passed. The species.filter.zip files contains the same information, but only for the subset of interactions that passed the biological and network-level filters. These files can be viewed in any spreadsheet program, such as Excel.</p> <p> </p> <p>Predictions were made for interactions between human and </p> <ol> <li>Mycobacterium tuberculosis: mtuber</li> <li>Mycobacterium leprae: mleprae</li> <li>Leishmania major: lmajor</li> <li>Trypanosoma brucei: tbrucei</li> <li>Trypanosoma cruzi: tcruzi</li> <li>Cryptosporidium hominis: chominis</li> <li>Cryptosporidium parvum: cparvum</li> <li>Plasmodium falciparum: pfalciparum</li> <li>Plasmodium vivax: pvivax</li> <li>Toxoplasma gondii: tgondi</li> </ol>
Predicted structures of SLC-protein complexes and controls
<p>Solute carrier (SLC) transporters form a major superfamily which transport a wide range of substrates across cellular and organellar membranes. To study their regulation on protein level and position SLCs in the human interactome, we conducted a large-scale interrogation of the protein-protein interactions (PPIs) of SLCs employing affinity purification combined with mass spectrometry (AP-MS). This study resulted in thousands of novel protein interactions of SLCs. For a subset of SLC protein complexes, we performed structural predictions using AlphaFold (v2.2, v2.3 and v3). The dataset attached contains the structures in PDB/CIF format. The structures were further discussed in the associated manuscript. In addition, an annotation table is provided, which summarizes the scores for each modelled complex. </p>
A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2
<p>Data produced and analyzed in the manuscript "A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2" by Pasquadibisceglie et al.</p> <p><br> If you include these data in your manuscript, please cite: Pasquadibisceglie A, Leccese A and Polticelli F (2022) A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2. <em>Front. Chem.</em> 10:1004815. doi: 10.3389/fchem.2022.1004815</p>
Simulated and experimental data distributed to the CASP13 participants in protein structure prediction assisted with sparse NMR data
<p>All simulated and experimental data distributed to the CASP participants in protein structure prediction assisted with sparse NMR data in CASP13.</p> <p>Also available at http://predictioncenter.org/casp13/results.cgi?view=targets&model=first&tr_type=others&sub_type=N&groups_id=</p> <p> </p>
Trypanosoma brucei predicted protein structures, part 2 of 2
<p>AlphaFold2-predicted protein structures for the <em>Trypanosoma brucei</em> (TREU927) proteome, predicted using input multiple sequence alignments optimised for the Discoba lineage in which <em>T. brucei </em>sits. The structure prediction methodology was exactly as described in <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p> <p>This deposition contains part 2 of 2. To get the full dataset, also download TbruceiTREU927_part1.zip from <a href="https://zenodo.org/record/7940748">doi:10.5281/zenodo.7940748</a>.</p> <p>Data are organised with one directory per <em>T. brucei </em>TREU927 gene ID (eg. Tb927.1.3600). Within each directory you will find:</p> <p><strong><gene id>_predmap.png</strong> A left to right representation of the linear protein sequence with one pixel per amino acid. Each horizontal bar represents one structure prediction, colour coded by pLDDT. For small proteins, there will likely be one prediction of the full-length protein. For large proteins, there may be many overlapping predictions.</p> <p><strong><gene id>_<start aa>-<end_aa> </strong>A directory containing structure prediction of that gene ID between the start and end amino acid. Within this directory you will find:</p> <p><strong><gene id>_<start aa>-<end_aa>.json</strong> The full data in a JSON format, including linear protein sequence and metadata, along with 5 predicted protein structures ranked from best to worst overall pAE. For each predicted protein structure, the structure (PDB format), its pLDDT per residue and pairwise pAE.</p> <p><strong><gene id>_<start aa>-<end_aa>_1.pdb</strong> The PDB file of the highest ranked structure.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-pae.png</strong> A plot of pAE, for the highest ranked structure, at one pixel per amino acid. Shade of green represents pAE for that amino acid pair, see below.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-plddt.png</strong> A plot of pLDDT, for the highest ranked structure, at one horizontal pixel per amino acid. Bar height and colour both represent pLDDT for that amino acid, see below.</p> <p>PDB structure and pAE/pLDDT of lower ranked models are embedded in the JSON file.</p> <p>All pLDDT and pAE plots use the colour scales used by https://alphafold.ebi.ac.uk/: For pLDDT: > 90 (dark blue), 90 > pLDDT > 70 (light blue), 70 > pLDDT > 50 (orange), < 50 (yellow) discontinuous. For pAE: 0 (white angstrom) to dark green (32 angstrom) continuous.</p> <p>If you use this resource, please cite this Zenodo deposition and <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p>
Trypanosoma brucei predicted protein structures, part 1 of 2
<p>AlphaFold2-predicted protein structures for the <em>Trypanosoma brucei</em> (TREU927) proteome, predicted using input multiple sequence alignments optimised for the Discoba lineage in which <em>T. brucei </em>sits. The structure prediction methodology was exactly as described in <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p> <p>This deposition contains part 1 of 2. To get the full dataset, also download TbruceiTREU927_part2.zip from <a href="http://zenodo.org/record/7948119">10.5281/zenodo.7948119</a>.</p> <p>Data are organised with one directory per <em>T. brucei </em>TREU927 gene ID (eg. Tb927.1.3600). Within each directory you will find:</p> <p><strong><gene id>_predmap.png</strong> A left to right representation of the linear protein sequence with one pixel per amino acid. Each horizontal bar represents one structure prediction, colour coded by pLDDT. For small proteins, there will likely be one prediction of the full-length protein. For large proteins, there may be many overlapping predictions.</p> <p><strong><gene id>_<start aa>-<end_aa> </strong>A directory containing structure prediction of that gene ID between the start and end amino acid. Within this directory you will find:</p> <p><strong><gene id>_<start aa>-<end_aa>.json</strong> The full data in a JSON format, including linear protein sequence and metadata, along with 5 predicted protein structures ranked from best to worst overall pAE. For each predicted protein structure, the structure (PDB format), its pLDDT per residue and pairwise pAE.</p> <p><strong><gene id>_<start aa>-<end_aa>_1.pdb</strong> The PDB file of the highest ranked structure.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-pae.png</strong> A plot of pAE, for the highest ranked structure, at one pixel per amino acid. Shade of green represents pAE for that amino acid pair, see below.</p> <p><strong><gene id>_<start aa>-<end_aa>_1-plddt.png</strong> A plot of pLDDT, for the highest ranked structure, at one horizontal pixel per amino acid. Bar height and colour both represent pLDDT for that amino acid, see below.</p> <p>PDB structure and pAE/pLDDT of lower ranked models are embedded in the JSON file.</p> <p>All pLDDT and pAE plots use the colour scales used by https://alphafold.ebi.ac.uk/: For pLDDT: > 90 (dark blue), 90 > pLDDT > 70 (light blue), 70 > pLDDT > 50 (orange), < 50 (yellow) discontinuous. For pAE: 0 (white angstrom) to dark green (32 angstrom) continuous.</p> <p>If you use this resource, please cite this Zenodo deposition and <a href="https://doi.org/10.1371/journal.pone.0259871">doi:10.1371/journal.pone.0259871</a>.</p>
Alphafold predicted structures of VPS13 proteins from model organisms
<p>This upload contains AlphaFold-predicted structures of VPS13 proteins from a variety of organisms. Given the large size of these proteins, only partial sequences were predicted with AlphaFold(1) and the resulting structures were aligned in PyMOL(2). A summary of the structures uploaded here is presented as a collection of domain cartoons in the "VPS13 domain organization across eukaryotic evolution.pdf" file. </p> <p>The structures were generated with AlphaFold v2.029 on the Yale High Performance Cluster. Each *.zip file contains the best ranked predictions (out of five) for each sequence (*.pdb files) and the PyMOL assembled full structure (*.pse file). In a few cases, where a good alignment was not possible due to long disordered regions in the C-terminal portions (mostly in proteins from <em>D. discoideum</em> and <em>A. thaliana</em>), the full structures were aligned manually in PyMOL based on the continuity of the lipid transfer groove. The structures in PyMOL can be colour-coded by the confidence value of AlphaFold predictions using the following prompt:</p> <p>set_color n0, [0.051, 0.341, 0.827]<br> set_color n1, [0.416, 0.796, 0.945]<br> set_color n2, [0.996, 0.851, 0.212]<br> set_color n3, [0.992, 0.490, 0.302]<br> color n0, b < 100; color n1, b < 90<br> color n2, b < 70; color n3, b < 50</p> <p>Considering that full length structures were assembled by aligning different protein fragments and in view of the presence of flexible loops with low prediction confidence scores, the relative positions of different folded domains are not necessarily correct.</p> <p> </p> <p><strong>References</strong></p> <p>1. J. Jumper, <em>et al.</em>, Highly accurate protein structure prediction with AlphaFold. <em>Nature</em> 596, 583–589 (2021).</p> <p>2. The PyMOL Molecular Graphics System, Version 2.0. Schrödinger LLC.</p>
Dataset for Peptide binder design with inverse folding and protein structure prediction
<p>Dataset for a paper on peptide design</p> <p> </p> <p><br> mutated_peptides - results for randomly intriduced mutations in protein-peptide complexes that can be predicted at 2 Å (Figure 1)<br> pdb_peptide - variation in the number of recycles (1-10) for 96 peptides (Figure 1)<br> minibinder - results for the minibinder set (Figure 2)<br> Pfam - results for the Pfam set (Figures 4+5)<br> protein_mpnn - results on protein_mpnn test set (Figure 6)</p> <p> </p> <p> </p>
DynamicBind: Predicting ligand-specific protein-ligand complex structure with a deep equivariant generative model.
<p>test and training data.</p>
Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)
<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>
Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)
<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>
predicted_protein_complex_structures_datasets
Open the record for dataset details and reuse information.
AlphaFold2 predicted structures of ThsA and ThsB proteins
<p>This Zenodo record contains the AlphaFold2 models described in the manuscript: Structural characterization of macro domain-containing Thoeris antiphage defense systems</p>
Fueling ab initio folding with oceanic metagenomics enables structure and function predictions of new protein families
<p>Code and protein sequence database to construct multiple sequence alignment from Tara Ocean data.</p>
Protein structure data for "AI-predicted protein deformation encodes energy landscape perturbation"
<p>AF2-predicted protein structures of WT and mutant proteins that have corresponding ddG measurements in the ThermoMutDB database of protein mutant stability measurements. PDB structures are compressed using <a href="https://github.com/steineggerlab/foldcomp/">FoldComp</a>, and saved in "structures.zip".</p> <p>Summary of the final dataset and results can be found in "results_summary.pkl".</p> <p>Code used to plot figures can be found in "code4figs.zip".</p>
Exploring zero-shot structure-based protein fitness prediction
<p>This repository contains data used in Exploring zero-shot structure-based protein fitness<br>prediction.</p> <p>Directions to use this data can be found on <a href="https://github.com/gitter-lab/benchmarking-structure-based-models">our GitHub repository</a>.</p> <ol> <li><code>experimental_struct_artifacts</code> contains the experimentally determined structures for ProteinGym assays used in our analysis along with the reference file needed to generate ESM inverse folding predictions for these structures in ProteinGym.</li> <li><code>results</code> contains the prediction results obtained by running SSEmb on the 216 ProteinGym assays being considered in this study.</li> <li><code>test.tar.gz</code> contains all the structures from ProteinGym as well as MSAs generated using mmseqs2. To use this directory: <ul> <li>Setup SSEmb as directed in its <a href="https://github.com/KULL-Centre/_2023_Blaabjerg_SSEmb">repository</a></li> <li>Download this file and extract it in the data folder.</li> </ul> </li> </ol> <p> </p> <p> </p>
Predicted structures of the periplasmic adaptor protein, CmeA
<p>Predicted structures for CmeA, sequence alignments, and a fully assembled model for CmeABC (Chapter 6 of Kahlan Newman's Doctoral Thesis). </p>
Structure prediction from SARS-CoV-2 accessory proteins ORF-6
<p>Structure prediction made with Collabfold for SARS-CoV-2 accessory protein ORF-6.</p> <p>The archive contains both the structure and</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.