Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “Protein structure”
Memory Effects in a Random Walk Description of Protein Structure Ensembles
<p>This <a href="https://www.activepapers.org/">ActivePaper </a>file contains all the code and data that was used in generating the figures for the article "Memory Effects in a Random Walk Description of Protein Structure Ensembles" by Gerald R Kneller and Konrad Hinsen, J. Chem. Phys. <strong>150</strong>, 064911 (2019); <a href="https://doi.org/10.1063/1.5054887">https://doi.org/10.1063/1.5054887</a></p>
Extended data for the paper "Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction"
<p>Extended data for the paper:<br> Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction</p> <p>Authors:<br> Shaun M Kandathil, Mario Garza-Fabre, Simon C Lovell and Julia Handl</p> <p>--------------------------------</p> <p>Contents of the zip file:</p> <p> </p> <p>Directory 'ECDFplots':<br> ----------------------<br> Data corresponding to Figure 3 for all targets, for the bilevel and ILS protocols. Data are available following stages 3 and 4 of the low-resolution protocol.</p> <p>Directory 'ScoreRMSDplots_3archivers':<br> --------------------------------------<br> Data corresponding to Figures 6 and 9 for all targets. Data corresponding to decoys obtained after low-resolution stages 3 and 4 can be found in subdirectories 'Stage3' and 'Stage4', respectively.<br> </p>
Figure 2 in Solution structure of the first RRM domain of human spliceosomal protein SF3b49
Figure 2. – Fresh specimen of Lutjanus madras (UPVMI 1084, 211.3 mm SL, Panay Island, Republic of the Philippines).
Figure 3 in Solution structure of the first RRM domain of human spliceosomal protein SF3b49
Figure 3. – Diagrams of the first gill arch of the right side of Lutjanus madras. A: UPVMI 1084; B: MUFS 46214 (226.0 mm SL, Beruwala, Sri Lanka). Arrowhead and arrows indicate soft flesh-like mass and rudimentary gill rakers, respectively.
DPCstruct Classification of AlphaFold2-Predicted Protein Structures
<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong> </span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us: </p> <p>federico.barone@areasciencepark.it</p>
Initial Structures of PKM1/M2 proteins for AMOEBA Molecular Dynamics studies (xyz Tinker format)
<p>Here are presented our initial structures of PKM1/M2 (solvated and neutralized) for the different states to initiate molecular dynamics in AMOEBA force field.</p> <p>Those are represented in xyz Tinker format and come from their respectives PDB crystal structure after extraction of the unwanted ligands :</p> <p>3SRF for PKM1,</p> <p>1ZJH for monomer PKM2,</p> <p>6B6U for dimer PKM2,</p> <p>3SRH for free-tetramer PKM2,</p> <p>3SRD for tetramer PKM2 bound to FBP,</p> <p>3U2Z for tetramer PKM2 bound to TEPP-46.</p>
AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes
<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span> </span>alanine<span> </span>sequence masking<span> </span>with shallow multiple sequence alignment<span> </span>subsampling to expand the conformational diversity of the predicted structural<span> </span>ensembles and<span> </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span> </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span> </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span> </span>the view </span><span> </span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span> </span>a significant <span>correlation </span>between conformational flexibility and <span> </span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span> </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span> </span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span> </span></p>
Data for "Training data composition affects performance of protein structure analysis algorithms" by A. Derry, K. A. Carpenter, & R. B. Altman
<p><strong>Description</strong></p> <p>This repository contains all data used in "Training data composition affects performance of protein structure analysis algorithms", published in the Pacific Symposium on Biocomputing 2022 by A. Derry, K. A. Carpenter, & R. B. Altman. </p> <p>The data consists of the following files:</p> <ul> <li>ema_zenodo_data.tar.gz: train, validation, and test splits for Estimation of Model Accuracy task, in LMDB format</li> <li>design_zenodo_data.tar.gz: train, validation, and test splits for Protein Sequence Design task, in JSON format</li> <li>enz_cat_res_zenodo_data.tar.gz: train, validation, and test splits for Catalytic Residue and Enzyme Prediction task, in TF record format</li> </ul> <p>Details on dataset construction can be found in our paper and dataloaders can be found in our <a href="https://github.com/awfderry/ml-structure-bias">Github repo</a>.</p> <p><strong>Reference</strong></p> <p>A. Derry*, K. A. Carpenter*, & R. B. Altman, "Training data composition affects performance of protein structure analysis algorithms", 2021.</p> <p><strong>Dataset References</strong></p> <p>Datasets used were derived from the following works:</p> <p>Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K., & Moult, J. (2019). Critical assessment of methods of protein structure prediction (CASP)—Round XIII. In <em>Proteins: Structure, Function and Bioinformatics</em> (Vol. 87, Issue 12, pp. 1011–1020). https://doi.org/10.1002/prot.25823</p> <p>Ingraham, J., Garg, V. K., Barzilay, R., & Jaakkola, T. (2019). <em>Generative Models for Graph-Based Protein Design</em>. https://openreview.net/pdf?id=SJgxrLLKOE</p> <p>Furnham, N., Holliday, G. L., de Beer, T. A. P., Jacobsen, J. O. B., Pearson, W. R., & Thornton, J. M. (2014). The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. <em>Nucleic Acids Research</em>, <em>42 </em>(Database issue), D485–D489.</p>
Protein structure model predictions for secreted fungal proteins
<p><strong>Dataset A - Alphafold2 prediction output data for 753 secreted proteins of <em>Rhizophagus irregularis </em>DAOM197198</strong>. Gene IDs are taken from the annotation by Yildirir et al. 2021, <a href="https://doi.org/10.1111/nph.17842">doi.org/10.1111/nph.17842</a></p> <p><strong>Dataset B - Alphafold2 prediction output data for 10 fungal effectors.</strong><strong> </strong>These are nine effectors from <em>Fusarium oxysporum</em> f. sp.<em> lycopersici</em> and RiSLM from <em>Rhizophagus irregularis</em> as well as their amino acid sequences. Signal peptides and sequences preceding a predicted Kex2 processing site were removed.</p> <p><strong>Dataset C - Alphafold2 prediction output data for 454 matches of a MycFOLD-HMM search</strong> across the Mycocosm genome database (<a href="https://mycocosm.jgi.doe.gov/mycocosm/home">https://mycocosm.jgi.doe.gov/mycocosm/home</a>) and 36 Glomeromycotina fungal genomes.</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases
<p>The files included here are a set of Galaxy workflows, starting structure files (PDB, mol2, and frcmod), and specialized force field files (ZAFF) for the simulation of coronavirus helicases in the apo and drug-bound state. The inhibitor molecules include those from virtual screening (FCID1 and thioguanine), as well as experimentally validated candidates (Lumacaftor and SSYA10-001).</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Flavivirus Helicases
<p>The files included here are a set of Galaxy workflows and starting structure files (PDB, mol2, and frcmod) for the simulation of flavivirus helicases in the apo and drug-bound state. The inhibitors include the 4th highest ranking compound from a virtual screening of more than 12.7 million drug-like molecules.</p>
Supplementary data and code to "An assessment of quaternary structure functionality in homomer protein complexes" by G. Abrusan and C. Foguet, https://doi.org/10.1093/molbev/msad070
<p>Scripts and high-level data to reproduce the figures and supplementary figures of "An assessment of quaternary structure functionality in homomer protein complexes" by G. Abrusan and C. Foguet, https://doi.org/10.1093/molbev/msad070</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases -- Output Files
<p>These are the output files generated using the input files and Galaxy workflows for coronavirus helicase simulations, from: </p> <pre>https://doi.org/10.5281/zenodo.7492987</pre>
Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates
<p>Solid-state NMR spectra (DARR and NCA) of rehydrated freeze-dried free TTR and TTR in the presence of Tafamidis and Taf-PTX</p> <p>Reference citation: Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates. Angew Chem Int Ed Engl. 2023 Jun 5:e202303202. doi: 10.1002/anie.202303202. PMID: 37276329.</p>
FireProtDB + PDB Structural Protein Stability Dataset
<p>Dataset compiled and curated for use in the ThermoMPNN paper: <a href="https://doi.org/10.1073/pnas.2314853121">https://doi.org/10.1073/pnas.2314853121</a>: </p> <p>Dataset for training models for prediction of thermodynamic stability changes (ddG) of protein point mutations given a wildtype protein structure (PDB) file. Data was assembled by matching sequence-based ddG measurements in <a href="https://loschmidt.chemi.muni.cz/fireprotdb/">FireProtDB</a> to structures from the <a href="https://www.rcsb.org/">RCSB Protein Data Bank </a>(PDB). For details, see the Methods section of our manuscript.</p> <p>Citing this work: If you choose to use this dataset for your own research, please cite this repository and the ThermoMPNN paper: <a href="https://doi.org/10.1073/pnas.2314853121">https://doi.org/10.1073/pnas.2314853121</a>.</p> <p> </p> <p>Contents:</p> <p>pdbs/ directory contains all PDB files</p> <p>csvs/ directory contains all CSVs with mutation data</p> <p>csvs/4_fireprotDB_bestpH.csv is the main (full) dataset file with 3,438 mutations across 100 proteins.</p> <p>csvs/fireprot_splits.pkl contains the dataset splits (train/val/test) used in our study</p> <p>csvs/splits/ contains csvs for each of the splits (train/val/test/homologue-free) indexed from the full dataset csv.</p> <p>Important CSV columns:</p> <ul> <li>pdb_id_corrected: corresponds to the PDB in the pdbs/ directory (after curation and disambiguation)</li> <li>ddG: ddG value for mutation (mutant - WT)</li> <li>wild_type: wild-type amino acid (1-letter code)</li> <li>mutation: mutant amino acid (1-letter code)</li> <li>pdb_position: 0-based index of the mutated residue in the PDB file (may be different from position in the original FireProtDB sequence entry)</li> </ul> <p> </p>
Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15
<p>Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15</p>
Protein Structure Datasets for Protein Workshop
<p>Raw + Processed Datasets used in the ProteinWorkshop Representation Learning Benchmark</p> <p> </p> <p>Includes datasets from:</p> <p>* The Antibody Developability dataset from Chen et al. (https://doi.org/10.1101/2020.06.18.159798)</p> <p>* CATH from Ingraham et al. (https://www.mit.edu/~vgarg/GenerativeModelsForProteinDesign.pdf)</p> <p>* CCPDB datasets from Agrawal et al. (https://doi.org/10.1093/database/bay142)</p> <p>* The Deep Sea Proteins dataset from Sieg et al. (https://doi.org/10.1002/prot.26337)</p> <p>* Reaction Class prediction from Hermosilla et al. (https://doi.org/10.48550/arXiv.2007.06252)</p> <p>* FoldClassification from Hou et al. (https://doi.org/10.1093/bioinformatics/btx780)</p> <p>* MaSIF Site dataset from Gainza et al. (https://doi.org/10.1038/s41592-019-0666-6)</p> <p>* Metal3d Dataset from Duerr et al. (https://doi.org/10.1038/s41467-023-37870-6)</p> <p>* Post-translational Modification Dataset from Yan et al. (https://doi.org/10.1016/j.crmeth.2023.100430)</p>
Automated benchmarking of combined protein structure and ligand conformation prediction
<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB. </p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>
Data from: The structure of an ancient genotype-phenotype map shaped the functional evolution of a protein family
Open the record for dataset details and reuse information.
FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.