Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

347

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

347 results for “Protein structure”

Learn how ShareScore rates datasets ↗
zenodo40/100

Memory Effects in a Random Walk Description of Protein Structure Ensembles

<p>This <a href="https://www.activepapers.org/">ActivePaper </a>file contains all the code and data that was used in generating the figures for the article &quot;Memory Effects in a Random Walk Description of Protein Structure Ensembles&quot; by Gerald R Kneller and Konrad Hinsen, J. Chem. Phys. <strong>150</strong>, 064911 (2019); <a href="https://doi.org/10.1063/1.5054887">https://doi.org/10.1063/1.5054887</a></p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Extended data for the paper "Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction"

<p>Extended data for the paper:<br> Reliable generation of native-like decoys limits predictive ability in fragment-based protein structure prediction</p> <p>Authors:<br> Shaun M Kandathil, Mario Garza-Fabre, Simon C Lovell and Julia Handl</p> <p>--------------------------------</p> <p>Contents of the zip file:</p> <p>&nbsp;</p> <p>Directory &#39;ECDFplots&#39;:<br> ----------------------<br> &nbsp;&nbsp; &nbsp;Data corresponding to Figure 3 for all targets, for the bilevel and ILS protocols. Data are available following stages 3 and 4 of the low-resolution protocol.</p> <p>Directory &#39;ScoreRMSDplots_3archivers&#39;:<br> --------------------------------------<br> &nbsp;&nbsp; &nbsp;Data corresponding to Figures 6 and 9 for all targets. Data corresponding to decoys obtained after low-resolution stages 3 and 4 can be found in subdirectories &#39;Stage3&#39; and &#39;Stage4&#39;, respectively.<br> &nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Figure 2 in Solution structure of the first RRM domain of human spliceosomal protein SF3b49

Figure 2. – Fresh specimen of Lutjanus madras (UPVMI 1084, 211.3 mm SL, Panay Island, Republic of the Philippines).

opencc-by-4.0Dec 2017View details →
zenodo40/100

Figure 3 in Solution structure of the first RRM domain of human spliceosomal protein SF3b49

Figure 3. – Diagrams of the first gill arch of the right side of Lutjanus madras. A: UPVMI 1084; B: MUFS 46214 (226.0 mm SL, Beruwala, Sri Lanka). Arrowhead and arrows indicate soft flesh-like mass and rudimentary gill rakers, respectively.

opencc-by-4.0Dec 2017View details →
zenodo40/100

DPCstruct Classification of AlphaFold2-Predicted Protein Structures

<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong>&nbsp;</span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us:&nbsp;</p> <p>federico.barone@areasciencepark.it</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Initial Structures of PKM1/M2 proteins for AMOEBA Molecular Dynamics studies (xyz Tinker format)

<p>Here are presented our initial structures of PKM1/M2 (solvated and neutralized) for the different states to initiate molecular dynamics in AMOEBA force field.</p> <p>Those are represented in xyz Tinker format and come from their respectives PDB crystal structure after extraction of the unwanted ligands :</p> <p>3SRF for PKM1,</p> <p>1ZJH for monomer PKM2,</p> <p>6B6U for dimer PKM2,</p> <p>3SRH for free-tetramer PKM2,</p> <p>3SRD for tetramer PKM2 bound to FBP,</p> <p>3U2Z for tetramer PKM2 bound to TEPP-46.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes

<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span>&nbsp;</span>alanine<span>&nbsp; </span>sequence masking<span>&nbsp; </span>with shallow multiple sequence alignment<span>&nbsp; </span>subsampling to expand the conformational diversity of the predicted structural<span>&nbsp; </span>ensembles and<span>&nbsp;&nbsp; </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span>&nbsp; </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span>&nbsp; </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span>&nbsp; </span>the view </span><span>&nbsp;</span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span>&nbsp; </span>a significant <span>correlation </span>between conformational flexibility and <span>&nbsp;</span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span>&nbsp; </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span>&nbsp;</span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span>&nbsp; </span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Data for "Training data composition affects performance of protein structure analysis algorithms" by A. Derry, K. A. Carpenter, & R. B. Altman

<p><strong>Description</strong></p> <p>This repository contains all data used in&nbsp;&quot;Training data composition affects performance of protein structure analysis algorithms&quot;, published in the Pacific Symposium on Biocomputing 2022 by A. Derry, K. A. Carpenter, &amp; R. B. Altman.&nbsp;</p> <p>The data consists of the following files:</p> <ul> <li>ema_zenodo_data.tar.gz: train, validation, and test&nbsp;splits for Estimation of Model Accuracy task, in LMDB format</li> <li>design_zenodo_data.tar.gz: train, validation, and test&nbsp;splits for Protein Sequence Design&nbsp;task, in JSON format</li> <li>enz_cat_res_zenodo_data.tar.gz:&nbsp;train, validation, and test&nbsp;splits for Catalytic Residue and Enzyme Prediction task, in TF record format</li> </ul> <p>Details on dataset construction can be found in our paper and dataloaders can be found in our&nbsp;<a href="https://github.com/awfderry/ml-structure-bias">Github repo</a>.</p> <p><strong>Reference</strong></p> <p>A. Derry*, K. A. Carpenter*, &amp; R. B. Altman, &quot;Training data composition affects performance of protein structure analysis algorithms&quot;, 2021.</p> <p><strong>Dataset References</strong></p> <p>Datasets used were derived from the following works:</p> <p>Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K., &amp; Moult, J. (2019). Critical assessment of methods of protein structure prediction (CASP)&mdash;Round XIII. In <em>Proteins: Structure, Function and Bioinformatics</em> (Vol. 87, Issue 12, pp. 1011&ndash;1020). https://doi.org/10.1002/prot.25823</p> <p>Ingraham, J., Garg, V. K., Barzilay, R., &amp; Jaakkola, T. (2019). <em>Generative Models for Graph-Based Protein Design</em>. https://openreview.net/pdf?id=SJgxrLLKOE</p> <p>Furnham, N., Holliday, G. L., de Beer, T. A. P., Jacobsen, J. O. B., Pearson, W. R., &amp; Thornton, J. M. (2014). The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. <em>Nucleic Acids Research</em>, <em>42&nbsp;</em>(Database issue), D485&ndash;D489.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Protein structure model predictions for secreted fungal proteins

<p><strong>Dataset A - Alphafold2 prediction output data for 753 secreted proteins of <em>Rhizophagus irregularis </em>DAOM197198</strong>.&nbsp;Gene IDs are taken from the annotation by Yildirir et al. 2021,&nbsp;<a href="https://doi.org/10.1111/nph.17842">doi.org/10.1111/nph.17842</a></p> <p><strong>Dataset B - Alphafold2 prediction output data for 10 fungal effectors.</strong><strong>&nbsp;</strong>These are nine effectors from&nbsp;<em>Fusarium oxysporum</em>&nbsp;f. sp.<em>&nbsp;lycopersici</em>&nbsp;and RiSLM from&nbsp;<em>Rhizophagus irregularis</em>&nbsp;as well as their amino acid sequences. Signal peptides and sequences preceding a predicted Kex2 processing site were removed.</p> <p><strong>Dataset C - Alphafold2 prediction output data for 454 matches of a MycFOLD-HMM search</strong>&nbsp;across the Mycocosm genome database (<a href="https://mycocosm.jgi.doe.gov/mycocosm/home">https://mycocosm.jgi.doe.gov/mycocosm/home</a>) and 36 Glomeromycotina fungal genomes.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases

<p>The files included here are a set of Galaxy workflows, starting structure files (PDB, mol2, and frcmod), and specialized force field files (ZAFF) for the simulation of coronavirus helicases in the apo and drug-bound state. The inhibitor molecules include those from virtual screening (FCID1 and thioguanine), as well as experimentally validated candidates (Lumacaftor and&nbsp;SSYA10-001).</p>

opencc-zeroDec 2022View details →
zenodo40/100

Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Flavivirus Helicases

<p>The files included here are a set of Galaxy workflows and starting structure files (PDB, mol2, and frcmod) for the simulation of flavivirus helicases&nbsp;in the apo and drug-bound state. The inhibitors include the 4th highest ranking compound from a virtual screening of more than 12.7 million drug-like molecules.</p>

opencc-zeroDec 2022View details →
zenodo40/100

Supplementary data and code to "An assessment of quaternary structure functionality in homomer protein complexes" by G. Abrusan and C. Foguet, https://doi.org/10.1093/molbev/msad070

<p>Scripts and high-level data to reproduce the figures and supplementary figures of &quot;An assessment of quaternary structure functionality in homomer protein complexes&quot; by G. Abrusan and C. Foguet, https://doi.org/10.1093/molbev/msad070</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases -- Output Files

<p>These are the output files generated using the input files and Galaxy workflows for coronavirus helicase simulations, from:&nbsp;</p> <pre>https://doi.org/10.5281/zenodo.7492987</pre>

opencc-zeroApr 2023View details →
zenodo40/100

Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates

<p>Solid-state NMR spectra (DARR and NCA)&nbsp;of rehydrated freeze-dried free TTR and TTR in the presence of Tafamidis and Taf-PTX</p> <p>Reference citation:&nbsp;&nbsp;Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates. Angew Chem Int Ed Engl. 2023 Jun 5:e202303202. doi: 10.1002/anie.202303202. PMID: 37276329.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

FireProtDB + PDB Structural Protein Stability Dataset

<p>Dataset compiled and curated for use in the ThermoMPNN paper: <a href="https://doi.org/10.1073/pnas.2314853121">https://doi.org/10.1073/pnas.2314853121</a>:&nbsp;</p> <p>Dataset for training models for prediction of thermodynamic stability changes (ddG) of protein point mutations given a wildtype protein structure (PDB) file. Data was assembled by matching sequence-based ddG measurements in <a href="https://loschmidt.chemi.muni.cz/fireprotdb/">FireProtDB</a> to structures from the <a href="https://www.rcsb.org/">RCSB Protein Data Bank </a>(PDB). For details, see the Methods section of our manuscript.</p> <p>Citing this work: If you choose to use this dataset for your own research, please cite this repository and the ThermoMPNN paper: <a href="https://doi.org/10.1073/pnas.2314853121">https://doi.org/10.1073/pnas.2314853121</a>.</p> <p>&nbsp;</p> <p>Contents:</p> <p>pdbs/ directory contains all PDB files</p> <p>csvs/ directory contains all CSVs with mutation data</p> <p>csvs/4_fireprotDB_bestpH.csv is the main (full) dataset file with 3,438 mutations across 100 proteins.</p> <p>csvs/fireprot_splits.pkl contains the dataset splits (train/val/test) used in our study</p> <p>csvs/splits/ contains csvs for each of the splits (train/val/test/homologue-free) indexed from the full dataset csv.</p> <p>Important CSV columns:</p> <ul> <li>pdb_id_corrected: corresponds to the PDB in the pdbs/ directory (after curation and disambiguation)</li> <li>ddG: ddG value for mutation (mutant - WT)</li> <li>wild_type: wild-type amino acid (1-letter code)</li> <li>mutation: mutant amino acid (1-letter code)</li> <li>pdb_position: 0-based index of the mutated residue in the PDB file (may be different from position in&nbsp; the original FireProtDB sequence entry)</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15

<p>Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Protein Structure Datasets for Protein Workshop

<p>Raw + Processed Datasets used in the ProteinWorkshop Representation Learning Benchmark</p> <p>&nbsp;</p> <p>Includes datasets from:</p> <p>* The Antibody Developability dataset from Chen et al. (https://doi.org/10.1101/2020.06.18.159798)</p> <p>* CATH from Ingraham et al. (https://www.mit.edu/~vgarg/GenerativeModelsForProteinDesign.pdf)</p> <p>* CCPDB datasets from Agrawal et al. (https://doi.org/10.1093/database/bay142)</p> <p>* The Deep Sea Proteins dataset from Sieg et al. (https://doi.org/10.1002/prot.26337)</p> <p>* Reaction Class prediction from Hermosilla et al. (https://doi.org/10.48550/arXiv.2007.06252)</p> <p>* FoldClassification from Hou et al. (https://doi.org/10.1093/bioinformatics/btx780)</p> <p>* MaSIF Site dataset from Gainza et al. (https://doi.org/10.1038/s41592-019-0666-6)</p> <p>* Metal3d Dataset from Duerr et al. (https://doi.org/10.1038/s41467-023-37870-6)</p> <p>* Post-translational Modification Dataset from Yan et al. (https://doi.org/10.1016/j.crmeth.2023.100430)</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Automated benchmarking of combined protein structure and ligand conformation prediction

<p>The prediction of protein-ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein-Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three-dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time-split test-set is inappropriate for comprehensive PLC evaluation, with state-of-the-art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC-prediction-specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein-ligand interactions.</p> <p>This repository contains:</p> <ol> <li> <p>all_validation_clustering_data.tsv - X-ray validation data and MMSeqs cluster identifiers at different sequence identities for over a million small molecule and ion-binding pockets in the PDB.&nbsp;</p> </li> <li> <p>hqr_dataset.tsv - PDB IDs and ligand information for the high quality representative (HQR) dataset described in the manuscript</p> </li> <li> <p>score_files.tar.gz - Full docking results for all detected pockets for the PDBBind time-split test-set, the HQR dataset, and the subsets of AF models created for both datasets. One file per tool benchmarked with the following columns: Tool, Complex, Pocket, Rank, lDDT-PLI, lDDT-LP, BiSyRMSD, Reference_Ligand, Tool-generated Score</p> </li> <li> <p>errors_all_sets.csv - Report of failures running the pipeline with the following columns: Process, Complex/Ligand/Receptor, Problem</p> </li> </ol>

opencc-by-4.0Sep 2023View details →
dryad40/100

Data from: The structure of an ancient genotype-phenotype map shaped the functional evolution of a protein family

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad40/100

FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling

Open the record for dataset details and reuse information.

publicJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record