Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “protein-protein docking”

Learn how ShareScore rates datasets ↗
zenodo36/100

Illustration of Protein-Protein Docking

<p>Conceptual representation of protein-protein docking. Two proteins, represented as surfaces, moving towards each other with a computer server rack in the background. Background image sourced from Wikimedia, original author Victorgrigas.</p>

openother-openNov 2020View details →
zenodo36/100

Protein-protein docking with significant backbone flexibility

<p>The trajectory of a single replica from the protein-protein docking simulation of barnase/barstar system. The movie shows the barnase receptor in surface representation and the barstar ligand in ribbon. The presented replica reached the model with interface RMSD value 1.9 Angstrom from the complex X-ray structure, shown as transparent ribbon.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Protein-Protein Docking for PROTAC discovery

<p>My goal is to test whether structure-based approaches can guide the design of PROTACs. As a first step, I am evaluating whether protein-protein docking tools can accurately predict the interface between an E3 ligase and its target, a question recently explore by Drummond et al.&nbsp;This step is necessary to define the relative orientation of the chemical moiety binding the E3 ligase and the chemical moiety binding the target protein. Once we have this information, the second step will be to design PROTACs that are compatible with the relative orientation of these 2 chemical moieties. For this first protein-protein docking step, I will compare predicted structures with crystal structures of three complexes: the first bromodomain of BRD4 (BRD4<sup>BD1</sup>) bound to the E3 ligases CRBN [pdb codes 6boy and 6bnb], the second bromodomain of BRD4 (BRD4<sup>BD2</sup>) bound to the E3 ligase VHL [pdb code 5t35] and, the bromodomain of SMARCA2<sup>BD</sup> bound to the E3 ligase VHL [pdb code 6hay]. I will be using three different protein-protein docking tools: HADDOCK, Rosetta, and ICM, which all performed among the best at past CAPRI protein docking competitions.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks

<p><strong>FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks.&nbsp;</strong></p> <p>FTDMP is a software system for running docking experiments and scoring/ranking multimeric models. This dataset contains FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks. The FTDMP framework itself is available at https://github.com/kliment-olechnovic/ftdmp.&nbsp;</p> <p>Every *.tar.gz file in this dataset contains two folders: results for unbound-unbound and bound-bound docking. These folders contain results for the benchmark cases:</p> <p>252 folders with results for the protein-protein docking benchmark cases [1].<br>47 folders with results for the protein-DNA docking benchmark cases [2].<br>42 folders with results for the protein-RNA docking benchmark cases [3-6].&nbsp;</p> <p>Every folder is named according to the PDB ID of the complex. The folders contain:</p> <p>1. A subfolder named <em>relaxed_top_complexes</em>. This subfolder contains 200 pdb files of relaxed [7] top docking models.<br>2. A text file named <em>scoring_results-ranks.txt</em>. It contains the names of the models (that are in the relaxed_top_complexes folder) in the ranked order. This means that the first model in the file is considered to be the best prediction by the FTDMP framework.<br>3. A text file named <em>cad_scores.txt</em>. It contains interface CAD-score and binding site CAD-score [8] results for every model.<br>4. A text file named <em>rmsd_results.txt</em>, which is available only for protein-DNA and protein-RNA cases. The file contains ligand-RMSD values for the models, where the DNA/RNA is considered as the ligand.<br>5. A text file named <em>DockQ_results.txt</em>, which is available only for the protein-protein docking cases. The file contains DockQ [9] results for every model, as well as model accuracy based on CAPRI criteria (Incorrect, Acceptable, Medium, High)<br>6. A text file named <em>binding_site_CAD-scores.txt</em>, which contains the binding site CAD-score from the <strong>protein</strong> side for RNA and DNA docking. This binding site CAD-score shows how accurately the ligand (DNA/RNA) was docked to the protein without taking the orientation of the ligand into consideration. In the case of protein-protein docking the binding site CAD-score file is available only for antibody-antigen docking targets and contains the binding site (epitope) CAD-score for the antigen. &nbsp;</p> <p>The ligand-RMSD, CAD-scores, and DockQ scores were all calculated by comparing the models to the corresponding targets. The target structures are available at <span>https://zenodo.org/records/10517524</span>. These target structures have the same residue numbering as the models available here.&nbsp;</p> <p>REFERENCES&nbsp;</p> <p>[1] Guest, J. D., et al. (2021). An expanded benchmark for antibody-antigen docking and affinity prediction reveals insights into antibody recognition determinants. Structure, 29(6), 606&ndash;621.e5.<br>[2] van Dijk, M., Bonvin, A.M. (2008). A protein-DNA docking benchmark. Nucleic Acids Res, 36, e88.&nbsp;<br>[3] Perez-Cano, L., et. Al. (2012). A protein-RNA docking benchmark (II): extended set from experimental and homology modeling data. Proteins, 80(7): 1872-1882.&nbsp;<br>[4] Huang, S.Y., Zou, X. (2013). A nonredundant structure dataset for benchmarking protein-RNA computational docking. J Comput Chem, 34(4): 311-318.&nbsp;<br>[5] Nithin, C., et. al. (2017). A non-redundant protein-RNA docking benchmark version 2.0. Proteins, 85(2) :256-267.&nbsp;<br>[6] Zheng, J., et al. (2020). P3DOCK: a protein-RNA docking webserver based on template-based and template-free docking. Bioinformatics, 36(1), 96&ndash;103.&nbsp;<br>[7] Eastman, P., et al.(2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Comp. Biol., 13(7): e1005659. &nbsp;<br>[8] Olechnovic, K., Venclovas, C. (2020). Contact area-based structural analysis of proteins and their complexes using CAD-score. Methods Mol Biol, 2112, 75.<br>[9] Basu, S., Wallner, B. (2016). DockQ: A Quality Measure for Protein-Protein Docking Models. PLoS ONE 11(8): e0161879.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Molecular dynamics simulations of 20 complexes from the Protein-Protein Docking Benchmark

<p>We selected 20 complexes from the Protein-Protein Docking Benchmark 5.0 dataset based on structure resolution and parameterization difficulty. For each complex, we conducted a standard 1 &micro;s-long molecular dynamics (MD) simulation in the NPT ensemble (at 1 atm and 300 K, following a 2 ns NVT equilibration) for the bound receptor, unbound receptor, bound ligand and unbound ligand. We set up all systems using Amber ff14SB<sup> </sup>and its recommended TIP3P water model, running MD simulations with Amber 16. For the 80 (single chain structure) MD, we sampled 500 frames for each simulation and computed the average prediction confidence.</p>

opencc-by-4.0Jul 2024View details →
zenodo28/100

Flexible Protein-Protein Docking Benchmark(FD1.0)

<p>To effectively assess the capabilities of various methods in flexible protein-protein docking, it is essential for a protein-protein docking dataset to encompass not only the structures of the heterodimer but also that of unbound monomers. Existing datasets such as DB5.5 and AB-Benchmark, while useful, are relatively limited in scale. In contrast, the Database of Interacting Protein Structures (DIPS) contains up to 42,826 binary protein complex structures but lacks the unbound state structures of the monomers. This limitation restricts its applicability to evaluations of rigid docking models rather than flexible ones. Consequently, the impact of large-scale docking datasets on methods for flexible protein-protein docking has not been thoroughly explored. To address this gap, we introduce the Flexible Protein-Protein Docking Benchmark (FD1.0), which, to our knowledge, is currently the largest dataset dedicated to flexible protein-protein docking. By providing a large and well-characterized dataset, FD1.0 aims to foster innovation in the development of flexible docking algorithms. It allows researchers to rigorously test and refine their methods, facilitating more accurate predictions of protein interactions, which are essential for understanding biological functions and designing therapeutic interventions.</p> <p>In our analysis of the DIPS dataset, we identified several critical issues: (1) Multiple three-dimensional structures correspond to a single protein sequence, introducing substantial noise and affecting fair comparisons among baselines, especially for models reliant on 3D structural data. (2) The DIPS training set, primarily consisting of homo-multimers, fails to capture the diversity of interface types fully. Moreover, protein-protein docking predictions are most valuable for elucidating mechanisms of protein-protein interactions (PPIs), which predominantly involve heterodimers. Homomers, often synthesized directly rather than through docking, do not accurately represent typical PPI scenarios. (3) A significant number of docking cases in DIPS involve the interaction of one polymeric protein with another, further complicating the dataset.</p> <p>As a cornerstone for the flexible docking dataset, it is imperative to acquire the structures of protein monomers in their unbound state. Specifically, this can be achieved through protein structure prediction methods, such as AlphaFold2, and the aggregation of structural data from sources including electron microscopy. Additionally, acknowledging the deficiencies of the DIPS dataset, several guidelines were established in the construction process of the FD1.0 dataset: (1) Each protein monomer is associated with a unique three-dimensional structure, reducing dataset noise. (2) We ensured that the similarity score (as determined by MMSeq) between docking monomers does not exceed 0.6, thereby filtering out homodimeric pairs from the dataset. (3) Unlike DIPS, a certain proportion of cases in the which dataset actually involve docking of two protein multimer. Current methods for predicting multimeric structures, such as AlphaFold Multimer, still do not achieve satisfactory results (AlphaFold3's license prohibits its use for docking purposes). However, current methods for predicting monomeric structures have reached a high level of accuracy. Therefore, we filtered out such cases, ensuring that each docking instance involves only protein monomers, guaranteeing the quality of the dataset. By adhering to these standardized construction criteria and through the collection, cleaning, and organization of data from various sources, including the Protein Data Bank and existing datasets, we compiled 3721 entries. Following the DIPS division ratio, these entries were divided into training, validation, and test sets of 3546, 98, and 77, respectively.</p>

opencc-by-4.0Oct 2024View details →
zenodo28/100

Mapping Synthetic Binding Proteins Epitopes on Diverse Protein Targets by Protein Structure Prediction and Protein-Protein Docking

<p>The predicted 3D structures of 145 SBPs and the 96 models of SBPs in complex with protein targets.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record