Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

641

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

641 results for “Protein interaction”

Learn how ShareScore rates datasets ↗
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"

<p><strong>README file for the project files provided as supporting information to the manuscript &quot;Ligand-protein interactions in lysozyme investigated through a dual-resolution model&quot;</strong></p> <p>February 12, 2020</p> <p>Authors: Raffaele Fiorentini, Kurt Kremer and Raffaello Potestio</p> <p>================================</p> <p>Overview</p> <p>The dataset&nbsp;is organised in three (compressed) subfolders (see the tree diagrams in each section):</p> <p>- annihilation<br> - decoupling<br> - density</p> <p>The figure deltaG_binding_ann_dec_comparison.png shows the results of binding free energy calculations comparing the values obtained both for annihilation and decoupling.</p> <p>The figure deltaG_binding_annih_gromacs_espp.png displays the results for Binding FE, comparing the values obtained in GROMACS and ESPResSo++.</p> <p>The README.pdf file contains detailed information about these folders and their content.</p> <p>================================</p> <p>The &quot;annihilation&quot; folder contains all results concerning the calculation of binding free energy in case of annihilation and it is divided in two parts:&nbsp;</p> <p>- complex<br> - ligand</p> <p>In &quot;complex&quot; are reported the results of Ligand-Protein FE both in ESPResSo++ and GROMACS. All simulations are fully-atomistic.&nbsp;</p> <p>In &quot;ligand&quot; are reported the results of ligand solvation free energy both in ESPResSo++ and GROMACS. All simulations are fully-atomistic.&nbsp;</p> <p>====</p> <p>The &quot;decoupling&quot; folder contains all results concerning the calculation of binding free energy in case of decoupling and it is divided in three parts:&nbsp;</p> <p>- complex-DualRes<br> - complex-FullyAT<br> - ligand</p> <p>In &quot;complex-DualRes&quot; are reported the results of Ligand-Protein FE only in ESPResSo++ (GROMACS cannot do decoupling). The system is simulated in Dual-Resolution. It is possible to find the trajectory files in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-30&quot;.</p> <p>In &quot;complex-fullyAT&quot; are reported the results of Ligand-Protein FE only in ESPResSo++. The system simulated is fully-atomistic. It is possible to find the trajectory file in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-30&quot;.</p> <p>In &quot;ligand&quot; are reported the results of ligand solvation free energy only in ESPResSo++. All simulations are fully-atomistic. It is possible to find the trajectory file in the sub-directories &quot;lambdaindex-0&quot; and &quot;lambdaindex-20&quot;.</p> <p>====</p> <p>The &quot;density&quot; folder contains the data for the tuning of the c parameter of the steric repulsion among residues. This parameter is tuned so that the water density attains the value computed in all-atom simulations.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Scored protein-protein interactions accompanying "A pan-plant protein complex map reveals deep conservation and novel assemblies"

<p><a href="http://plants.proteincomplexes.org/static/data/panplant_cfms_scores_annot.txt.gz">All scored pairwise protein-protein interactions with CF-MS scores (3,076,999 unique pairwise interactions)</a></p> <ul> <li>Description: Scores between Orthogroups with the corresponding CF-MS score and eggNOG generated orthogroup descriptions.</li> <li>Note: Only the highest scoring pairs are considered significant. A CF-MS score &gt;= 0.509 corresponds to 10% FDR, &gt;= 0.207 corresponds to 50% FDR</li> <li>Format: OrthogroupID1 [tab] OrthogroupID2 [tab] Score [tab] Annotation1 [tab] Annotation2</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Human Pleckstrin Homology domain Interacting Protein (PHIP); A Target Enabling Package

<p>SGC Oxford has expressed, purified and crystallized the second bromodomain of PHIP as part of the probe programme. Fragment screening and X-ray crystallography identified binders, some of which optimised to uM affinity. However, molecules with probe properties were not obtained. Consequently it has been decided to put the information generated into the public domain.</p>

opencc-by-4.0Jun 2016View details →
zenodo44/100

Amino Acids Modulate Liquid-Liquid Phase Separation in vitro and in vivo by Regulating Protein-Protein Interactions

<p>The metadata, plots and microscopy images for the manuscript "Amino Acids Modulate Liquid-Liquid Phase Separation in vitro and in vivo by Regulating Protein-Protein Interactions".</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

RNA-Protein Interaction Prediction Using Network-Guided Deep Learning

<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div>&nbsp;</div>

openmit-licenseJul 2024View details →
zenodo44/100

Data for "Detection of metabolite-protein interactions in complex biological samples by high-resolution relaxometry: towards interactomics by NMR"

<p>Raw NMR data for relaxometry experiments, divided by donor sample. For every donor sample 2 or 3 different samples were used in order to record data at 19 different magnetic fields.</p> <p>Data from fast field-cycling relaxometry. All the data is&nbsp;in one xlsx file, divided by donor sample.</p> <p>Relaxometry results for alanine, lactate, creatinine and glutamine, obtained from the fitting of their relaxation decays recorded at 19 different fields, divided by donor sample.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Paramecium Polycomb Repressive Complex 2 physically interacts with the small RNA binding PIWI protein to repress transposable elements

<p>Polycomb Repressive Complex 2 (PRC2) maintains transcriptionally silent genes in a repressed state via deposition of histone H3 K27 trimethyl (me3) marks. PRC2 has also been implicated in silencing transposable elements (TEs), yet how PRC2 is targeted to TEs remains unclear. To address this question, we identified proteins that physically interact with the <em>Paramecium</em> Enhancer-of-zeste Ezl1 enzyme, which catalyzes H3K9me3 and H3K27me3 deposition at TEs. We show that the <em>Paramecium</em> PRC2 core complex comprises four subunits, each required <em>in vivo</em> for catalytic activity. We also identify PRC2 cofactors, including the RNA interference (RNAi) effector Ptiwi09, which are necessary to target H3K9me3 and H3K27me3 to TEs. We find that the physical interaction between PRC2 and the RNAi pathway is mediated by a RING finger protein and that small RNA recruitment of PRC2 to TEs is analogous to the small RNA recruitment of H3K9 methylation SU(VAR)3-9 enzymes.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Analysis of the interacting residues between wild type SARS-CoV-2 spike protein and natural ligand hACE2, as well as three engineered alternative ligands

<p>The analysis of residue interactions between the SARS-CoV-2 spike protein and its natural (hACE2 <sup>1</sup>) and engineered binders P17 Fab <sup>2</sup>, Ty1 VHH <sup>3</sup> and LCB1 peptide <sup>4</sup> reveals that glutamine, serine and especially tyrosine residues on the ligand side are more frequent and influence spike binding efficiency, and that spike residues Glu484, Phe486, Tyr489 and Gln493 are more recurrent targets for interactions with ligands. The list of residues establishing contacts between the wild type structure of the SARS-CoV-2 spike protein and the binders defined above are described in Table 1. In Figure 1, the frequency and type of amino acids that interact with each spike residue is illustrated.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Intermolecular interactions in G protein-coupled receptor allosteric sites at the membrane interface from molecular dynamics simulations and quantum chemical calculations

<p>Allosteric modulators are called to be promising candidates in G protein-coupled receptor (GPCR) drug development by displaying target selectivity and fewer side effects. Among the allosteric sites known to date, extrahelical cavities represent an uncharacteristic binding location that raises many questions about the ligand interactions and stability; the binding site structure, and how all of these are affected by lipid molecules. In this work, we analyze the dynamics and interactions in the PAR2, C5aR1, and GCGR receptors unbound and bound to allosteric modulators at the receptor-lipid interface using molecular dynamics simulations in three lipid compositions. In addition, we performed quantum chemical calculations to further explore electrostatic interactions and the strength of atom pairwise contacts in the stabilization of the ligand-receptor complexes. We show that besides classical hydrogen bonds weak polar interactions such as O-HC, O-Br, and S-HC contacts and aromatic interactions contribute to the binding of allosteric modulators at the extrahelical sites in the middle of the membrane. The allosteric cavities are open and detectable in various membrane compositions but not always predicted as druggable. &nbsp;The availability of polar atoms for interactions in such cavities can be assessed by water molecules from the simulations. Although ligand-lipid interactions are weak, the lipid tails play a role in sizing and shaping the large part of the allosteric cavity.&nbsp;</p> <p>You will find the following files:</p> <ul> <li>Input files of the equilibration and production protocols of MD simulations (MD_simulations_inputs.zip)</li> <li>Input files and coordinate files of F-SAPT and NCIPLOT calculations (quantum_chemical_coordiates_inputs.zip)</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Protein Protein Interaction Prdiction Datasets of H.Pylori and S.cerevisiae Species

<p><br> Protein protein interaction prediction datasets related to 2 different species. These datasets have been comprehensively used in published literature to assess the performance of protein-protein interaction predictors.<br> &nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex

<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar&nbsp;-- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo44/100

A complete map of specificity encoding for a partially fuzzy protein interaction

<p>All data required to run analyses for "A complete map of specificity encoding for a partially fuzzy protein interaction". Please see&nbsp;<a href="https://github.com/lehner-lab/fuzzy_specificity">https://github.com/lehner-lab/fuzzy_specificity</a> for instructions.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

PhasAGE Training School 1 - Structure and protein interactions of repeated and low complexity regions - LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction

<p>This dataset contains replication data for the paper titled &quot;DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction&quot;. The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test&nbsp;filename lists for cross-validation, that can be used to develop and evaluate&nbsp;protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users&#39; convenience (e.g. raw MSAs and profile HMMs for each protein complex).&nbsp;Our GitHub repository linked in the &quot;Additional notes&quot; metadata section below provides more details on how we parsed through these files to create our cross-validation&nbsp;datasets. The GitHub repository&nbsp;for DIPS-Plus&nbsp;also includes scripts that can be used&nbsp;to impute missing feature values and convert the&nbsp;final &quot;raw&quot; complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of&nbsp;residue&nbsp;embeddings, the final representation of each complex can easily be adapted to fit the users&#39; needs (e.g. feeding a complex&#39;s&nbsp;2D residue feature tensors into a convolutional neural network).</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Nuclear Magnetic resonance Dataset of 2D spectra of S100B and Tau to study their protein-protein interaction

<p>Nuclear Magnetic resonance dataset of 2D spectra corresponding to raw data of research published in Nature Communication in a communication entitled &quot;Dynamic interactions and Ca2+ 1 -binding modulate the holdase-type chaperone activity of S100B preventing tau&nbsp;aggregation and seeding&quot; by Moreira G. et al.</p> <p>Dataset corresponds to</p> <p>raw data files in Bruker format of NMR 2D spectra (ser), associated with&nbsp;files of acquisition parameters and processing parameters (pdata),</p> <p>files in .ucsf format that can be read with NMRFAM-Sparky (free download) of 2D spectra (in sub-directory pdata/1)</p> <p>files of chemical shift value lists that can be read as text files or in NMRFAM sparky together with the corresponding ucsf files.</p> <p>physico-chemical conditions are found in title in pdata\1</p> <p>Data were acquired on a Bruker 900-MHz spectrometer equipped with a triple-resonance cryogenic probe (Bruker, Karlsruhe, Germany)</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction (Supplementary Data)

<p>This dataset contains supplementary replication data for the paper titled &quot;DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction&quot;. In particular, it contains a new version of our `final_raw_dips.tar.gz` protein pair representations which now contain (1) residue-level&nbsp;annotations for intrinsic disorder regions (IDRs) as well as (2) a copy of each protein pair representation in the HDF5 file format for programming language-agnostic read capabilities. In addition, this record also contains (3) raw MSAs (in HDF5 file format)&nbsp;generated for each protein pair using Jackhmmer and AlphaFold&#39;s small version of the Big Fantastic Database (BFD). Lastly, this record contains (4) PDB metadata derived for each DIPS-Plus complex using Graphein&#39;s PDBManager&nbsp;API&nbsp;as well as (5) structure-based (i.e., FoldSeek-based) training and validation splits of the dataset&#39;s complexes in the form of respective text files containing the file paths of complexes assigned to each split.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

DATASET: Exploring the Residue-Level Interactions between the R2ab Protein and Polystyrene Nanoparticles

<p>This file contains the raw hydrogen-deuterium exchange data used in the manuscript. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD044469. 202308</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction

<p>The code, dataset, and model weights are described in the paper "Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction."</p> <p>&nbsp;</p> <p><strong>experiment_results.zip:</strong> Contains generated results that can reproduce the result from the reported paper.</p> <p><strong>benchmark.zip:</strong> Contains docking and affinity input data of the interformer. You can use the source code to make predictions and reproduce the number of the reported paper.</p> <p><strong>checkpoints.zip: </strong>Contains one weight for the Energy and four PoseScore and Affinity models.</p> <p><strong>source_code_1.0.zip:</strong> Contains the initial version of the source code.</p> <p><strong>interformer_train.tar.gz:</strong> Contains prepared training data for interformer. poses/ contains all structure need for training, poses/ligand contains the re-docking poses generated by interformer energy, poses/ligand/rcsb contains the conformation of reference ligand, poses/pocket contains all pocket extract by raw PDB from rcsb, poses/uff contains all ligand conformation minimized using UFF from reference ligand, and train/ contains the training csv.</p> <p><strong>baseline_results.tar.gz:</strong>&nbsp; Contains the predictions from three methods: Interformer, DiffDock, and DeepDock. The results align with the exact numbers reported in the paper. For further details, please refer to the <em>eda/ </em>directory.</p> <p>&nbsp;</p> <p>You can also find the newest version of the source code at <a href="https://github.com/tencent-ailab/Interformer" target="_blank" rel="noopener">https://github.com/tencent-ailab/Interformer</a></p> <p>&nbsp;</p>

openapache2.0Mar 2024View details →
zenodo40/100

Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record