Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
641
datasets available to search
ShareScore release 0.9.0
Dataset results
641 results for “Protein interaction”
Bioactivity deep learning for structure-free compound-protein interaction
<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p> </p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p> </p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p> </p>
Project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"
<p><strong>README file for the project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"</strong></p> <p>February 12, 2020</p> <p>Authors: Raffaele Fiorentini, Kurt Kremer and Raffaello Potestio</p> <p>================================</p> <p>Overview</p> <p>The dataset is organised in three (compressed) subfolders (see the tree diagrams in each section):</p> <p>- annihilation<br> - decoupling<br> - density</p> <p>The figure deltaG_binding_ann_dec_comparison.png shows the results of binding free energy calculations comparing the values obtained both for annihilation and decoupling.</p> <p>The figure deltaG_binding_annih_gromacs_espp.png displays the results for Binding FE, comparing the values obtained in GROMACS and ESPResSo++.</p> <p>The README.pdf file contains detailed information about these folders and their content.</p> <p>================================</p> <p>The "annihilation" folder contains all results concerning the calculation of binding free energy in case of annihilation and it is divided in two parts: </p> <p>- complex<br> - ligand</p> <p>In "complex" are reported the results of Ligand-Protein FE both in ESPResSo++ and GROMACS. All simulations are fully-atomistic. </p> <p>In "ligand" are reported the results of ligand solvation free energy both in ESPResSo++ and GROMACS. All simulations are fully-atomistic. </p> <p>====</p> <p>The "decoupling" folder contains all results concerning the calculation of binding free energy in case of decoupling and it is divided in three parts: </p> <p>- complex-DualRes<br> - complex-FullyAT<br> - ligand</p> <p>In "complex-DualRes" are reported the results of Ligand-Protein FE only in ESPResSo++ (GROMACS cannot do decoupling). The system is simulated in Dual-Resolution. It is possible to find the trajectory files in the sub-directories "lambdaindex-0" and "lambdaindex-30".</p> <p>In "complex-fullyAT" are reported the results of Ligand-Protein FE only in ESPResSo++. The system simulated is fully-atomistic. It is possible to find the trajectory file in the sub-directories "lambdaindex-0" and "lambdaindex-30".</p> <p>In "ligand" are reported the results of ligand solvation free energy only in ESPResSo++. All simulations are fully-atomistic. It is possible to find the trajectory file in the sub-directories "lambdaindex-0" and "lambdaindex-20".</p> <p>====</p> <p>The "density" folder contains the data for the tuning of the c parameter of the steric repulsion among residues. This parameter is tuned so that the water density attains the value computed in all-atom simulations.</p>
Scored protein-protein interactions accompanying "A pan-plant protein complex map reveals deep conservation and novel assemblies"
<p><a href="http://plants.proteincomplexes.org/static/data/panplant_cfms_scores_annot.txt.gz">All scored pairwise protein-protein interactions with CF-MS scores (3,076,999 unique pairwise interactions)</a></p> <ul> <li>Description: Scores between Orthogroups with the corresponding CF-MS score and eggNOG generated orthogroup descriptions.</li> <li>Note: Only the highest scoring pairs are considered significant. A CF-MS score >= 0.509 corresponds to 10% FDR, >= 0.207 corresponds to 50% FDR</li> <li>Format: OrthogroupID1 [tab] OrthogroupID2 [tab] Score [tab] Annotation1 [tab] Annotation2</li> </ul>
Human Pleckstrin Homology domain Interacting Protein (PHIP); A Target Enabling Package
<p>SGC Oxford has expressed, purified and crystallized the second bromodomain of PHIP as part of the probe programme. Fragment screening and X-ray crystallography identified binders, some of which optimised to uM affinity. However, molecules with probe properties were not obtained. Consequently it has been decided to put the information generated into the public domain.</p>
Amino Acids Modulate Liquid-Liquid Phase Separation in vitro and in vivo by Regulating Protein-Protein Interactions
<p>The metadata, plots and microscopy images for the manuscript "Amino Acids Modulate Liquid-Liquid Phase Separation in vitro and in vivo by Regulating Protein-Protein Interactions".</p>
RNA-Protein Interaction Prediction Using Network-Guided Deep Learning
<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div> </div>
Data for "Detection of metabolite-protein interactions in complex biological samples by high-resolution relaxometry: towards interactomics by NMR"
<p>Raw NMR data for relaxometry experiments, divided by donor sample. For every donor sample 2 or 3 different samples were used in order to record data at 19 different magnetic fields.</p> <p>Data from fast field-cycling relaxometry. All the data is in one xlsx file, divided by donor sample.</p> <p>Relaxometry results for alanine, lactate, creatinine and glutamine, obtained from the fitting of their relaxation decays recorded at 19 different fields, divided by donor sample.</p>
Paramecium Polycomb Repressive Complex 2 physically interacts with the small RNA binding PIWI protein to repress transposable elements
<p>Polycomb Repressive Complex 2 (PRC2) maintains transcriptionally silent genes in a repressed state via deposition of histone H3 K27 trimethyl (me3) marks. PRC2 has also been implicated in silencing transposable elements (TEs), yet how PRC2 is targeted to TEs remains unclear. To address this question, we identified proteins that physically interact with the <em>Paramecium</em> Enhancer-of-zeste Ezl1 enzyme, which catalyzes H3K9me3 and H3K27me3 deposition at TEs. We show that the <em>Paramecium</em> PRC2 core complex comprises four subunits, each required <em>in vivo</em> for catalytic activity. We also identify PRC2 cofactors, including the RNA interference (RNAi) effector Ptiwi09, which are necessary to target H3K9me3 and H3K27me3 to TEs. We find that the physical interaction between PRC2 and the RNAi pathway is mediated by a RING finger protein and that small RNA recruitment of PRC2 to TEs is analogous to the small RNA recruitment of H3K9 methylation SU(VAR)3-9 enzymes.</p>
Analysis of the interacting residues between wild type SARS-CoV-2 spike protein and natural ligand hACE2, as well as three engineered alternative ligands
<p>The analysis of residue interactions between the SARS-CoV-2 spike protein and its natural (hACE2 <sup>1</sup>) and engineered binders P17 Fab <sup>2</sup>, Ty1 VHH <sup>3</sup> and LCB1 peptide <sup>4</sup> reveals that glutamine, serine and especially tyrosine residues on the ligand side are more frequent and influence spike binding efficiency, and that spike residues Glu484, Phe486, Tyr489 and Gln493 are more recurrent targets for interactions with ligands. The list of residues establishing contacts between the wild type structure of the SARS-CoV-2 spike protein and the binders defined above are described in Table 1. In Figure 1, the frequency and type of amino acids that interact with each spike residue is illustrated.</p>
Intermolecular interactions in G protein-coupled receptor allosteric sites at the membrane interface from molecular dynamics simulations and quantum chemical calculations
<p>Allosteric modulators are called to be promising candidates in G protein-coupled receptor (GPCR) drug development by displaying target selectivity and fewer side effects. Among the allosteric sites known to date, extrahelical cavities represent an uncharacteristic binding location that raises many questions about the ligand interactions and stability; the binding site structure, and how all of these are affected by lipid molecules. In this work, we analyze the dynamics and interactions in the PAR2, C5aR1, and GCGR receptors unbound and bound to allosteric modulators at the receptor-lipid interface using molecular dynamics simulations in three lipid compositions. In addition, we performed quantum chemical calculations to further explore electrostatic interactions and the strength of atom pairwise contacts in the stabilization of the ligand-receptor complexes. We show that besides classical hydrogen bonds weak polar interactions such as O-HC, O-Br, and S-HC contacts and aromatic interactions contribute to the binding of allosteric modulators at the extrahelical sites in the middle of the membrane. The allosteric cavities are open and detectable in various membrane compositions but not always predicted as druggable. The availability of polar atoms for interactions in such cavities can be assessed by water molecules from the simulations. Although ligand-lipid interactions are weak, the lipid tails play a role in sizing and shaping the large part of the allosteric cavity. </p> <p>You will find the following files:</p> <ul> <li>Input files of the equilibration and production protocols of MD simulations (MD_simulations_inputs.zip)</li> <li>Input files and coordinate files of F-SAPT and NCIPLOT calculations (quantum_chemical_coordiates_inputs.zip)</li> </ul>
Protein Protein Interaction Prdiction Datasets of H.Pylori and S.cerevisiae Species
<p><br> Protein protein interaction prediction datasets related to 2 different species. These datasets have been comprehensively used in published literature to assess the performance of protein-protein interaction predictors.<br> </p>
Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex
<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar -- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>
A complete map of specificity encoding for a partially fuzzy protein interaction
<p>All data required to run analyses for "A complete map of specificity encoding for a partially fuzzy protein interaction". Please see <a href="https://github.com/lehner-lab/fuzzy_specificity">https://github.com/lehner-lab/fuzzy_specificity</a> for instructions. </p>
PhasAGE Training School 1 - Structure and protein interactions of repeated and low complexity regions - LECTURE
<p>The Training School 1 <strong>“Computational Methods to Study Protein Phase Separation”</strong> is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of <strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide <strong>an overview of the available computational resources</strong> to navigate this knowledge. Participants will have <strong>hands-on training</strong> in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction
<p>This dataset contains replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test filename lists for cross-validation, that can be used to develop and evaluate protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users' convenience (e.g. raw MSAs and profile HMMs for each protein complex). Our GitHub repository linked in the "Additional notes" metadata section below provides more details on how we parsed through these files to create our cross-validation datasets. The GitHub repository for DIPS-Plus also includes scripts that can be used to impute missing feature values and convert the final "raw" complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of residue embeddings, the final representation of each complex can easily be adapted to fit the users' needs (e.g. feeding a complex's 2D residue feature tensors into a convolutional neural network).</p>
Nuclear Magnetic resonance Dataset of 2D spectra of S100B and Tau to study their protein-protein interaction
<p>Nuclear Magnetic resonance dataset of 2D spectra corresponding to raw data of research published in Nature Communication in a communication entitled "Dynamic interactions and Ca2+ 1 -binding modulate the holdase-type chaperone activity of S100B preventing tau aggregation and seeding" by Moreira G. et al.</p> <p>Dataset corresponds to</p> <p>raw data files in Bruker format of NMR 2D spectra (ser), associated with files of acquisition parameters and processing parameters (pdata),</p> <p>files in .ucsf format that can be read with NMRFAM-Sparky (free download) of 2D spectra (in sub-directory pdata/1)</p> <p>files of chemical shift value lists that can be read as text files or in NMRFAM sparky together with the corresponding ucsf files.</p> <p>physico-chemical conditions are found in title in pdata\1</p> <p>Data were acquired on a Bruker 900-MHz spectrometer equipped with a triple-resonance cryogenic probe (Bruker, Karlsruhe, Germany)</p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction (Supplementary Data)
<p>This dataset contains supplementary replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". In particular, it contains a new version of our `final_raw_dips.tar.gz` protein pair representations which now contain (1) residue-level annotations for intrinsic disorder regions (IDRs) as well as (2) a copy of each protein pair representation in the HDF5 file format for programming language-agnostic read capabilities. In addition, this record also contains (3) raw MSAs (in HDF5 file format) generated for each protein pair using Jackhmmer and AlphaFold's small version of the Big Fantastic Database (BFD). Lastly, this record contains (4) PDB metadata derived for each DIPS-Plus complex using Graphein's PDBManager API as well as (5) structure-based (i.e., FoldSeek-based) training and validation splits of the dataset's complexes in the form of respective text files containing the file paths of complexes assigned to each split.</p>
DATASET: Exploring the Residue-Level Interactions between the R2ab Protein and Polystyrene Nanoparticles
<p>This file contains the raw hydrogen-deuterium exchange data used in the manuscript. The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the dataset identifier PXD044469. 202308</p>
Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction
<p>The code, dataset, and model weights are described in the paper "Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction."</p> <p> </p> <p><strong>experiment_results.zip:</strong> Contains generated results that can reproduce the result from the reported paper.</p> <p><strong>benchmark.zip:</strong> Contains docking and affinity input data of the interformer. You can use the source code to make predictions and reproduce the number of the reported paper.</p> <p><strong>checkpoints.zip: </strong>Contains one weight for the Energy and four PoseScore and Affinity models.</p> <p><strong>source_code_1.0.zip:</strong> Contains the initial version of the source code.</p> <p><strong>interformer_train.tar.gz:</strong> Contains prepared training data for interformer. poses/ contains all structure need for training, poses/ligand contains the re-docking poses generated by interformer energy, poses/ligand/rcsb contains the conformation of reference ligand, poses/pocket contains all pocket extract by raw PDB from rcsb, poses/uff contains all ligand conformation minimized using UFF from reference ligand, and train/ contains the training csv.</p> <p><strong>baseline_results.tar.gz:</strong> Contains the predictions from three methods: Interformer, DiffDock, and DeepDock. The results align with the exact numbers reported in the paper. For further details, please refer to the <em>eda/ </em>directory.</p> <p> </p> <p>You can also find the newest version of the source code at <a href="https://github.com/tencent-ailab/Interformer" target="_blank" rel="noopener">https://github.com/tencent-ailab/Interformer</a></p> <p> </p>
Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.