Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
80
datasets available to search
ShareScore release 0.9.0
Dataset results
80 results for “binding affinity”
Binding Affinity Prediction Workflow - Simulation Input Files and Absolute Binding Free Energies
<p>The Binding Affinity Prediction (BAP) workflow calculates absolute binding free energies for protein-ligand complexes by taking their crystal structures, converting them into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks, and analysing the resulting trajectories with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the free-energy estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the BAP workflow was run on the PDBbind 2020 (http://www.pdbbind.org.cn/index.php) refined set. This entry contains the MD simulation input files (BAPSimulationInputFiles.tar.gz) and the ABFE estimates (BAPBindingFreeEnergyEstimates.csv) obtained from four 250 ns trajectories for each complex. The MD simulations for more than 4000 complexes were run on the Leonardo supercomputer while the implicit-solvent calculations were carried out on Galileo, both operated by Cineca (Italy). The MD trajectories will be stored at Cineca for approx. 1 year after publication of this entry; contact Cineca's user support if you are interested in the trajectories.</p> <p>The README file describes how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Binding-Affinity-Prediction-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>
P2PXML Dataset: Deep Geometric Framework to Predict Antibody-Antigen Binding Affinity
<p>In drug development, the efficacy of an antibody depends on how the antibody interacts with the target antigen. The strength of these interactions indicates how successful an antibody is in neutralizing an antigen. Therefore, the strength, measured by “binding affinity”, is a critical aspect of antibody engineering. In theory, the higher the binding affinity, the higher the chances are that the antibody is successful against the target antigen. Currently, techniques such as molecular docking and molecular dynamics are utilized in quantifying the binding affinity. However, owing to the computational complexity of the aforementioned techniques, running simulations for large antibodies/antigens remains a daunting task. Despite the commendable improvements in deep learning-based binding affinity prediction, such approaches are highly dependent on the quality of the antibody-antigen structures and they tend to overlook the importance of capturing the evolutionary details of proteins upon mutation. Further, most of the existing datasets for the task only include antibody-antigen pairs related to one antigen variant and, thus, are not suitable for developing comprehensive data-driven approaches. To circumvent the said complexities, we first curate the largest and most generalized datasets for antibody-antigen binding affinity prediction, consisting of both protein sequences and structures. Subsequently, we propose a deep geometric neural network comprising a structure-based model and a sequence-based model that considers both atomistic and evolutionary details when predicting the binding affinity. The proposed framework exhibited a 10% improvement in mean absolute error compared to the state-of-the-art models while showing a strong correlation between the predictions and target values. We release the datasets and code publicly https://drug-discovery-entc.github.io/p2pxml/ to support the development of antibody-antigen binding affinity prediction frameworks for the benefit of science and society. </p>
Phenoxytacrine derivatives: Low-toxicity neuroprotectants exerting affinity to ifenprodil-binding site and cholinesterase inhibition
<p><a title="Learn more about Tacrine from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/chemistry/tacrine">Tacrine</a> (THA), a long withdrawn drug, is still a popular scaffold used in medicinal <a title="Learn more about chemistry from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/chemistry/chemistry">chemistry</a>, mainly for its good reactivity and multi-targeted effect. However, THA-associated hepatotoxicity is still an issue and must be considered in <a title="Learn more about drug discovery from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/drug-discovery">drug discovery</a> based on the THA scaffold. Following our previously identified hit compound 7-phenoxytacrine (7-PhO-THA), we systematically explored the chemical space with 30 novel derivatives, with a focus on low hepatotoxicity, <a title="Learn more about anticholinesterase from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/cholinesterase-inhibitor">anticholinesterase</a> action, and antagonism at the GluN1/GluN2B subtype of the <a title="Learn more about NMDA receptor from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/n-methyl-dextro-aspartic-acid-receptor">NMDA receptor</a>. Applying the down-selection process based on <em>in vitro</em> and <em>in vivo</em> <a title="Learn more about pharmacokinetic from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/pharmacokinetics">pharmacokinetic</a> data, two candidates, <strong>I-52</strong> and <strong>II-52,</strong> selective GluN1/GluN2B inhibitors thanks to the interaction with the ifenprodil-binding site, have entered <em>in vivo</em> <a title="Learn more about pharmacodynamic from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/biochemistry-genetics-and-molecular-biology/pharmacodynamics">pharmacodynamic</a> studies. Finally, compound <strong>I-52,</strong> showing only minor affinity to <a title="Learn more about AChE from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/acetylcholinesterase">AChE</a>, was identified as a lead candidate with favorable behavioral and <a title="Learn more about neuroprotective from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/pharmacology-toxicology-and-pharmaceutical-science/neuroprotective-agent">neuroprotective</a> effects using open-field and <a title="Learn more about prepulse inhibition from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/biochemistry-genetics-and-molecular-biology/prepulse-inhibition">prepulse inhibition</a> tests, along with scopolamine-based behavioral and NMDA-induced hippocampal lesion models. Our data show that compound <strong>I-52</strong> exhibits low toxicity often associated with <a title="Learn more about NMDA receptor from ScienceDirect's AI-generated Topic Pages" href="https://www.sciencedirect.com/topics/biochemistry-genetics-and-molecular-biology/nmda-receptor">NMDA receptor</a> ligands, and low hepatotoxicity, often related to THA-based compounds.</p>
Data from: Residues neighboring an SH3-binding motif participate in determining affinity and specificity in vivo
<p>In signaling networks, many protein-protein interactions are mediated by modular domains that bind short linear motifs. The motifs' sequences modulate many factors, among them affinity and specificity, or the ability to bind strongly and to bind the appropriate partners. Previous studies have proposed a trade-off between affinity and specificity, suggesting that motifs with high affinity are less capable of differentiating between domains with similar sequences and structures. Using Deep Mutational Scanning to create a mutant library of a well characterized binding motif, and protein complementation assays to measure protein-protein interactions, we tested this trade-off in vivo for the first time. We measured the binding strength and specificity of a library of mutants of a binding motif on the MAP kinase kinase Pbs2, which binds the SH3 domain of the osmosensor protein Sho1 in Saccharomyces cerevisiae. We find that many mutations in the region surrounding the binding motif modify binding strength, but that few mutations have a strong impact on specificity. Moreover, we find no systematic relationship between affinity and specificity as measured in vivo. Interestingly, all Pbs2 mutations which increase affinity or specificity are situated outside of the Pbs2 residues that interact with the canonical SH3-binding pocket, suggesting that other surfaces on Sho1 contribute to binding. We use predicted structures to propose a model of binding which involves residues neighboring the core Pbs2 motif binding outside of the canonical SH3-binding pocket, allowing affinity and specificity to be determined by a broader range of sequences than what has previously been considered.</p>
Cooperative action of separate interaction domains promotes high-affinity DNA binding of Arabidopsis thaliana ARF transcription factors
<p>The repository contains the smFRET and SAXS data presented in the preprint https://doi.org/10.1101/2022.11.16.516730 (BioRxiv)</p>
Target-ligand binding affinity from single point enthalpy calculation and elemental composition
<p>This repository contains supporting files for the manuscript entitled: Target-ligand binding affinity from single point enthalpy calculation and elemental composition.</p>
Simulations for "Molecular Dynamics-Based Identification of Binding Pathways and Two Distinct High-Affinity Sites for Succinate in the Succinate Receptor 1 SUCNR1/GPR91"
Open the record for dataset details and reuse information.
PPB-Affinity: Protein-Protein Binding Affinity dataset for AI-based protein drug discovery
<p>Prediction of protein-protein binding (PPB) affinity plays an important role in large-molecular drug discovery. Deep learning (DL) has been adopted to predict the changes of PPB binding affinities upon mutations, but there was a scarcity of studies predicting the PPB affinity itself. The major reason is the paucity of open-source dataset with PPB affinity data. To address this gap, the current study introduced a large comprehensive PPB affinity (PPB-Affinity) dataset. The PPB-Affinity dataset contains key information such as crystal structures of protein-protein complexes (with or without protein mutation patterns), PPB affinity, receptor protein chain, ligand protein chain, etc. To the best of our knowledge, this is the largest publicly available PPB affinity dataset, and we believe it will significantly advance drug discovery by streamlining the screening of potential large-molecule drugs. We also developed a deep-learning benchmark model with this dataset to predict the PPB affinity, providing a foundational comparison for the research community.</p> <p>Codes for PPB-Affinity database preparation is disclosed at <a title="https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow" href="https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow">https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow</a>.<br>Codes for the benchmark algorithm is disclosed at <a href="https://github.com/ChenPy00/PPB-Affinity">https://github.com/ChenPy00/PPB-Affinity</a>.</p> <p>The article related to the PPB-Affinity dataset can be found at <a href="https://doi.org/10.1038/s41597-024-03997-4">https://doi.org/10.1038/s41597-024-03997-4</a>.</p> <p>Files are orginized as follows:</p> <p><a href="13054646" target="_blank" rel="noopener noreferrer">- PPB-Affinity.xlsx</a></p> <p><a href="13054646" target="_blank" rel="noopener noreferrer">- samples_deleted.zip</a></p> <p><a href="../api/records/13067409/draft/files/PPB-Affinity-AF.zip/content" target="_blank" rel="noopener noreferrer">- PPB-Affinity-AF.zip</a></p> <p>- PDB/</p> <p> - Affinity Benchmark v5.5/</p> <p> - file1.pdb</p> <p> - file2.pdb</p> <p> - ...</p> <p> - filek.pdb</p> <p> - ATLAS/</p> <p> - PDBbind v2020/</p> <p> - SAbDab/</p> <p> - SKEMPIv2.0/</p>
Insect Odorant Binding Protein Dataset of Binding Affinities against Volatile Organic Compounds
<p>This is an archival version of the initial iOBPdb dataset of insect odorant binding protein binding affinities against volatile organic compounds. It contains 215 functional studies containing 381 unique OBPs from 91 insect species.</p>
Data for "Binding Affinity of Monoalkyl Phosphinic Acid Ligands toward Nanocrystal Surfaces".
<p>Data of the figures in the publication "<strong>Binding Affinity of Monoalkyl Phosphinic Acid Ligands toward Nanocrystal Surfaces</strong>".</p> <p>The <em>.pxp</em> documents contain the experimental data of the figures in the manuscript and they can be opened/edited with the software IGOR Pro 6.3 or higher.</p> <p>Table of contents:</p> <p><strong>Figure 1.</strong> (A) General reaction scheme toward monoclinic HfO<sub>2</sub>/oleate NCs. (B) TEM image and (C) DOSY NMR spectrum in C<sub>6</sub>D<sub>6</sub> of HfO<sub>2</sub>/oleate NCs. (D) General reaction scheme toward zinc blend CdSe/oleate NCs. (E) TEM image and (F) DOSY NMR spectrum in C<sub>6</sub>D<sub>6</sub> of CdSe/oleate NCs.</p> <p><strong>Figure 2.</strong> (Left) Titration of HfO<sub>2</sub>/oleate with 6-(hexyloxy)hexylphosphinic acid. (A) General reaction scheme. (B) <sup>1</sup>H NMR spectra of the titration. (C) <sup>31</sup>P NMR after 0.95 and 1.15 equiv of phosphinic acid is added. (D) Quantification of the different compounds as a function of the added equivalents. (Right) Titration of CdSe/oleate with 6-(hexyloxy)hexylphosphinic acid. (E) General reaction scheme. (F) <sup>1</sup>H NMR spectra of the titration. (G) <sup>31</sup>P NMR after 1.0 and 1.2 equiv of phosphinic acid is added. (H) Quantification of the different compounds as a function of the added equivalents.</p> <p><strong>Figure 3.</strong> (A) <sup>1</sup>H NMR spectrum of oleate-capped HfO<sub>2</sub> NCs in C<sub>6</sub>D<sub>6</sub>. (B) <sup>1</sup>H NMR spectrum of HfO<sub>2</sub> in C<sub>6</sub>D<sub>6</sub> after ligand exchange for 6-(hexyloxy)hexylphosphinic acid and purification. The inset shows the <sup>31</sup>P NMR spectrum.</p> <p><strong>Figure 4.</strong> (Left) Titration of HfO<sub>2</sub>/[6-(hexyloxy)hexyl]phosphinate with oleylphosphonic acid. (A) General reaction scheme. (B) <sup>1</sup>H NMR spectra of the titration. (C) Quantification of the different compounds as a function of the added equivalents. Note that the small amount of residual oleic acid present at the start of the titration in (B) is due to a challenging purification (high solubility of the HfO<sub>2</sub>/[6-(hexyloxy)hexyl]phosphinate NCs). This small signal was integrated and subtracted from the spectra for the quantification in (C). (Right) Titration of CdSe/[6-(hexyloxy)hexyl]phosphinate with oleylphosphonic acid. (D) General reaction scheme. (E) <sup>1</sup>H NMR spectra of the titration. (F) Quantification of the different compounds as a function of the added equivalents.</p> <p><strong>Figure 5.</strong> Mole fraction of bound oleylphosphonate in the ligand shell, χ<sub>phosphon (bound)</sub>, as a function of the overall mole fraction of oleylphosphonic acid (=unbound phosphonic acid and bound phosphonate), χ<sub>phosphon (total)</sub>, during the titrations of HfO<sub>2</sub> (blue), CdSe (red), and ZnS (green) NCs. The full lines represent different calculated equilibrium constants.</p> <p><strong>Figure 6.</strong> Mole fraction of a bound incoming ligand (= titrating ligand) in the ligand shell, χ<sub>incomingligand (bound)</sub>, as a function of the overall mole fraction of the total incoming ligand (=unbound and bound), χ<sub>incomingligand (total)</sub>, during the titrations of ZnS NCs stabilized with a 50/50 mixture of <em>n</em>-hexyl/<em>n</em>-octadecylphosphinate with oleylphosphonic acid (gray), ZnS NCs stabilized with 6-(hexyloxy)hexylphosphinate with oleylphosphonic acid (green) and ZnS NCs stabilized with oleate with a 50/50 mixture of <em>n</em>-hexyl/<em>n</em>-octadecylcarboxylic acids (orange). The full lines represent the different calculated equilibrium constants.</p> <p><strong>Figure S1.</strong> (A) HfO2/oleate NCs, (B) 1H NMR spectrum in C6D6, and (C) DOSY decay curve of the alkene region.</p> <p><strong>Figure S2.</strong> (A) CdSe/oleate NCs, (B) 1H NMR spectrum in C6D6, and (C) UV-vis absorption spectrum, (D) DOSY decay curve of the alkene region, and (E) DOSY decay curve of the methylene region.</p> <p><strong>Figure S3.</strong> Synthesis of zinc blende ZnS/oleate NCs. (A) General reaction scheme, (B) TEM image, and (C) 1H NMR spectrum in C6D6, (D) DOSY NMR spectrum in C6D6, (E) DOSY decay curve of the alkene region, and (F) UV-vis absorption spectrum.</p> <p><strong>Figure S4.</strong> 1H NMR of (top, black line) the supernatant of the HfO2/oleate NCs after titration until 1.0 equivalent 6-(hexyloxy)hexyl phosphinic acid is added, and (bottom grey line) reference spectrum of oleic acid in CDCl3.</p> <p><strong>Figure S5.</strong> 1H NMR of (top, black line) the supernatant of the CdSe/oleate NCs after titration until 1.0 equivalent 6-(hexyloxy)hexyl phosphinic acid is added, and (bottom grey line) reference spectrum of oleic acid in CDCl3</p> <p><strong>Figure S6.</strong> Titration of ZnS/oleate NCs with 6-(hexyloxy)hexylphosphinic acid. (A) General reaction scheme, (B) 1H NMR spectra of the titration, (C) 31P NMR after 1.0 and 1.6 equivalent phosphinic acid is added, and (D) quantification of the different compounds as a function of added equivalents.</p> <p><strong>Figure S7.</strong> 1H and 31P NMR spectra of (top gray line) n-tetradecylphosphinic acid dehydrated<br>with dicyclohexylcarbodiimide (DCC) to form n-tetradecylphosphinic anhydride, and (bottom<br>black line) n-tetradecylphosphinic acid reference, both in C6D6. </p> <p><strong>Figure S8.</strong> (A) 1H and (B) 31P NMR of (top black line) the supernatant of the ZnS/oleate NCs after titration until 1.6 equivalent 6-(hexyloxy)hexylphosphinic acid is added, and reference spectra of (red line) oleic acid and (blue line) 6-(hexyloxy)hexylphosphinic acid in CDCl3.</p> <p><strong>Figure S9.</strong> Purified HfO2/[6-(hexyloxy)hexyl]phosphinate NCs. (A) HfO2/phosphinate NCs. (B) 1H NMR spectrum in C6D6 with zoom inset of the broadened P-H resonance, (C) 31P NMR spectrum in C6D6, (D) DOSY NMR spectrum in C6D6, (E) DOSY decay curve of the ether region.</p> <p><strong>Figure S10.</strong> Purified CdSe/[6-(hexyloxy)hexyl]phosphinate NCs. (A) CdSe/phosphinate NCs. (B) 1H NMR spectrum in C6D6, (C) 31P NMR spectrum in C6D6, (D) DOSY NMR spectrum in C6D6, (E) DOSY decay curve of the ether region, (F) DOSY decay curve of the methylene region, and (G) UV-vis absorption spectrum.</p> <p><strong>Figure S11.</strong> Purified ZnS/[6-(hexyloxy)hexyl]phosphinate NCs. (A) ZnS/phosphinate NCs. (B) 1H NMR spectrum in C6D6 with zoom inset of the broadened P-H resonance, (C) 31P NMR spectrum in C6D6, (D) DOSY NMR spectrum in C6D6, (E) DOSY decay curve of the ether region, and (F) UV-vis absorption spectrum.</p> <p><strong>Figure S12.</strong> (A) 1H and (B) 31P NMR of the supernatant (black line) of the HfO2/phosphinate NCs after titration until 2.0 equivalent oleylphosphonic acid is added, and reference spectra of (green line) oleylphosphonic acid and (blue line) 6-(hexyloxy)hexylphosphinic acid in C6D6</p> <p><strong>Figure S13.</strong> (A) 1H and (B) 31P NMR of the supernatant (black line) of the CdSe/phosphinate NCs after titration until 2.0 equivalent oleylphosphonic acid is added, and reference spectra of (green line) oleylphosphonic acid and (blue line) 6-(hexyloxy)hexylphosphinic acid in C6D6. </p> <p><strong>Figure S14.</strong> Titration of ZnS/[(6-hexyloxy)hexyl]phosphinate with oleylphosphonic acid. (A) General reaction scheme, (B) 1H NMR spectra of the titration, and (C) quantification of the different compounds as a function of added equivalents.</p> <p><strong>Figure S15.</strong> (A) 1H and (B) 31P NMR of the supernatant (black line) of the ZnS/phosphinate NCs after titration until 2.0 equivalent oleylphosphonic acid is added, and reference spectra of (green line) oleylphosphonic acid and (blue line) 6-(hexyloxy)hexyl phosphinic acid in C6D6.</p> <p><strong>Figure S16.</strong> Titration of HfO2 NCs stabilized with a mixed ligand shell consistent of 6-(hexyloxy)hexylphosphinate and oleylphosphonate. (A) General reaction scheme, (B) 1H NMR spectrum of the purified NCs prior to titration (at 0.0 added equivalents of 6-(hexyloxy)hexylphosphinic acid), (C) 1H NMR spectra of the titration, and (D) quantification of the different compounds as a function of added equivalents 6-(hexyloxy)hexylphosphinic acid.</p> <p><strong>Figure S17.</strong> Titration of CdSe NCs stabilized with a mixed ligand shell consistent of 6-(hexyloxy)hexylphosphinate and oleylphosphonate. (A) General reaction scheme, (B) 1H NMR spectrum of the purified NCs prior to titration (at 0.0 added equivalents of 6-(hexyloxy)hexylphosphinic acid), (C) 1H NMR spectra of the titration, and (D) quantification of the different compounds as a function of added equivalents 6-(hexyloxy)hexylphosphinic acid.</p> <p><strong>Figure S18.</strong> Titration of ZnS NCs stabilized with a mixed ligand shell consistent of 6-(hexyloxy)hexylphosphinate acid and oleylphosphonate. (A) General reaction scheme, (B) 1H NMR spectrum of the purified NCs prior to titration (at 0.0 added equivalents of 6-(hexyloxy)hexylphosphinic acid) (C) 1H NMR spectra of the titration, and (D) quantification of the different compounds as a function of added equivalents 6-(hexyloxy)hexylphosphinic acid.</p> <p><strong>Figure S19.</strong> The mole fraction of bound oleylphosphonate in the ligand shell, 𝜒𝑝ℎ𝑜𝑠𝑝ℎ𝐨𝑛 (𝑏𝑜𝑢𝑛𝑑), as a function of the overall mole fraction of oleylphosphonic acid (= unbound phosphonic acid and bound phosphonate), 𝜒𝑝ℎ𝑜𝑠𝑝ℎ𝐨𝑛 (𝑡𝑜𝑡𝑎𝑙) , during the titrations of (A) HfO2 (blue), (B) CdSe (red), and (C) ZnS (green) NCs. The full lines represent different calculated equilibrium constants.</p> <p><strong>Figure S20.</strong> Changes in chemical shift during the ligand exchange of phosphinate for phosphonate for HfO2, CdSe, and ZnS NCs</p> <p><strong>Figure S21.</strong> Purified ZnS/n-alkylphosphinate NCs with a 50/50 mixture of n-hexyl/noctadecylphosphinate. (A) ZnS/phosphinate NCs. (B) 1H NMR spectrum in C6D6 with zoom inset of the broadened P-H resonance, (C) 31P NMR spectrum in C6D6, (D) DOSY NMR spectrum in C6D6, (E) DOSY decay curve of the alkane region, and (F) UV-vis absorption spectrum.</p> <p><strong>Figure S22.</strong> Titration of ZnS/n-alkylphosphinate with a 50/50 mixture of n-hexyl/noctadecylphosphinate with oleylphosphonic acid. (A) General reaction scheme, (B) 1H NMR spectra of the titration (zoom of the alkene resonance and the adjacent methylene groups), and (C) quantification of the different compounds as a function of added equivalents.</p> <p><strong>Figure S23.</strong> Comparative ligand exchange experiments where oleate capped ZnS NCs are titrated with a 50/50 mixture of n-hexyl/n-octadecylcarboxylic, -phosphinic, or -phosphonic acids. (A) General reaction scheme for the 3 separate titrations with carboxylic, phosphinic, or phosphonic acids. (B) Bound fraction of oleate on ZnS NCs as a function of the added equivalents of the titrating acid mixture, including the theoretical expected quantitative and random exchange development.</p> <p><strong>Figure S24.</strong> Titration of ZnS/oleate NCs with a 50/50 mixture of n-hexyl, and n-octadecyl carboxylic, phosphinic, and phosphonic acids in C6D6. (A) General reaction scheme. (B) 1H NMR spectra of the titration with carboxylic acids. (C) 1H NMR spectra of the titration with phosphinic acids. (D) 1H NMR spectra of the titration with phosphonic acids (added from a concentrated solution in THF-d8).</p> <p> </p> <p> </p> <p> </p>
ESM-2 embeddings for TCR-Epitope Binding Affinity Prediction Task
<p>This is the accompanying dataset that was generated by the GitHub project: <a href="https://github.com/tonyreina/tdc-tcr-epitope-antibody-binding">https://github.com/tonyreina/tdc-tcr-epitope-antibody-binding</a>. In that repository I show how to create a machine learning models for predicting if a T-cell receptor (TCR) and protein epitope will bind to each other.</p> <p>A model that can predict how well a TCR bindings to an epitope can lead to more effective treatments that use immunotherapy. For example, in anti-cancer therapies it is important for the T-cell receptor to bind to the protein marker in the cancer cell so that the T-cell (actually the T-cell's friends in the immune system) can kill the cancer cell.</p> <div> <div>[HuggingFace](https://huggingface.co/facebook/esm2_t36_3B_UR50D) provides a "one-stop shop" to train and deploy AI models. In this case, we use Facebook's open-source [Evolutionary Scale Model (ESM-2)](https://github.com/facebookresearch/esm). These embeddings turn the protein sequences into a vector of numbers that the computer can use in a mathematical model.</div> <div> </div> To load them into Python use the Pandas library:</div> <pre><code>import pandas as pd train_data = pd.read_pickle("train_data.pkl") validation_data = pd.read_pickle("validation_data.pkl") test_data = pd.read_pickle("test_data.pkl")</code></pre> <p>The <strong>epitope_aa</strong> and the <strong>tcr_full</strong> columns are the protein (peptide) sequences for the epitope and the T-cell receptor, respectively. The letters correspond to the <a href="https://en.wikipedia.org/wiki/DNA_and_RNA_codon_tables">standard amino acid codes</a>.</p> <p>The <strong>epitope_smi</strong> column is the <a href="https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system">SMILES</a> notation for the chemical structure of the epitope. We won't use this information. Instead, the ESM-1b embedder should be sufficient for the input to our binary classification model.</p> <p>The <strong>tcr</strong> column is the CDR3 hyperloop. It's the part of the TCR that actually binds (assuming it binds) to the epitope.</p> <p>The <strong>label</strong> column is whether the two proteins bind. 0 = No. 1 = Yes.</p> <p>The <strong>tcr_vector</strong> and <strong>epitope_vector</strong> columns are the bio-embeddings of the TCR and epitope sequences generated by the Facebook ESM-1b model. These two vectors can be used to create a machine learning model that predicts whether the combination will produce a successful protein binding.</p> <p>From the TDC website:</p> <blockquote> <p>T-cells are an integral part of the adaptive immune system, whose survival, proliferation, activation and function are all governed by the interaction of their T-cell receptor (TCR) with immunogenic peptides (epitopes). A large repertoire of T-cell receptors with different specificity is needed to provide protection against a wide range of pathogens. This new task aims to predict the binding affinity given a pair of TCR sequence and epitope sequence.</p> <p>Weber et al.</p> </blockquote> <p>Dataset Description: The dataset is from Weber et al. who assemble a large and diverse data from the VDJ database and ImmuneCODE project. It uses human TCR-beta chain sequences. Since this dataset is highly imbalanced, the authors exclude epitopes with less than 15 associated TCR sequences and downsample to a limit of 400 TCRs per epitope. The dataset contains amino acid sequences either for the entire TCR or only for the hypervariable CDR3 loop. Epitopes are available as amino acid sequences. Since Weber et al. proposed to represent the peptides as SMILES strings (which reformulates the problem to protein-ligand binding prediction) the SMILES strings of the epitopes are also included. 50% negative samples were generated by shuffling the pairs, i.e. associating TCR sequences with epitopes they have not been shown to bind.</p> <blockquote> <p>Task Description: Binary classification. Given the epitope (a peptide, either represented as amino acid sequence or as SMILES) and a T-cell receptor (amino acid sequence, either of the full protein complex or only of the hypervariable CDR3 loop), predict whether the epitope binds to the TCR.</p> <p>Dataset Statistics: 47,182 TCR-Epitope pairs between 192 epitopes and 23,139 TCRs.</p> <p>References:</p> </blockquote> <ol> <li>Weber, Anna, Jannis Born, and María Rodriguez Martínez. “TITAN: T-cell receptor specificity prediction with bimodal attention networks.” Bioinformatics 37.Supplement_1 (2021): i237-i244.</li> <li>Bagaev, Dmitry V., et al. “VDJdb in 2019: database extension, new analysis infrastructure and a T-cell receptor motif compendium.” Nucleic Acids Research 48.D1 (2020): D1057-D1062.</li> <li>Dines, Jennifer N., et al. “The immunerace study: A prospective multicohort study of immune response action to covid-19 events with the immunecode™ open access database.” medRxiv (2020).</li> </ol> <blockquote> <p>Dataset License: CC BY 4.0.</p> <p>Contributed by: Anna Weber and Jannis Born.</p> </blockquote> <p> </p> <div>The Facebook ESM-2 model has the MIT license and was published in:</div> <div> </div> <div>* Zeming Lin et al, Evolutionary-scale prediction of atomic-level protein structure with a language model, Science (2023). DOI: 10.1126/science.ade2574 https://www.science.org/doi/10.1126/science.ade2574</div> <div> </div> <div>HuggingFace has several versions of the trained model.</div> <div> </div> <div> <table> <tbody> <tr> <td>Checkpoint name</td> <td>Number of layers</td> <td>Number of parameters</td> </tr> <tr> <td>esm2_t48_15B_UR50D</td> <td>48</td> <td>15B</td> </tr> <tr> <td>esm2_t36_3B_UR50D</td> <td>36</td> <td>3B</td> </tr> <tr> <td>esm2_t33_650M_UR50D</td> <td>33</td> <td>650M</td> </tr> <tr> <td>esm2_t30_150M_UR50D</td> <td>30</td> <td>150M</td> </tr> <tr> <td>esm2_t12_35M_UR50D</td> <td>12</td> <td>35M</td> </tr> <tr> <td>esm2_t6_8M_UR50D</td> <td>6</td> <td>8M</td> </tr> </tbody> </table> </div>
Improving generalisability of 3D binding affinity models in low data regimes
<p>Structures of the PDBBind dataset (general protein-ligand) prepared with CCDC protein preparation software. After preparation, 18310 structures out of the total 19443 remained (1133 failed).</p>
Consensus machine-learning models for protein-ligand binding affinity estimation
<p><strong>Motivation:</strong> In structure-based virtual screening, machine learning based scoring function gained popularity in the last few years as they outperformed classical scoring function. The protein-ligand system can be encoded by a set of orthogonal descriptor spaces, which are then mined by machine learning algorithms to find a relationship with the binding affinity experimental value.</p> <p><strong> </strong></p> <p><strong>Results:</strong> In this work we propose our modelling approach to derive a new scoring function, derived from a combination of multiple descriptor spaces coupled with machine learning algorithms ensembled in consensus. The SF has been trained on the PDBbind v.2019 data and has been extensively internally and externally validated on a large set of complexes. When benchmarked on the PDBbind core set, it achieved better performance than state-of-the-art counterparts, scoring: R<sub>Pearson </sub>= 0.85-0.86 r<sup>2</sup> = 0.70-0.72 and RMSE = 1.15-1.21. As highlights: (i) an applicability domain definition has been implemented to delimit the SF’s application boundaries, and (ii) a mechanistic interpretation is proposed by investigating the contribution of each protein-ligand atom pairs in the prediction of the binding affinity, which could provide a support in the lead-optimization process.</p> <p><strong> </strong></p> <p><strong>Availability and implementation:</strong> Our scoring function is freely available through the webportal: <a href="https://predictor.exscalate.eu/">https://predictor.exscalate.eu/</a></p>
Insect Odorant Binding Protein Dataset of Binding Affinities against Volatile Organic Compounds
<p>This is an archival version of the initial iOBPdb dataset of insect odorant binding protein binding affinities against volatile organic compounds. It contains 181 functional studies containing 382 unique OBPs from 91 insect species for 622 individual VOC targets.</p>
PIGNet2: A versatile deep learning-based protein-ligand interaction prediction model for accurate binding affinity scoring and virtual screening
<p>Training and test datasets of the paper "Improving the versatility of deep learning-based protein-ligand interaction prediction for accurate binding affinity scoring and virtual screening".</p>
LMW Binding Affinity Screen for TBXT
<p>Uploaded excel spreadsheets contain the raw data for a SPR campaign targeting TBXT. The data was collected by HDBioscience and is made public by UNC-SGC, The Chordoma Foundation, The Mark Foundation for Cancer Research.</p> <p>If prompted "HDB" is the password for any Excel file.</p> <p>The "ID with Smiles" spreadsheet contains the corresponding SMILES for all the compounds screened. </p> <p>----- </p> <p>More SPR data has been uploaded from 20230103. This dataset has a tab with the SMILES.</p>
Supplementary Data for "Using AlphaFold and Experimental Structures for the Prediction of the Structure and Binding Affinities of GPCR Complexes via Induced Fit Docking and Free Energy Perturbation"
<p>Supplementary data for publication "Using AlphaFold and Experimental Structures for the Prediction of the Structure and Binding Affinities of GPCR Complexes via Induced Fit Docking and Free Energy Perturbation".</p><p>Includes:</p><ul><li>All input structures used in the the retrospective benchmark dataset as well as the (at most) 5 best scoring output models.</li><li>Input structures and output models for IFD-MD predictions of SSTR2, SSTR4, and SSTR5 complexes.</li><li>Output FEP+ maps (in fmp format) for SSTR2, SSTR4, and SSTR5 best models (representative runs shown in publication).</li></ul>
Data for: Accurate sequence-to-affinity models for SH2 domains from multi-round peptide binding assays coupled with free-energy regression
Open the record for dataset details and reuse information.
GAABind: A Geometry-Aware Attention-Based Network for Accurate Protein-Ligand Binding Pose and Binding Affinity Prediction
<p>The preprocessed dataset for paper "GAABind: A Geometry-Aware Attention-Based Network for Accurate Protein-Ligand Binding Pose and Binding Affinity Prediction" with associated code at https://github.com/Mercuryhs/GAABind.</p><p>The dataset files are saved as .pkl file for the convenience of use.</p><p><strong>Paper Abstract</strong>:</p><p>Protein-ligand interactions are increasingly profiled at high-throughput, playing a vital role in lead compound discovery and drug optimization. Accurate prediction of binding pose and binding affinity constitutes a pivotal challenge in advancing our computational understanding of protein-ligand interactions. However, inherent limitations still exist, including high computational cost for conformational search sampling in traditional molecular docking tools, and the unsatisfactory molecular representation learning and intermolecular interaction modeling in deep learning-based methods. Here we propose a geometry-aware attention-based deep learning model, GAABind, which effectively predicts the pocket- ligand binding pose and binding affinity within a multi-task learning framework. Specifically, GAABind comprehensively captures the geometric and topological properties of both binding pockets and ligands, and employs expressive molecular representation learning to model intramolecular interactions. Moreover, GAABind proficiently learns the intermolecular many-body interactions and simulates the dynamic conformational adaptations of the ligand during its interaction with the protein through meticulously designed networks. We trained GAABind on the PDBbindv2020 and evaluated it on the CASF2016 dataset, the results indicate that GAABind achieves state-of-the-art performance in binding pose prediction and shows comparable binding affinity prediction performance. Notably, GAABind achieves a success rate of 82.8% in binding pose prediction, and the Pearson correlation between predicted and experimental binding affinities reaches up to 0.803. Additionally, we assessed GAABind's performance on the SARS-CoV-2 main protease cross-docking dataset. In this evaluation, GAABind demonstrates a notable success rate of 76.5% in binding pose prediction and achieves the highest Pearson correlation coefficient in binding affinity prediction compared with all baseline methods.</p>
Molecular docking analysis was performed to assess the affinity of onalespib for their targets LOX, elucidating binding poses, protein interactions, and associated binding energies.
<p>To analyze the binding affinities and interaction modes between the drug candidates and their targets, we employed the Autodock Vina software [21]. Molecular structures of the candidate drugs and targets of hub genes were retrieved from Pubchem (https://pubchem.ncbi.nlm.nih.gov/) and Protein Data Bank database (http://www.rcsb.org/), respectively. <span>In the analysis of docking, the files for all proteins and molecules were converted to PDBQT format. Water molecules were removed and polar hydrogen atoms were added. The grid box was positioned at the center to encompass the protein domain, allowing for unrestricted movement of molecules.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.