Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
80
datasets available to search
ShareScore release 0.9.0
Dataset results
80 results for “binding affinity”
Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction
<p>This Zenodo repository provides comprehensive resources for the paper titled "Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction" published on <a href="https://academic.oup.com/bioinformatics/article/41/8/btaf429/8238154">Bioinformatics</a>. We created a dataset of 63,000 molecular dynamics simulations by performing 10 simulations of 10 ns on 6,300 complexes. Neural networks were developed to learn from this data in order to predict the binding affinities of protein-ligand complexes. The implementation of these neural networks are available on <a href="https://github.com/ICOA-SBC/MD_DL_BA" target="_blank" rel="noopener">github</a>. Our collection includes training/benchmark datasets, trained statistical models, and results on test sets (CSV & PDF files).</p> <p> </p> <p><strong>Training/benchmark datasets:</strong></p> <p>Training, validation and test sets are provided to train and evaluate the following neural networks:</p> <ul> <li>Pafnucy, Proli and Densenucy without MD data augmentation (dataset file names contain "initial")</li> <li>Pafnucy, Proli and Densenucy with MD data augmentation (dataset file names contain "MDDA")</li> <li>Pafnucy with/without MD data augmentation and Proli and Densenucy with MD data augmentation were also evaluated on the fep test set (test set file name contain "fep")</li> <li>Timenucy and Videonucy using spatiotemporal learning methods (dataset file names contain "4D")</li> <li>Pafnucy without MD data augmentation and on a reduced training set (dataset file names contain "reduced")</li> </ul> <p>For each training methodology (MD data augmentation and spatiotemporal learning), we provide the data for the whole complex, only the ligand or only the protein. Additionally for spatiotemporal learning, we provide the data with only the ligand using the tracking mode.</p> <p> </p> <p><strong>Statistical models:</strong></p> <p>We provide the models trained with Pafnucy, Proli, Densenucy, Timenucy and Videonucy. Each models were trained in 10 replicates. </p> <p>For Pafnucy, Proli, Densenucy, we provide the models trained with random and systematic rotations, as well as with or without MD data augmentation.</p> <p>For Proli, Densenucy, Timenucy and Videonucy, we provide the models trained on the whole complex, only the ligand or only the protein.</p> <p>For Pafnucy we also provide the models trained on the reduced set (5932 complexes).</p> <p> </p> <p><strong>Results on test sets (CSV & PDF files):</strong></p> <p>We provide the predictions on the PDBbind v.2016 core set.</p> <ul> <li>For spatiotemporal learning methods (Timenucy and Videonucy), there are predictions for only 83 complexes, as we did not perform simulations on the whole test set.</li> <li>For models trained with MD DA, predictions were carried on the crystallographic structures as well as on the frames extracted from the simulations performed on the test set (augmented test).</li> </ul> <p>Results on the FEP dataset are also provided for Pafnucy, Proli and Densenucy.</p> <p> </p> <p>The Raw MD data (~4.5 To) are stored, and can be visualized/downloaded, on the <a href="https://mdposit.mddbr.eu/#/browse?search=MDBind">MDDB</a>.</p> <p>This work was performed using HPC resources from GENCI-IDRIS (Grant 2021-A0100712496 & 2022-AD011013521) and CRIANN (Grant 2021002).</p>
Data and Code from "Structure-based prediction of Ras-effector binding affinities and design of 'branchegetic' interface mutations"
<p>Data, data generation and data analysis for manuscript "Structure-based prediction of Ras-effector binding affinities and design of ‘branchegetic’ interface mutations", currently available as a preprint <a href="https://doi.org/10.1101/2022.09.04.506480">here</a>.</p> <p>Contains the following directories:</p> <ul> <li>01_models: Contains all scripts for model generation and selection, as well as some of the generated and selected models. <ul> <li>01_inputs: The different inputs for the homology modelling pipeline. This includes AlphaFold single and complex templates, PDB templates and sequence alignments.</li> <li>02_validation: Model generation and initial selection for validation models, based on AF2 single models and PDB complex models.</li> <li>03_production1: Model generation and initial selection for Ras effector complexes, based on AF2 single models and PDB complex models.</li> <li>04_production2: Model generation and initial selection for Ras effector complexes, based on AF2 single models and AF2 complex models.</li> <li>05_selection_optics: Code and analysis for selection by unsupervised learning using OPTICS.</li> </ul> </li> <li>02_selected_models: The three representative models selected for each complex.</li> <li>03_affinity_prediction: Contains code and data for the prediction of binding affinities for Ras effector complexes.</li> <li>04_branch_pruning: Contains code and data for branch pruning analysis.</li> <li>05_systems_analysis: Contains code and data for the analysis of Ras effector systems based on affinities derived from affinity prediction and branch pruning analysis.</li> <li>06_visualization: Information on where in the raw data the panels for the figures in the manuscript can be found.</li> </ul>
DOX_BDW: Incorporating Solvation and Desolvation Effects of Cavity Water into Nonfitting Protein–Ligand Binding Affinity Prediction
<p><strong>structures.zip:</strong> including the coordinates of all optimized proteinligand complex structure obtained by DOX_BDW calculation. (compressed PDB file). These pdb files could also be used as input for the binding energy calculation,as illustrated in SI section 8. </p> <p><strong>mdinput.zip:</strong> Including the input files,parameter files, topology files needed to run MD simulation for water mapping, as illustrated in SI section 8. Note that all of the parameter files and topology files would be automatically generated using the RUNMD program we uploaded with the example file. </p> <p><strong>example.zip:</strong> The programs and input files needed to run an example, as illustrated in SI section 9. And all the output files except MD trajectories are in there,too.</p>
LILBID laser dissociation curves: a mass spectrometry-based method for the quantitative assessment of dsDNA binding affinities
<p>The data and data analysis scripts in this dataset are referenced in the manuscript, "LILBID laser dissociation curves: a mass spectrometry-based method for the quantitative assessment of dsDNA binding affinities", which is in preparation for publication. The contents of this dataset are as follows:</p> <p>1) raw data from UV melting curves<br> 2) settings, concentrations, and both raw and processed data from ITC experiments<br> 3) raw spectrum and imaging data from qLILBID experiments<br> 4) programming scripts used to process the qLILBID data</p> <p><em>Notes on the ITC data:</em><br> <em>The iTC200 microcalorimeter (Malvern Panalytical, Malvern, UK) used in the ITC experiments produces .itc files to be opened and analyzed in Origin (Originlab, Northampton, MA, US) using an add-on. The resulting Origin files, including data interpretation and figures, are provided here. Raw data and interpreted data have been gathered from the .itc files and the Origin files and assembled into tab-separated .dat files, so that the data are also accessible to users who do not have Origin.</em></p> <p><em>The Origin files can be understood as follows. After data collection, the ITC raw data are loaded into the Origin-based software. Initially, the baseline is created (Data1Coeff worksheet) and the data plotted in µcal/second as a function of time (minutes), shown in the mRawITC (graph) and the Data1RAW (data) windows. The peaks are integrated (area in µcal) and then plotted in units of kcal/mole of injectant as a function of molar ratio (injected ligand per molecule in the cell), shown in the DeltaH window. The first injection is negligible and therefore always deleted. According to the data points in the DeltaH window, a curve is fitted to obtain the molar ratio (N), Ka, ΔH and ΔS, the data is shown in the Data1 worksheet. The Data1 worksheet hereby contains the following information: DH: heat change resulting from the given injection (µcal/injection); INJV: volume of the injection; Xt: concentration of injected ligand in the cell before next injection; Mt: concentration of molecule in the cell before next injection; volume corrected; XMt: molar ratio of ligand per molecule in the cell after the injection as displayed in the DeltaH window; NDH: Normalized DH in kcal/moles of injectant as displayed in the DeltaH window, Fit: data points of the fitted curve. In the end, the results are presented in the ITCFINAL window (final figure). Additional information can be found in the MicroCal iTC200 System User Manual.</em></p> <p><em>The same data labeling system has been used for the tab-separated .dat files.</em></p>
Dynamic profiling and binding affinity prediction of NBTI antibacte-rials against DNA gyrase enzyme by multidimensional machine learning and molecular dynamics simulations
<p>The chemical libraries used in this study comprised of 199 and 133 structurally diverse novel bacterial topoisomerase inhibitors (<em>alias</em> NBTIs), with experimentally determined <em>in vitro</em> antibacterial potencies against <em>Staphylococcus aureus</em> DNA gyrase (IC<sub>50</sub>=0.007-50 µM) and <em>Escherichia coli</em> DNA gyrase (IC<sub>50</sub>=0.020-100 µM), respectively (named as NBTI<em><sub>SA</sub></em> and NBTI<em><sub>EC</sub></em>), were compiled from the literature as *.sdf file format. The chemical structures comprising both NBTI libraries were initially sketched by using ChemDraw Professional 20.1.1 suite and subsequently energetically minimized utilizing Discovery Studio’s integrated Merck Molecular Force Field (MMFF) module. Moreover 4D ligands ensembles of both libraries ready to be used for multidimensional QSAR modeling are available, as well.</p>
A Folding–Docking–Affinity framework for protein–ligand binding affinity prediction
Open the record for dataset details and reuse information.
AlphaFold2 Modeling and Molecular Dynamics Simulations of the Conformational Ensembles for the SARS-CoV-2 Spike Omicron JN.1, KP.2 and KP.3 Variants : Mutational Profiling of Binding Energetics Reveals Epistatic Drivers of the ACE2 Affinity and Escape Hotspots of Antibody Resistance
Open the record for dataset details and reuse information.
Machine Learning on the Impacts of Mutations in the SARS-CoV-2 Spike RBD on Binding Affinity to Human ACE2 based on Deep Mutational Scanning Data
Open the record for dataset details and reuse information.
Supplementary Information: Learning Protein-Ligand Binding Affinity with Atomic Environment Vectors
<p>Supplementary Information: Learning Protein-Ligand Binding Affinity with Atomic Environment Vectors</p>
Underlying data for 'Discordant results among MHC binding affinity prediction tools',
<p>Underlying data for ‘Discordant results among MHC binding affinity prediction tools’,</p>
Data from: Facile immobilisation of glucose oxidase onto gold nanostars with enhanced binding affinity and optimal function
Open the record for dataset details and reuse information.
Endogenous miRNA and target concentrations determine susceptibility to potential ceRNA competition, based on hierarchical binding affinities
GEO Series GSE61348. Mus musculus. 9 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing; Other.
Low-affinity SPL binding sites contribute to subgenome expression divergence in allohexaploid wheat [ATAC-seq]
GEO Series GSE188696. Triticum aestivum. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Low-affinity SPL binding sites contribute to subgenome expression divergence in allohexaploid wheat [RNA-seq]
GEO Series GSE188723. Triticum aestivum. 16 samples. Type: Expression profiling by high throughput sequencing.
Administration of anti-HIV-1 broadly neutralizing monoclonal antibodies with increased binding affinity to Fcg receptors during acute SHIV infection shapes innate and adaptive cellular immunity [set1]
GEO Series GSE254781. Macaca mulatta. 175 samples. Type: Expression profiling by high throughput sequencing.
MicroRNA 3'-compensatory pairing occurs through two binding modes, with affinity shaped by nucleotide identity and position [2]
GEO Series GSE196457. Mus musculus. 8 samples. Type: Other.
Low-affinity SPL binding sites contribute to subgenome expression divergence in allohexaploid wheat [DAP-seq]
GEO Series GSE188699. Triticum aestivum. 48 samples. Type: Other.
Chromosomal High Affinity Binding Sites for the Drosophila Dosage Compensation Complex.
GEO Series GSE12292. Drosophila melanogaster. 20 samples. Type: Genome binding/occupancy profiling by genome tiling array.
BANCseq for determination of genome-wide apparent transcription factor binding affinities [ATAC-Seq]
GEO Series GSE219034. Mus musculus. 6 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
BANCseq for determination of genome-wide apparent transcription factor binding affinities [CUT&Run]
GEO Series GSE193553. Homo sapiens; Mus musculus. 72 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.