Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.7.1
Dataset results
30 results for “protein-ligand”
Integrated Protein-Ligand Interaction Database
<p>Computational prediction of genome-wide protein-ligand interactions plays a key role in drug discovery, toxicology, and in many other applications. Despite recent advances in <em>deep learning</em>, the large quantity of high-quality data required for training and evaluating models has impeded its applications in computational drug development. To obviate this problem, we have developed an <em>integrated Protein-Ligand Interaction Database </em>(<strong>IPLID</strong>). <strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. <strong>IPLID</strong> is enabled with search functionalities specifically designed for machine learning, particularly deep learning projects. Users can retrieve numerically or binary labeled (e.g. pki, pkd, or binary) protein-ligand interaction data for different classes of proteins (e.g. GPCRs, kinases, FDA-approved targets, protein products of cancer-related genes, etc.). To facilitate the development of benchmarks for training and testing of machine learning algorithms, it also provides chemical-chemical structure similarity scores calculated by a well-established method, Tanimoto coefficient (Jaccard similarity) of two Extended Connectivity Fingerprint (ECFP4) molecular representations. Protein sequence similarities by BLAST score comparison and position-specific scoring matrices against UniRef50 sequence database are also available for more complicated protein-ligand interaction modeling projects. We believe our database can facilitate projects in <em>machine learning or deep learning-based drug development</em> and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, <strong>IPLID</strong> offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <ul> <li>Activities are in <em>tab-delimited</em> text file formats (.tsv).</li> <li>Binary activities are under '<em>binary_activity</em>' directory, and numerical activities are under '<em>numerical_activity</em>' directory.</li> <li>File names are in "(<strong>targets</strong>)_(<strong>activity_type</strong>).tsv"</li> <li>Long target names are abbreviated; abbreviations listed below.</li> <li>Ligand-ligand similarity scores are under '<em>ligand_info</em>' directory.</li> <li>Protein-protein similarity scores and position-specific scoring matrices are under '<em>protein_info</em>' directory.</li> <li>Primary ligand-id and protein-id are <em>InChIKey</em> and <em>UniProt ID</em>, respectively.</li> </ul> <p>*<strong>Abbreviations</strong>: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p>
Assessing interaction recovery of predicted protein-ligand poses
<p>We provide the following data:</p> <ol> <li>PDB files containing protein structures from the PoseBusters dataset prepared with OpenEye's SPRUCE protein preparation software</li> <li>SDF files containing the corresponding PoseBusters ligand in its crystal pose</li> </ol> <p>This data is provided for the 256 PoseBuster targets used in our paper "Assessing interaction recovery of predicted protein-ligand poses" [1] with associated code at https://github.com/Exscientia/plif_validity.</p> <p> </p> <h3>References</h3> <p>[1] Errington D, Schneider C, Bouysset C, Dreyer FA, Assessing interaction recovery of predicted protein-ligand poses, arXiv; 2024. Available from: https://arxiv.org/abs/2409.20227 </p>
The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction
<p>Datasets, features, and some representative scripts utilized in the paper "The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction".</p>
Supplementary Information: Learning Protein-Ligand Binding Affinity with Atomic Environment Vectors
<p>Supplementary Information: Learning Protein-Ligand Binding Affinity with Atomic Environment Vectors</p>
PDBscreen with multiple data augmentation strategies suitable for training protein-ligand interaction prediction methods
<p>PDBscreen with multiple data augmentation strategies suitable for training protein-ligand interaction prediction methods.</p> <p>PDBscreen is the training dataset for EquiScore.</p>
APObind core set for KarmaDock (229 protein-ligand complexes).
<p>APObind core set for KarmaDock (229 protein-ligand complexes). </p>
Structure prediction of protein-ligand complexes from sequence information with Umol
<p>posebusters_benchmark_set.tar.zst - files for the prediction (features to Umol) and scoring of the pose busters benchmark </p><p>posebusters_pred_native.tar.zst - pdb and sdf files of proteins and ligands. Includes native structures, predicted structures and relaxed predicted structures with plDDT in the B factor column.</p><p>posebusters_scores.csv - contains ligand RMSD and other metrics for the unrelaxed structures predicted with Umol.</p><p>PDBBind_processed.tar.zst - files for the training (features to Umol) using PDBbind version 2020</p><p> </p><p> </p>
Data for "Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions"
<p>Data used in "<em>Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions</em>"<br> by F. Pellicani, D. Dal Ben, A. Perali, S. Pilati</p> <p>If you use these data or the python script for your research or other activities, please cite the corresponding journal article.</p> <p> </p> <p>====================</p> <p>Uncompressing the zipped file <em>DataSFUnicam.zip</em> provies the following files and folders:</p> <p><br> <strong>DataSFUnicam/</strong></p> <p> </p> <p> ExperimentalDataPDBFiles/<br> <em>This folder contains 2408 .pdb files of experimental complex structures. The files are named with a univocal code corresponding to the protein-ligand complex.</em></p> <p> </p> <p> ExperimentalDataXLSXFile.xlsx<br> <em>This Excel file reports the experimental protein-ligand chemical information. In the sheet named “Foglio1”, the first column contains the univocal code of the protein-ligand complex, the second column contains the experimentally measured pK_d.</em></p> <p> </p> <p> SyntheticDataPDBFiles/<br> <em>This folder contains the .pdb files of the synthetic complex structures. The .pdb files are grouped in 17 folders according to just as many target proteins. The folders are named after the corresponding protein. Each folder contains the .pdb files for the best position of each protein-ligand pair according to the MOE docking score. The files are named with a univocal code.</em></p> <p> </p> <p> SyntheticDataXLSXFiles/<br> <em> The folder contains 17 Excel files with the chemical information of the synthetic protein-ligand complexes. The files are named after the corresponding target protein. In the sheet named “Foglio1” of each .xlsx file, the first column contains a univocal code of the protein-ligand complex in each conformation, the second column contains an auxiliary numerical code corresponding to the protein-ligand pair, the third column contains the experimentally measured pK_i, and the fourth column contains the docking score provided by the MOE software.</em></p> <p>====================</p> <p>USER GUIDE FOR THE PYTHON SCRIPT</p> <p>Download and uncompress the zipped file "<em>SFUnicam.zip</em>" with a command like "<em>unzip SFUnicam.zip</em>". </p> <p>The following file structure is created:</p> <p><em>SFUnicam/</em></p> <p> <em>ComplexToBePredictedFolder/4ey5_30.pdb <br> MaxAssMatrix.npy<br> my_model<br> devStndSynt.npy<br> mediaSynt.npy<br> UnicamSF13prot.py<br> README.txt</em><br> <br> The subfolder "<em>ComplexToBePredictedFolder/</em>" contains the example PDB file "<em>4ey5_30.pdb</em>".</p> <p>-) To execute the script "<em>UnicamSF13prot.py</em>", Python 3 should be installed with the following libraries and sublibraries:<br> <em>Keras:<br> Regularizers<br> Sequential (keras.models)<br> Conv1D, Dense, MaxPooling1D, GlobalMaxPooling1D, GlobalAveragePooling1D, AveragePooling1D (keras.layers)<br> Adam (keras.optimizers)<br> Numpy</em><br> <em>Tensorflow</em></p> <p>Operation:<br> -) Copy the .pdb file related to the protein-ligand complex whose affinity is to be predicted in the subfolder “<em>ComplexToBePredictedFolder/</em>”.<br> -) Make sure the following files are in the same folder where the python script is:<br> <em>MaxAssMatrix.npy<br> mediaSynt.npy<br> devStndSynt.npy<br> my_model</em><br> -) Run the code using Python 3 with a command like "<em>python3.x UnicamSF13prot.py</em>".<br> -) Enter the name of the protein-ligand PDB file whose affinity is to be predicted (excluding the extension ".pdb").<br> -) Read the predicted affinity from screen.<br> </p> <p> </p>
High-throughput diversification of protein-ligand surfaces to discover chemical inducers of proximity
GEO Series GSE278582. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.
Deep Learning-Ready Voxel Representation of Protein-Ligand Complexes from an Enhanced PBDbind v.2020 Dataset
<p>A critical aspect of successful deep learning (DL) modelling in computer-aided drug discovery (CADD) is the representation of biomolecular data. Voxel grid representations have emerged as a straightforward method for depicting 3D molecular structures of protein-ligand complexes. Proper structural preparation of these complexes is also crucial, particularly in models where the orientation of hydrogen atoms and the accurate assignment of protonation/tautomeric states are vital. The PDBbind, a widely used dataset, can be improved in this regard. This work presents an enhanced version of the PDBbind v.2020 refined set concerning structural preparation, a voxel representation of these structures suitable for DL model training and a diverse set of docking-generated poses that could be used to develop new scoring functions for pose prediction. With this dataset, we aim to provide the CADD community with high-quality, accessible resources to facilitate the development of DL models for drug discovery.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.