Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.7.1
Dataset results
30 results for “protein-ligand”
Associated Data: RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features
<p>Additional digital data to "RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features" (ChemRxiv preprint:<a href="https://doi.org/10.26434/chemrxiv.12636704.v1">https://doi.org/10.26434/chemrxiv.12636704</a>).</p> <p>Associated code can be found at: <a href="https://github.com/HITS-MCM/RASPDplus">https://github.com/HITS-MCM/RASPDplus</a></p> <p>Files:</p> <ul> <li>weights.tar.gz: contains the model weights of one random dataset split and its associated crossvalidation folds. Used for standard RASPD+ evaluation.</li> <li>additional_model_replicates.tar.gz: contains the remaining models trained on the full set of descriptors.</li> <li>external_test_sets.tar.gz: contains the descriptor tables for all external test sets used</li> <li>dude.tar.gz: contains the descriptor tables for and several identifier lists for evaluation on the Directory of Useful Decoys - Enhanced (DUD-E)</li> <li>run_outputs.tar.gz: Performance metric data and predicted values created during the model training and evaluation runs. Basis for the figures and metrics in the manuscript.</li> </ul> <p> </p>
Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction
<p>The code, dataset, and model weights are described in the paper "Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction."</p> <p> </p> <p><strong>experiment_results.zip:</strong> Contains generated results that can reproduce the result from the reported paper.</p> <p><strong>benchmark.zip:</strong> Contains docking and affinity input data of the interformer. You can use the source code to make predictions and reproduce the number of the reported paper.</p> <p><strong>checkpoints.zip: </strong>Contains one weight for the Energy and four PoseScore and Affinity models.</p> <p><strong>source_code_1.0.zip:</strong> Contains the initial version of the source code.</p> <p><strong>interformer_train.tar.gz:</strong> Contains prepared training data for interformer. poses/ contains all structure need for training, poses/ligand contains the re-docking poses generated by interformer energy, poses/ligand/rcsb contains the conformation of reference ligand, poses/pocket contains all pocket extract by raw PDB from rcsb, poses/uff contains all ligand conformation minimized using UFF from reference ligand, and train/ contains the training csv.</p> <p><strong>baseline_results.tar.gz:</strong> Contains the predictions from three methods: Interformer, DiffDock, and DeepDock. The results align with the exact numbers reported in the paper. For further details, please refer to the <em>eda/ </em>directory.</p> <p> </p> <p>You can also find the newest version of the source code at <a href="https://github.com/tencent-ailab/Interformer" target="_blank" rel="noopener">https://github.com/tencent-ailab/Interformer</a></p> <p> </p>
PSnpBind: A database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow
<p>A key concept in drug design is how natural variants, especially the ones occurring in the binding site of drug targets, affect the inter-individual drug response and efficacy by altering binding affinity. These effects have been studied on very limited and small datasets while, ideally, a large dataset of binding affinity changes due to binding site single-nucleotide polymorphisms (SNPs) is needed for evaluation. However, to the best of our knowledge, such a dataset does not exist. Thus, a reference dataset of ligands binding affinities to proteins with all their reported binding sites’ variants was constructed using a molecular docking approach. Having a large database of protein-ligand complexes covering a wide range of binding pocket mutations and a large small molecules’ landscape is of great importance for several types of studies. For example, developing machine learning algorithms to predict protein-ligand affinity or a SNP effect on it requires an extensive amount of data. In this work, we present PSnpBind: A large database of mutated binding site protein-ligand complexes constructed using a multithreaded virtual screening workflow. It provides a web interface to explore and visualize the protein-ligand complexes and a REST API to programmatically access the different aspects of the database contents. PSnpBind is freely available at <a href="https://psnpbind.org">https://psnpbind.org</a>.<strong> </strong>The source code of the tools used in constructing PSnpBind is available on <a href="https://github.com/ammar257ammar/PSnpBind-Build">GitHub</a>.</p>
Protein-ligand interactions between mAMCase and chitin.
<p>This directory contains all files to analyze mouse AMCase-ligand interactions as presented in <strong>Supplemental Figure 6</strong> in the manuscript <a href="https://www.biorxiv.org/content/10.1101/2023.06.03.542675">Díaz et al.<em> </em>(2023)</a>.</p> <p>Structure models were analyzed in PyMOL. Figures were compiled using Adobe Illustrator.</p> <p> </p> <p>Contact:</p> <p>Roberto Efraín Díaz, robertoefrain.diaz@ucsf.edu</p> <p>James Fraser, jfraser@fraserlab.com</p>
FABind: Fast and Accurate Protein-Ligand Binding
<p>The preprocessed PDBbind2020 dataset for paper "FABind: Fast and Accurate Protein-Ligand Binding" with associated code at <a href="https://github.com/QizhiPei/FABind">https://github.com/QizhiPei/FABind</a>.</p><p>The dataset files are saved as .pt and lmdb file for the convenience of use.</p><p>We follow the same preprocessing as TankBind.<br><br><strong>Paper Abstract:</strong></p><p>Modeling the interaction between proteins and ligands and accurately predicting their binding structures is a critical yet challenging task in drug discovery. Recent advancements in deep learning have shown promise in addressing this challenge, with sampling-based and regression-based methods emerging as two prominent approaches. However, these methods have notable limitations. Sampling-based methods often suffer from low efficiency due to the need for generating multiple candidate structures for selection. On the other hand, regression-based methods offer fast predictions but may experience decreased accuracy. Additionally, the variation in protein sizes often requires external modules for selecting suitable binding pockets, further impacting efficiency. In this work, we propose FABind, an end-to-end model that combines pocket prediction and docking to achieve accurate and fast protein-ligand binding. FABind incorporates a unique ligand-informed pocket prediction module, which is also leveraged for docking pose estimation. The model further enhances the docking process by incrementally integrating the predicted pocket to optimize protein-ligand binding, reducing discrepancies between training and inference. Through extensive experiments on benchmark datasets, our proposed FABind demonstrates strong advantages in terms of effectiveness and efficiency compared to existing methods. Our code is available at <a href="https://github.com/QizhiPei/FABind">https://github.com/QizhiPei/FABind</a>.</p>
Exploring Data-Driven Chemical SMILES Tokenization Approaches to Identify Key Protein-Ligand Binding Moieties
<p>This repository contains materials for the paper, "Exploring Data-Driven Chemical SMILES Tokenization Approaches to Identify Key Protein-Ligand Binding Moieties", published in <a href="https://onlinelibrary.wiley.com/doi/10.1002/minf.202300249">Molecular Informatics.</a></p> <p>`data.zip` contains vocabulary and dataset files for identifying chemical vocabularies and key chemical words associated with protein ligand binding. </p> <p>`results.zip` comprises outputs specific to vocabularies and datasets, as well as various related statistics.</p> <p> </p> <p> </p>
The dataset used in the article "A point cloud graph neural network for protein-ligand binding site prediction"
Open the record for dataset details and reuse information.
Data Set "Protein-Ligand Interaction Energies from Quantum-Chemical Fragmentation Methods: Upgrading the MFCC-Scheme with Many-Body Contributions"
<p>This data set accompanies the publication "Protein-Ligand Interaction Energies from Quantum-Chemical Fragmentation Methods: Upgrading the MFCC-Scheme with Many-Body Contributions" by Johannes Vornweg and Christoph R. Jacob (TU Braunschweig, Germany) </p> <p>It contains the following files:</p> <p><br>Directory 1_structures:</p> <p> PDB files of all structures used for the test calculations. <br> The PDB files correspond to the protonated structures obtained <br> as described in the main text.</p> <p><br>Directory 02_figure_scripts:</p> <p> Jupyter Notebook for generating all plots included in the manuscript, <br> including raw numerical data.</p> <p><br>Directory 03_input_scripts:</p> <p> - min_congrad.mdp: input file for partial optimization of protonated <br> protein--ligand complexes with Gromacs</p> <p> PyADF input scripts:</p> <p> - sp_single.pyadf: single-point calculations of separate protein and ligand<br> - sp_complex.pyadf: single-point calculation of protein-ligand complex<br> - mfccmbe3.pyadf: MFCC and MFCC-MBE(2) calculations of P-L interaction energy</p> <p> These scripts can be used with PyADF v1.5 (DOI: 10.5281/zenodo.13236550)</p>
Targeting protein-ligand neosurfaces with a generalizable deep learning tool
<p>Molecular recognition events between proteins drive biological processes in living systems. However, higher levels of mechanistic regulation have emerged, where protein-protein interactions are conditioned to small molecules. Despite recent advances, computational tools for the design of novel chemically-induced protein interactions have remained a challenging task for the field. Here, we present a computational strategy for the design of proteins that target neosurfaces, i.e. surfaces arising from protein-ligand complexes. To do so, we leveraged a geometric deep learning approach based on learned molecular surface representations and experimentally validated binders against three drug-bound protein complexes: Bcl2:Venetoclax, DB3:Progesterone and PDF1:Actinonin. All binders demonstrated high affinities and accurate specificities assessed by mutational and structural characterization. Remarkably, surface fingerprints previously trained only on proteins can be applied to neosurfaces emerging from small molecules, serving as a powerful demonstration of generalizability that is uncommon in other deep learning approaches. We anticipate that the designed chemically-induced protein interactions hold the potential to expand the sensing repertoire and the assembly of new synthetic pathways in engineered cells for innovative drug-controlled cell-based therapies</p>
Consensus machine-learning models for protein-ligand binding affinity estimation
<p><strong>Motivation:</strong> In structure-based virtual screening, machine learning based scoring function gained popularity in the last few years as they outperformed classical scoring function. The protein-ligand system can be encoded by a set of orthogonal descriptor spaces, which are then mined by machine learning algorithms to find a relationship with the binding affinity experimental value.</p> <p><strong> </strong></p> <p><strong>Results:</strong> In this work we propose our modelling approach to derive a new scoring function, derived from a combination of multiple descriptor spaces coupled with machine learning algorithms ensembled in consensus. The SF has been trained on the PDBbind v.2019 data and has been extensively internally and externally validated on a large set of complexes. When benchmarked on the PDBbind core set, it achieved better performance than state-of-the-art counterparts, scoring: R<sub>Pearson </sub>= 0.85-0.86 r<sup>2</sup> = 0.70-0.72 and RMSE = 1.15-1.21. As highlights: (i) an applicability domain definition has been implemented to delimit the SF’s application boundaries, and (ii) a mechanistic interpretation is proposed by investigating the contribution of each protein-ligand atom pairs in the prediction of the binding affinity, which could provide a support in the lead-optimization process.</p> <p><strong> </strong></p> <p><strong>Availability and implementation:</strong> Our scoring function is freely available through the webportal: <a href="https://predictor.exscalate.eu/">https://predictor.exscalate.eu/</a></p>
PIGNet2: A versatile deep learning-based protein-ligand interaction prediction model for accurate binding affinity scoring and virtual screening
<p>Training and test datasets of the paper "Improving the versatility of deep learning-based protein-ligand interaction prediction for accurate binding affinity scoring and virtual screening".</p>
Data set for validation of a Python script for computation of Protein-Ligand Interaction Fingerprints
<p><strong>1. Data set for for validation of the Protein-Ligand Interaction Fingerprints, which includes examples of protein structures (original PDB and equilibrated) and molecular dynamics trajectories (equilibration and ligand dissociation generated using Random Acceleration MD simulations, RAMD)</strong></p> <p><strong>mdifp_validation_data.tar.gz - </strong>archive that contains benchmark dataset for evaluation of the protein-ligand IFP protocol (PDB structures of protonated complexes, ligands, and MOL2 files of ligands) published in D. B. Kokha, B. Doser, S. Richter, F. Ormersbach, X. Cheng, R. C. Wade "A Workflow for Exploring Ligand Dissociation from a Macromolecule: Efficient Random Acceleration Molecular Dynamics Simulation and Interaction Fingerprints Analysis of Ligand Trajectories" J. Chem. Phys. <strong>153</strong>, 125102 (2020); <a href="https://doi.org/10.1063/5.0019088">https://doi.org/10.1063/5.0019088</a></p> <p>(2020) <a href="https://arxiv.org/abs/2006.11066">arXiv:2006.11066</a> </p> <p><strong>2YKI </strong>- protein-ligand complex , PDB ID 2YKI<br> - 2yki_MOE.pdb complex with hydrogen added and energy minimized using MOE software (https://www.chemcomp.com/)<br> - ligand_2yki_MOE.mol2 and ligand_2yki_MOE.pdb - ligand structure with hydrogens prepered by MOE software (https://www.chemcomp.com/)</p> <p><strong>6EI5</strong> - MD trajectory of the protein-ligand complex generated from PDB ID 6EI5<br> - ref-min.pdb minimized structure<br> - ref.prmtop topology file<br> - moe.mol2 - ligand structure in mol2 format<br> - amber2namd2.dcd generated MD trajectory </p> <p><strong>SAD_3-RAMD-03-2020.pkl </strong>- a pkl dataset with IFPs generated from RAMD dissociation trajectory of the complex PDB ID: 5LQ9 (trajectories from the paper Front. Mol. Biosci., 2019 DOI:10.3389/fmolb.2019.00036)</p> <p><strong>HSP90_Gromacs.zip </strong>- an archive that contains three pkl data sets of protein-ligand IFPs (for three HSP90 complexes; PDB ID: 5J64, 5J86, 5LQ9) generated from RAMD dissociation trajectories simulated using new Gromacs-RAMD engine (https://github.com/HITS-MCM/gromacs-ramd)</p> <p>The rest of the files contains data obtained from simulation of the complex of <strong>GPCR muscarinic receptor M2 (PDB ID:4MQT);</strong> immersed in a mixed membrane: 50% CHL, 30% POPC, 20% POPE) with a small molecule agonist iperoxo. <br> - <strong>IXO.pdb and moe.mol2 </strong>- PDBand MOL2 structure of iperoxo<br> - <strong>AMBER_eq.tar.gz</strong> - structure of the equilibrated complex generated using AMBER software<br> -<strong> NAMD_eq.tar.gz </strong>- two equilibration trajectories in dcd format generated using NAMD software <br> - <strong>RAMD_eq.tar.gz </strong>- dissociation tarjectoris of iprtoxo from the M2 protein generated from the last snapshot of two NAMD equilibration trajectories (for each case 2 RAMD dissociaiton trajectories are available) </p> <p>( *csv files were added erroneously and do not belong to the project)</p>
GAABind: A Geometry-Aware Attention-Based Network for Accurate Protein-Ligand Binding Pose and Binding Affinity Prediction
<p>The preprocessed dataset for paper "GAABind: A Geometry-Aware Attention-Based Network for Accurate Protein-Ligand Binding Pose and Binding Affinity Prediction" with associated code at https://github.com/Mercuryhs/GAABind.</p><p>The dataset files are saved as .pkl file for the convenience of use.</p><p><strong>Paper Abstract</strong>:</p><p>Protein-ligand interactions are increasingly profiled at high-throughput, playing a vital role in lead compound discovery and drug optimization. Accurate prediction of binding pose and binding affinity constitutes a pivotal challenge in advancing our computational understanding of protein-ligand interactions. However, inherent limitations still exist, including high computational cost for conformational search sampling in traditional molecular docking tools, and the unsatisfactory molecular representation learning and intermolecular interaction modeling in deep learning-based methods. Here we propose a geometry-aware attention-based deep learning model, GAABind, which effectively predicts the pocket- ligand binding pose and binding affinity within a multi-task learning framework. Specifically, GAABind comprehensively captures the geometric and topological properties of both binding pockets and ligands, and employs expressive molecular representation learning to model intramolecular interactions. Moreover, GAABind proficiently learns the intermolecular many-body interactions and simulates the dynamic conformational adaptations of the ligand during its interaction with the protein through meticulously designed networks. We trained GAABind on the PDBbindv2020 and evaluated it on the CASF2016 dataset, the results indicate that GAABind achieves state-of-the-art performance in binding pose prediction and shows comparable binding affinity prediction performance. Notably, GAABind achieves a success rate of 82.8% in binding pose prediction, and the Pearson correlation between predicted and experimental binding affinities reaches up to 0.803. Additionally, we assessed GAABind's performance on the SARS-CoV-2 main protease cross-docking dataset. In this evaluation, GAABind demonstrates a notable success rate of 76.5% in binding pose prediction and achieves the highest Pearson correlation coefficient in binding affinity prediction compared with all baseline methods.</p>
DynamicBind: Predicting ligand-specific protein-ligand complex structure with a deep equivariant generative model.
<p>test and training data.</p>
Datasets for "Physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data"
<div> <p>Datasets used for implementing the <a href="https://github.com/huankoh/PSICHIC">PSICHIC</a> experiments shown in the <a href="https://doi.org/10.1101/2023.09.17.558145">manuscript</a>.</p> <p> </p> </div>
Targeting protein-ligand neosurfaces using a generalizable deep learning approach [benchmark dataset]
<p>PDB files and processed surface meshes for the binder recovery benchmark. For the larger PDBbind decoy set only the PDB files are provided.</p>
Prepared protein-ligand complex from MOAD
<p>This is a protein-ligand data set. </p> <p>Each compressed (tar.gz) represents a collection of PDB IDs starting with the same initial ID, for example, 1.tar.gz conatins all PDB IDs that start with 1.</p> <p>With each subfolder, protein.pdb is a prepared protein structure file. This file has been properly added with correct hydrogen atoms, protonation information, and completion of missing heavy atoms and residues. The JSON and SDF files with the same name record a ligand, named according to the format: residue name_chain id_residue number. The JSON file contains information about the center coordinates of the ligand's heavy atoms and a box of dimensions +/- 10 angstroms around the ligand, which can be directly used as grid parameters in molecular docking. The SDF file provides the ligand atom coordinates with correct bond order information.</p>
MDD-Molecular Dynamics Dataset: Collection of protein-ligand complex simulations
<p>Dataset is part of the paper: https://chemrxiv.org/engage/chemrxiv/article-details/664c73f6418a5379b0de8152.</p> <p>This dataset consists of molecular dynamics (MD) simulations of 862 unique protein-ligand complexes, covering a wide range of protein families and diverse chemical classes of ligands. It is derived from publicly available repositories and represents the largest single source of MD simulations to date.</p> <p>All protein-ligand complexes included in the dataset were prepared following a standardized protocol. Missing atoms in the protein structures were added using the PDBFixer tool. The protein targets were parameterized using the AMBER99SB-ILDN force field, while ligands were parameterized with the ANTECHAMBER module within the ACPYPE tool. Ligand partial charges were determined to match the quantum-mechanically generated electrostatic potential via the Restrained Electrostatic Potential (RESP) method, and the remaining parameters were set using the GAFF2 force field. The molecular dynamics simulations were performed using GROMACS. The simulations were configured in a cubic simulation box with periodic boundary conditions and employed a TIP3P water model within an electrostatically neutral environment. The simulation protocol included an initial minimization cycle, followed by temperature equilibration in the NVT ensemble and pressure equilibration in the NPT ensemble. Production simulations were conducted over a period of 200 ns, with a timestep of 100 ps.</p> <p>Constructing a large, representative set of MD simulations poses challenges due to the high computational costs and complexities associated with preparing molecular systems. Moreover, given the limited number of suitable training examples (complexes) and the large volume of MD data from each simulation, careful filtering and feature selection are crucial. This dataset is valuable for exploring how molecular dynamics simulation data can be integrated with protein-ligand binding affinity prediction tasks, an essential component of in silico drug discovery pipelines. MD simulations, in particular, offer a dynamic view by illustrating the temporal interactions within protein-ligand complexes, potentially providing additional insights for affinity and specificity estimates.</p>
Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction
<p>This Zenodo repository provides comprehensive resources for the paper titled "Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction" published on <a href="https://academic.oup.com/bioinformatics/article/41/8/btaf429/8238154">Bioinformatics</a>. We created a dataset of 63,000 molecular dynamics simulations by performing 10 simulations of 10 ns on 6,300 complexes. Neural networks were developed to learn from this data in order to predict the binding affinities of protein-ligand complexes. The implementation of these neural networks are available on <a href="https://github.com/ICOA-SBC/MD_DL_BA" target="_blank" rel="noopener">github</a>. Our collection includes training/benchmark datasets, trained statistical models, and results on test sets (CSV & PDF files).</p> <p> </p> <p><strong>Training/benchmark datasets:</strong></p> <p>Training, validation and test sets are provided to train and evaluate the following neural networks:</p> <ul> <li>Pafnucy, Proli and Densenucy without MD data augmentation (dataset file names contain "initial")</li> <li>Pafnucy, Proli and Densenucy with MD data augmentation (dataset file names contain "MDDA")</li> <li>Pafnucy with/without MD data augmentation and Proli and Densenucy with MD data augmentation were also evaluated on the fep test set (test set file name contain "fep")</li> <li>Timenucy and Videonucy using spatiotemporal learning methods (dataset file names contain "4D")</li> <li>Pafnucy without MD data augmentation and on a reduced training set (dataset file names contain "reduced")</li> </ul> <p>For each training methodology (MD data augmentation and spatiotemporal learning), we provide the data for the whole complex, only the ligand or only the protein. Additionally for spatiotemporal learning, we provide the data with only the ligand using the tracking mode.</p> <p> </p> <p><strong>Statistical models:</strong></p> <p>We provide the models trained with Pafnucy, Proli, Densenucy, Timenucy and Videonucy. Each models were trained in 10 replicates. </p> <p>For Pafnucy, Proli, Densenucy, we provide the models trained with random and systematic rotations, as well as with or without MD data augmentation.</p> <p>For Proli, Densenucy, Timenucy and Videonucy, we provide the models trained on the whole complex, only the ligand or only the protein.</p> <p>For Pafnucy we also provide the models trained on the reduced set (5932 complexes).</p> <p> </p> <p><strong>Results on test sets (CSV & PDF files):</strong></p> <p>We provide the predictions on the PDBbind v.2016 core set.</p> <ul> <li>For spatiotemporal learning methods (Timenucy and Videonucy), there are predictions for only 83 complexes, as we did not perform simulations on the whole test set.</li> <li>For models trained with MD DA, predictions were carried on the crystallographic structures as well as on the frames extracted from the simulations performed on the test set (augmented test).</li> </ul> <p>Results on the FEP dataset are also provided for Pafnucy, Proli and Densenucy.</p> <p> </p> <p>The Raw MD data (~4.5 To) are stored, and can be visualized/downloaded, on the <a href="https://mdposit.mddbr.eu/#/browse?search=MDBind">MDDB</a>.</p> <p>This work was performed using HPC resources from GENCI-IDRIS (Grant 2021-A0100712496 & 2022-AD011013521) and CRIANN (Grant 2021002).</p>
Integrated Protein-Ligand Interaction Database
<p><strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. Our database can facilitate projects in <em>machine learning or deep learning-based drug development </em>and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, IPLID offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <p>Data statistics</p> <p>Target (data type) Activities | Unique chemicals | Unique proteins | File name</p> <p>All (binary) 96318 | 18107 | 3107 | integrated_binary_activity.tsv</p> <p>All (numerical) 2798365 | 683009 | 5876 | integrated_continuous_activity.tsv</p> <p>CYP450 (binary) 67552 | 17273 | 47 | integrated_cyp450_binary.tsv</p> <p>CRT (binary) 4152 | 1219 | 412 | integrated_cancer_related_targets_binary.tsv</p> <p>CDT (binary) 519 | 349 | 88 | integrated_cardio_targets_binary.tsv</p> <p>DRT (binary) 4433 | 1325 | 852 | integrated_disease_related_targets_binary.tsv</p> <p>FDA (binary) 6217 | 1521 | 592 | integrated_fda_approved_targets_binary.tsv </p> <p>GPCR (binary) 1958 | 545 | 129 | integrated_gpcr_binary.tsv</p> <p>NR (binary) 1335 | 657 | 264 | integrated_nr_binary.tsv</p> <p>PDT (binary) 1469 | 674 | 404 | integrated_potential_drug_targets_binary.tsv</p> <p>TF (binary) 1966 | 998 | 304 | integrated_tf_binary.tsv</p> <p> </p> <p>*Abbreviations: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p> <p><a href="https://github.com/XieResearchGroup/DrugTargetInteraction/blob/master/iplid/IPLID_data_stat.png">IPLID data statistics</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.