Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
873
datasets available to search
ShareScore release 0.9.0
Dataset results
873 results for “ligands”
Integrated Protein-Ligand Interaction Database
<p><strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. Our database can facilitate projects in <em>machine learning or deep learning-based drug development </em>and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, IPLID offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <p>Data statistics</p> <p>Target (data type) Activities | Unique chemicals | Unique proteins | File name</p> <p>All (binary) 96318 | 18107 | 3107 | integrated_binary_activity.tsv</p> <p>All (numerical) 2798365 | 683009 | 5876 | integrated_continuous_activity.tsv</p> <p>CYP450 (binary) 67552 | 17273 | 47 | integrated_cyp450_binary.tsv</p> <p>CRT (binary) 4152 | 1219 | 412 | integrated_cancer_related_targets_binary.tsv</p> <p>CDT (binary) 519 | 349 | 88 | integrated_cardio_targets_binary.tsv</p> <p>DRT (binary) 4433 | 1325 | 852 | integrated_disease_related_targets_binary.tsv</p> <p>FDA (binary) 6217 | 1521 | 592 | integrated_fda_approved_targets_binary.tsv </p> <p>GPCR (binary) 1958 | 545 | 129 | integrated_gpcr_binary.tsv</p> <p>NR (binary) 1335 | 657 | 264 | integrated_nr_binary.tsv</p> <p>PDT (binary) 1469 | 674 | 404 | integrated_potential_drug_targets_binary.tsv</p> <p>TF (binary) 1966 | 998 | 304 | integrated_tf_binary.tsv</p> <p> </p> <p>*Abbreviations: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p> <p><a href="https://github.com/XieResearchGroup/DrugTargetInteraction/blob/master/iplid/IPLID_data_stat.png">IPLID data statistics</a></p>
Integrated Protein-Ligand Interaction Database
<p>Computational prediction of genome-wide protein-ligand interactions plays a key role in drug discovery, toxicology, and in many other applications. Despite recent advances in <em>deep learning</em>, the large quantity of high-quality data required for training and evaluating models has impeded its applications in computational drug development. To obviate this problem, we have developed an <em>integrated Protein-Ligand Interaction Database </em>(<strong>IPLID</strong>). <strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. <strong>IPLID</strong> is enabled with search functionalities specifically designed for machine learning, particularly deep learning projects. Users can retrieve numerically or binary labeled (e.g. pki, pkd, or binary) protein-ligand interaction data for different classes of proteins (e.g. GPCRs, kinases, FDA-approved targets, protein products of cancer-related genes, etc.). To facilitate the development of benchmarks for training and testing of machine learning algorithms, it also provides chemical-chemical structure similarity scores calculated by a well-established method, Tanimoto coefficient (Jaccard similarity) of two Extended Connectivity Fingerprint (ECFP4) molecular representations. Protein sequence similarities by BLAST score comparison and position-specific scoring matrices against UniRef50 sequence database are also available for more complicated protein-ligand interaction modeling projects. We believe our database can facilitate projects in <em>machine learning or deep learning-based drug development</em> and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, <strong>IPLID</strong> offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <ul> <li>Activities are in <em>tab-delimited</em> text file formats (.tsv).</li> <li>Binary activities are under '<em>binary_activity</em>' directory, and numerical activities are under '<em>numerical_activity</em>' directory.</li> <li>File names are in "(<strong>targets</strong>)_(<strong>activity_type</strong>).tsv"</li> <li>Long target names are abbreviated; abbreviations listed below.</li> <li>Ligand-ligand similarity scores are under '<em>ligand_info</em>' directory.</li> <li>Protein-protein similarity scores and position-specific scoring matrices are under '<em>protein_info</em>' directory.</li> <li>Primary ligand-id and protein-id are <em>InChIKey</em> and <em>UniProt ID</em>, respectively.</li> </ul> <p>*<strong>Abbreviations</strong>: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p>
Supplemental Data: Differential roles of kinetic on- and off-rates in T-cell receptor signal integration revealed with a modified Fab'-DNA ligand
<p>Microscopy data and associated code for analysis. For more information refer to https://doi.org/10.1101/2024.04.01.587588.</p>
PanDDA analysis of ligand screen against the NSP3 macrodomain of SARS-CoV-2: ligands from FrankenROCS fragment-linking pipeline and subsequent optimization of AVI-313
<p>This deposition contains the X-ray diffraction data used to run PanDDA in the ligand screen against the NSP3 macrodomain of SARS-CoV-2 described in Correy et al. 2024 (doi: https://doi.org/10.1101/2024.08.25.609621). Compounds were from fragment linking using FrankenROCS and subsequent optimization of AVI-313. </p> <p>frankenROCS_mac1.tar.gz contains structure factor intensities, PanDDA input/output and refined models/maps.</p> <p>frankenROCS_mac1_ligand-bound-states.tar.gz contains the ligand-bound states extracted from the multi-state PDB files.</p>
Preserving Precise Choreography of Bonds in Z-Stereoretentive Olefin Metathesis Even at Elevated Temperature by Using Quinoxaline-2,3-dithiolate Ligand
<p>Data confirming the structure of the new compounds obtained within the project, published in Nature Communications. The research was performed within OPUS project and was funded by National Science Centre, Poland, grant number UMO-2019/33/B/ST4/00874.</p>
Dataset 3 for paper: "Estimation of free energy of ligand binding using Multi-eGO"
<p>Dataset 3 contains Abeta42 dataset:</p> <ul> <li>APO: reference and multi-eGO simulations</li> <li>HOLO: reference and multi-eGO simualtions</li> <li>Titration at multiple concentrations with two different multi-eGO parameters</li> </ul>
Dataset 2 for paper "Esimation of free energy of ligand binding using Multi-eGO"
<p>Dataset2 contains Kinase dataset and part of Lysozyme-Benzene dataset:</p> <p>LYZ-BNZ:</p> <ul> <li>Unbiased binding simulations</li> </ul> <p>Kinase:</p> <ul> <li>APO trainng, reference, and multi-eGO simulations</li> <li>HOLO training, reference and multi-eGO simulations of both Dasatinib and PP1</li> <li>Thermodynamic integration of both Dasatinib and PP1</li> </ul> <p> </p> <p> </p>
Dataset 1 for paper: "Estimation of free energy of ligand binding using Multi-eGO"
<p>The dataset 1 contains Lysozyme-Benzene dataset:</p> <ul> <li>APO training, reference and multi-eGO simulations</li> <li>HOLO training, reference and multi-eGO simulations</li> <li> Thermodynamic integration calculations, Volume based metadynamics</li> </ul>
Synthesis of β-(Hetero)aryl Ketones via Ligand-Enabled Nondirected C–H Alkylation
<p>This folder contains the DFT-optimized geometries (in .xyz format together with the gas-phase energy, E) accompanying the paper</p> <p>"Synthesis of β-(Hetero)aryl Ketones via Ligand-Enabled Nondirected C–H Alkylation"</p>
[PART 3] A twist of fate: the helix-turn-helix motif in Pseudomonas aeruginosa ExsA can allosterically stabilize the ligand-binding domain
<p>1273- ExsA - Dimer1 without DNA - alphaFold compact - 230ns/day 10x500ns<br>*1274- ExsA - Dimer2 without DNA - xtal extended (like 1026) - 10x1µs (not everything is 1 µs)<br>1275- ExsA - Dimer3 without DNA - made by hand / 150ns/day (reps 1-5 broken, reps 6-10 correct)<br>*1276- ExsA- monomer - Site1-PexoT - 240 ns/day - 10x1µs (some replicas have pieces missing)<br>*1277- ExsA- monomer - Site2-PexoT - 10x1µs<br>*1278- ExsA - Dimer2 with broken DNA - xtal extended (like 1026) - 5x200ns?</p>
[PART 2] A twist of fate: the helix-turn-helix motif in Pseudomonas aeruginosa ExsA can allosterically stabilize the ligand-binding domain
<p>1273- ExsA - Dimer1 without DNA - alphaFold compact <br>1274- ExsA - Dimer2 without DNA - xtal extended <br>1275- ExsA - Dimer3 without DNA - made by hand (reps 1-5 broken, reps 6-10 correct)<br>1276- ExsA- monomer - Site1-PexoT <br>1277- ExsA- monomer - Site2-PexoT<br>1278- ExsA - Dimer2 with broken DNA</p>
Assessing interaction recovery of predicted protein-ligand poses
<p>We provide the following data:</p> <ol> <li>PDB files containing protein structures from the PoseBusters dataset prepared with OpenEye's SPRUCE protein preparation software</li> <li>SDF files containing the corresponding PoseBusters ligand in its crystal pose</li> </ol> <p>This data is provided for the 256 PoseBuster targets used in our paper "Assessing interaction recovery of predicted protein-ligand poses" [1] with associated code at https://github.com/Exscientia/plif_validity.</p> <p> </p> <h3>References</h3> <p>[1] Errington D, Schneider C, Bouysset C, Dreyer FA, Assessing interaction recovery of predicted protein-ligand poses, arXiv; 2024. Available from: https://arxiv.org/abs/2409.20227 </p>
ligand-target dataset from NicheNet
Open the record for dataset details and reuse information.
Data from: 'Venus trapped, Mars transits': Cu and Fe redox chemistry, cellular topography and in situ ligand binding in terrestrial isopod hepatopancreas
Woodlice efficiently sequester copper (Cu) in 'cuprosomes' within hepatopancreatic 'S' cells. Binuclear 'B' cells in the hepatopancreas form iron (Fe) deposits; these cells apparently undergo an apocrine secretory diurnal cycle linked to nocturnal feeding. Synchrotron-based m-focus X-ray spectroscopy undertaken on thin sections was used to characterize the ligands binding Cu and Fe in S and B cells of Oniscus asellus (Isopoda). Main findings were: (i) morphometry confirmed a diurnal B-cell apocrine cycle; (ii) X-ray fluorescence (XRF) mapping indicated that Cu was co-distributed with sulfur (mainly in S cells), and Fe was co-distributed with phosphate (mainly in B cells); (iii) XRF mapping revealed an intimate morphological relationship between the basal regions of adjacent S and B cells; (iv) molecular modelling and Fourier transform analyses indicated that Cu in the reduced Cuþ state is mainly coordinated to thiol-rich ligands (Cu–S bond length 2.3 A˚ ) in both cell types, while Fe in the oxidized Fe3þ state is predominantly oxygen coordinated (estimated Fe–O bond length of approx. 2 A˚ ), with an outer shell of Fe scatterers at approximately 3.05 A˚ ; and (v) no significant differences occur in Cu or Fe speciation at key nodes in the apocrine cycle. Findings imply that S and B cells form integrated unit-pairs; a functional role for secretions from these cellular units in the digestion of recalcitrant dietary components is hypothesized.
Delineating the ligand-receptor interactions that lead to biased signaling at the mu-opioid receptor
<p>Dataset for the manuscript entitled "Delineating the ligand-receptor interactions that lead to biased signaling at the mu-opioid receptor."</p>
Ligand docking models of vanillic acid decarboxylase (VdcCD)
<p>Docking of substrates and intermediates to the vanillic acid decarboxylase, VdcCD, in models of domain closure.</p>
The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction
<p>Datasets, features, and some representative scripts utilized in the paper "The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction".</p>
Ligand binding remodels protein side chain conformational heterogeneity
<p>While protein conformational heterogeneity plays an important role in many aspects of biological function, including ligand binding, its impact has been difficult to quantify. Macromolecular X-ray diffraction is commonly interpreted with a static structure, but it can provide information on both the anharmonic and harmonic contributions to conformational heterogeneity. Here, through multiconformer modeling of time- and space-averaged electron density, we measure conformational heterogeneity of 743 stringently matched pairs of crystallographic datasets that reflect unbound/apo and ligand-bound/holo states. When comparing the conformational heterogeneity of side chains, we observe that when binding site residues become more rigid upon ligand binding, distant residues tend to become more flexible, especially in non-solvent exposed regions. Among ligand properties, we observe increased protein flexibility as the number of hydrogen bonds decrease and relative hydrophobicity increases. Across a series of 13 inhibitor bound structures of CDK2, we find that conformational heterogeneity is correlated with inhibitor features and identify how conformational changes propagate differences in conformational heterogeneity away from the binding site. Collectively, our findings agree with models emerging from NMR studies suggesting that residual side chain entropy can modulate affinity and point to the need to integrate both static conformational changes and conformational heterogeneity in models of ligand binding.</p>
Fragment screening using biolayer interferometry reveals ligands targeting the SHP-motif binding site of the AAA+ ATPase p97
<p>Biosensor techniques have become increasingly important for fragment-based drug discovery during the last years. The AAA+ ATPase p97 is an essential protein with key roles in protein homeostasis and a possible target for cancer chemotherapy. Currently available p97 inhibitors address its ATPase activity and globally impair p97-mediated processes. In contrast, inhibition of cofactor binding to the N-domain by a protein-protein-interaction inhibitor would enable the selective targeting of specific p97 functions. Here, we describe a biolayer interferometry-based fragment screen targeting the N-domain of p97 and demonstrate that a region known as SHP-motif binding site can be targeted with small molecules. Guided by molecular dynamics simulations, the binding sites of selected screening hits were postulated and experimentally validated using protein- and ligand-based NMR techniques, as well as X-ray crystallography, ultimately resulting in the first structure of a small molecule in complex with the N-domain of p97. The identified fragments provide insights into how this region could be targeted and present first chemical starting points for the development of a protein-protein interaction inhibitor preventing the binding of selected cofactors to p97.</p> <p><br> Here we publish the primary data of the biolayer-interferometry (BLI), STD-NMR, HSQC-NMR and mixed-solvent MD simulation results.<br> For BLI, the sensorgrams of 10 binding fragments identified in the screening are given. Measurements were conducted with the N-domain and the ND1-construct of p97. Additionaly, the measurements with ADP as positive control (n=6) are provided.<br> Four of the 10 fragments were selected as most promising hits for further investigation: TROLL2, TROLL7, TROLL8 and TROLL12. For these fragemnts, the STD-NMR build-up data, the HSQC-NMR spectra and the representative poses from the mixed solvent MD simulations used for CORCEMA predictions are provided. For TROLL2 the anomalous map of the bromine atom of the crystal structure (PDB entry: 7PUX) is given.</p>
Dataset for: Kinetics of free and ligand-bound atacicept in human serum
<p>Dataset for the publication "Kinetics of free and ligand-bound atacicept in human serum", Eslami et al, Frontiers in Immunology (2022), containing:</p> <p>a) Pictures of Coomassie blue gels and Western blots used in the study</p> <p>b) An Excel file containing data used to generate graphs of the publication. The file contains several Tabs referring to the indicated figure panels.</p> <p>c) Maps of plasmids used in this study.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.