Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
316
datasets available to search
ShareScore release 0.9.0
Dataset results
316 results for “Drug Target”
Datasets for "Advancing Drug-Target Interactions Prediction: Leveraging a Large-Scale Dataset with a Rapid and Robust Chemogenomic Algorithm"
<p>All datasets required to reproduce the results of publication "Drug-Target Interactions Prediction at Scale: the Komet Algorithm with the LCIdb Dataset"</p>
Phenome-wide association studies across large population cohorts support drug target validation
<p>Summary-level data generated by Genomics plc as presented in:<br> Diogo, D. et al. Phenome-wide association studies across large population cohorts support drug target validation. Nat. Commun. 9, 4285 (2018). https://doi.org/10.1038/s41467-018-06540-3</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at <a href="mailto:research@genomicsplc.com">research@genomicsplc.com</a></p> <p>NOTES<br> -----------------------------<br> These analyses were carried out using the interim UK Biobank imputation data release. Analyses were restricted to a subset of "white-British" unrelated samples with a maximum sample size of 112,337 individuals. </p> <p>Case control phenotypes were defined based on categorical datafields as listed in the accompanying file. <br> Quantitative phenotypes were either rank-normalised before analysis, or beta/se values were standardised after analysis using the variance of the phenotype. The normalisation value is indicated in the accompanying file.<br> <br> All analyses included Age at assessment, sex, genotyping chip, and 10 principal components as covariates. </p> <p>We used plink1.9 linear/logistic regression as appropriate. For chromosome X variants males were treated as having 0 or 2 alternative alleles. </p> <p>The results are not adjusted for genomic control.</p> <p>DATA FILE CONTENT DESCRIPTION<br> -----------------------------<br> CHR - Chromosome<br> SNP - Variant rsID<br> ALT - Alternative allele (effect allele)<br> REF - Reference Allele (non-effect allele)<br> BP - Position in base pairs (b37, 1-based)<br> NMISS - Number of samples with non-missing genotypes<br> BETA - Effect size (log odds ratio or standardised effect size)<br> SE - Standard error<br> P - P-value<br> F_MISS - genotype missing rate<br> P_hwe - Hardy-weinberg p-value<br> MAF - ALT allele frequency</p>
Drug Mechanism of Action Target Database
<p>First release contains data from March 2016 release of DrugCentral database.</p>
SUPPLEMENTARY (For MD) An integrative pan-genome and subtractive proteomics approach for the identification of potential novel therapeutic drug target against antibiotic resistant honeybee pathogen Paenibacillus larvae
<p><strong>Parameters</strong></p><p>Force field: AMBER ff19SB</p><p>Water type: TIP3P</p><p>Ions: NaCl </p><p>Ligand topology force field: GAFF2</p><p>Temperature: 298k</p><p>Pressure: 1 bar</p><p>minimization step: 20000 on 5 nanoseconds</p><p>initial velocity is changed by changing "ntx" and "ig"</p><p>C2: ntx = 5 , ig = 8</p><p>C3: ntx = 2 , ig = 5</p><p> </p><p><strong>Uploads</strong>- </p><p>1. Zip file of all 3 main files</p><p>2. Unzip file of C1 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>3. Zip file of C1</p><p>4. Unzip file of C2 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>5. Zip file of C2</p><p>6. Unzip file of C3 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>7. Zip file of C3</p><p>8. Zip and unzip file of <strong>Initial</strong> PDB of complex prior to MD simulation with <strong>Post</strong> MD PDB (C1, C2, C3)</p><p>9. Zip file of <strong>topology</strong> files for C1, C2, and C3</p>
Exploiting Pretrained Biochemical Language Models for Targeted Drug Design
<p>This repository contains materials for the paper,<em> Exploiting Pretrained Biochemical Language Models for Targeted Drug Design, </em>which<em> </em>has been accepted for publication in <em>Bioinformatics</em> Published by Oxford University Press.</p> <p><em>data.zip</em> contains vocabulary files for the pretrained models, additional information regarding proteins (PFAM family, protein similarity) and interactions filtered from <a href="https://www.bindingdb.org/bind/index.jsp">BindingDB</a> which are further split into train, validation and test sets and used to train target specific molecule generation models. </p> <p><em>models.zip </em>includes files for the models trained in this study. </p> <p><em>predictions.zip </em>comprises the compounds generated with the targeted models and the result of their evaluation with respect to benchmarking metrics. </p> <p><em>docking.zip </em>contains <em>targets/ </em>including PDB files of the test proteins selected for docking evaluation, <em>ligands/ </em>including SDF files for molecules generated with the targeted models and two decoding strategies (i.e. beam search and sampling) and <em>complex/ </em>including docking outputs. </p> <p> </p> <p> </p>
Industry-scale Application and Evaluation of Deep Learning for Drug Target Prediction
<p>Artificial intelligence (AI) is undergoing a revolution thanks to the breakthroughs of machine learning algorithms in computer vision, speech recognition, natural language processing and generative modelling. Recent works on publicly available pharmaceutical data showed that AI methods are highly promising for Drug Target prediction. However, the quality of public data might be different than that of industry data due to different labs reporting measurements, different measurement techniques, fewer samples and less diverse and specialized assays. As part of a European funded project (ExCAPE), that brought together expertise from pharmaceutical industry, machine learning, and high-performance computing, we investigated how well machine learning models obtained from public data can be transferred to internal pharmaceutical industry data. Our results show that machine learning models trained on public data can indeed maintain their predictive power to a large degree when applied to industry data. Moreover, we observed that deep learning derived machine learning models outperformed comparable models, which were trained by other machine learning algorithms, when applied to internal pharmaceutical company datasets. To our knowledge, this is the first large-scale study evaluating the potential of machine learning and especially deep learning directly at the level of industry-scale settings and moreover investigating the transferability of publicly learned target prediction models towards industrial bioactivity prediction pipelines.</p>
Large curated dataset for drug target interaction
<p>A large data curation from the PubChem, ChEMBL and BindingDB public sources. The curated dataset includes samples of pairs of small molecules and protein targets, with information about their binding interactions. The data is stored in an efficient tables format, decoupling entity IDs from their string representations, to avoid redundancy. The curation also includes meaningful splits of the dataset into train, validation and test sets for the purpose of utilizing it for learning based affinity prediction models.</p>
Retinal proteome profiling of inherited retinal degeneration across three different mouse models suggests common drug targets in retinitis pigmentosa
Open the record for dataset details and reuse information.
Comparative assessment of line-probe assays and targeted next-generation sequencing in drug-resistant tuberculosis diagnosis
Open the record for dataset details and reuse information.
Bespoke plant glycoconjugates for gut microbiota-mediated drug targeting
Open the record for dataset details and reuse information.
Nano MOFs as targeted drug delivery agents to combat antibiotic resistant bacterial infections
<p>The drug resistance of bacteria is a significant threat to human civilization while the action of antibiotics against drug-resistant bacteria is severely limited due to the hydrophobic nature of drug molecules, which unquestionably inhibit its permanency for clinical applications. The antibacterial action of nanomaterials offers major modalities to combat drug resistance of bacteria. The current work reports, the use of nano MOFs encapsulating drug molecules to enhance its antibacterial activity against model drug-resistant free living bacteria and biofilm of the bacteria. We have attached rifampicin (RF), a well-documented antituberculosis drug with tremendous pharmacological significance, into the pore surface of zeolitic imidazolate framework 8 (ZIF8) by a <span><span>simple synthetic procedure</span></span><span>.</span> The synthesized ZIF8 has been characterized using X-ray diffraction (XRD) method before and after drug encapsulation. The electron microscopic strategies such as scanning electron microscope (SEM) and transmission electron microscope (TEM) methods was performed to characterize the binding between ZIF8 and RF. We have also performed picosecond resolved fluorescence spectroscopy to validate the formation of the ZIF8-RF nanohybrids (NHs). The drug release profile experiment demonstrates that ZIF8-RF depicts pH-responsive drug delivery and ideal for targeting bacterial disease corresponding to its inherent acidic nature. Most remarkably, ZIF8-RF gives enhanced antibacterial activity against methicillin-resistant <i>S. aureus</i> (MRSA) bacteria and also prompts entire damage of structurally robust bacterial biofilms. Overall, the present study depicts a detailed physical insight for manufactured antibiotic-encapsulated NHs presenting tremendous antimicrobial activity that can be beneficial for manifold practical applications.</p>
The dataset used in the article "Evidential Deep Learning for Guided Drug-Target Interaction Prediction"
<p>The file consists of two parts, the first part is the feature file used for model training and the second part is the saved model.</p> <p>bond_angle file is the feature file for the drugbank dataset, KIBA file is the feature file for the kiba dataset, davis is the feature file for the davis dataset. case_drugbank is the dungbank feature file for the independent test set and case_len_drug is the feature file for the patent validation dataset in the TKIs feature file.len_vs is the feature file for TKIs virtual screening dataset.</p> <p>In the runs folder, the models saved by random division on the Drugbank, KIBA, and Davis datasets are stored.</p> <p>The <a href="https://zenodo.org/api/records/14056305/draft/files/bert_weights_encoderMedium_10.h5/content" target="_blank" rel="noopener noreferrer">bert_weights_encoderMedium_10.h5</a> is the parameter for the small molecule pre-training model, MG-BERT.</p>
Gene prioritization scores and drug target information
<p>This data folder accompanies the article Sadler MC, Auwerx C, Deelen P, Kutalik Z. Multi-layered genetic approaches to identify approved drug targets. Cell Genomics 3, no. 7 (July 2023): 100341(https://doi.org/10.1016/j.xgen.2023.100341).</p><p>It contains disease-drug-target links as well as gene prioritisation scores of all the assessed methods.</p><p> </p>
Efficient Drug-Target Interactions Prediction Framework via Transferable Knowledge Fusing
<p><span>Drug-target integrations (DTI) prediction is a niche in drug discovery, streamlining the search for potential drugs. Computer-aided drug discovery (CADD) has gained traction for its precise predictions, efficiency, and adaptability across various situations. Yet, the computational demands of current top CADD models hinder their practical use due to heavy resource needs.</span></p> <p><span>In this research, we introduce TransFusE DTI, an effective framework for predicting DTIs that leverages pre-trained knowledge to construct models that optimize predictive accuracy while minimizing computational demands. The encoder uses a pre-extracted embedding vector from ProtBERT to reduce computational load and adapts a smaller ProtBERT model. It also includes target-related functional text to boost predictive accuracy. We evaluate the performance of TransFusE DTI using three widely-recognized benchmark datasets: BIOSNAP, DAVIS, and BindingDB, and compare its results to prior studies.</span></p> <p><span>Our results demonstrate that TransFusE DTI exhibits superior predictive performance on the BIOSNAP and BindingDB datasets. Notably, the model's parameter count is only 60% of that of the previous top-performing model by </span><span><a href="https://www.mdpi.com/1999-4923/14/8/1710"><span>Kang et al. (2022)</span></a></span><span>, and it operates efficiently with a learning rate of 26%. Furthermore, the model's video memory requirement is 11.2 GB, rendering it suitable for use on general-purpose Graphics Processing Units (GPUs). </span></p>
Expanding drug targets for 112 chronic diseases using a machine learning-assisted genetic priority score
<h2>ML-GPS: Machine Learning-Assisted Genetic Priority Score</h2> <p>This Zenodo repository contains data and code associated with the publication:</p> <p>Chen R, Duffy Á, Petrazzini BO, Vy HM, Stein D, Mort M, Park JK, Schlessinger A, Itan Y, Cooper DN, Jordan DM, Rocheleau G, Do R. Expanding drug targets for 112 chronic diseases using a machine learning-assisted genetic priority score. Nat Commun. 2024 Oct 15;15(1):8891. doi: <a href="https://doi.org/10.1038/s41467-024-53333-y">10.1038/s41467-024-53333-y</a>.</p> <h3>Important notes</h3> <ul> <li>You can interactively view the top 10% of ML-GPS predictions without download at <a href="https://rstudio-connect.hpc.mssm.edu/mlgps/">https://rstudio-connect.hpc.mssm.edu/mlgps/</a>.</li> <li>For running Jupyter notebooks, please follow the instructions in the README of the GitHub repository at <a href="https://github.com/robchiral/ML-GPS">https://github.com/robchiral/ML-GPS</a>.</li> </ul> <h3>Repository contents</h3> <p>Files needed to train ML-GPS and ML-GPS DOE:</p> <ul> <li><strong>Files needed for Jupyter notebooks.zip</strong>: Data files required for preprocessing and training.</li> <li><strong>Jupyter notebooks.zip</strong>: Notebooks for cleaning data, training models, and generating predictions.</li> </ul> <h3>Other files:</h3> <ul> <li><strong>Predictions for all gene-phecode pairs.zip</strong>: ML-GPS and ML-GPS DOE scores for all analyzed gene-phecode pairs.</li> <li><strong>Summary statistics.zip</strong>: Genetic association summary statistics for all tested gene-phecode pairs.</li> </ul> <h3>Updated performance metrics</h3> <table> <tbody> <tr> <td><strong>Model</strong></td> <td><strong>Open Targets AUPRC</strong></td> <td><strong>SIDER AUPRC</strong></td> </tr> <tr> <td>ML-GPS (non-DOE)</td> <td>0.074</td> <td>0.080</td> </tr> <tr> <td>ML-GPS DOE (activator predictions)</td> <td>0.029</td> <td>0.042</td> </tr> <tr> <td>ML-GPS DOE (inhibitor predictions)</td> <td>0.067</td> <td>0.064</td> </tr> </tbody> </table> <h3>Zenodo versions</h3> <ul> <li><strong>Version 4: </strong>Updated notebooks and external data to use Open Targets 2024.9; summary statistics are unchanged</li> <li><strong>Version 3: </strong>Corrected error where DOE for rare and ultrarare variants was incorrectly incorporated</li> <li><strong>Version 2: </strong>Original release accompanying the publication</li> </ul>
PIGNet: A physics-informed deep learning model toward generalized drug-target interaction predictions
<p>Training, test datasets of the paper "PIGNet: A physics-informed deep learning model toward generalized drug-target interaction predictions".</p>
BIOMETSCO progress presentation 2022: Identification of Novel Biomarkers and Drug Targets for the Detection and Elimination of Occult Metastases in Colon Cancer
<p>This is a recorded talk with head of the BIOMETSCO project - Lasse Sommer Kristensen - where he presents the latest project progress.</p> <p>The talk was given at one of ODIN's (the Open Discovery Innovation Network) Knowledge Sharing Events in May 2022.</p> <p> </p> <p> </p>
Molecular Dynamics Simulation and Docking Studies Reveals Inhibition of NF-kB signaling as a Promising Therapeutic Drug Target for reduction in Cytokines Storms
<p><span>The complexes of the top identified molecules with NF-kB-kB site, as well as all the designed molecules used in the screening process. </span></p>
A novel preclinical secondary pharmacology resource illuminates target-adverse drug reaction associations of marketed drugs - Supplementary Material
<p>All Supplementary materials and source data for the manuscript.</p> <p>The reported analyses and results can be reproduced via the described IPython notebooks which can be found under www.github.com/Novartis/SPD</p>
Data for: A cell surface-binding antibody atlas nominates a MUC18-directed antibody-drug conjugate for targeting melanoma
<p><span>Recent advances in targeted therapy and immunotherapy have substantially improved the treatment of melanoma. However, therapeutic strategies are still needed for unresponsive or treatment-relapsed melanoma patients. To discover antibody-drug conjugate (ADC)-tractable cell surface targets for melanoma, we developed an atlas of melanoma cell surface binding antibodies (pAbs) using a proteome-scale antibody array platform (PETAL). Target identification of pAbs led to development of melanoma cell killing ADCs against LGR6, TRPM1, ASAP1, and MUC18, among others. MUC18 was overexpressed in both tumor cells and tumor-infiltrating blood vessels across major melanoma subtypes, making it a potential dual-compartment and universal melanoma therapeutic target. AMT-253, an MUC18-directed ADC based on topoisomerase I inhibitor exatecan and a self-immolative T moiety, had a higher therapeutic index compared to its microtubule inhibitor-based counterpart and favorable pharmacokinetics and tolerability in monkeys. AMT-253 exhibited MUC18-specific cytotoxicity through DNA damage and apoptosis and a strong bystander killing effect, leading to potent antitumor activities against melanoma cell line and patient-derived xenograft models. Tumor vasculature-targeting by a mouse MUC18-specific antibody-T1000-exatecan conjugate inhibited tumor growth in human melanoma xenografts. Combination therapy of AMT-253 with an anti-angiogenic agent generated higher efficacy than single agent in a mucosal melanoma model. Beyond melanoma, AMT-253 was also efficacious in a wide range of MUC18-expressing solid tumors. Efficient target/antibody discovery in combination with the T moiety-exatecan linker-payload exemplified here may facilitate discovery of new ADC to improve cancer treatment</span><span>.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.