Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

131

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

131 results for “Drug discovery”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset for manuscript titled "Exploring SureChEMBL from a drug discovery perspective".

<p>This is the data directory for running the code available on the GitHub repository for the manuscript titled "Exploring SureChEMBL from a drug discovery perspective<strong></strong>". The GitHub repository is available at <a href="https://github.com/Fraunhofer-ITMP/patent-clinical-candidate-characteristics">https://github.com/Fraunhofer-ITMP/patent-clinical-candidate-characteristics.</a></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Supporting data for: Low-cost anti-mycobacterial drug discovery using engineered E. coli

<p>Supporting data for: Low-cost anti-mycobacterial drug discovery using engineered E. col</p> <p>This dataset pertains to our work developing TESEC Mtb ALR, a genetically engineered strain of E. coli expressing the enzyme ALR derived from Mtb. We used the TESEC Mtb ALR strain in a high-throughput drug screen and identified benazepril as targeted inhibitor of the ALR enzyme. We then performed additional experiments to characterize the activity of benazepril against E. coli, Mtb and purified enzymes. Finally, we tested the extensibility of the platform by constructing and screening against similar strains for additional targets: Asd, CysH, DapB, and TrpD.</p> <p>These files include growth measurements, biochemical assays and other forms of biological data. They are packaged together with scripts used to analyze the data and present them in figures. Our goal in creating this archive was to present our complete analysis pipeline in the spirit of open science. It is not intended to be a stand-alone resource. Consult the associated manuscript for protocols, units of measurement and other essential technical context.</p> <p>Our scripts were written for Python 3.8. The raw data is presented as human-readable .csv files intended to be imported as Pandas DataFrames. Some data is also packaged as Python dictionaries saved with the Pickle package.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Improving the drug discovery process by using multiple classifier systems

<p>High-quality dataset gathered from ChEMBL version 22 based on UniProt accession P34972. Regarding to activity data potential, duplicates were ignored, no activity or data validity comments were allowed, only data from binding assays and with a pCheMBL value were kept. This led to a dataset composed of 3925 chemical compounds (instances) represented using 2132 features. The first 2048 features epitomize different chemical structures fingerprints (represented using FCFP_6 notation), while the remaining 84 are associated with several physicochemical descriptors (such as Fractional Polar Surface Area, Rotatable Bonds&nbsp;or Molecular Weight). Finally, the set was transformed into a binary classification set where the activity cut-off was defined at a pChEMBL value &gt; 7 and written to a tab-delimited text file. The final set contained 1977 active compounds and 1948 inactive compounds. Table 3 shows the codification of each feature grouped by type.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

MISATO - Machine learning dataset for structure-based drug discovery

<p>Developments in Artificial Intelligence (AI) have had an enormous impact on scientific research in recent years. Yet, relatively few robust methods have been reported in the field of structure-based drug discovery. To train AI models to abstract from structural data, highly curated and precise biomolecule-ligand interaction datasets are urgently needed. We present MISATO, a curated dataset of almost 20000 experimental structures of protein-ligand complexes, associated molecular dynamics traces, and electronic properties. Semi-empirical quantum mechanics was used to systematically refine protonation states of proteins and small molecule ligands. Molecular dynamics traces for protein-ligand complexes were obtained in explicit water. The dataset is made readily available to the scientific community via simple python data-loaders. AI baseline models are provided for dynamical and electronic properties. This highly curated dataset is expected to enable the next-generation of AI models for structure-based drug discovery. Our vision is to make MISATO the first step of a vibrant community project for the development of powerful AI-based drug discovery tools.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

MBC and ECBL Libraries: outstanding tools for drug discovery

<p><strong>UPDATE</strong>.&nbsp;New in this revision: python scripts to process DBs and calculate the percentage of molecules&nbsp;which pass the&nbsp;Veber and Ghose filters. Two new DBs&nbsp;were also added and considered for the analysis.</p> <p>Data and scripts to reproduce all the graphics reported in the Manuscript entitled: &quot;MBC and ECBL Libraries: outstanding tools for drug discovery&quot;.</p> <p><strong>List of analyzed DBs:</strong></p> <ol> <li>MBC2016 (Total entries: 1,096 cmpds; 7.39% excluded from properties analysis - QikProp failure).</li> <li>MBC2022 (Total entries: 2,577 cmpds; 3.14% excluded from properties analysis - QikProp failure).</li> <li>ECBL (Total entries: 101,021 cmpds; 0.20% excluded from properties analysis - QikProp failure).</li> <li>ChEMBL v.31 (Total entries 1,908,325 cmpds; 2.97% excluded from properties analysis - QikProp failure).</li> <li>DrugBank v.5.0 (Total entries 10,981 cmpds; 4.13% excluded from properties analysis - QikProp failure).</li> <li>ZINC20 (Total entries 10,723,360 cmpds; 0.61% excluded from properties analysis - QikProp failure).</li> <li>NuBBE (Total entries 2,223 cmpds) - <strong>NEW</strong></li> <li>Approved drugs (Total entries: 3,140 cmpds) - <strong>NEW</strong></li> </ol> <p><strong>Files:</strong></p> <p><em>QikProp_properties.docx</em>: doc file&nbsp;containing the full list of QikProp properties calculated for each analyzed DB.</p> <p><em>DATA_comparison.xlsx</em>: excel file containing data used to reproduce plots in&nbsp;<strong>Figure 4</strong>&nbsp;of the MS.</p> <ul> <li><em>Murcko_scaffold_percentages</em>: distribution (%) of the first 50 most populated Murcko scaffolds for MBC2016, MBC2022 and ECBL.</li> <li><em>Murcko_scaffolds_comparison</em>: distribution (count) of the first 94 common Murcko scaffolds for MBC2016, MBC2022 and ECBL.</li> </ul> <p>QikProp properties for all the analyzed DBs (8&nbsp;files; CSV format).</p> <p>SMILES codes for all the analyzed DBs (8&nbsp;files; SMI format).&nbsp;</p> <p><em>joinplots.py</em>: python script to generate the 2D plots in&nbsp;<strong>Figure 2</strong>&nbsp;of the MS.</p> <p><em>fingerprint_similarity.py</em>: python script to run and generate the Tanimoto similarity plots in&nbsp;<strong>Figure 3</strong>&nbsp;of the MS.</p> <p><em>calc_kde.py</em>: python script to run kernel density analysis reported in&nbsp;<strong>Figure 5&nbsp;</strong>of the MS.</p> <p><em>Veber_filter.py: python&nbsp;script to generate&nbsp;</em>data presented&nbsp;in <strong>Table 1 </strong>of the MS. <strong>(NEW)</strong></p> <p><em>Ghose filter.py: &nbsp;python&nbsp;script to generate&nbsp;</em>data presented&nbsp;in <strong>Table 1 </strong>of the MS.&nbsp;<strong>(NEW)</strong></p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Raw data for the SciPipe machine learning in drug discovery case study

<p>Accompanying raw data for the machine learning in&nbsp;drug discoverycase studies for SciPipe [1] available at&nbsp;https://github.com/pharmbio/scipipe-demo&nbsp;</p> <p>[1]&nbsp;http://scipipe.org</p>

opencc-by-4.0Jul 2018View details →
zenodo40/100

Data source for "A universal programmable Gaussian Boson Sampler for drug discovery"

<p>Source Data for the figures in paper.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery

<p>All of the individual docking data and ESSENCE-Dock consensus results for 21 diverse DUD-E targets as presented in the paper "ESSENCE-Dock: A Consensus-Based Approach to Enhance Virtual Screening Enrichment in Drug Discovery".</p> <p>The data is sorted per DUD-E target. It contains the prepared data that was used for the docking calculations (in the Undocked directory), as well as our docking results. Finally, our ESSENCE-Dock Consensus results are included as well</p> <p>Docking calculations were performed using:</p> <ul> <li><a href="https://github.com/bio-hpc/metascreener">Metascreener (V1.1)</a> (Gnina and LeadFinder Calculations; prefix VS_GN_ and VS_LF_ respectively)</li> <li><a href="https://github.com/Jnelen/DiffDockHPC/tree/DiffDockHPCv1.0">DiffDockHPC (v1.0)</a> (DiffDock calculations; prefix VS_DD_ )</li> </ul> <p>The consensus calculations were performed using ESSENCE-Dock, available via <a href="https://github.com/bio-hpc/metascreener">Metascreener </a>as well.</p> <p>The whole methodology and all of the details are described in the ESSENCE-Dock paper: <a href="https://doi.org/10.1021/acs.jcim.3c01982">https://doi.org/10.1021/acs.jcim.3c01982</a></p> <p><strong>Paper Abstract</strong></p> <p>Drug development is a complex, costly, and time-consuming endeavor. While high-throughput screening (HTS) plays a critical role in the discovery stage, it is one of many factors contributing to these challenges. In certain contexts, virtual screening can complement HTS, potentially offering a more streamlined approach in the initial stages of drug discovery. Molecular docking is an example of a popular virtual screening technique that is often used for this purpose, however, its effectiveness can vary greatly. This has led to the use of consensus docking approaches, which combine results from different docking methods to improve the identification of active compounds and reduce the occurrence of false positives. However, many of these methods do not fully leverage the latest advancements in molecular docking.<br>In response, we present ESSENCE-Dock (Effective Structural Screening ENrichment ConsEnsus Dock), a new consensus docking workflow aimed at decreasing false positives and increasing the discovery of active compounds. By utilizing a combination of novel docking algorithms, we improve the selection process for potential active compounds. ESSENCE-Dock has been made to be user-friendly, requiring only a few simple commands to perform a complete screening, while also being designed for use in high-performance computing (HPC) environments.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

A Comprehensive Dataset of protein-protein interactions and Ligand Binding Pockets for Advancing Drug Discovery

<p>This dataset presents a comprehensive collection of structural data related to protein-protein interactions (PPIs) and ligand binding pockets. The dataset includes high-quality structural information that can aid researchers in the fields of bioinformatics, structural biology, and drug discovery. It encompasses a diverse set of PPI complexes and associated ligands, enabling detailed investigations into molecular interactions at the atomic level. This article introduces an indispensable resource designed to unlock the full potential of PPIs while pioneering a novel metric for pocket similarity for repurposing protein partners.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Predicting transcriptional responses to novel chemical perturbations using deep generative model for drug discovery

<p>Understanding transcriptional responses to chemical perturbations is central to drug discovery, but exhaustive experimental screening of diseasecompound combinations is unfeasible. To overcome this limitation, here we introduce PRnet, a perturbation-conditioned deep generative model that predicts transcriptional responses to novel chemical perturbations that have never experimentally perturbed at bulk and single-cell levels. Evaluations indicate that PRnet outperforms alternative methods in predicting responses across novel compounds, pathways, and cell lines. PRnet enables gene-level response interpretation and in-silico drug screening for diseases based on gene signatures. PRnet further identifies and experimentally validates novel compound candidates against small cell lung cancer and colorectal cancer. Lastly, PRnet generates a large-scale integration atlas of perturbation profiles, covering 88 cell lines, 52 tissues, and various compound libraries. PRnet provides a robust and scalable candidate recommendation workflow and successfully recommends drug candidates for 233 diseases. Overall, PRnet is an effective and valuable tool for gene-based therapeutics screening.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

PPB-Affinity: Protein-Protein Binding Affinity dataset for AI-based protein drug discovery

<p>Prediction of protein-protein binding (PPB) affinity plays an important role in large-molecular drug discovery. Deep learning (DL) has been adopted to predict the changes of PPB binding affinities upon mutations, but there was a scarcity of studies predicting the PPB affinity itself. The major reason is the paucity of open-source dataset with PPB affinity data. To address this gap, the current study introduced a large comprehensive PPB affinity (PPB-Affinity) dataset. The PPB-Affinity dataset contains key information such as crystal structures of protein-protein complexes (with or without protein mutation patterns), PPB affinity, receptor protein chain, ligand protein chain, etc. To the best of our knowledge, this is the largest publicly available PPB affinity dataset, and we believe it will significantly advance drug discovery by streamlining the screening of potential large-molecule drugs. We also developed a deep-learning benchmark model with this dataset to predict the PPB affinity, providing a foundational comparison for the research community.</p> <p>Codes for PPB-Affinity database preparation is &nbsp;disclosed at <a title="https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow" href="https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow">https://github.com/Huatsing-Lau/PPB-Affinity-DataPrepWorkflow</a>.<br>Codes for the benchmark algorithm is disclosed at&nbsp;<a href="https://github.com/ChenPy00/PPB-Affinity">https://github.com/ChenPy00/PPB-Affinity</a>.</p> <p>The article related to the PPB-Affinity dataset can be found at&nbsp;<a href="https://doi.org/10.1038/s41597-024-03997-4">https://doi.org/10.1038/s41597-024-03997-4</a>.</p> <p>Files are orginized as follows:</p> <p><a href="13054646" target="_blank" rel="noopener noreferrer">- PPB-Affinity.xlsx</a></p> <p><a href="13054646" target="_blank" rel="noopener noreferrer">- samples_deleted.zip</a></p> <p><a href="../api/records/13067409/draft/files/PPB-Affinity-AF.zip/content" target="_blank" rel="noopener noreferrer">- PPB-Affinity-AF.zip</a></p> <p>- PDB/</p> <p>&nbsp; - Affinity Benchmark v5.5/</p> <p>&nbsp; &nbsp; - file1.pdb</p> <p>&nbsp; &nbsp; - file2.pdb</p> <p>&nbsp; &nbsp; - ...</p> <p>&nbsp; &nbsp; - filek.pdb</p> <p>&nbsp; - ATLAS/</p> <p>&nbsp; - PDBbind v2020/</p> <p>&nbsp; - SAbDab/</p> <p>&nbsp; - SKEMPIv2.0/</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

CARA: Benchmarking Compound Activity Prediction for Real-World Drug Discovery Applications

<p>Identifying active compounds for target proteins is fundamental in early drug discovery.&nbsp;Recently, data-driven computational methods have demonstrated promising potential in predicting compound activities.&nbsp;However, there lacks a well-designed benchmark to comprehensively evaluate these methods from a practical perspective.&nbsp;To fill this gap, we propose a benchmark, named CARA.Through carefully distinguishing assay types, designing train-test splitting schemes and selecting evaluation metrics, CARA can consider the biased distribution of current real-world compound activity data and avoid overestimation of model performances. We observed that current models can make successful predictions for certain proportions of assays, while the performances varied across different assays. In addition, evaluation of several few-shot training strategies demonstrated different performances related to task types. Overall, we provide a high-quality dataset for developing and evaluating compound activity prediction models, and the analyses in this work may inspire better applications of data-driven models in drug discovery.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery

<p>This contains data described in detail in our paper, &quot;Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery&quot;, where we develop a novel&nbsp;algorithm called RPath that prioritizes drugs for a given disease by reasoning over causal paths in a knowledge graph (KG), guided by both drug-perturbed as well as disease-specific transcriptomic signatures.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Ensembles of knowledge graph embedding models improve predictions for drug discovery

<p>This contains data described in detail in our paper, &quot;Ensembles of knowledge graph embedding models improve predictions for drug discovery&quot;. The metadata involves the different trained models that were used for prediction analysis as well as all the predictions from the trained models.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Workshop Material - 3D-e-Chem Structural Cheminformatics Workflows for Computer-Aided Drug Discovery

<p>The workshop at the KNIME user meeting (Berlin 9th of March 2018) &nbsp;is set up to stimulate participants with varying degrees of experience in cheminformatics to learn and apply the different structural cheminformatics tools and workflows developed within the context of the 3D-e-Chem project. You will learn how to construct and apply integrated cheminformatics workflows using the 3D-e-Chem KNIME nodes for the exploitation of G protein-coupled receptor and kinase data (two important pharmaceutical target classes) to obtain useful information for drug discovery.</p> <p>Information on the 3D-e-Chem KNIME nodes and workflows can be found online:</p> <p>3D-e-Chem GitHub website: <a href="http://3d-e-chem.github.io/">http://3d-e-chem.github.io/</a></p>

openapache2.0Mar 2018View details →
zenodo36/100

Computational biology and how this field supports new drug discovery

<p>PechaKucha 20x20 is a simple presentation format where you show 20 images, each for 20 seconds. The images advance automatically and this presentation will explore the use of computational modeling supports the early stage of drug discovery. https://www.pechakucha.org/presentations/drug-discovery-and-computational-biology</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Traversing Chemical Space with Active Deep Learning for Low-data Drug Discovery

<p>Raw and processed data from LitPCBA used in the paper "Traversing Chemical Space with Active Deep Learning for Low-data Drug Discovery"</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Supplement tables of Facilitating Antiviral Drug Discovery by Genetic and Evolutionary Knowledge

<p>Table S1: 36&nbsp;human targets for approved antiviral drugs downloaded from Drugbank,</p> <p>Table S2: Supplementary information for Table 2,</p> <p>Table S3: Pockdrug druggability prediction results.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

SARS-CoV-2–host proteome interactions for antiviral drug discovery

<p>Images and datasets used in Fig 6 and corresponding supplementary material Image analysis was performed with Harmony 4.9 software (PerkinElmer) with feature extraction and linear classification of N-protein positive cells from the total population as presented in the PlateResults file. Prism files (GraphPad Software) include the calculations for curve fits of drug testing data (4PL logistic regression) as well as area under the curve (AUC) of the fitted curves..</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Exploiting Vector Pattern Diversity of Molecular Scaffolds for Cheminformatics Tasks in Drug Discovery

<p>Data and code to accompany the paper:&nbsp;<em>Exploiting Vector Pattern Diversity of Molecular Scaffolds for Cheminformatics Tasks in Drug Discovery.&nbsp;</em></p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record