Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

35 results for “drug design”

Learn how ShareScore rates datasets ↗
zenodo48/100

Molecular datasets from "SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design"

<p>Herein find the molecular datasets from &quot;<a href="https://chemrxiv.org/articles/SMILES-Based_Deep_Generative_Scaffold_Decorator_for_De-Novo_Drug_Design/11638383">SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design</a>&quot;. These were generated with&nbsp;SMILES-based scaffold decorator generative models&nbsp;trained with two training sets (DRD2 and ChEMBL). These generative models require a partially-built molecule (scaffold) as input and output several possible completions for each scaffold. Each dataset corresponds to a model trained with the&nbsp; ChEMBL or DRD2&nbsp;sets, wither multi-step (ms) or single-step (ss) and the provenance of the scaffolds (validation set, or non-dataset).</p> <p>The molecules generated are annotated with a set of descriptors. The DRD2 datasets have the predicted probability of each molecule to be active&nbsp;on DRD2 (p)&nbsp;obtained from a Random Forest model. The ChEMBL model&#39;s descriptors are related to the synthesizability of the molecules (see manuscript). Also, the datasets decorated from validation set scaffolds are annotated whether they are part of the validation set (in_validation).</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

A consensus compound/bioactivity dataset for data-driven drug design and chemogenomics

<p>This is the <strong>updated version</strong> of the dataset from <strong>10.5281/zenodo.6320761</strong></p> <p><strong>Information</strong></p> <p>The diverse publicly available compound/bioactivity databases constitute a key resource for data-driven applications in chemogenomics and drug design. Analysis of their coverage of compound entries and biological targets revealed considerable differences, however, suggesting benefit of a consensus dataset. Therefore, we have combined and curated information from five esteemed databases (ChEMBL, PubChem, BindingDB, IUPHAR/BPS and Probes&amp;Drugs) to assemble a consensus compound/bioactivity dataset comprising <a href="tel:1144648">1144648</a> compounds with 10915362 bioactivities on 5613 targets (including defined macromolecular targets as well as cell-lines and phenotypic readouts). It also provides simplified information on assay types underlying the bioactivity data and on bioactivity confidence by comparing data from different sources. We have unified the source databases, brought them into a common format and combined them, enabling an ease for generic uses in multiple applications such as chemogenomics and data-driven drug design.</p> <p>The consensus dataset provides increased target coverage and contains a higher number of molecules compared to the source databases which is also evident from a larger number of scaffolds. These features render the consensus dataset a valuable tool for machine learning and other data-driven applications in (de novo) drug design and bioactivity prediction. The increased chemical and bioactivity coverage of the consensus dataset may improve robustness of such models compared to the single source databases. In addition, semi-automated structure and bioactivity annotation checks with flags for divergent data from different sources may help data selection and further accurate curation.</p> <p>This dataset belongs to the publication:&nbsp;<a href="https://doi.org/10.3390/molecules27082513">https://doi.org/10.3390/molecules27082513</a><br> &nbsp;</p> <p><strong>Structure and content of the dataset</strong></p> <table align="left"> <caption><strong>Dataset structure</strong></caption> <thead> <tr> <th scope="col"> <p>ChEMBL</p> <p>ID</p> </th> <th scope="col"> <p>PubChem</p> <p>ID</p> </th> <th scope="col"> <p>IUPHAR</p> <p>ID</p> </th> <th scope="col">Target</th> <th scope="col"> <p>Activity</p> <p>type</p> </th> <th scope="col">Assay type</th> <th scope="col">Unit</th> <th scope="col">Mean C (0)</th> <th scope="col">...</th> <th scope="col">Mean PC (0)</th> <th scope="col">...</th> <th scope="col">Mean B (0)</th> <th scope="col">...</th> <th scope="col">Mean I (0)</th> <th scope="col">...</th> <th scope="col">Mean PD (0)</th> <th scope="col">...</th> <th scope="col">Activity check annotation</th> <th scope="col">Ligand names</th> <th scope="col">Canonical SMILES C</th> <th scope="col">...</th> <th scope="col">Structure check (Tanimoto)</th> <th scope="col">Source</th> </tr> </thead> <tbody> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>The dataset was created using the Konstanz Information Miner (KNIME) (https://www.knime.com/) and was exported as a CSV-file and a compressed CSV-file.</p> <p>Except for the canonical SMILES columns, all columns are filled with the datatype &lsquo;string&rsquo;. The datatype for the canonical SMILES columns is the smiles-format. We recommend the <strong>File Reader</strong> node for using the dataset in KNIME. With the help of this node the data types of the columns can be adjusted exactly. In addition, only this node can read the compressed format.</p> <p>Column content:</p> <ul> <li>ChEMBL ID, PubChem ID, IUPHAR ID: chemical identifier of the databases</li> <li>Target: biological target of the molecule expressed as the HGNC gene symbol</li> <li>Activity type: for example, pIC<sub>50</sub></li> <li>Assay type: Simplification/Classification of the assay into cell-free, cellular, functional and unspecified</li> <li>Unit: unit of bioactivity measurement</li> <li>Mean columns of the databases: mean of bioactivity values or activity comments denoted with the frequency of their occurrence in the database, e.g. Mean C = 7.5 *(15) -&gt; the value for this compound-target pair occurs 15 times in ChEMBL database</li> <li>Activity check annotation: a bioactivity check was performed by comparing values from the different sources and adding an activity check annotation to provide automated activity validation for additional confidence <ul> <li>no comment: bioactivity values are within one log unit;</li> <li>check activity data: bioactivity values are not within one log unit;</li> <li>only one data point: only one value was available, no comparison and no range calculated;</li> <li>no activity value: no precise numeric activity value was available;</li> <li>no log-value could be calculated: no negative decadic logarithm could be calculated, e.g., because the reported unit was not a compound concentration</li> </ul> </li> <li>Ligand names: all unique names contained in the five source databases are listed</li> <li>Canonical SMILES columns: Molecular structure of the compound from each database</li> <li>Structure check (Tanimoto): To denote matching or differing compound structures in different source databases <ul> <li>match: molecule structures are the same between different sources;</li> <li>no match: the structures differ. We calculated the Jaccard-Tanimoto similarity coefficient from Morgan Fingerprints to reveal true differences between sources and reported the minimum value;</li> <li>1 structure: no structure comparison is possible, because there was only one structure available;</li> <li>no structure: no structure comparison is possible, because there was no structure available.</li> </ul> </li> <li>Source: From which databases the data come from</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Hybrid quantum-classical machine learning for generative chemistry and drug design: Generated molecules

<p>Deep generative chemistry models emerge as powerful tools to expedite drug discovery. How- ever, the immense size and complexity of the structural space of all possible drug-like molecules pose significant obstacles, which could be overcome with hybrid architectures combining quantum computers with deep classical networks.&nbsp;As the first step toward this goal, we built a compact discrete variational autoencoder (DVAE) with a Restricted Boltzmann Machine (RBM) of reduced size in its latent layer. The size of the proposed model was small enough to fit on a state-of-the-art D-Wave quantum annealer and allowed training on a subset of the ChEMBL dataset of biologically active compounds. Finally, we generated 2331 novel chemical structures with medicinal chemistry and synthetic accessibility properties in the ranges typical for molecules from ChEMBL.&nbsp;The pre- sented results demonstrate the feasibility of using already existing or soon-to-be-available quantum computing devices as testbeds for future drug discovery applications.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

PDEStrIAn: A phosphodiesterase structure and ligand interaction annotated database as a tool for structure-based drug design

<p>A systematic analysis is presented of the 220 phosphodiesterase (PDE) catalytic domain crystal structures present in the Protein Data Bank (PDB) with a focus on PDE-ligand interactions. The consistent structural alignment of 57 PDE ligand binding site residues enables the systematic analysis of PDE-ligand Interaction FingerPrints (IFPs), the identification of subtype-specific PDE-ligand interaction features, and the classification of ligands according to their binding modes. We illustrate how systematic mining of this phosphodiesterase structure and ligand interaction annotated (PDEStrIAn) database provides new insights into how conserved and selective PDE interaction hot spots can accommodate the large diversity of chemical scaffolds in PDE ligands. A substructure analysis of the co-crystalized PDE ligands in combination with those in the ChEMBL database provides a toolbox for scaffold hopping and ligand design. These analyses lead to an improved understanding of the structural requirements of PDE binding that will be useful in future drug discovery studies.</p>

opencc-zeroFeb 2016View details →
zenodo40/100

Exploiting Pretrained Biochemical Language Models for Targeted Drug Design

<p>This repository contains materials for the paper,<em> Exploiting Pretrained Biochemical Language Models for Targeted Drug Design, </em>which<em>&nbsp;</em>has been accepted for publication in <em>Bioinformatics</em> Published by Oxford University Press.</p> <p><em>data.zip</em>&nbsp;contains vocabulary files for the pretrained models, additional information regarding proteins (PFAM family, protein similarity) and&nbsp;interactions filtered from <a href="https://www.bindingdb.org/bind/index.jsp">BindingDB</a>&nbsp;which are further split into train, validation and test sets and used to train target specific molecule generation models.&nbsp;</p> <p><em>models.zip&nbsp;</em>includes files for the models trained in this study. &nbsp;&nbsp;&nbsp;&nbsp;</p> <p><em>predictions.zip&nbsp;</em>comprises the compounds generated with the targeted models and the result of their evaluation with respect to benchmarking metrics.&nbsp;</p> <p><em>docking.zip&nbsp;</em>contains <em>targets/ </em>including PDB files of the test proteins selected for docking evaluation, <em>ligands/ </em>including SDF files for molecules generated with the targeted models and two decoding strategies (i.e. beam search and sampling)&nbsp;and <em>complex/&nbsp;</em>including docking outputs.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates

<p>Solid-state NMR spectra (DARR and NCA)&nbsp;of rehydrated freeze-dried free TTR and TTR in the presence of Tafamidis and Taf-PTX</p> <p>Reference citation:&nbsp;&nbsp;Combining Solid-State NMR with Structural and Biophysical Techniques to Design Challenging Protein-Drug Conjugates. Angew Chem Int Ed Engl. 2023 Jun 5:e202303202. doi: 10.1002/anie.202303202. PMID: 37276329.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Computer-Aided Drug Design (CADD): To Screen Potential Antibiotics Against Klebsiella Pneumoniae Beta-Lactamase Enzyme

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo32/100

Table A6 images from Computer-aided drug design (CADD) to de-orphanise marine molecules: Finding potential therapeutic agents for neurodegenerative and cardiovascular diseases

<p>High Reslution Images from Table A6</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Design of Modular Autoproteolytic Gene Switches Responsive to Anti-Coronavirus Drug Candidates

<p>Data underlying the figures in the publication &ldquo;Design of modular autoproteolytic gene switches responsive to anti-coronavirus drug candidates&rdquo;, published in <em>Nat. Commun.</em>, <strong>2021</strong>, 12, 6786.</p> <p><a href="https://doi.org/10.1038/s41467-021-27072-3">https://doi.org/10.1038/s41467-021-27072-3</a></p> <p>&nbsp;</p> <p>Table of contents:</p> <p><strong>1. Supplementary Information.pdf</strong>: Document containing the Supplementary Tables 1-7 and the Supplementary Figures 1-10.</p> <p><strong>2. Dataset 1</strong>: Raw data points used to create the main and supplementary figures of the paper arranged in worksheets panel by panel.</p> <p>&nbsp;</p> <p>Data availability</p> <p>Sequence data of plasmids encoding PLpro-TAGS and Mpro-TAGS have been deposited in GenBank under accession codes OK425851, OK425852, and OK425853. Original plasmids are available upon request. All vector information is provided in Supplementary Table 4. Detailed statistical analysis is provided in Supplementary Table 7. Source data is provided in the Source data file. Source data are provided with this paper.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Binding networks identify targetable protein pockets for mechanism-based drug design

<p>This repository contains supporting files&nbsp;for the manuscript entitled:&nbsp;Binding networks identify targetable protein pockets for mechanism-based drug design.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Pharmulator™ Module: Quantifying Functional Groups and Its Applications in Drug Design

<p><strong>MS-Excel-1.xlsx:</strong>&nbsp;It contains SMILES codes of the training and test sets along with their GHS classification&nbsp;</p> <p><strong>MS-Excel-2.xlsx: </strong>It contains<strong>&nbsp;</strong>SMILES codes of the 8993 chemicals and their experimental aqueous solubility parameters and FGs, Morgan and MACCS-based structural descriptor values</p> <p><strong>MS-Excel-3.xlsx:&nbsp;</strong>It contains&nbsp;the DrugBank IDs, SMILES codes and number of FG occurrences (binary string) of the 2356 drug molecules</p> <p><strong>MS-Excel-4.xlsx:</strong>&nbsp;It contains&nbsp;SMILES codes&nbsp;as well as FGs descriptors for the entire training sets, balanced training sets and final test sets</p> <p><strong>MS-Excel-5.xlsx:</strong>&nbsp;It contains&nbsp;SMILES codes&nbsp;as well as Morgan descriptors for the entire training sets, balanced training sets and final test sets</p> <p><strong>MS-Excel-6.xlsx:</strong>&nbsp;It contains&nbsp;SMILES codes&nbsp;as well as MACCS descriptors for the entire training sets, balanced training sets and final test sets</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Fig. 1 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold

Fig. 1. Chemical structures of the targeted NP diterpenoid scaffolds (1–2), methylated products (3–4) and the amide library (5–16).

opennotspecifiedOct 2021View details →
zenodo32/100

Fig. 4 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold

Fig. 4. Representative flow cytometry examples of GFP expression (i.e., HIV production) in J-Lat 10.6 cells in the absence of stimulation (A), treatment with 50 nM PMA (B) and PMA plus 100 μM compound 2 (C). Dose response profiles (D) of compounds on GFP expression induced by 50 nM PMA in J-Lat 10.6 cells. Data denote mean ± s.d. from three independent experiments.

opennotspecifiedOct 2021View details →
zenodo32/100

Fig. 3 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold

Fig. 3. Microscopic images of the skinny (Skn) phenotype in L4s induced by compound 3 compared with the wild-type (WT) phenotype control (cultured in culture medium + 0.25% DMSO); the bottom panels enlarged from the dashed/ boxed areas in the upper panels.

opennotspecifiedOct 2021View details →
ClinicalTrials.gov32/100

Equivalence Among Antiepileptic Drug Generic and Brand Products in People With Epilepsy: Chronic-Dose 4-Period Replicate Design

ClinicalTrials.gov study NCT01713777. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Efficacy and Safety of New Generation Drug Eluting Stents Associated With an Ultra Short Duration of Dual Antiplatelet Therapy. Design of the Short Duration of Dual antiplatElet Therapy With SyNergy I

ClinicalTrials.gov study NCT02099617. IPD Sharing: NO. Countries: 9. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Equivalence Among Antiepileptic Drug Generic and Brand Products in People With Epilepsy: Single-Dose 6-Period Replicate Design (EQUIGEN Single-Dose Study)

ClinicalTrials.gov study NCT01733394. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

An Exploratory Pilot Study in Healthy Volunteers to Assess the Parameters for the Design of Bioequivalence Studies on Moderately Lipophilic, Moderately to Highly Protein Bound Drugs Using Dermal Open

ClinicalTrials.gov study NCT03613207. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Multi-domain Distribution Learning for De Novo Drug Design

<p>Model checkpoints, processed dataset and samples.</p>

opencc-by-4.0Sep 2024View details →
zenodo28/100

DigFrag as a digital fragmentation method used for artificial intelligence-based drug design

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record