Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

27 results for “ChEMBL”

Learn how ShareScore rates datasets ↗
zenodo40/100

Analog series from ChEMBL, PubChem, and DrugBank

<p>The datasets consist&nbsp;of analog series and key compounds extracted from ChEMBL, PubChem, and DrugBank. For each compound structural and activity information is provided. &nbsp;</p>

opencc-zeroAug 2016View details →
zenodo40/100

Analog series-based scaffolds from ChEMBL with associated activity information

<p>Reported is the activity information for the 12,294 analog series-based (ASB) scaffolds extracted from ChEMBL database. For each ASB scaffold structural and activity information for all analogs comprising the analog series is provoded. </p>

opencc-by-4.0Sep 2016View details →
zenodo40/100

784 promiscuity cliffs from ChEMBL

<p>Reported are the 784 promiscuity cliffs formed by compounds from medicinal chemistry sources. Compounds forming cliffs are provided as SMILES. For each compound its promiscuity degree (PD) and the list of ChEMBL target IDs are provided. </p>

opencc-by-4.0Dec 2016View details →
zenodo40/100

286 new target pairs based on shared compounds from ChEMBL

<p>Reported is the list of 286 compound-based target pairs between distantly related or unrelated pharmaceutical target proteins with shared common compounds, derived from ChEMBL22 high-confidence data. For each target, the corresponding UniProt ID is provided. In addition, for each given target pair, the number of shared compounds and structures by SMILES notations are added as well.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

A new ChEMBL dataset for FastTargetPred and target fishing for an exhaustive list of linear tetrapeptides

<p>A ChEMBL-v29 dataset was generated to be used with the ligand-based similarity search target prediction engine FastTargetPred (https://github.com/ludovicchaput/FastTargetPred). Using this new dataset, attempts to predict macromolecular targets for a published dataset of 160,000 tetrapeptides was performed.</p> <p>The dbchembl29 directory contains all the files for FastTargetPred. This command line tool compares using different types of fingerprints, a file containing small query molecules in SDF format (it can be 1 molecule or a collection) to molecules extracted from ChEMBL29. If a match is found, this suggests that your query molecule is similar to a ChEMBL compound and as the ChEMBL compound has bioactivity data against one or more macromolecular target, then this suggests that your query compound could bind to targets that interact with compounds that are similar to the query molecule. The so-called similarity principle in chemistry. Fingerprints are computed with http://www.mayachemtools.org/, a collection of Perl and Python scripts for Chemoinformatics and (Structural) Bioinformatics.</p> <p>To run FastTargetPred with the new ChEMBL 29 data, you just need to unzip the dbchembl29 directory into the FastTargetPred main directory.</p> <p>The default FastTargetPred commands (e.g.,&nbsp;python3 FastTargetPred.py rivaroxaban.sdf, the default command uses ECFP4 fingerprints and a Tanimoto coef of 0.6, rivaroxaban here is the query compound, it is provided in the extra_data directory) will use the data present in the default db directory and thus a curated version of ChEMBL-25 release. It was the version of the ChEMBL database available when FastTargetPred was developed. Since then, many new molecules have been added and this is why we generated the ChEMBL-29 dataset (last ChEMBL version at the time of writing).</p> <p>To use the new ChEMBL-29 data, you can run the following command:</p> <p>python3 FastTargetPred.py rivaroxaban.sdf -fp MACCS -tc 0.9 -db dbchembl29/chembl29_active</p> <p>This applies a similarity search for the query compound (here rivaroxaban, you can for instance move this SDF file in the directory containing the file FastTargetPred.py) using MACCS fingerprints, a Tanimoto coefficient threshold of 0.9 and the -db option forces the system to look at the ChEMBL29 curated data and not the default ChEMBL-25 data. This 0.9 value means to focus on molecules very similar to rivaroxaban present in the ChEMBL data. If one is looking for more distantly related compounds, then a value of 0.7 can be used. Users can try different values or try consensus scoring...See FastTargetPred: a program enabling the fast prediction of putative protein targets for input chemical databases. Chaput et al., Bioinformatics. 2020 Aug 15;36(14):4225-4226</p> <p>The chembl29 directory also contains 714,780 compounds (canonical SMILES strings) extracted from ChEMBL29 that have bioactivity data. Fingerprints could not be computed for 19 molecules that have unusual chemistry. We selected the following thresholds (eg, binding assays, activity against targets less than 20 micro-molar, ChEMBL confidence_score = 6 or above, maximum = 9).</p> <p>With this new dataset, we attempted to predict potential targets with&nbsp;FastTargetPred for 160,000 input query peptides (4 amino acids, combination should be 20 x 20 x 20 x 20) previously reported by Dewi Prasasty and Perdana Istyastono, Data in brief 27 (2019) 104607. The peptides for which a putative target was predicted are shown in two DataWarrior files with the amino acid sequence of the query peptide, the compounds found to be similar in the ChEMBL29 dataset (fingerprints = ECFP4, Tanimoto 0.6) and thus the compound IDs, the target ChEMBL IDs, mapping to the UniProt database when available, information about disease involvements, Reactome pathway database identifiers. These two DataWarrior files are searchable, can be sorted and hyperlinks to the ChEMBL, UniProt and Reactome databases have been inserted.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Collection of analog series-based (ASB) scaffolds shared between ZINC, ChEMBL, and PubChem

<p>Analog series-based (ASB) scaffolds shared between ZINC and ChEMBL (version 22), ZINC and PubChem and all the three databases are provided as three separate files. For each ASB scaffold, the SMILES representation of ZINC compounds is provided. In addition, the number of ZINC compounds, the number and the list of targets it was annotated with is reported. A README file is also given.</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Selectivity profiling of multi-kinase inhibitors across the Human Kinome from ChEMBL

<p>Reported is the list of 596 protein kinase pairs, consisting of 141 kinases and selectivity profiles of 10,060 multi-kinase inhibitors found in ChEMBL23 high-confidence data. For each of the reported protein kinase pairs, UniProt IDs defining the kinase forming a pair is provided, as well as the shared inhibitors and their selectivity profiles. For each target within the pair, potency value for each compound is reported as pIC50 value, as well as the absolute potency difference used to assess the selectivity profiles.</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

ChEMBL data against CHEMBL367, CHEMBL368 and CHEMBL612348

<p>Data from ChEMBL compounds reported with an activity against one of the following targets: CHEMBL367 :&nbsp;<em>Leishmania donovani,&nbsp;</em>CHEMBL368 : <em>Trypanosoma cruzi, </em>and&nbsp;CHEMBL612348 :&nbsp;<em>Trypanosoma brucei rhodesiense</em>.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

31 ChEMBL data sets for regression modeling

<p>From ChEMBL version 17, 31 compound data sets have been selected for regression modeling. Compounds had to be active against human targets in a direct inhibition/binding assay with highest ChEMBL confidence score and Ki values below 100 micromolar.&nbsp;Multiple Ki values for the same compound were averaged if they fell into the same order of magnitude, or else they were disregarded. Duplicates,&nbsp;known pan-assay interference, and other reactive molecules were removed.&nbsp;Only sets with at least 500 compounds were considered.</p> <p>&nbsp;</p> <p>Note:&nbsp;The SD files contain a field &quot;pKi&quot;; note however that this field contains the Ki value in nM units, not the logarithmic value.</p>

opencc-zeroJan 2015View details →
zenodo36/100

Sets of ChEMBL compounds with high or low confidence activity data

<p>Two sets of compounds assembled from ChEMBL release 20 that were annotated with high or low confidence activity data were provided in separate files. For each compound in a file, the unique compound identifier (i.e., molregno), the number of targets in individual years (from 1976 to 2014) and the list of target annotations (if any) was given.</p>

opencc-zeroMay 2015View details →
zenodo36/100

Currently available scaffolds and MMP cores in ChEMBL 20

<p>All target-based BM scaffolds (K<sub>i</sub>: 35,872; IC<sub>50</sub>: 74,379), CSKs (K<sub>i</sub>: 23,056; IC<sub>50</sub>: 49,216), MMP cores (K<sub>i</sub>: 42,104; IC<sub>50</sub>: 73,616), and retrosynthetic MMP cores (K<sub>i</sub>: 19,040; IC<sub>50</sub>: 32,382) for high confident bioactive compounds extracted from ChEMBL version 20 are provided herein.&nbsp;</p>

opencc-zeroDec 2015View details →
zenodo36/100

Compound activity records associated with original publications in ChEMBL 21

<p>Provided are two sets of compound activity records (set 1 and set 2) that were traced back to original publications and assembled from ChEMBL release 21. For each compound-target combination, the corresponding potency measurements and publications are provided. In addition, the list of unique publications is given for both sets 1 and 2.</p>

opencc-zeroMay 2016View details →
zenodo36/100

Compound activity data sets for 15 biological targets compiled from the ChEMBL and PubChem databases.

<p>Compound activity data sets for the 15 biological targets are&nbsp;deposited, along with structure-activity relationship matrices IDs. Active compounds were extracted from the ChEMBL database and inactive were from the PubChem database. Details of the data sets are described in the original publication.&nbsp;and the summary of the data sets is given in the readme.txt file.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

PubChem and ChEMBL-series processed dataset used in Exhaustive local chemical space exploration using a transformer model

<p>PubChem and ChEMBL-series processed dataset used in&nbsp;<span>Exhaustive local chemical space exploration using </span><span>a transformer model</span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Compound activity classes from ChEMBL for machine learning analysis

<p>Ten activity classes are provided that were extracted from ChEMBL version 24 for machine learning studies. Compounds are given in SMILES representations. The following selection criteria were applied. Compounds were required to be tested in a direct binding assay against a single human protein with a ChEMBL assay confidence score of 9. In addition, K<sub>i</sub>&nbsp;measurements had to be available. If multiple K<sub>i</sub>&nbsp;values were available for a compound and did not fall within the same order of magnitude, the compound was not selected. Furthermore only compounds with (mean) pK<sub>i</sub>&nbsp;of at least 5 were considered. Moreover, activity classes had to contain at least 200 compounds belonging to at least 50 computationally determined analog series.&nbsp;The 10 deposited classes consist of 243 to 955&nbsp;compounds and 57 to 216 analog series.</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Two ChEMBL-34 subsets (lead-like and drug-like molecules)

<p>Two subsets of molecules from ChEMBL-34[1].</p> <p>Those molecular datasets might be useful to people training molecular generators.</p> <p>After decompression, you will get:<br>chembl34_stable_ES_OA_LL.smi: 585,272 molecules.<br>chembl34_stable_ES_OA_DL.smi: 756,420 molecules.</p> <p>stable=non-reactive molecules (filtered-out reactive functional groups from [5]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py</a></p> <p>ES=Easy Synthesis (SAscore &lt;= 3.0) [2].<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py</a></p> <p>OA=Orally Available (according to a classifier trained on the dataset from [6]).</p> <p>LL=Lead-Like (almost the definition from [3]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py</a></p> <p>DL=Drug-Like (definition from [4]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py</a></p> <div> <h1>Bibliography</h1> <a href="https://github.com/UnixJunkie/chembl34_subsets#bibliography"></a></div> <ol> <li> <p>Zdrazil, B., Felix, E., Hunter, F., Manners, E. J., Blackshaw, J., Corbett, S., ... &amp; Leach, A. R. (2024). The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic acids research, 52(D1), D1180-D1192. <a href="https://doi.org/10.1093/nar/gkad1004" rel="nofollow">https://doi.org/10.1093/nar/gkad1004</a></p> </li> <li> <p>Ertl, P., &amp; Schuffenhauer, A. (2009). Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics, 1, 1-11. <a href="https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8" rel="nofollow">https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8</a></p> </li> <li> <p>Hann, M. M., &amp; Oprea, T. I. (2004). Pursuing the leadlikeness concept in pharmaceutical research. Current opinion in chemical biology, 8(3), 255-263. <a href="https://doi.org/10.1016/j.cbpa.2004.04.003" rel="nofollow">https://doi.org/10.1016/j.cbpa.2004.04.003</a></p> </li> <li> <p>Tran-Nguyen, V. K., Jacquemard, C., &amp; Rognan, D. (2020). LIT-PCBA: an unbiased data set for machine learning and virtual screening. Journal of chemical information and modeling, 60(9), 4263-4273. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155</a></p> </li> <li> <p>Lisurek, M., Rupp, B., Wichard, J., Neuenschwander, M., von Kries, J. P., Frank, R., ... &amp; K&uuml;hne, R. (2010) Design of chemical libraries with potentially bioactive molecules applying a maximum common substructure concept. Molecular diversity, 14, 401-408. <a href="https://link.springer.com/article/10.1007/s11030-009-9187-z" rel="nofollow">https://link.springer.com/article/10.1007/s11030-009-9187-z</a></p> </li> <li> <p>Falcon-Cano, G., Molina, C., &amp; Cabrera-Perez, M. A. (2020). ADME prediction with KNIME: development and validation of a publicly available workflow for the prediction of human oral bioavailability. Journal of chemical information and modeling, 60(6), 2660-2667. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019</a></p> </li> </ol>

opencc-by-3.0Sep 2024View details →
zenodo36/100

Raw data extracted from ChEMBL

<p>Raw data files extracted from ChEMBL for the MELLODDY project.</p>

opencc-by-3.0Jun 2021View details →
zenodo36/100

Chembl Filtered Dataset for TorchDrug

<p>A preprocessed ChEMBL dataset containing 456K molecules with 1310 kinds of diverse and extensive biochemical assays from paper &quot;Strategies for Pre-training Graph Neural Networks&quot; used by TorchDrug library.</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

cardiotoxic_chembl

<p>Cardiotoxic and not compounds based on Chembl-filtered dataset</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

ChemBioSim Recalibration: Twelve Preprocessed ChEMBL Data Sets

<p><strong>Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data</strong></p> <p><strong>Project description</strong></p> <p>Machine learning models are powerful tools for the prediction of molecular properties or the biological activity of chemical compounds. However, to make these models useful and applicable, the confidence in the predictions should also be specified. For that purpose, models may be integrated in a conformal prediction (CP) framework that adds a calibration step to estimate the confidence of the predictions. CP models offer the advantage of ensuring a predefined error rate, as long as the test and training sets are exchangeable.</p> <p>In cases where the test data presents a drift from the descriptor space of the training data, or where assay setups change, this assumption may not be fulfilled and the models are not guaranteed to be valid.&nbsp;</p> <p>In this study, the performance of internally valid CP models was evaluated upon application to either newer time-split data or to external data. More specifically, temporal data drifts were analysed based on time-splits of twelve toxicity-related datasets from the ChEMBL database. Moreover, models trained on publicly available data for liver toxicity and MNT in vivo were applied on proprietary data to evaluate the discrepancies. In general it was observed that the training and (holdout) test sets were not exchangeable in the studied set-ups, and the models were therefore not applicable (i.e. non-valid CP models).</p> <p>To recover the validity of the models on the holdout test set, a strategy for updating the calibration set with data more similar to the holdout set was investigated. Restored validity is the main requisite for applying the CP models with confidence. However, this comes at the cost of decreased model efficiency, as more predictions are identified as inconclusive.</p> <p>&nbsp;</p> <p><strong>Dataset</strong></p> <p>The uploaded file contains the ChEMBL data used in the work for the manuscript &ldquo;Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data&rdquo;.</p> <p>Twelve preprocessed datasets containing molecule chembl ID, SMILES, binary activity (i.e. 1 if active, 0 if inactive), publication year, and CHEMBIO descriptors are available for the following ChEMBL endpoints, extracted from ChEMBL Version 26:</p> <ul> <li> <p>CHEMBL220: Acetylcholinesterase (human), 2673 compounds</p> </li> <li> <p>CHEMBL4078: Acetylcholinesterase (fish), 3811 compounds</p> </li> <li> <p>CHEMBL5763: Cholinesterase, 2755 compounds</p> </li> <li> <p>CHEMBL203: EGFR erbB1, 4059 compounds</p> </li> <li> <p>CHEMBL206: Estrogen receptor alpha, 1416 compounds</p> </li> <li> <p>CHEMBL279: VEGFR 2, 5174 compounds</p> </li> <li> <p>CHEMBL230: Cyclooxygenase-2, 2020 compounds</p> </li> <li> <p>CHEMBL340: Cytochrome P450 3A4, 3316 compounds</p> </li> <li> <p>CHEMBL240: HERG, 4976 compounds</p> </li> <li> <p>CHEMBL2039: Monoamine oxidase B, 2534 compounds</p> </li> <li> <p>CHEMBL222: Norepinephrine transporter, 1566 compounds</p> </li> <li> <p>CHEMBL228: Serotonin transporter, 2111 compounds</p> </li> </ul> <p>&nbsp;</p> <p><strong>Usage</strong></p> <p>This dataset can be used as input&nbsp; to run the notebooks available at&nbsp;&nbsp;</p> <p><a href="https://github.com/volkamerlab/CPRecalibration_manuscript_SI">https://github.com/volkamerlab/CPRecalibration_manuscript_SI</a></p> <ol> <li> <p>Clone the GitHub repository.</p> </li> <li> <p>Download the dataset provided here.</p> </li> <li> <p>Copy the dataset (don&rsquo;t extract) into the data folder of the cloned GitHub repository.</p> </li> <li> <p>Follow the instructions on GitHub.</p> </li> </ol>

opengpl-3.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record