Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.9.0
Dataset results
27 results for “ChEMBL”
Analog series from ChEMBL, PubChem, and DrugBank
<p>The datasets consist of analog series and key compounds extracted from ChEMBL, PubChem, and DrugBank. For each compound structural and activity information is provided. </p>
Analog series-based scaffolds from ChEMBL with associated activity information
<p>Reported is the activity information for the 12,294 analog series-based (ASB) scaffolds extracted from ChEMBL database. For each ASB scaffold structural and activity information for all analogs comprising the analog series is provoded. </p>
784 promiscuity cliffs from ChEMBL
<p>Reported are the 784 promiscuity cliffs formed by compounds from medicinal chemistry sources. Compounds forming cliffs are provided as SMILES. For each compound its promiscuity degree (PD) and the list of ChEMBL target IDs are provided. </p>
286 new target pairs based on shared compounds from ChEMBL
<p>Reported is the list of 286 compound-based target pairs between distantly related or unrelated pharmaceutical target proteins with shared common compounds, derived from ChEMBL22 high-confidence data. For each target, the corresponding UniProt ID is provided. In addition, for each given target pair, the number of shared compounds and structures by SMILES notations are added as well.</p>
A new ChEMBL dataset for FastTargetPred and target fishing for an exhaustive list of linear tetrapeptides
<p>A ChEMBL-v29 dataset was generated to be used with the ligand-based similarity search target prediction engine FastTargetPred (https://github.com/ludovicchaput/FastTargetPred). Using this new dataset, attempts to predict macromolecular targets for a published dataset of 160,000 tetrapeptides was performed.</p> <p>The dbchembl29 directory contains all the files for FastTargetPred. This command line tool compares using different types of fingerprints, a file containing small query molecules in SDF format (it can be 1 molecule or a collection) to molecules extracted from ChEMBL29. If a match is found, this suggests that your query molecule is similar to a ChEMBL compound and as the ChEMBL compound has bioactivity data against one or more macromolecular target, then this suggests that your query compound could bind to targets that interact with compounds that are similar to the query molecule. The so-called similarity principle in chemistry. Fingerprints are computed with http://www.mayachemtools.org/, a collection of Perl and Python scripts for Chemoinformatics and (Structural) Bioinformatics.</p> <p>To run FastTargetPred with the new ChEMBL 29 data, you just need to unzip the dbchembl29 directory into the FastTargetPred main directory.</p> <p>The default FastTargetPred commands (e.g., python3 FastTargetPred.py rivaroxaban.sdf, the default command uses ECFP4 fingerprints and a Tanimoto coef of 0.6, rivaroxaban here is the query compound, it is provided in the extra_data directory) will use the data present in the default db directory and thus a curated version of ChEMBL-25 release. It was the version of the ChEMBL database available when FastTargetPred was developed. Since then, many new molecules have been added and this is why we generated the ChEMBL-29 dataset (last ChEMBL version at the time of writing).</p> <p>To use the new ChEMBL-29 data, you can run the following command:</p> <p>python3 FastTargetPred.py rivaroxaban.sdf -fp MACCS -tc 0.9 -db dbchembl29/chembl29_active</p> <p>This applies a similarity search for the query compound (here rivaroxaban, you can for instance move this SDF file in the directory containing the file FastTargetPred.py) using MACCS fingerprints, a Tanimoto coefficient threshold of 0.9 and the -db option forces the system to look at the ChEMBL29 curated data and not the default ChEMBL-25 data. This 0.9 value means to focus on molecules very similar to rivaroxaban present in the ChEMBL data. If one is looking for more distantly related compounds, then a value of 0.7 can be used. Users can try different values or try consensus scoring...See FastTargetPred: a program enabling the fast prediction of putative protein targets for input chemical databases. Chaput et al., Bioinformatics. 2020 Aug 15;36(14):4225-4226</p> <p>The chembl29 directory also contains 714,780 compounds (canonical SMILES strings) extracted from ChEMBL29 that have bioactivity data. Fingerprints could not be computed for 19 molecules that have unusual chemistry. We selected the following thresholds (eg, binding assays, activity against targets less than 20 micro-molar, ChEMBL confidence_score = 6 or above, maximum = 9).</p> <p>With this new dataset, we attempted to predict potential targets with FastTargetPred for 160,000 input query peptides (4 amino acids, combination should be 20 x 20 x 20 x 20) previously reported by Dewi Prasasty and Perdana Istyastono, Data in brief 27 (2019) 104607. The peptides for which a putative target was predicted are shown in two DataWarrior files with the amino acid sequence of the query peptide, the compounds found to be similar in the ChEMBL29 dataset (fingerprints = ECFP4, Tanimoto 0.6) and thus the compound IDs, the target ChEMBL IDs, mapping to the UniProt database when available, information about disease involvements, Reactome pathway database identifiers. These two DataWarrior files are searchable, can be sorted and hyperlinks to the ChEMBL, UniProt and Reactome databases have been inserted.</p> <p> </p>
Collection of analog series-based (ASB) scaffolds shared between ZINC, ChEMBL, and PubChem
<p>Analog series-based (ASB) scaffolds shared between ZINC and ChEMBL (version 22), ZINC and PubChem and all the three databases are provided as three separate files. For each ASB scaffold, the SMILES representation of ZINC compounds is provided. In addition, the number of ZINC compounds, the number and the list of targets it was annotated with is reported. A README file is also given.</p>
Selectivity profiling of multi-kinase inhibitors across the Human Kinome from ChEMBL
<p>Reported is the list of 596 protein kinase pairs, consisting of 141 kinases and selectivity profiles of 10,060 multi-kinase inhibitors found in ChEMBL23 high-confidence data. For each of the reported protein kinase pairs, UniProt IDs defining the kinase forming a pair is provided, as well as the shared inhibitors and their selectivity profiles. For each target within the pair, potency value for each compound is reported as pIC50 value, as well as the absolute potency difference used to assess the selectivity profiles.</p>
ChEMBL data against CHEMBL367, CHEMBL368 and CHEMBL612348
<p>Data from ChEMBL compounds reported with an activity against one of the following targets: CHEMBL367 : <em>Leishmania donovani, </em>CHEMBL368 : <em>Trypanosoma cruzi, </em>and CHEMBL612348 : <em>Trypanosoma brucei rhodesiense</em>. </p>
31 ChEMBL data sets for regression modeling
<p>From ChEMBL version 17, 31 compound data sets have been selected for regression modeling. Compounds had to be active against human targets in a direct inhibition/binding assay with highest ChEMBL confidence score and Ki values below 100 micromolar. Multiple Ki values for the same compound were averaged if they fell into the same order of magnitude, or else they were disregarded. Duplicates, known pan-assay interference, and other reactive molecules were removed. Only sets with at least 500 compounds were considered.</p> <p> </p> <p>Note: The SD files contain a field "pKi"; note however that this field contains the Ki value in nM units, not the logarithmic value.</p>
Sets of ChEMBL compounds with high or low confidence activity data
<p>Two sets of compounds assembled from ChEMBL release 20 that were annotated with high or low confidence activity data were provided in separate files. For each compound in a file, the unique compound identifier (i.e., molregno), the number of targets in individual years (from 1976 to 2014) and the list of target annotations (if any) was given.</p>
Currently available scaffolds and MMP cores in ChEMBL 20
<p>All target-based BM scaffolds (K<sub>i</sub>: 35,872; IC<sub>50</sub>: 74,379), CSKs (K<sub>i</sub>: 23,056; IC<sub>50</sub>: 49,216), MMP cores (K<sub>i</sub>: 42,104; IC<sub>50</sub>: 73,616), and retrosynthetic MMP cores (K<sub>i</sub>: 19,040; IC<sub>50</sub>: 32,382) for high confident bioactive compounds extracted from ChEMBL version 20 are provided herein. </p>
Compound activity records associated with original publications in ChEMBL 21
<p>Provided are two sets of compound activity records (set 1 and set 2) that were traced back to original publications and assembled from ChEMBL release 21. For each compound-target combination, the corresponding potency measurements and publications are provided. In addition, the list of unique publications is given for both sets 1 and 2.</p>
Compound activity data sets for 15 biological targets compiled from the ChEMBL and PubChem databases.
<p>Compound activity data sets for the 15 biological targets are deposited, along with structure-activity relationship matrices IDs. Active compounds were extracted from the ChEMBL database and inactive were from the PubChem database. Details of the data sets are described in the original publication. and the summary of the data sets is given in the readme.txt file. </p>
PubChem and ChEMBL-series processed dataset used in Exhaustive local chemical space exploration using a transformer model
<p>PubChem and ChEMBL-series processed dataset used in <span>Exhaustive local chemical space exploration using </span><span>a transformer model</span></p>
Compound activity classes from ChEMBL for machine learning analysis
<p>Ten activity classes are provided that were extracted from ChEMBL version 24 for machine learning studies. Compounds are given in SMILES representations. The following selection criteria were applied. Compounds were required to be tested in a direct binding assay against a single human protein with a ChEMBL assay confidence score of 9. In addition, K<sub>i</sub> measurements had to be available. If multiple K<sub>i</sub> values were available for a compound and did not fall within the same order of magnitude, the compound was not selected. Furthermore only compounds with (mean) pK<sub>i</sub> of at least 5 were considered. Moreover, activity classes had to contain at least 200 compounds belonging to at least 50 computationally determined analog series. The 10 deposited classes consist of 243 to 955 compounds and 57 to 216 analog series.</p>
Two ChEMBL-34 subsets (lead-like and drug-like molecules)
<p>Two subsets of molecules from ChEMBL-34[1].</p> <p>Those molecular datasets might be useful to people training molecular generators.</p> <p>After decompression, you will get:<br>chembl34_stable_ES_OA_LL.smi: 585,272 molecules.<br>chembl34_stable_ES_OA_DL.smi: 756,420 molecules.</p> <p>stable=non-reactive molecules (filtered-out reactive functional groups from [5]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py</a></p> <p>ES=Easy Synthesis (SAscore <= 3.0) [2].<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py</a></p> <p>OA=Orally Available (according to a classifier trained on the dataset from [6]).</p> <p>LL=Lead-Like (almost the definition from [3]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py</a></p> <p>DL=Drug-Like (definition from [4]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py</a></p> <div> <h1>Bibliography</h1> <a href="https://github.com/UnixJunkie/chembl34_subsets#bibliography"></a></div> <ol> <li> <p>Zdrazil, B., Felix, E., Hunter, F., Manners, E. J., Blackshaw, J., Corbett, S., ... & Leach, A. R. (2024). The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic acids research, 52(D1), D1180-D1192. <a href="https://doi.org/10.1093/nar/gkad1004" rel="nofollow">https://doi.org/10.1093/nar/gkad1004</a></p> </li> <li> <p>Ertl, P., & Schuffenhauer, A. (2009). Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics, 1, 1-11. <a href="https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8" rel="nofollow">https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8</a></p> </li> <li> <p>Hann, M. M., & Oprea, T. I. (2004). Pursuing the leadlikeness concept in pharmaceutical research. Current opinion in chemical biology, 8(3), 255-263. <a href="https://doi.org/10.1016/j.cbpa.2004.04.003" rel="nofollow">https://doi.org/10.1016/j.cbpa.2004.04.003</a></p> </li> <li> <p>Tran-Nguyen, V. K., Jacquemard, C., & Rognan, D. (2020). LIT-PCBA: an unbiased data set for machine learning and virtual screening. Journal of chemical information and modeling, 60(9), 4263-4273. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155</a></p> </li> <li> <p>Lisurek, M., Rupp, B., Wichard, J., Neuenschwander, M., von Kries, J. P., Frank, R., ... & Kühne, R. (2010) Design of chemical libraries with potentially bioactive molecules applying a maximum common substructure concept. Molecular diversity, 14, 401-408. <a href="https://link.springer.com/article/10.1007/s11030-009-9187-z" rel="nofollow">https://link.springer.com/article/10.1007/s11030-009-9187-z</a></p> </li> <li> <p>Falcon-Cano, G., Molina, C., & Cabrera-Perez, M. A. (2020). ADME prediction with KNIME: development and validation of a publicly available workflow for the prediction of human oral bioavailability. Journal of chemical information and modeling, 60(6), 2660-2667. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019</a></p> </li> </ol>
Raw data extracted from ChEMBL
<p>Raw data files extracted from ChEMBL for the MELLODDY project.</p>
Chembl Filtered Dataset for TorchDrug
<p>A preprocessed ChEMBL dataset containing 456K molecules with 1310 kinds of diverse and extensive biochemical assays from paper "Strategies for Pre-training Graph Neural Networks" used by TorchDrug library.</p>
cardiotoxic_chembl
<p>Cardiotoxic and not compounds based on Chembl-filtered dataset</p>
ChemBioSim Recalibration: Twelve Preprocessed ChEMBL Data Sets
<p><strong>Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data</strong></p> <p><strong>Project description</strong></p> <p>Machine learning models are powerful tools for the prediction of molecular properties or the biological activity of chemical compounds. However, to make these models useful and applicable, the confidence in the predictions should also be specified. For that purpose, models may be integrated in a conformal prediction (CP) framework that adds a calibration step to estimate the confidence of the predictions. CP models offer the advantage of ensuring a predefined error rate, as long as the test and training sets are exchangeable.</p> <p>In cases where the test data presents a drift from the descriptor space of the training data, or where assay setups change, this assumption may not be fulfilled and the models are not guaranteed to be valid. </p> <p>In this study, the performance of internally valid CP models was evaluated upon application to either newer time-split data or to external data. More specifically, temporal data drifts were analysed based on time-splits of twelve toxicity-related datasets from the ChEMBL database. Moreover, models trained on publicly available data for liver toxicity and MNT in vivo were applied on proprietary data to evaluate the discrepancies. In general it was observed that the training and (holdout) test sets were not exchangeable in the studied set-ups, and the models were therefore not applicable (i.e. non-valid CP models).</p> <p>To recover the validity of the models on the holdout test set, a strategy for updating the calibration set with data more similar to the holdout set was investigated. Restored validity is the main requisite for applying the CP models with confidence. However, this comes at the cost of decreased model efficiency, as more predictions are identified as inconclusive.</p> <p> </p> <p><strong>Dataset</strong></p> <p>The uploaded file contains the ChEMBL data used in the work for the manuscript “Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data”.</p> <p>Twelve preprocessed datasets containing molecule chembl ID, SMILES, binary activity (i.e. 1 if active, 0 if inactive), publication year, and CHEMBIO descriptors are available for the following ChEMBL endpoints, extracted from ChEMBL Version 26:</p> <ul> <li> <p>CHEMBL220: Acetylcholinesterase (human), 2673 compounds</p> </li> <li> <p>CHEMBL4078: Acetylcholinesterase (fish), 3811 compounds</p> </li> <li> <p>CHEMBL5763: Cholinesterase, 2755 compounds</p> </li> <li> <p>CHEMBL203: EGFR erbB1, 4059 compounds</p> </li> <li> <p>CHEMBL206: Estrogen receptor alpha, 1416 compounds</p> </li> <li> <p>CHEMBL279: VEGFR 2, 5174 compounds</p> </li> <li> <p>CHEMBL230: Cyclooxygenase-2, 2020 compounds</p> </li> <li> <p>CHEMBL340: Cytochrome P450 3A4, 3316 compounds</p> </li> <li> <p>CHEMBL240: HERG, 4976 compounds</p> </li> <li> <p>CHEMBL2039: Monoamine oxidase B, 2534 compounds</p> </li> <li> <p>CHEMBL222: Norepinephrine transporter, 1566 compounds</p> </li> <li> <p>CHEMBL228: Serotonin transporter, 2111 compounds</p> </li> </ul> <p> </p> <p><strong>Usage</strong></p> <p>This dataset can be used as input to run the notebooks available at </p> <p><a href="https://github.com/volkamerlab/CPRecalibration_manuscript_SI">https://github.com/volkamerlab/CPRecalibration_manuscript_SI</a></p> <ol> <li> <p>Clone the GitHub repository.</p> </li> <li> <p>Download the dataset provided here.</p> </li> <li> <p>Copy the dataset (don’t extract) into the data folder of the cloned GitHub repository.</p> </li> <li> <p>Follow the instructions on GitHub.</p> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.