Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

226

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

226 results for “Active compounds”

Learn how ShareScore rates datasets ↗
OpenNeuro44/100

Differences in Chemo-signaling Compound-Evoked Brain Activity in Male and Female Young Adults: A Pilot Study in the Role of Sexual Dimorphism in Olfactory Chemo-Signaling

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo44/100

Systematic Data Analysis and Diagnostic Machine Learning Reveal Differences between Compounds with Single- and Multitarget Activity

<p>The deposited files contain balanced data sets of multi-target (MT) and single-target (ST) compounds (CPDs) used for machine learning studies (https://dx.doi.org/10.1021/acs.molpharmaceut.0c00901).&nbsp; The first file (st_mt_data.tsv) contains 15,142 MT- and 15,081 ST-CPDs and the second (st_dt_data.tsv)&nbsp; 1828 DT- and 1776 ST-CPDs. For each CPD, a nonstereo_aromatic_SMILES representation, the original ChEMBL_cid, UniProt (target) IDs, and CPD category (CPD_CAT) (i.e. DT/MT/ST) is provided. DT stands for &#39;diverse-target&#39; and denotes a subset of MT-CPDs (as detailed in the publication). In addition, a CPD is tagged &ldquo;Y&rdquo; if it continued to be present in the data set after removal of 50% randomly selected CPDs or 50%&nbsp; CPD nearest neighbors (NN), respectively.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Publication & Supplementary data: Pro-health compounds and antioxidant activity of 65 potato cultivars

<p>Metadata (climatic conditions, list of varieties, field plan, etc.) and data (carotenoid content, vitamin C content, radical scavenging activity DPPH and FRAP, yellow index) related to the paper of Tatarowska et al., &quot;The content of total carotenoids, vitamin C and antioxidant properties of 65 potato cultivars characterised under the European project ECOBREED&quot; published in Int. J. Mol. Sci. 24 (2023), 11716.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Bioactive compounds with no structural analogs (high-confidence activity data)

<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis&nbsp;to be published in &#39;Medicinal Chemistry Communications&#39;. &nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-zeroNov 2015View details →
zenodo40/100

Systematic Design of Analogs of Active Compounds Covering More than 1000 Targets

<p>The analog database consisting of 1,297,204 virtual compounds is provided. Virtual compounds are reported in SMILES representation. In addition, for each virtual compound all available ChEMBL analogs (CHEMBL_COMPOUND_ID) and their activities (CHEMBL_TARGET_IDs) are given.</p>

opencc-zeroFeb 2016View details →
zenodo40/100

Morphing libraries, QSAR models, and compounds predicted to be active on the Glucocorticoid receptor (GR)

<p>This repository contains datasets and files related to the computational drug discovery project of the chemical space exploration of the Glucocorticoid receptor. The accompanying Python code is freely available in the GitHub repository (<a title="https://github.com/Iagea/GRML_analyses" href="https://github.com/Iagea/GRML_analyses" target="_blank" rel="noreferrer noopener">https://github.com/Iagea/GRML_analyses</a>).</p> <p><strong>Morphing Libraries:</strong></p> <ul> <li><strong>GRML_library.csv:</strong> The GRML library is the collection of 999,015 virtual compounds generated by Molpher [1-2] starting from GR ligands with unique Bemis-Murcko scaffolds collected from the ChEMBL17 and IMG libraries.</li> <li><strong>RML_library.csv:</strong> The RML library is the collection of 1,346,310 virtual compounds generated by Molpher starting from compounds with unique Bemis-Murcko scaffolds&nbsp; randomly selected from the ZINC database.</li> </ul> <p><strong>IMG library:</strong></p> <ul> <li><strong>IMG_non_proprietary.csv</strong>: The non-proprietary IMG library subset containing 12,956 compounds and their corresponding B-scores from the primary screen.</li> </ul> <p><strong>Molpher inputs:</strong></p> <ul> <li><strong>GR_inputs.csv</strong>: The GR inputs are the ligands used to create the GRML library, 204 compounds from ChEMBL17 (95 compounds) and the non-proprietary dataset from IMG (109 compounds).</li> <li><strong>Random_inputs.csv</strong>: The random inputs are 249 random ZINC compounds used to create the Random library.</li> </ul> <p><strong>Model's training sets:</strong></p> <ul> <li><strong>Model33_training_set.csv</strong>: Random forest classification model training set, it includes 865 compounds; known GR actives and inactives from ChEMBL33 (738 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>Model17_training_set.csv</strong>: Random forest classification model training set, it includes 601 compounds; known GR actives and inactives from ChEMBL17 (474 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>RFR_training_set.csv</strong>: Random forest regression model training set, it includes 89 compounds; known GR actives and inactives from ChEMBL33 that fit into the GR pharmacophore with the four features we describe in our paper.</li> </ul> <p><strong>Models:</strong></p> <ul> <li><strong>Model33.pkl:</strong> Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL33 and IMG libraries.</li> <li><strong>Model17.pkl</strong>: Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL17 and IMG libraries.</li> <li><strong>RFR_models.pkl</strong><em>: </em>Python pickle file containing the 100 trained random forest regression models used to rank the proposed active morphs. These models were trained with the RFR_training_set.csv.</li> </ul> <p><strong>Active predicted morphs:</strong></p> <ul> <li><strong>all_morphs_actives</strong><em><strong>_</strong></em><strong>predicted.xlsx:</strong> An Excel spreadsheet containing two sheets. 1) All 22,524 GRML active predicted morphs. 2) All 4,341 RML active predicted morphs. The QED, NIBR Severity Score, and Molskill Score are given for each morph.</li> </ul> <p><strong>Proposed GR active ligands:</strong></p> <ul> <li><strong>designed_ligands.xlsx</strong>: An Excel spreadsheet containing two sheets. 1) All 54 designed GR ligands with their QED, NIBR severity score, MolSkill score, consensus ranking from the 100 RFR models, and the result of the manual annotation and remarks, if available. 2) The structure of the 54 ligands based on their manual annotation and presence or not in ChEMBL33 database.</li> </ul> <p>Researchers and professionals in the field of drug discovery and cheminformatics may find these resources useful for further analysis and investigations.</p> <p><strong>Bibliography</strong></p> <p>[1] Hoksza, D., &Scaron;koda, P., Vor&scaron;il&aacute;k, M. <em>et al.</em> Molpher: a software framework for systematic chemical space exploration. <em>J Cheminform</em> <strong>6</strong>, 7 (2014). https://doi.org/10.1186/1758-2946-6-7</p> <p>[2] <a href="https://github.com/lich-uct/molpher-lib">https://github.com/lich-uct/molpher-lib</a></p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Dataset: Consistent release of volatile organic compounds across an actively degrading permafrost peatland

<p>Here, we conducted in situ measurements of soil and pond VOC emissions across an actively degrading permafrost peatland in subarctic Norway. We used a permafrost thaw gradient that covered bare soil and vegetated palsa plateaus, underlain by intact permafrost, and increasingly degraded permafrost landscapes: thaw slumps, thaw ponds, and vegetated thaw ponds.</p> <p>This dataset includes two excel files: 1) the first one &quot;Finnmark_source_data&quot; is the source data for figures&nbsp;in the publication <a href="https://doi.org/10.1016/j.geoderma.2023.116355">https://doi.org/10.1016/j.geoderma.2023.116355</a>. ii) the second one &quot;Rawdata_of_emission_rate&quot; is the emission rate of the 210 VOC species identified in this study.</p> <p>Results showed that every peatland landscape type was an important and consistent source of atmospheric VOCs, with a large variety species, such as methanol, acetone, monoterpenes, sesquiterpenes, isoprene, hydrocarbons, oxygenated VOCs, etc. VOC composition varied considerably across the measurement period and across the permafrost thaw gradient. We observed enhanced terpenoid emissions following thaw slump degradation, highlighting the potential atmospheric impact of permafrost thaw, due to the high chemical reactivities of terpenoid compounds. Overall, our study demonstrates that VOCs are being emitted in significant quantities and with largely similar composition upon permafrost thawing, inundation, and subsequent vegetation development, despite major differences in microclimate, hydrological regime, vegetation, and permafrost occurrence.</p> <p>Should you have any questions regarding the dataset, please free feel to contact Yi jiao at yi.jiao@bio.ku.dk or the PI of this project Prof. Rinnan at riikkar@bio.ku.dk</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Analog Series of Compounds with High Frequency of Activity in Screening Assays

<p>A set of 6941 analog series and associated data are provided. These series exclusively consist of compounds that are most frequently active across public screening assays.&nbsp; &nbsp;</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Data for Non-Equilibrium Sensing of Volatile Compounds Using Active and Passive Analyte Delivery

<blockquote> <p>Version 2: Added missing files to&nbsp;<code>sniffing_data.zip</code></p> </blockquote> <p>See GitHub repository for data processing functions and examples: <a href="https://github.com/soerenbrandt/sniffing-sensor">https://github.com/soerenbrandt/sniffing-sensor</a></p> <p><strong>Abstract</strong>:<br>Sensor technologies have allowed us to outperform the human senses of sight, hearing, and touch; however, the development of artificial noses is significantly behind their biological counterparts. This is largely due to the complexity of natural olfaction, as it incorporates complex fluid dynamics within the nasal anatomy together with the response patterns of hundreds to thousands of unique molecular-scale receptors for odor interpretation. We designed a sensing approach to identify volatiles that exploits time-dependent information from a single sensor (here, the reflectance spectra from a mesoporous one-dimensional photonic crystal) by augmenting and accentuating differences in the non-equilibrium mass-transport dynamics of vapors stemming from their distinct physicochemical properties, thus obviating the need for a large sensor array. By training a machine learning algorithm on the sensor output, we clearly identify polar and nonpolar volatile organic compounds, determine the mixing ratios of binary mixtures, and accurately predict the boiling point, flash point, vapor pressure, and viscosity of several volatile liquids within those used for training as well as compounds unknown to the model. We further implement a bioinspired active sniffing approach, in which the fluid dynamics and patterns of analyte delivery are controlled, enabling an additional modality of differentiation and reducing the duration of data collection and analysis to seconds. These results outline a strategy to build accurate and rapid artificial noses for volatile liquids that can provide useful information on chemicals such as their composition and properties, and can be applied in a variety of fields, including disease diagnosis, hazardous waste management, and healthy building monitoring.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Dataset for "Electrodeposited p-Cu2O Films – Role of Redox-Active Compounds Under Photoelectrochemical Operation Revisited"

<p>The p-type semiconducting copper oxides CuO and Cu2O are of interest for the conversion of solar energy due to their medium wide bandgap. The position of their conduction band should allow for reductive processes in junctions with electrolytes under irradiation. In this work, on Cu2O, the efficiency of several such processes in competition with self-reduction was studied.<br>Thin films of Cu2O were synthesised via potentiostatic electrodeposition on FTO (fluorine doped SnO2 on glass) using aqueous electrolytes. Deposition parameters (electrolyte composition, electrode potential, temperature) had a great influence on electrode properties (phase purity, preferential orientation). Electrodeposited Cu2O films were p-type. In junctions with aqueous electrolytes, under electrical bias, cathodic and photocathodic currents passed which increased dramatically when reducible redox compounds were added. The influence of various redox couples (O2, H2O2, methylviologen) and their concentration in the electrolyte on the stability of the electrodes was studied.<br>Long time experiments showed that oxygen saturated and H2O2 containing electrolytes gave rise to constant photocurrents and no alteration of the electrodes was found by XRD. With the investigated redox couples, no product of interest was produced.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Can metal organic frameworks outperform adsorptive removal of harmful phenolic compound 2-chlorophenol by activated carbon?

<p>Dataset supporting publication. High resolution images, and full data set as produced in manuscript figures.</p> <p><strong>Preprint</strong>: <a href="https://doi.org/10.26434/chemrxiv.10320752.v1">https://doi.org/10.26434/chemrxiv.10320752.v1 </a></p> <p><strong>Published article:</strong> <a href="https://doi.org/10.1016/j.cherd.2020.03.017">https://doi.org/10.1016/j.cherd.2020.03.017 </a></p> <p><strong>Abstract:</strong> Removal of persistent organic compounds from aqueous solutions is generally achieved using adsorbent like activated carbon (AC) but it suffers from limited adsorption capacity due to low surface area. This paper describes a pioneering work on the adsorption of an organic pollutant, 2-chlorophenol (2-CP) by two MOFs with high surface area and water stability; MIL-101 and its amino-derivative, MIL-101-NH<sub>2</sub>. Although MOFs have higher surface area than AC, the latter was proven better having the highest equilibrium 2-CP uptake (345 mg.g<sup>-1</sup>), followed by MIL-101 (121 mg.g<sup>-1</sup>) and MIL-101-NH<sub>2</sub> (84 mg.g<sup>-1</sup>). Used MIL-101 could be easily regenerated multiple times by washing with ethanol and even showed improved adsorption capacity after each washing cycle. These results can open the doors to meticulous adsorbent selection for treating 2-CP-contaminated water. &nbsp;&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

In silico design, docking simulation, and ANN-QSAR model for predicting the anticoagulant activity of thiourea isosteviol compounds as FXa inhibitors

<p>The present work combined molecular modeling and docking approach for searching and designing novel thiourea isosteviol-based compounds as potential FXa inhibitors. Elaborated regression model establishes the relationships between experimentally determined anticoagulant activity and molecular descriptors and enables the prediction of FXa inhibitory activity for novel compounds. The obtained results proved that the Artificial Neural Network algorithm facilitates the search for the most promising isosteviol derivatives incorporating thiourea fragments as FXa inhibitors. Moreover, docking simulation confirms the prominent binding of the newly in silico designed molecules with the active sites of the protein, which may be the lead molecules and can be further optimized for the efficient pharmacodynamic and pharmacokinetic profiles.&nbsp;The enclosed files are representations of&nbsp;molecular structures of thiourea isosteviol compounds with experimentally tested FXa inhibitory activity (i20-i39) geometrically optimized in hyperchem, newly in silico designed thiourea isosteviol compounds geometrically optimized in hyperchem (e1-e11), one file contains molecular descriptors for optimized structures calculated in Dragon and there is also a code for ANN QSAR model for predicting activity of novel thiourea isosteviol compounds.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Detailed data sets of MMP-cliffs, SAR transfer series, RECAP-MMPs and compound activities

<p>An up-to-date version of three MMP-based data sets derived from compounds included in the latest release of ChEMBL is presented. These data sets include activity cliffs, structure-activity relationship (SAR) transfer series, and second generation MMPs based upon retrosynthetic rules. The structural data and information are provided in eight different files comprising the data sets of MMP-cliffs, SAR transfer series, and RECAP-MMPs. Compound activities are incorporated in files of RECAP-MMPs. For transfer series, substituted fragments are also provided. All MMP-cliffs, SAR transfer series with approximate or regular potency progression, and RECAP-MMPs are provided in canonical SMILES representation on a per-target basis separately for the Ki and IC50 subsets. The corresponding files are clearly designated.</p>

opencc-zeroFeb 2014View details →
zenodo36/100

Follow-up: Prospective compound design using the ‘SAR Matrix’ method and matrix-derived conditional probabilities of activity

<p>Details of the conditional probability calculations on exemplary the matrix provided in Figure 3 of the publication (see Gupta-Ostermann, Hirose, Odagami &amp; Bajorath, Follow-up: Prospective compound design using the &lsquo;SAR Matrix&rsquo; method and matrix-derived conditional probabilities of activity, F1000Research 2015, 4:75 , DOI:&nbsp;10.12688/f1000research.6271.1 ) is provided in an excel sheet.&nbsp;</p> <p>Informative SARMs from the PRISM library are included. Due to proprietary issues the structural information of compounds is not included. Key and value fragments of SARMs are provided with an identifier.</p>

opencc-zeroApr 2015View details →
zenodo36/100

Currently available 3D activity cliffs and 2D-analogs of 3D-cliff compounds

<p>Three dimensional activity cliffs (3D-cliffs) were systematically determined based on currently available X-ray structures in PDB. The list of all 236, 292, 595 3D-cliffs that were identified from the K<sub>i</sub>, IC<sub>50</sub>, and K<sub>i</sub>/IC<sub>50</sub> sets, respectively, is provided. In addition,&nbsp;on the basis of matched molecular pairs, the 2D structural analogs of 3D-cliff compounds identified from ChEMBL database (release 19) are given.</p>

opencc-zeroMay 2015View details →
zenodo36/100

Sets of ChEMBL compounds with high or low confidence activity data

<p>Two sets of compounds assembled from ChEMBL release 20 that were annotated with high or low confidence activity data were provided in separate files. For each compound in a file, the unique compound identifier (i.e., molregno), the number of targets in individual years (from 1976 to 2014) and the list of target annotations (if any) was given.</p>

opencc-zeroMay 2015View details →
zenodo36/100

Compounds with multi-target activities forming target cliffs and selectivity cliffs

<p>On the basis of high-confidence activity data assembled from ChEMBL release 20, compounds active against one or more pairs of targets are provided (in the file &#39;Compounds_In_Target_Pairs.xlsx&#39;) for the K<sub>i</sub> and IC<sub>50</sub> value-based data sets, respectively. Target cliffs formed by selective compounds are listed in the file &#39;Target_Cliffs.xlsx&#39;. In addition, selectivity cliffs formed by structurally analogous compounds with significantly different selectivity are given in the file &#39;Selectivity_Cliffs.xlsx&#39;. The content of individual files is detailed in the README file.</p>

opencc-zeroAug 2015View details →
zenodo36/100

Compound activity records associated with original publications in ChEMBL 21

<p>Provided are two sets of compound activity records (set 1 and set 2) that were traced back to original publications and assembled from ChEMBL release 21. For each compound-target combination, the corresponding potency measurements and publications are provided. In addition, the list of unique publications is given for both sets 1 and 2.</p>

opencc-zeroMay 2016View details →
zenodo36/100

NMR data for "Application of a Hydrophobic Polyglutamate Bearing a Triphenylphosphine Group for the Orientation of Pharmaceutically Active Compounds and the Measurement of Residual Dipolar Couplings"

<p>NMR raw data for the work titled:</p> <p>"Application of a Hydrophobic Polyglutamate Bearing a Triphenylphosphine Group for the Orientation of Pharmaceutically Active Compounds and the Measurement of Residual Dipolar Couplings"</p> <p>The archive consists of NMR spectra of all compounds synthesized and the NMR spectra used for the determination of RDCs (isotropic and anisotropic).</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

CARA: Benchmarking Compound Activity Prediction for Real-World Drug Discovery Applications

<p>Identifying active compounds for target proteins is fundamental in early drug discovery.&nbsp;Recently, data-driven computational methods have demonstrated promising potential in predicting compound activities.&nbsp;However, there lacks a well-designed benchmark to comprehensively evaluate these methods from a practical perspective.&nbsp;To fill this gap, we propose a benchmark, named CARA.Through carefully distinguishing assay types, designing train-test splitting schemes and selecting evaluation metrics, CARA can consider the biased distribution of current real-world compound activity data and avoid overestimation of model performances. We observed that current models can make successful predictions for certain proportions of assays, while the performances varied across different assays. In addition, evaluation of several few-shot training strategies demonstrated different performances related to task types. Overall, we provide a high-quality dataset for developing and evaluating compound activity prediction models, and the analyses in this work may inspire better applications of data-driven models in drug discovery.</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record