Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

30

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

30 results for “bioactivity data”

Learn how ShareScore rates datasets ↗
zenodo48/100

Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data

<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed &ldquo;similarity-based merger models&rdquo; which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC&gt;0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p>&nbsp;</p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version&nbsp;</strong>data was incorrectly&nbsp;duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this&nbsp;fixed 05.5 version</strong>&nbsp;the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p>&nbsp;</p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p>&nbsp;</p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers&rsquo; time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>

opencc-by-sa-4.0Aug 2022View details →
zenodo40/100

Bioactive compounds with no structural analogs (high-confidence activity data)

<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis&nbsp;to be published in &#39;Medicinal Chemistry Communications&#39;. &nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-zeroNov 2015View details →
zenodo40/100

Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"

<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used&nbsp; to calculate the embeddings.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis

<p>We carried out a systematic evaluation of target selectivity profiles across three recent large-scale biochemical assays of kinase inhibitors and further compared these standardized bioactivity assays with data reported in the widely used databases ChEMBL and STITCH. Our comparative evaluation revealed relative benefits and potential limitations among the bioactivity types, as well as pinpointed biases in the database curation processes. Ignoring such issues in data heterogeneity and representation may lead to biased modeling of drugs' polypharmacological effects as well as to unrealistic evaluation of computational strategies for the prediction of drug-target interaction networks. Toward making use of the complementary information captured by the various bioactivity types, including IC50, K(i), and K(d), we also introduce a model-based integration approach, termed KIBA, and demonstrate here how it can be used to classify kinase inhibitor targets and to pinpoint potential errors in database-reported drug-target interactions. An integrated drug-target bioactivity matrix across 52,498 chemical compounds and 467 kinase targets, including a total of 246,088 KIBA scores, has been made freely available.</p> <p>Please cite:&nbsp;</p> <p>https://pubmed.ncbi.nlm.nih.gov/24521231/&nbsp;</p> <p>https://pubs.acs.org/doi/10.1021/ci400709d</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Data underlying the article: 3DDPDs: Describing protein dynamics for proteochemometric bioactivity prediction. A case for (mutant) G protein-coupled receptors

<p>This repository contains the datasets and results supporting the conclusions of the manuscript &quot;<strong>3DDPDs: Describing protein dynamics for proteochemometric bioactivity prediction. A case for (mutant) G protein-coupled receptors</strong>&quot;.&nbsp;</p> <p>Publicly available data is not included in this repository. The source code to generate the results gathered here can be found on GitHub (https://github.com/CDDLeiden/3ddpd).&nbsp;</p>

opencc-by-4.0May 2023View details →
dryad36/100

Umbrella review data of colour-associated bioactive pigments found in fruit and vegetables, compared to placebo or low intakes, on human health outcomes relevant to public health

<p>This dataset comprises data extracted from 86 publications included in an umbrella review which compared the effect of colour-associated bioactive pigments found in fruit and vegetables (carotenoids, flavonoids, betalains and chlorophyll) on human health outcomes relevant to public health.  Meta-analysed data from 83 systematic literature reviews were available for 17 different bioactive pigments spanning all colours of fruit and vegetables except green, with additional data from two randomised controlled trials and one cohort study for chlorophyll. This dataset represents 2,847 original research studies and data from over 37 million participants. There were 449 meta-analysed health outcomes extracted from the 83 systematic literature reviews. Extracted health outcomes were categorised according to pigment, comparator type, broad health outcome, study design, age group, source of pigment, risk of bias, and confidence in the estimated effect.  Data extracted were dose, intervention duration, sample size, number of original studies, effect estimate, 95% confidence intervals, I2 statistics, publication bias and p-value. Rigorous analysis, including estimations of a common effect size, stratification of the evidence, study level sensitivity analyses and reporting on heterogeneity and potential biases, may be carried out using these data, to further evaluate the effect of colour-associated bioactive pigments in fruit and vegetables on human health.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Data underlying the article: "Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling"

<p>This repository contains the data underlying the article: &ldquo;Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling&rdquo; available as a preprint on ChemRxiv.</p> <p>Main authors: Marina Gorostiola Gonz&aacute;lez &amp; Olivier J.M. B&eacute;quignon (Leiden University)</p> <p>Senior author: Gerard J.P. van Westen (Leiden University)</p> <p>This analysis was performed using the code available at <a href="https://github.com/CDDLeiden/chembl_variants" target="_blank" rel="noopener">https://github.com/CDDLeiden/chembl_variants</a></p>

openmit-licenseMay 2024View details →
zenodo36/100

Data described in the article "Three-step enzymatic remodeling of chitin into bioactive chitooligomers"

<p>The Supporting Information file contains the following data: <span>PCR primers and reaction conditions for the generation of <em>Tf</em>Chit mutants; HPLC chromatogram of chitin hydrolysate; detailed sequence alignment; computational data; additional figures for modeling of <em>Tf</em>Chit;</span><span> <span>MALDI-TOF MS analysis of a mixture of insoluble chitooligomers prepared by Y445N <em>Ao</em>Hex, and <em>m/z</em> values monitored by HPLC-MS of the deacetylation reactions.</span></span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Primary data for Manuscript provisionally titled Stereoselective access to bioactive cyclopropanes (+)-PPCC and (1R,2S)-2-aminomethyl-1-arylcyclopropane-1-carboxamides from (−)-levoglucosenone

<p><span>Contains HRMS and FID data for compounds described in the manuscript titled, "<span>Stereoselective access to bioactive cyclopropanes (+)-PPCC and (1<em><span>R</span></em>,2<em><span>S</span></em>)-2-aminomethyl-1-arylcyclopropane-1-carboxamides from (&minus;)-levoglucosenone"</span></span></p> <p><span>FIDs can be opened using SpinWorks or Topspin programs.</span></p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

Data from: Identification of anti-fungal bioactive terpenoids from the bioenergy crop switchgrass (Panicum virgatum)

<p>Plant derived bioactive small molecules have attracted attention of scientists across fundamental and applied scientific disciplines. We seek to understand the influence of these phytochemicals on functional phytobiomes. Increased knowledge of specialized metabolite bioactivities could inform strategies for sustainable crop production. We hypothesized that – consistent with accumulating evidence that switchgrass genotype impacts microbiome assembly – differential terpenoid accumulation contributes to switchgrass ecotype-specific microbiome composition. An initial in vitro plate-based disc diffusion screen of 18 switchgrass root derived fungal isolates revealed differential responses to upland- and lowland-isolated metabolites. To identify specific fungal growth-modulating metabolites, we tested fractions from root extracts on three ecologically important fungal isolates – <em>Linnemania elongata</em>, <em>Trichoderma</em> sp. and <em>Fusarium</em> sp. Saponins and diterpenoids were identified as the most prominent antifungal metabolites. Finally, analysis of liquid chromatography-purified terpenoids revealed fungal inhibition structure – activity relationships (SAR). Saponin antifungal activity was primarily determined by the number of sugar moieties – saponins glycosylated at a single core position were inhibitory whereas saponins glycosylated at two core positions were inactive. Saponin core hydroxylation and acetylation were also associated with reduced activity. Diterpenoid activity required the presence of an intact furan ring for strong fungal growth inhibition.</p>

opencc-zeroJun 2023View details →
dryad36/100

Umbrella review data of colour-associated bioactive pigments found in fruit and vegetables, compared to placebo or low intakes, on human health outcomes relevant to public health

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad36/100

Data from: Identification of anti-fungal bioactive terpenoids from the bioenergy crop switchgrass (Panicum virgatum)

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad32/100

Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis

Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad28/100

Data from: Evolutionary origins of a bioactive peptide buried within preproalbumin

The de novo evolution of proteins is now considered a frequented route for biological innovation, but the genetic and biochemical processes that lead to each newly created protein are often poorly documented. The common sunflower (Helianthus annuus) contains the unusual gene PawS1 (Preproalbumin with SFTI-1) that encodes a precursor for seed storage albumin; however, in a region usually discarded during albumin maturation, its sequence is matured into SFTI-1, a protease-inhibiting cyclic peptide with a motif homologous to unrelated inhibitors from legumes, cereals, and frogs. To understand how PawS1 acquired this additional peptide with novel biochemical functionality, we cloned PawS1 genes and showed that this dual destiny is over 18 million years old. This new family of mostly backbone-cyclic peptides is structurally diverse, but the protease-inhibitory motif was restricted to peptides from sunflower and close relatives from its subtribe. We describe a widely distributed, potential evolutionary intermediate PawS-Like1 (PawL1), which is matured into storage albumin, but makes no stable peptide despite possessing residues essential for processing and cyclization from within PawS1. Using sequences we cloned, we retrodict the likely stepwise creation of PawS1's additional destiny within a simple albumin precursor. We propose that relaxed selection enabled SFTI-1 to evolve its inhibitor function by converging upon a successful sequence and structure.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Green approach for synthesis of bioactive Hantzsch 1,4-dihydropyridine derivatives based on thiophene moiety via multicomponent reaction

A novel green and efficient one-pot multicomponent reaction of dihydropyridine derivatives was reported as having good to excellent yield. In the presence of the catalyst ceric ammonium nitrate (CAN), different 1,3-diones and same starting materials as 5-bromothiophene-2-carboxaldehyde and ammonium acetate were used at room temperature under solvent-free condition for the Hantzsch pyridine synthesis within a short period of time. All compounds were evaluated for their in vitro antibacterial and antifungal activity and, interestingly, we found that 5(b–f) show excellent activity compared with Ampicillin, whereas only the 5e compound shows excellent antifungal activity against Candida albicans compared with griseofulvin. The cytotoxicity of all compounds has been assessed against breast tumour cell lines (BT-549), but no activity was found. The X-ray structure of one such compound, 5a, viewed as a colourless block crystal, corresponded accurately to a primitive monoclinic cell.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Random sequences are an abundant source of bioactive RNAs or peptides

It is generally assumed that new genes arise through duplication and/or recombination of existing genes. The probability that a new functional gene could arise out of random non-coding DNA is so far considered to be negligible, as it seems unlikely that such an RNA or protein sequence could have an initial function that influences the fitness of an organism. Here, we have tested this question systematically, by expressing clones with random sequences in Escherichia coli and subjecting them to competitive growth. Contrary to expectations, we find that random sequences with bioactivity are not rare. In our experiments we find that up to 25% of the evaluated clones enhance the growth rate of their cells and up to 52% inhibit growth. Testing of individual clones in competition assays confirms their activity and provides an indication that their activity could be exerted by either the transcribed RNA or the translated peptide. This suggests that transcribed and translated random parts of the genome could indeed have a high potential to become functional. The results also suggest that random sequences may become an effective new source of molecules for studying cellular functions, as well as for pharmacological activity screening.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Identification of reactive intermediate formation and bioactivation pathways in Abemaciclib metabolism by LC–MS/MS: in vitro metabolic investigation

Abemaciclib (Verzenio®) is approved as tyrosine kinase inhibitor (TKI) for breast cancer treatment. In this study, in vitro phase I metabolic profiling of Abemaciclib (ABC) was done using rat liver microsomes (RLMs). We checked the formation of reactive intermediates in ABC metabolism using (RLMs) in the presence of potassium cyanide (KCN) that was used as capturing agent for iminium reactive intermediates forming a stable complex that can be characterized by LC-MS/MS. Nine in vitro phase I metabolites and three cyano adducts were identified. The metabolic reactions involved in formation of these metabolites and adducts are reduction, oxidation, hydroxylation and cyanide addition. The bioactivation pathway was also proposed. Knowing the electrodeficient bioactive centre in ABC structure helped in making targeted modifications to improve its safty and retain its efficacy. Blocking or isosteric replacement of α carbon to the tertiary nitrogen atoms of piperazine ring can aid in reducing toxic side effects of ABC. No previous articles were found about in vitro metabolic profiling for ABC or structural identification of the formed reactive metabolites for ABC.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Green approach for synthesis of bioactive Hantzsch 1,4-dihydropyridine derivatives based on thiophene moiety via multicomponent reaction

Open the record for dataset details and reuse information.

publicMay 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record