Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.9.0
Dataset results
30 results for “bioactivity data”
Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data
<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed “similarity-based merger models” which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC>0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>
Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions
<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p> </p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version </strong>data was incorrectly duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this fixed 05.5 version</strong> the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p> </p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p> </p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>
Bioactive compounds with no structural analogs (high-confidence activity data)
<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis to be published in 'Medicinal Chemistry Communications'. </p>
Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"
<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used to calculate the embeddings.</p>
Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis
<p>We carried out a systematic evaluation of target selectivity profiles across three recent large-scale biochemical assays of kinase inhibitors and further compared these standardized bioactivity assays with data reported in the widely used databases ChEMBL and STITCH. Our comparative evaluation revealed relative benefits and potential limitations among the bioactivity types, as well as pinpointed biases in the database curation processes. Ignoring such issues in data heterogeneity and representation may lead to biased modeling of drugs' polypharmacological effects as well as to unrealistic evaluation of computational strategies for the prediction of drug-target interaction networks. Toward making use of the complementary information captured by the various bioactivity types, including IC50, K(i), and K(d), we also introduce a model-based integration approach, termed KIBA, and demonstrate here how it can be used to classify kinase inhibitor targets and to pinpoint potential errors in database-reported drug-target interactions. An integrated drug-target bioactivity matrix across 52,498 chemical compounds and 467 kinase targets, including a total of 246,088 KIBA scores, has been made freely available.</p> <p>Please cite: </p> <p>https://pubmed.ncbi.nlm.nih.gov/24521231/ </p> <p>https://pubs.acs.org/doi/10.1021/ci400709d</p>
Data underlying the article: 3DDPDs: Describing protein dynamics for proteochemometric bioactivity prediction. A case for (mutant) G protein-coupled receptors
<p>This repository contains the datasets and results supporting the conclusions of the manuscript "<strong>3DDPDs: Describing protein dynamics for proteochemometric bioactivity prediction. A case for (mutant) G protein-coupled receptors</strong>". </p> <p>Publicly available data is not included in this repository. The source code to generate the results gathered here can be found on GitHub (https://github.com/CDDLeiden/3ddpd). </p>
Umbrella review data of colour-associated bioactive pigments found in fruit and vegetables, compared to placebo or low intakes, on human health outcomes relevant to public health
<p>This dataset comprises data extracted from 86 publications included in an umbrella review which compared the effect of colour-associated bioactive pigments found in fruit and vegetables (carotenoids, flavonoids, betalains and chlorophyll) on human health outcomes relevant to public health. Meta-analysed data from 83 systematic literature reviews were available for 17 different bioactive pigments spanning all colours of fruit and vegetables except green, with additional data from two randomised controlled trials and one cohort study for chlorophyll. This dataset represents 2,847 original research studies and data from over 37 million participants. There were 449 meta-analysed health outcomes extracted from the 83 systematic literature reviews. Extracted health outcomes were categorised according to pigment, comparator type, broad health outcome, study design, age group, source of pigment, risk of bias, and confidence in the estimated effect. Data extracted were dose, intervention duration, sample size, number of original studies, effect estimate, 95% confidence intervals, I2 statistics, publication bias and p-value. Rigorous analysis, including estimations of a common effect size, stratification of the evidence, study level sensitivity analyses and reporting on heterogeneity and potential biases, may be carried out using these data, to further evaluate the effect of colour-associated bioactive pigments in fruit and vegetables on human health.</p>
Data underlying the article: "Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling"
<p>This repository contains the data underlying the article: “Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling” available as a preprint on ChemRxiv.</p> <p>Main authors: Marina Gorostiola González & Olivier J.M. Béquignon (Leiden University)</p> <p>Senior author: Gerard J.P. van Westen (Leiden University)</p> <p>This analysis was performed using the code available at <a href="https://github.com/CDDLeiden/chembl_variants" target="_blank" rel="noopener">https://github.com/CDDLeiden/chembl_variants</a></p>
Data described in the article "Three-step enzymatic remodeling of chitin into bioactive chitooligomers"
<p>The Supporting Information file contains the following data: <span>PCR primers and reaction conditions for the generation of <em>Tf</em>Chit mutants; HPLC chromatogram of chitin hydrolysate; detailed sequence alignment; computational data; additional figures for modeling of <em>Tf</em>Chit;</span><span> <span>MALDI-TOF MS analysis of a mixture of insoluble chitooligomers prepared by Y445N <em>Ao</em>Hex, and <em>m/z</em> values monitored by HPLC-MS of the deacetylation reactions.</span></span></p>
Primary data for Manuscript provisionally titled Stereoselective access to bioactive cyclopropanes (+)-PPCC and (1R,2S)-2-aminomethyl-1-arylcyclopropane-1-carboxamides from (−)-levoglucosenone
<p><span>Contains HRMS and FID data for compounds described in the manuscript titled, "<span>Stereoselective access to bioactive cyclopropanes (+)-PPCC and (1<em><span>R</span></em>,2<em><span>S</span></em>)-2-aminomethyl-1-arylcyclopropane-1-carboxamides from (−)-levoglucosenone"</span></span></p> <p><span>FIDs can be opened using SpinWorks or Topspin programs.</span></p> <p> </p>
Data from: Identification of anti-fungal bioactive terpenoids from the bioenergy crop switchgrass (Panicum virgatum)
<p>Plant derived bioactive small molecules have attracted attention of scientists across fundamental and applied scientific disciplines. We seek to understand the influence of these phytochemicals on functional phytobiomes. Increased knowledge of specialized metabolite bioactivities could inform strategies for sustainable crop production. We hypothesized that – consistent with accumulating evidence that switchgrass genotype impacts microbiome assembly – differential terpenoid accumulation contributes to switchgrass ecotype-specific microbiome composition. An initial in vitro plate-based disc diffusion screen of 18 switchgrass root derived fungal isolates revealed differential responses to upland- and lowland-isolated metabolites. To identify specific fungal growth-modulating metabolites, we tested fractions from root extracts on three ecologically important fungal isolates – <em>Linnemania elongata</em>, <em>Trichoderma</em> sp. and <em>Fusarium</em> sp. Saponins and diterpenoids were identified as the most prominent antifungal metabolites. Finally, analysis of liquid chromatography-purified terpenoids revealed fungal inhibition structure – activity relationships (SAR). Saponin antifungal activity was primarily determined by the number of sugar moieties – saponins glycosylated at a single core position were inhibitory whereas saponins glycosylated at two core positions were inactive. Saponin core hydroxylation and acetylation were also associated with reduced activity. Diterpenoid activity required the presence of an intact furan ring for strong fungal growth inhibition.</p>
Umbrella review data of colour-associated bioactive pigments found in fruit and vegetables, compared to placebo or low intakes, on human health outcomes relevant to public health
Open the record for dataset details and reuse information.
Data from: Identification of anti-fungal bioactive terpenoids from the bioenergy crop switchgrass (Panicum virgatum)
Open the record for dataset details and reuse information.
Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis
Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.
Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis
Open the record for dataset details and reuse information.
Data from: Evolutionary origins of a bioactive peptide buried within preproalbumin
The de novo evolution of proteins is now considered a frequented route for biological innovation, but the genetic and biochemical processes that lead to each newly created protein are often poorly documented. The common sunflower (Helianthus annuus) contains the unusual gene PawS1 (Preproalbumin with SFTI-1) that encodes a precursor for seed storage albumin; however, in a region usually discarded during albumin maturation, its sequence is matured into SFTI-1, a protease-inhibiting cyclic peptide with a motif homologous to unrelated inhibitors from legumes, cereals, and frogs. To understand how PawS1 acquired this additional peptide with novel biochemical functionality, we cloned PawS1 genes and showed that this dual destiny is over 18 million years old. This new family of mostly backbone-cyclic peptides is structurally diverse, but the protease-inhibitory motif was restricted to peptides from sunflower and close relatives from its subtribe. We describe a widely distributed, potential evolutionary intermediate PawS-Like1 (PawL1), which is matured into storage albumin, but makes no stable peptide despite possessing residues essential for processing and cyclization from within PawS1. Using sequences we cloned, we retrodict the likely stepwise creation of PawS1's additional destiny within a simple albumin precursor. We propose that relaxed selection enabled SFTI-1 to evolve its inhibitor function by converging upon a successful sequence and structure.
Data from: Green approach for synthesis of bioactive Hantzsch 1,4-dihydropyridine derivatives based on thiophene moiety via multicomponent reaction
A novel green and efficient one-pot multicomponent reaction of dihydropyridine derivatives was reported as having good to excellent yield. In the presence of the catalyst ceric ammonium nitrate (CAN), different 1,3-diones and same starting materials as 5-bromothiophene-2-carboxaldehyde and ammonium acetate were used at room temperature under solvent-free condition for the Hantzsch pyridine synthesis within a short period of time. All compounds were evaluated for their in vitro antibacterial and antifungal activity and, interestingly, we found that 5(b–f) show excellent activity compared with Ampicillin, whereas only the 5e compound shows excellent antifungal activity against Candida albicans compared with griseofulvin. The cytotoxicity of all compounds has been assessed against breast tumour cell lines (BT-549), but no activity was found. The X-ray structure of one such compound, 5a, viewed as a colourless block crystal, corresponded accurately to a primitive monoclinic cell.
Data from: Random sequences are an abundant source of bioactive RNAs or peptides
It is generally assumed that new genes arise through duplication and/or recombination of existing genes. The probability that a new functional gene could arise out of random non-coding DNA is so far considered to be negligible, as it seems unlikely that such an RNA or protein sequence could have an initial function that influences the fitness of an organism. Here, we have tested this question systematically, by expressing clones with random sequences in Escherichia coli and subjecting them to competitive growth. Contrary to expectations, we find that random sequences with bioactivity are not rare. In our experiments we find that up to 25% of the evaluated clones enhance the growth rate of their cells and up to 52% inhibit growth. Testing of individual clones in competition assays confirms their activity and provides an indication that their activity could be exerted by either the transcribed RNA or the translated peptide. This suggests that transcribed and translated random parts of the genome could indeed have a high potential to become functional. The results also suggest that random sequences may become an effective new source of molecules for studying cellular functions, as well as for pharmacological activity screening.
Data from: Identification of reactive intermediate formation and bioactivation pathways in Abemaciclib metabolism by LC–MS/MS: in vitro metabolic investigation
Abemaciclib (Verzenio®) is approved as tyrosine kinase inhibitor (TKI) for breast cancer treatment. In this study, in vitro phase I metabolic profiling of Abemaciclib (ABC) was done using rat liver microsomes (RLMs). We checked the formation of reactive intermediates in ABC metabolism using (RLMs) in the presence of potassium cyanide (KCN) that was used as capturing agent for iminium reactive intermediates forming a stable complex that can be characterized by LC-MS/MS. Nine in vitro phase I metabolites and three cyano adducts were identified. The metabolic reactions involved in formation of these metabolites and adducts are reduction, oxidation, hydroxylation and cyanide addition. The bioactivation pathway was also proposed. Knowing the electrodeficient bioactive centre in ABC structure helped in making targeted modifications to improve its safty and retain its efficacy. Blocking or isosteric replacement of α carbon to the tertiary nitrogen atoms of piperazine ring can aid in reducing toxic side effects of ABC. No previous articles were found about in vitro metabolic profiling for ABC or structural identification of the formed reactive metabolites for ABC.
Data from: Green approach for synthesis of bioactive Hantzsch 1,4-dihydropyridine derivatives based on thiophene moiety via multicomponent reaction
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.