Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

606

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

606 results for “Bioactivation”

Learn how ShareScore rates datasets ↗
zenodo52/100

Bioactivity of small-molecule compounds against Haemonchus contortus

<div> <div> <div> <p>This dataset of small-molecule compounds and their effects on <em>H. contortus </em>was assembled based on the results obtained from screening two compound libraries (Medicines for Malaria Venture Pathogen Box, Compounds Australia Open Scaffolds set) to assess the effect of compounds on the motility of exsheathed third-stage larvae (xL3) of <em>H. contortus </em>(Preston et al., 2016, 2017). Additionally, select literature data were included to augment the in-house generated data.</p> </div> </div> </div>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Dataset for: Stereorandomization as a Method to Probe Peptide Bioactivity

<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM, Serum stability, Vesicle leakage assay, Cytotoxicity, Hemolysis.</p>

opencc-by-4.0Jun 2021View details →
zenodo52/100

Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.

<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">&quot;Predicting compound activity from phenotypic profiles and chemical structures&quot;</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper&#39;s GitHub repository</a>&nbsp;for reproduction.</p> <p>Folders and files&nbsp;and are described&nbsp;below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data

<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed &ldquo;similarity-based merger models&rdquo; which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC&gt;0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions

<div> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="https://doi.org/10.26434/chemrxiv-2021-1rxhk">https://doi.org/10.26434/chemrxiv-2021-1rxhk</a>.</p> <p>&nbsp;</p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers&rsquo; time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p> </div>

opencc-by-sa-4.0Apr 2022View details →
zenodo44/100

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p>&nbsp;</p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version&nbsp;</strong>data was incorrectly&nbsp;duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this&nbsp;fixed 05.5 version</strong>&nbsp;the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p>&nbsp;</p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p>&nbsp;</p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers&rsquo; time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>

opencc-by-sa-4.0Aug 2022View details →
zenodo44/100

Bioactive fatty-acid derivated ovalicin from Pseudallescheria boydii

<p>Original data for the mMolecular network with t-SNE representation of crude extracts of <em>P. boydii</em>. SNB-CN71, -CN73, -CN81 and -CN85</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Dataset - Papyrus 2024 - A large scale curated dataset aimed at bioactivity predictions

<p><strong>This update of release 2024.1 fixes the following:</strong></p> <ul> <li>Metadata in the columns <em>type_IC50</em>, <em>type_EC50</em>, <em>type_KD</em>, <em>type_Ki</em>, and <em>type_other</em> did not contain multiple values when multiple pChEMBL values where available but reported only a single value. This fix ensures all values are reported.</li> <li>Molecules were incorrectly standardized and mixtures were included in the dataset. Standardization (using the&nbsp;<a href="https://github.com/OlivierBeq/papyrus_structure_pipeline" target="_blank" rel="noopener">papyrus_structure_pipeline</a>) is now correctly enforced and mixtures have been removed.</li> </ul> <p><strong>Changes since version 05.6</strong></p> <ul> <li>ChEMBL data was updated to ChEMBL version 34</li> <li>data from the IUPHAR/BPS Guide to PHARMACOLOGY has been included</li> <li>data from Pickett et al.'s publication on MMP-12 has been included (<a href="https://doi.org/10.1021/ml100191f">ACS Med Chem Lett. 2011 Jan 13; 2(1): 28&ndash;33. DOI: 10.1021/ml100191f</a>)</li> </ul> <p><strong>Papyrus++:</strong></p> <p>Previous versions mistakenly considered a deviation of 2 log units around compound-target pairs to determine the reproducibility of assays (see published article for more details). This has been fixed to 0.5 log units to ensure data points fall within a maximum range of 1 log unit. As a result, the number of entries in the Papyrus++ set from this release has drastically reduced compared to previous releases.</p>

opencc-by-sa-4.0Dec 2023View details →
zenodo44/100

The effect of dietary bioactive on gut microbiome diversity (DIME) – a pilot study

<p>The DIME study consists of a randomised 2x2 cross-over human intervention where healthy participants (n = 20) are subjected to a diet high in bioactive-rich food for two weeks and a diet low in bioactive-rich food. There is a four-week washout between the two interventions.&nbsp;</p> <p>The continuous glucose monitoring was achieved using the Abbott freesylte libre flash glucose device. The baseline of the participants were determined 7 days before the start of the intervention, followed by the first arm and second arm. The period between the two arms (washout) was not recorded.</p> <p>We also included&nbsp;sleep data which consists of the amount of time spent in bed and during that time the amount of time&nbsp;spent in light, deep and rem in all 20 participants during the course of the dietary intervention, both the high and low bioactive diet.&nbsp;&nbsp;that was captured using Fitbit wearables during both stages of the dietary intervention,</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Bioactive compounds with no structural analogs (high-confidence activity data)

<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis&nbsp;to be published in &#39;Medicinal Chemistry Communications&#39;. &nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-zeroNov 2015View details →
zenodo40/100

Comprehensive 16s rRNA sequencing and metabolomics to investigate the effect of anticancer bioactive peptides combined with oxaliplatin on gastric cancer

<p>背景: 胃癌的发生、发展与肠道菌群密切相关。既往研究发现抗癌生物活性肽(ACBP)与奥沙利铂(OXA)联合对胃癌有显着的治疗作用,但ACBP-OXA对肠道菌群的影响仍不清楚。</p><p><strong>Methods:</strong> We established a nude mouse model of ACBP-OXA combined therapy for gastric cancer, the diversity of gut microbiota and fecal metabolomics were studied, and the correlation between gut microbiota and metabolites was analyzed.</p><p><strong>Results:&nbsp;</strong>ACBP-OXA联合疗法对肠道菌群具有很强的调节作用。16s rRNA研究发现,在门中,ACBP-OXA处理后,厚壁菌门和拟杆菌门的相对丰度发生显着变化,厚壁菌门的相对丰度下降,拟杆菌门的相对丰度增加。属内,ACBP-OXA组中毛螺菌科NK4AB6组的相对丰度降低,odpribacter和拟杆菌属的相对丰度增加。ACBP组乳酸菌相对丰度增加,ACBP-OXA和OXA组葡萄球菌相对丰度下降。GO和KEGG研究发现联合治疗机制与代谢和免疫有关。通过代谢组学研究,本研究发现差异代谢物与Benzenoids、Ligans、neoligans、其中脂质和脂类大多参与酪氨酸代谢、不饱和脂肪酸生物合成、苯丙氨酸代谢α-生物过程。将代谢组学与16s rRNA长寿素相结合,发现氨基酸相关代谢物与Jetgalilicus、Staphylococcus、Proteiniphilum等细菌属相关。</p><p>结论: &nbsp; ACBP与ACBP-OXA联合治疗可能通过改变肠道菌群的分布多样性和菌群结构来改善和恢复胃癌裸鼠的肠道菌群,这可能是抑制胃癌发生、发展的关键。该研究为进一步研究ACBP-OXA在胃癌治疗中的应用提供了新的方向。</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"

<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used&nbsp; to calculate the embeddings.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Figure 4 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 4: Isolate UE-5 (Aspergillus terreus). (A) Surface of colony, on potato dextrose agar after 6 days culture at 28 °C. (B) Reverse of colony. (C) Mycelia, conidiophores and conidia after 5 days culture. (D) Conidiophore with conidia. (E) Phylogenetic tree inferred from internal transcribed spacer sequences using maximum likelihood method.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 3 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 3: Isolates UE-3 (Curvularia sp., A–C) and UE-4 (Curvularia moringae, D–H). (A) Surface of colony, on potato dextrose agar (PDA) after 6 days culture at 28 °C. (B) Reverse of colony. (C) Mycelia and conidia after 7 days culture. (D) Surface of colony, on PDA after 12 days culture at 28 °C. (E) Reverse of colony. (F) Mycelia and conidia after 7 days culture. (G) Conidium. (H) Phylogenetic tree inferred from internal transcribed spacer sequences using maximum likelihood method.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 6 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 6: Antimicrobial activity of the crude extracts obtained from marine endophytic fungi associated with Ulva sp. against five bacteria (Staphylococcus aureus, Bacillus megaterium, Escherichia coli, Salmonella typhi, Pseudomonas aeruginosa) and one fungus (Aspergillus flavus). Values are mean ± standard deviation, n = 3. Bars with different letters are significantly different according to Tukey's post hoc test at p = 0.05. Note: The solvent control (dichloromethane) showed no inhibition (0 mm). The strongest inhibitory effects were observed with the positive controls kanamycin (S1) and ketoconazole (S2).

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 2 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 2: Isolate UE-2 (Nigrospora magnoliae). (A) Surface of colony, on potato dextrose agar after 6 days culture at 28 °C. (B) Reverse of colony. (C) Mycelia and conidia after 21 days culture. (D) Conidiogenus cells with conidia. (E) Phylogenetic tree inferred from internal transcribed spacer sequences using maximum likelihood method.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 5 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 5: Isolate UE-6 (Collariella sp.). (A) Surface of colony, on potato dextrose agar after 12 days culture at 28 °C. (B) Reverse of colony. (C) Terminal ascomatal hairs with ascospores after 45 days culture. (D) Phylogenetic tree inferred from internal transcribed spacer sequences using maximum likelihood method.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 1 in Bioactivity and chemical screening of endophytic fungi associated with the seaweed Ulva sp. of the Bay of Bengal, Bangladesh

Figure 1: Isolate UE-1 (Chaetomium globosum). (A) Surface of colony, on potato dextrose agar after 6 days culture at 28 °C. (B) Reverse of colony. (C) Ascomata after 28 days culture. (D) Asci.(E) Ascus with ascospores. (F) Ascospores. (G) Phylogenetic tree inferred from internal transcribed spacer sequences using maximum likelihood method.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 1 in Bioactivity of aqueous extract of Jacaranda spp. (Bignoniaceae) on Plutella xylostella L. 1758 (Lepidoptera: Plutellidae)

Figure 1. Food preference index effect of aqueous extracts of Jacaranda decurrens and J. mimosifolia at 10% on P. xylostella. Table 1. Larval duration (days), larval survival (%) and egg survival (%) of Plutella xylostella fed with aqueous extract of Jacaranda spp.

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record