Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “cell painting”

Learn how ShareScore rates datasets ↗
zenodo52/100

Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.

<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">&quot;Predicting compound activity from phenotypic profiles and chemical structures&quot;</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper&#39;s GitHub repository</a>&nbsp;for reproduction.</p> <p>Folders and files&nbsp;and are described&nbsp;below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Cell Painting Dataset of EUbOPEN compounds

<p>This parquet file is a dataset of cell painting profiles generated using a technique of morphological profiling known as the Cell Painting Assay (CPA). The CPA aims at deeply characterizing modulator from their &lsquo;fingerprint&rsquo; in human cells as seen in microscopic images.</p> <p>Cells subjected to treatments are fixed, stained and imaged. These images are analysed and more than 7500 features are extracted using a cell profiler pipeline. Single cell level profiles are aggregated into well-level profiles which can be used to compare and contrast chemical treatments. </p> <p>Morphological profiling using cell painting begins with the culture of bone sarcoma U2OS cells in control and perturbation conditions for 24 to 48 hours. These cells are then subjected to 4% PFA fixation and staining with 6 inorganic dyes (Hoechst 33342 Nuclear Stain, Concanavalin A 488, Nucleic Acid Stain, WGA 555, Phalloidin 568, Mitotracker Mitochondrial stain 641) in order to mark 8 cellular components (nucleus, endoplasmic reticulum, mitochondria, actin, Golgi, plasma membrane, nucleolus and RNA). Images at 20x magnification are then captured using a high content microscope Yokogawa Cell Voyager CV8000. These images go through a rigorous quality control process&nbsp; are then passed through the profiling pipeline using the open-source software Cell Profiler. Thousands of features are extracted from each image and value of each feature is computed so as to detect even the slightest phenotypic variation occurring in cells in an extremely sensitive manner.</p> <p>This dataset consists of profiles of 595 compounds from the EUbOPEN consortium (cell painting first batch) in 106 plates in the 384 well format. DMSO and positive controls are present in columns 1,2, 23 and 24. There are 39206 rows, each row corresponding to one well, and 7721 columns. The treatments have been done in a range of 8 doses and the highest dose treatment has been done at 2 time points.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

AI4Life-MDC24 Challenge data: JUMP Cell Painting Datasets

<p>This is a subset of <em>the cpg0016-jump</em> dataset (<em>Chandrasekaran et al., 2022), available from the Cell Painting Gallery on the Registry of Open Data on AWS (<a href="https://registry.opendata.aws/cellpainting-gallery/" rel="nofollow">https://registry.opendata.aws/cellpainting-gallery/</a>).&nbsp;<br><br></em>The selected <strong>subset</strong> contains 517 images with four channels in the form of a single multi-channel tiff file. Gaussian noise is applied to each image to simulate the detector noise.</p> <p>The GitHub repository with the details about the original dataset is available at <a href="https://github.com/jump-cellpainting/datasets">https://github.com/jump-cellpainting/datasets</a>. The preprint describing the original dataset is&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2023.03.23.534023v1">available on BioRxiv</a><br><br></p> <p>AI4Life has received funding from the European Union&rsquo;s Horizon Europe research and innovation programme under grant agreement number 101057970. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"

<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used&nbsp; to calculate the embeddings.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

broadinstitute/lincs-cell-painting: Full release of LINCS Cell Painting dataset

<p>Summary</p> <p>This release contains finalized profile and metadata information for the CellProfiler-derived features from the LINCS Cell Painting dataset.</p> <p>We assayed two batches of data and applied an image analysis pipeline using CellProfiler to segment cells and extract morphology features. Using the CellProfiler output, we applied an image-based profiling pipeline to process the morphology profiles into an analysis-ready form.</p> <p>In this repository, we provide the full image-based profiling pipeline we used to process the data, as well as most of the intermediate data.</p> <p>Data counts</p> <ul> <li>Two batches (&quot;2016_04_01_a549_48hr_batch1&quot; and &quot;2017_12_05_Batch2&quot;)</li> <li>159,717,488 unique cells total</li> </ul> <p>2016_04_01_a549_48hr_batch1</p> <ul> <li>52,223 profiles (batch corrected)</li> <li>10,752 unique perturbations</li> <li>1,571 unique compounds</li> <li>1 time point (48 hour)</li> <li>7 doses</li> <li>1 cell line (A549)</li> <li>110,012,425 unique cells</li> </ul> <p>2017_12_05_Batch2</p> <ul> <li>51,447 profiles (batch corrected)</li> <li>10,368 unique perturbations</li> <li>349 unique compounds</li> <li>3 time points (6 hour, 24 hour, 48 hour)</li> <li>6 doses</li> <li>3 cell lines (A549, MCF7, U2OS)</li> <li>49,705,063 unique cells</li> </ul>

openother-openJun 2021View details →
zenodo36/100

Unbiased single-cell morphology with self-supervised vision transformers -- Cell Painting

<p>The data necessary to reproduce the Cell Painting results in the paper&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

An Interpretable Framework to Characterize Compound Treatments on Filamentous Fungi using Cell Painting and Deep Metric Learning

<p>This deposit contains:</p> <ul> <li>{train,test,val}.csv:&nbsp; &nbsp; Meta-data files</li> <li>images.tar.gz:&nbsp; &nbsp; &nbsp;Archive of images</li> <li>checkpoint.pth.tar:&nbsp; &nbsp; &nbsp;Model weights</li> <li>code.tar.gz:&nbsp; &nbsp; Code necessary to reproduce our experiments.</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Cell Painting Botryris

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Cell Painting Data of Botrytis cinerea

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo28/100

CellDeathPred: a deep learning framework for ferroptosis and apoptosis prediction based on cell painting

<p><strong>Data Description:</strong></p> <ul> <li>Two cell death modalities are considered in this study: apoptosis and ferroptosis. The treatments and their cell death types used in these experiments are shown in Figure 1 in the paper. In total we have seven apoptosis and seven ferroptosis inducers.&nbsp;They target specific cell organelles and components. Besides, active treatments there are negative controls DMSO and untreated wells.</li> <li>Every image has a resolution of 1360 x 1024. These four channels&nbsp;were imaged in the experiments. For every well multiple images (fields) are taken typically&nbsp;ten fields.</li> <li>Three plates with four different settings were imaged. 20x non-confocal (very fast [1.5h]), 20x confocal (fast [2.5h]), 40x non-confocal (slow [5h]), 40x confocal (very slow [8h]). All for one time point 28h (incubation time).</li> </ul>

opencc-by-4.0Jul 2023View details →
geo24/100

Assessing the Effects of Silver Nanoparticles on ARPE-19 Cells via High-Throughput Phenotypic Profiling with the Cell Painting Assay.

GEO Series GSE297342. Homo sapiens. 88 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record