Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “cell painting”
Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.
<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">"Predicting compound activity from phenotypic profiles and chemical structures"</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper's GitHub repository</a> for reproduction.</p> <p>Folders and files and are described below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p> </p>
Cell Painting Dataset of EUbOPEN compounds
<p>This parquet file is a dataset of cell painting profiles generated using a technique of morphological profiling known as the Cell Painting Assay (CPA). The CPA aims at deeply characterizing modulator from their ‘fingerprint’ in human cells as seen in microscopic images.</p> <p>Cells subjected to treatments are fixed, stained and imaged. These images are analysed and more than 7500 features are extracted using a cell profiler pipeline. Single cell level profiles are aggregated into well-level profiles which can be used to compare and contrast chemical treatments. </p> <p>Morphological profiling using cell painting begins with the culture of bone sarcoma U2OS cells in control and perturbation conditions for 24 to 48 hours. These cells are then subjected to 4% PFA fixation and staining with 6 inorganic dyes (Hoechst 33342 Nuclear Stain, Concanavalin A 488, Nucleic Acid Stain, WGA 555, Phalloidin 568, Mitotracker Mitochondrial stain 641) in order to mark 8 cellular components (nucleus, endoplasmic reticulum, mitochondria, actin, Golgi, plasma membrane, nucleolus and RNA). Images at 20x magnification are then captured using a high content microscope Yokogawa Cell Voyager CV8000. These images go through a rigorous quality control process are then passed through the profiling pipeline using the open-source software Cell Profiler. Thousands of features are extracted from each image and value of each feature is computed so as to detect even the slightest phenotypic variation occurring in cells in an extremely sensitive manner.</p> <p>This dataset consists of profiles of 595 compounds from the EUbOPEN consortium (cell painting first batch) in 106 plates in the 384 well format. DMSO and positive controls are present in columns 1,2, 23 and 24. There are 39206 rows, each row corresponding to one well, and 7721 columns. The treatments have been done in a range of 8 doses and the highest dose treatment has been done at 2 time points. </p> <p> </p>
AI4Life-MDC24 Challenge data: JUMP Cell Painting Datasets
<p>This is a subset of <em>the cpg0016-jump</em> dataset (<em>Chandrasekaran et al., 2022), available from the Cell Painting Gallery on the Registry of Open Data on AWS (<a href="https://registry.opendata.aws/cellpainting-gallery/" rel="nofollow">https://registry.opendata.aws/cellpainting-gallery/</a>). <br><br></em>The selected <strong>subset</strong> contains 517 images with four channels in the form of a single multi-channel tiff file. Gaussian noise is applied to each image to simulate the detector noise.</p> <p>The GitHub repository with the details about the original dataset is available at <a href="https://github.com/jump-cellpainting/datasets">https://github.com/jump-cellpainting/datasets</a>. The preprint describing the original dataset is <a href="https://www.biorxiv.org/content/10.1101/2023.03.23.534023v1">available on BioRxiv</a><br><br></p> <p>AI4Life has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement number 101057970. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p> <p> </p>
Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"
<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used to calculate the embeddings.</p>
broadinstitute/lincs-cell-painting: Full release of LINCS Cell Painting dataset
<p>Summary</p> <p>This release contains finalized profile and metadata information for the CellProfiler-derived features from the LINCS Cell Painting dataset.</p> <p>We assayed two batches of data and applied an image analysis pipeline using CellProfiler to segment cells and extract morphology features. Using the CellProfiler output, we applied an image-based profiling pipeline to process the morphology profiles into an analysis-ready form.</p> <p>In this repository, we provide the full image-based profiling pipeline we used to process the data, as well as most of the intermediate data.</p> <p>Data counts</p> <ul> <li>Two batches ("2016_04_01_a549_48hr_batch1" and "2017_12_05_Batch2")</li> <li>159,717,488 unique cells total</li> </ul> <p>2016_04_01_a549_48hr_batch1</p> <ul> <li>52,223 profiles (batch corrected)</li> <li>10,752 unique perturbations</li> <li>1,571 unique compounds</li> <li>1 time point (48 hour)</li> <li>7 doses</li> <li>1 cell line (A549)</li> <li>110,012,425 unique cells</li> </ul> <p>2017_12_05_Batch2</p> <ul> <li>51,447 profiles (batch corrected)</li> <li>10,368 unique perturbations</li> <li>349 unique compounds</li> <li>3 time points (6 hour, 24 hour, 48 hour)</li> <li>6 doses</li> <li>3 cell lines (A549, MCF7, U2OS)</li> <li>49,705,063 unique cells</li> </ul>
Unbiased single-cell morphology with self-supervised vision transformers -- Cell Painting
<p>The data necessary to reproduce the Cell Painting results in the paper <a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>. </p>
An Interpretable Framework to Characterize Compound Treatments on Filamentous Fungi using Cell Painting and Deep Metric Learning
<p>This deposit contains:</p> <ul> <li>{train,test,val}.csv: Meta-data files</li> <li>images.tar.gz: Archive of images</li> <li>checkpoint.pth.tar: Model weights</li> <li>code.tar.gz: Code necessary to reproduce our experiments.</li> </ul>
Cell Painting Botryris
Open the record for dataset details and reuse information.
Cell Painting Data of Botrytis cinerea
Open the record for dataset details and reuse information.
CellDeathPred: a deep learning framework for ferroptosis and apoptosis prediction based on cell painting
<p><strong>Data Description:</strong></p> <ul> <li>Two cell death modalities are considered in this study: apoptosis and ferroptosis. The treatments and their cell death types used in these experiments are shown in Figure 1 in the paper. In total we have seven apoptosis and seven ferroptosis inducers. They target specific cell organelles and components. Besides, active treatments there are negative controls DMSO and untreated wells.</li> <li>Every image has a resolution of 1360 x 1024. These four channels were imaged in the experiments. For every well multiple images (fields) are taken typically ten fields.</li> <li>Three plates with four different settings were imaged. 20x non-confocal (very fast [1.5h]), 20x confocal (fast [2.5h]), 40x non-confocal (slow [5h]), 40x confocal (very slow [8h]). All for one time point 28h (incubation time).</li> </ul>
Assessing the Effects of Silver Nanoparticles on ARPE-19 Cells via High-Throughput Phenotypic Profiling with the Cell Painting Assay.
GEO Series GSE297342. Homo sapiens. 88 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.