Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

25

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

25 results for “high-content screening”

Learn how ShareScore rates datasets ↗
zenodo36/100

Transitive prediction of small molecule function through alignment of high-content screening resources

<p>This dataset supports the development of CLIP&lt;sup&gt;n&lt;/sup&gt;, a contrastive-learning framework designed to align heterogeneous high-content screening (HCS) profile datasets.</p> <p><strong>GitHub link</strong>: https://github.com/AltschulerWu-Lab/CLIPn</p> <h2>Directory Structure</h2> <h3>Data Files</h3> <ul> <li>HCS_datasets.pkl: Contains 13 high-content screening (HCS) datasets from multiple studies across 20 years.</li> <li>Hypoxia.pkl: Contains 8 profile datasets using different assays and treated under diverse hypoxia durations.</li> <li>Expression.pkl: Contains 2 transcriptional profile datasets and 6 image profile datasets for multimodal analysis.</li> </ul> <h3>Folders</h3> <p><strong>raw_profiles</strong>:<br>HCS13/<br>- Contains raw data from 13 high-content screening (HCS) datasets. Each dataset includes meta and feature files.&nbsp;</p> <p>L1000/<br>- CDRP_feature_exp.csv: Raw L1000 expression data from the CDRP dataset.<br>- CDRP_meta_exp.csv: Metadata associated with the CDRP expression data.<br>- LINCS_feature_exp.csv: Raw L1000 expression data from the LINCS dataset.<br>- LINCS_meta_exp.csv: Metadata associated with the LINCS expression data.</p> <p>RxRx3/<br>- RxRx3_feature_final.csv: Profile data from the RxRx3 dataset.<br>- RxRx3_meta_final.csv: Metadata from the RxRx3 dataset.</p> <p>Uncharacterized_compounds/<br>- NCI_cpnData.csv: Feature data for uncharacterized compounds from the NCI dataset.<br>- NCI_cpnInfo.csv: Information about uncharacterized compounds in the NCI dataset.<br>- Prestwick_UTSW_cpnData.csv: Feature data for uncharacterized compounds from the Prestwick UTSW dataset.<br>- Prestwick_UTSW_cpnInfo.csv: Information about uncharacterized compounds from the Prestwick UTSW dataset.</p> <h2><br>Usage</h2> <p><br><br><code>import pickle</code><br><code>with open('data.pkl', 'rb') as f:</code><br><code>&nbsp; &nbsp; data = pickle.load(f)</code></p> <p><code>X = data['X']</code><br><code>y = data['y']</code><br><br></p> <h2>Data Reference</h2> <p><br>For raw datasets from 13 HCS database, data and analysis pipeline for dataset 1 was obtained from https://www.science.org/doi/suppl/10.1126/science.1100709/suppl_file/perlman.som.zip; for datasets 2-3, data were shared by authors; For datasets 4-5, analysis code was downloaded from https://static-content.springer.com/esm/art%3A10.1038%2Fnbt.3419/MediaObjects/41587_2016_BFnbt3419_MOESM21_ESM.zip and data were shared by authors; For datasets 6-7, processed dataset was downloaded from AWS following instructions from https://github.com/carpenter-singh-lab/2022_Haghighi_NatureMethods, and replicate_level_cp_normalized.csv.gz features were used. For project datasets 8-13, datasets and analysis results were downloaded from https://zenodo.org/records/7352487. For RxRx3, dataset was obtained from https://www.rxrx.ai/rxrx3. L1000 transcript datasets were downloaded using the same link as datasets 6-7 and the processed transcript data files (named &ldquo;replicate_level_l1k.csv&rdquo;) were used.&nbsp;</p>

openmit-licenseOct 2024View details →
zenodo28/100

Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry_01

<p>Pooled CRISPR screening of high-content cellular phenotypes by ghost cytometry (NFkB NGS data, conventional flow cytometer data, Amnis data)</p>

openApr 2023View details →
zenodo28/100

Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry_02

<p>Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry (Macrophage NGS data)</p>

openApr 2023View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [ChIPseq]

GEO Series GSE210491. Homo sapiens. 24 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

A kinome-wide high-content siRNA screen identifies MEK5-ERK5 signaling as critical for breast cancer cell EMT and metastasis

GEO Series GSE100403. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2018View details →
geo24/100

High-content screen identifies drugs that restrict tumor cell extravasation across the endothelial barrier

GEO Series GSE160915. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2020View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [Pilot scRNA-seq]

GEO Series GSE212396. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [RNAseq]

GEO Series GSE210522. Homo sapiens. 52 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

An in vitro high-content screening unveils miR-429 as a protective molecule in photoreceptor degeneration models

GEO Series GSE272674. Mus musculus. 18 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2025View details →
geo24/100

Deep learning detects cardiotoxicity in a high-content screen with induced pluripotent stem cell-derived cardiomyocytes

GEO Series GSE172181. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2021View details →
geo24/100

Integrated time-series analysis and high-content CRISPR screening delineate the dynamics of macrophage immune regulation [CROP-seq KO150]

GEO Series GSE263761. Mus musculus. 27 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenJul 2025View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [ATACseq]

GEO Series GSE210489. Homo sapiens. 30 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

Integrated time-series analysis and high-content CRISPR screening delineate the dynamics of macrophage immune regulation [CROP-seq KO15]

GEO Series GSE263760. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenJul 2025View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs

GEO Series GSE210523. Homo sapiens. 185 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing; Other.

openGEO-OpenNov 2022View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [Bulk RNA-seq]

GEO Series GSE232400. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2023View details →
geo24/100

High-Content Imaging-Based Pooled CRISPR Screens in Mammalian Cells

GEO Series GSE156623. Homo sapiens. 26 samples. Type: Other.

openGEO-OpenJan 2021View details →
geo24/100

High-content, targeted RNA-seq screening in organoids for drug discovery in colorectal cancer

GEO Series GSE157167. Mus musculus. 1249 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2021View details →
geo24/100

A high-content arrayed CRISPR screen reveals genetic requirements for dynein-based trafficking

GEO Series GSE218249. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2024View details →
geo24/100

High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [scRNAseq]

GEO Series GSE210681. Homo sapiens. 40 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenNov 2022View details →
zenodo24/100

High-Content Live-Cell Multiplex Screen for Chemogenomic Compound Annotation based on Nuclear Morphology - Supplemental Material

<p>High-Content Live-Cell Multiplex Screen for Chemogenomic Compound Annotation based on Nuclear Morphology</p> <p>Supplementary Material of linked STAR-protocol</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record