Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “high-content screening”
Transitive prediction of small molecule function through alignment of high-content screening resources
<p>This dataset supports the development of CLIP<sup>n</sup>, a contrastive-learning framework designed to align heterogeneous high-content screening (HCS) profile datasets.</p> <p><strong>GitHub link</strong>: https://github.com/AltschulerWu-Lab/CLIPn</p> <h2>Directory Structure</h2> <h3>Data Files</h3> <ul> <li>HCS_datasets.pkl: Contains 13 high-content screening (HCS) datasets from multiple studies across 20 years.</li> <li>Hypoxia.pkl: Contains 8 profile datasets using different assays and treated under diverse hypoxia durations.</li> <li>Expression.pkl: Contains 2 transcriptional profile datasets and 6 image profile datasets for multimodal analysis.</li> </ul> <h3>Folders</h3> <p><strong>raw_profiles</strong>:<br>HCS13/<br>- Contains raw data from 13 high-content screening (HCS) datasets. Each dataset includes meta and feature files. </p> <p>L1000/<br>- CDRP_feature_exp.csv: Raw L1000 expression data from the CDRP dataset.<br>- CDRP_meta_exp.csv: Metadata associated with the CDRP expression data.<br>- LINCS_feature_exp.csv: Raw L1000 expression data from the LINCS dataset.<br>- LINCS_meta_exp.csv: Metadata associated with the LINCS expression data.</p> <p>RxRx3/<br>- RxRx3_feature_final.csv: Profile data from the RxRx3 dataset.<br>- RxRx3_meta_final.csv: Metadata from the RxRx3 dataset.</p> <p>Uncharacterized_compounds/<br>- NCI_cpnData.csv: Feature data for uncharacterized compounds from the NCI dataset.<br>- NCI_cpnInfo.csv: Information about uncharacterized compounds in the NCI dataset.<br>- Prestwick_UTSW_cpnData.csv: Feature data for uncharacterized compounds from the Prestwick UTSW dataset.<br>- Prestwick_UTSW_cpnInfo.csv: Information about uncharacterized compounds from the Prestwick UTSW dataset.</p> <h2><br>Usage</h2> <p><br><br><code>import pickle</code><br><code>with open('data.pkl', 'rb') as f:</code><br><code> data = pickle.load(f)</code></p> <p><code>X = data['X']</code><br><code>y = data['y']</code><br><br></p> <h2>Data Reference</h2> <p><br>For raw datasets from 13 HCS database, data and analysis pipeline for dataset 1 was obtained from https://www.science.org/doi/suppl/10.1126/science.1100709/suppl_file/perlman.som.zip; for datasets 2-3, data were shared by authors; For datasets 4-5, analysis code was downloaded from https://static-content.springer.com/esm/art%3A10.1038%2Fnbt.3419/MediaObjects/41587_2016_BFnbt3419_MOESM21_ESM.zip and data were shared by authors; For datasets 6-7, processed dataset was downloaded from AWS following instructions from https://github.com/carpenter-singh-lab/2022_Haghighi_NatureMethods, and replicate_level_cp_normalized.csv.gz features were used. For project datasets 8-13, datasets and analysis results were downloaded from https://zenodo.org/records/7352487. For RxRx3, dataset was obtained from https://www.rxrx.ai/rxrx3. L1000 transcript datasets were downloaded using the same link as datasets 6-7 and the processed transcript data files (named “replicate_level_l1k.csv”) were used. </p>
Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry_01
<p>Pooled CRISPR screening of high-content cellular phenotypes by ghost cytometry (NFkB NGS data, conventional flow cytometer data, Amnis data)</p>
Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry_02
<p>Pooled CRISPR Screening of High-content Cellular Phenotypes by Ghost Cytometry (Macrophage NGS data)</p>
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [ChIPseq]
GEO Series GSE210491. Homo sapiens. 24 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A kinome-wide high-content siRNA screen identifies MEK5-ERK5 signaling as critical for breast cancer cell EMT and metastasis
GEO Series GSE100403. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.
High-content screen identifies drugs that restrict tumor cell extravasation across the endothelial barrier
GEO Series GSE160915. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [Pilot scRNA-seq]
GEO Series GSE212396. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [RNAseq]
GEO Series GSE210522. Homo sapiens. 52 samples. Type: Expression profiling by high throughput sequencing.
An in vitro high-content screening unveils miR-429 as a protective molecule in photoreceptor degeneration models
GEO Series GSE272674. Mus musculus. 18 samples. Type: Expression profiling by high throughput sequencing.
Deep learning detects cardiotoxicity in a high-content screen with induced pluripotent stem cell-derived cardiomyocytes
GEO Series GSE172181. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.
Integrated time-series analysis and high-content CRISPR screening delineate the dynamics of macrophage immune regulation [CROP-seq KO150]
GEO Series GSE263761. Mus musculus. 27 samples. Type: Expression profiling by high throughput sequencing; Other.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [ATACseq]
GEO Series GSE210489. Homo sapiens. 30 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Integrated time-series analysis and high-content CRISPR screening delineate the dynamics of macrophage immune regulation [CROP-seq KO15]
GEO Series GSE263760. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing; Other.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs
GEO Series GSE210523. Homo sapiens. 185 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing; Other.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [Bulk RNA-seq]
GEO Series GSE232400. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.
High-Content Imaging-Based Pooled CRISPR Screens in Mammalian Cells
GEO Series GSE156623. Homo sapiens. 26 samples. Type: Other.
High-content, targeted RNA-seq screening in organoids for drug discovery in colorectal cancer
GEO Series GSE157167. Mus musculus. 1249 samples. Type: Expression profiling by high throughput sequencing.
A high-content arrayed CRISPR screen reveals genetic requirements for dynein-based trafficking
GEO Series GSE218249. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.
High-content CRISPR screens link coronary artery disease genes to endothelial cell programs [scRNAseq]
GEO Series GSE210681. Homo sapiens. 40 samples. Type: Expression profiling by high throughput sequencing; Other.
High-Content Live-Cell Multiplex Screen for Chemogenomic Compound Annotation based on Nuclear Morphology - Supplemental Material
<p>High-Content Live-Cell Multiplex Screen for Chemogenomic Compound Annotation based on Nuclear Morphology</p> <p>Supplementary Material of linked STAR-protocol</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.