Skip to main content
zenodoopen

Transitive prediction of small molecule function through alignment of high-content screening resources

<p>This dataset supports the development of CLIP&lt;sup&gt;n&lt;/sup&gt;, a contrastive-learning framework designed to align heterogeneous high-content screening (HCS) profile datasets.</p> <p><strong>GitHub link</strong>: https://github.com/AltschulerWu-Lab/CLIPn</p> <h2>Directory Structure</h2> <h3>Data Files</h3> <ul> <li>HCS_datasets.pkl: Contains 13 high-content screening (HCS) datasets from multiple studies across 20 years.</li> <li>Hypoxia.pkl: Contains 8 profile datasets using different assays and treated under diverse hypoxia durations.</li> <li>Expression.pkl: Contains 2 transcriptional profile datasets and 6 image profile datasets for multimodal analysis.</li> </ul> <h3>Folders</h3> <p><strong>raw_profiles</strong>:<br>HCS13/<br>- Contains raw data from 13 high-content screening (HCS) datasets. Each dataset includes meta and feature files.&nbsp;</p> <p>L1000/<br>- CDRP_feature_exp.csv: Raw L1000 expression data from the CDRP dataset.<br>- CDRP_meta_exp.csv: Metadata associated with the CDRP expression data.<br>- LINCS_feature_exp.csv: Raw L1000 expression data from the LINCS dataset.<br>- LINCS_meta_exp.csv: Metadata associated with the LINCS expression data.</p> <p>RxRx3/<br>- RxRx3_feature_final.csv: Profile data from the RxRx3 dataset.<br>- RxRx3_meta_final.csv: Metadata from the RxRx3 dataset.</p> <p>Uncharacterized_compounds/<br>- NCI_cpnData.csv: Feature data for uncharacterized compounds from the NCI dataset.<br>- NCI_cpnInfo.csv: Information about uncharacterized compounds in the NCI dataset.<br>- Prestwick_UTSW_cpnData.csv: Feature data for uncharacterized compounds from the Prestwick UTSW dataset.<br>- Prestwick_UTSW_cpnInfo.csv: Information about uncharacterized compounds from the Prestwick UTSW dataset.</p> <h2><br>Usage</h2> <p><br><br><code>import pickle</code><br><code>with open('data.pkl', 'rb') as f:</code><br><code>&nbsp; &nbsp; data = pickle.load(f)</code></p> <p><code>X = data['X']</code><br><code>y = data['y']</code><br><br></p> <h2>Data Reference</h2> <p><br>For raw datasets from 13 HCS database, data and analysis pipeline for dataset 1 was obtained from https://www.science.org/doi/suppl/10.1126/science.1100709/suppl_file/perlman.som.zip; for datasets 2-3, data were shared by authors; For datasets 4-5, analysis code was downloaded from https://static-content.springer.com/esm/art%3A10.1038%2Fnbt.3419/MediaObjects/41587_2016_BFnbt3419_MOESM21_ESM.zip and data were shared by authors; For datasets 6-7, processed dataset was downloaded from AWS following instructions from https://github.com/carpenter-singh-lab/2022_Haghighi_NatureMethods, and replicate_level_cp_normalized.csv.gz features were used. For project datasets 8-13, datasets and analysis results were downloaded from https://zenodo.org/records/7352487. For RxRx3, dataset was obtained from https://www.rxrx.ai/rxrx3. L1000 transcript datasets were downloaded using the same link as datasets 6-7 and the processed transcript data files (named &ldquo;replicate_level_l1k.csv&rdquo;) were used.&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4