Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14,866
datasets available to search
ShareScore release 0.9.0
Dataset results
14,866 results for “cancer cell”
pix2pix_HUVEC_nuclei_cancer_cells_dataset
<p>This repository contains a Pix2Pix deep learning model to generate synthetic nuclear staining from brightfield images. The model was trained on 258 paired brightfield and fluorescent microscopy images of circulating cancer cells perfused over an endothelial cell monolayer. To improve performance, the dataset was augmented computationally by a factor of 8. The model was trained over 400 epochs using a patch size of 512x512, a batch size of 1, and a vanilla GAN loss function. The final model was selected based on its performance metrics and visual fidelity when compared to ground truth images, achieving an average SSIM score of 0.755 and an LPIPS score of 0.120.</p> <h3>Specifications</h3> <ul> <li> <p>Model: Pix2Pix for generating synthetic nuclear staining from brightfield images</p> </li> <li> <p>Training Dataset:</p> </li> <ul> <li> <p>Cancer Cells: 258 paired brightfield and fluorescent microscopy images</p> </li> <li> <p>Microscope: Nikon Eclipse Ti2-E, brightfield/fluorescence microscope with a 20x objective</p> </li> <li> <p>Data Type: Brightfield and fluorescent microscopy images</p> </li> <li> <p>File Format: TIFF (.tif), 16-bit</p> </li> <li> <p>Image Size: 1024 x 1022 pixels (Pixel size: 650 nm)</p> </li> </ul> <li> <p>Training Parameters:</p> </li> <ul> <li> <p>Epochs: 400</p> </li> <li> <p>Patch Size: 512 x 512 pixels</p> </li> <li> <p>Batch Size: 1</p> </li> <li> <p>Loss Function: Vanilla GAN loss function</p> </li> </ul> <li> <p>Model Performance:</p> </li> <ul> <li> <p>Circulating Cancer Cells:</p> </li> <ul> <li> <p>SSIM Score: 0.755</p> </li> <li> <p>LPIPS Score: 0.120</p> </li> </ul> </ul> <li> <p>Model Selection: Models were selected based on quality metric scores and visual inspection compared to ground truth images.</p> </li> <li> <p>Model Training: Conducted using ZeroCostDL4Mic (<a href="https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki/pix2pix">https://github.com/HenriquesLab/ZeroCostDL4Mic/wiki</a>)</p> </li> </ul> <div> <h3>Reference</h3> <div><strong>Fast label-free live imaging reveals key roles of flow dynamics and CD44-HA interaction in cancer cell arrest on endothelial monolayers</strong></div> </div> <div>Gautier Follain, Sujan Ghimire, Joanna W. Pylvänäinen, Monika Vaitkevičiūtė, Diana Wurzinger, Camilo Guzmán, James RW Conway, Michal Dibus, Sanna Oikari, Kirsi Rilla, Marko Salmi, Johanna Ivaska, Guillaume Jacquemet</div> <div>bioRxiv 2024.09.30.615654; doi: <a href="https://www.biorxiv.org/content/10.1101/2024.09.30.615654v1">https://doi.org/10.1101/2024.09.30.615654</a></div> <p> </p>
A stratification system for breast cancer based on basoluminal tumor cells and spatial tumor architecture (IF/mIF data)
<p>This repository contains all <strong>raw whole-slide immunofluorescence (IF) data</strong> for the breast cancer study from Meyer et al., 2025. The code that was used to process and analyze the data is available at <a href="https://github.com/BodenmillerGroup/TNBC_publication">https://github.com/BodenmillerGroup/TNBC_publication</a>. </p> <p><strong>Structure:</strong><br>BasoLum.zip - Contains whole-slide IF data for CK5, CK7, CK19 stainings of selected TNBC patients (related to Figure 4)</p> <p><strong>NOTE:</strong> Raw <strong>multiplexed whole-slide immunofluorescence (mIF) data</strong> (related to Figure 5) for this study is available from the corresponding author upon reasonable request and has not been uploaded to Zenodo due to the large data size (~ 50 GB per image). </p>
Source data sets_Figure 1, 2, 6_Salmonella cancer therapy metabolically disrupts tumours at the collateral cost of T cell immunity
<p>Flow cytometry data files associated with Copland <em>et al., </em><strong><em><span>Salmonella </span></em></strong><strong><span>cancer therapy metabolically disrupts tumours at the collateral cost of T cell immunity.</span></strong></p> <p><span>Data associated to Figures 1, 2 and 6. <br></span></p>
Single-cell RNA sequencing reveals immunosuppressive pathways associated with metastatic breast cancer
Open the record for dataset details and reuse information.
Utilizing DNA pooling to predict cancer eye, ocular squamous cell carcinoma, in Hereford cattle
<p>Files include genotypes and dye intensity data from BovineHD bead array from a small closed population of Hereford cattle. 42 animals are individually genotyped and 10 pools of 50 animals are genotyped. Random regression of individual dye intensity data were regressed on pool data to estimate genetic connections between animals and pools. Covariances among animals for pool contributions were also estimated.</p>
Pan-Cancer T cell atlas from "The combined use of scRNA-seq and network propagation highlights key features of pan-cancer Tumor-Infiltrating T cells" (https://doi.org/10.1371/journal.pone.0315980)
<p>The scRNA-seq data were collected from previously published datasets (GSE140228, GSE139555, GSE155698, GSE121636, and GSE139324), adhering to the following selection criteria: 1) presence of T cells, 2) treatment-naïve patients, 3) solid tumors, and 4) inclusion of at least tumor and blood samples.<br>Each scRNA-seq dataset underwent separate preprocessing in R (v4.0.2). We filtered out cells from the original count matrices that had fewer than 200 genes detected or more than 10% mitochondrial UMI counts and we only kept genes detected in at least 3 cells. Then, we applied Seurat (v4.0.5) with default parameters for count data normalization and scaling. Each cell was assigned a cell cycle score using the CellCycleScoring function and we computed the difference between the G2M and S phase scores. This approach allows for the separation of non-cycling from cycling cells while minimizing the differences in cell cycle phase among proliferating cells. The SelectIntegrationFeatures function was ran with the nfeatures parameter set to 3,000 before merging all samples from each dataset. These integration features were then used for Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP). Clustering was performed using the Louvain algorithm with the resolution parameter set to 2.0 for all datasets. Finally, T cells were isolated based on CD3D and CD3G genes expression (CD3D or CD3G expression level > 0).</p> <p>To integrate heterogeneous data from different sources, a two-step procedure was applied. We first concatenated all datasets together and ran the scaling and PCA steps based on the top 3,000 highly variable genes identified by the FindVariableFeatures function with the “vst” method. Harmony was applied for batch effect correction then UMAP and clustering using the Louvain algorithm with the resolution parameter set to 2.0 were performed on the harmony reduction. Examining the result from the first clustering run, we identified contamination clusters and clusters that arose from unwanted factors: we removed the contamination clusters including low quality cells highly expressing marker genes associated with apoptosis and tissue dissociation operation, pancreatic acinar cells (expressing PRSS1, CLPS, PNLIP and CTRB1 among others), myeloid cells (expressing CD68) and B cells (expressing CD79A). Then, we performed the second run of integration and clustering excluding immunoglobulin, ribosome-protein-coding, and T cell receptor (TCR) genes (gene symbol with string pattern "^IGK|^IGH|^IGL|^IGJ|^IGS|^IGD|IGFN1", "^RP([0–9]+-|[LS])", and "^TRA|^TRB|^TRG" respectively) from the top 3,000 highly variable genes and regressing out the cell cycle difference effect as well as the percentage of mitochondrial UMI counts. Harmony (v0.1.0) was applied again for batch effect correction and UMAP was performed on the harmony reduction.<br>T cell subtypes identification and annotation was performed by clustering cells using the Louvain algorithm with the resolution parameter set to 4.1 after iterative testing from 3.5 to 5.0 by 0.1 (more granular than default), computing clusters signatures based on differential gene expression using the FindAllMarkers function with the “MAST” method and interrogating known gene markers expression. A resolution value of 4.1 was notably found to be the lowest resolution value enabling the correct separation of proliferating CD4+ T cells from proliferating CD8+ T cells.</p>
scRNA-seq dataset from "Glucose deprivation and identification of TXNIP as an immunometabolic modulator of T cell activation in cancer"
<p>Gene expression profiling analysis of single cell RNA-seq data from MLR, anti-CD3/anti-CD28 treated and paired untreated CD4+ T cells samples under high glucose (11 mM) and low glucose (1 mM) conditions.</p> <p>For each sample, cells suspensions in culture medium were recovered, washed once with 0.04% BSA in 1X PBS and processed through 10x Cell Multiplexing Oligo Labeling protocol (10x Genomics, USA) according to the manufacturer’s instructions. ~1,600 cells/µl pooled cell suspensions were prepared with equal number of cells per sample: one for MLR samples, one for non-stimulated T cells samples, and one for anti-CD3/anti-CD28-stimulated T cells samples. Libraries were prepared using the 10x Chromium Single-Cell 3’ v3.1 protocol with Feature Barcode (10x Genomics, USA), according to the manufacturer’s instructions. Sequencing was performed on a NovaSeq 6000 sequencer (Illumina, USA).</p> <p>Cell Ranger (v6.0.1, 10x Genomics Inc) was applied for demultiplexing, reads mapping against the GRCh38 human reference genome, and UMI counting. Seurat package (v4.4.0) was used to generate Seurat objects. Only genes detected in at least 3 cells were kept. Cells with fewer than 200 genes detected or >15% mitochondrial UMI counts were filtered out. Samples were merged in a unique Seurat object then count data normalization and scaling was performed using Seurat with default parameters. The 2000 most highly variable genes were used for Principal Component Analysis (PCA). Harmony (v0.1.1) was applied for batch effect correction then Uniform Manifold Approximation and Projection (UMAP) and clustering using the Louvain algorithm were performed on the harmony reduction. Non-T or -MoDC clusters were removed for further analysis.</p>
Dataset for "Luminal breast epithelial cells from wildtype and BRCA mutation carriers harbor copy number alterations commonly associated with breast cancer"
<p>Processed single cell whole genome sequencing data from:</p> <p>Luminal breast epithelial cells from wildtype and BRCA mutation carriers harbor copy number alterations commonly associated with breast cancer Williams, Vinci Oliphant et al 2024</p> <p>Included are processed copy number profiles for all cells included in the study.</p>
Fig. 9 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 9. Binding interactions of active site residues of CAIX (PDB ID 3IAI) with 4-hydroxy-5-methylcoumarin. The ligand is black in colour while the amino acid residues are green in colour. Purple dash represents hydrophobic interactions while green dash depicts hydrogen bonding. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)
Fig. 8 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 8. Binding interactions of active site residues of CAIX (PDB ID 3IAI) with 5-methylcoumarin-4-β-glucoside. The ligand is black in colour while the amino acid residues are green in colour. Purple dash represents hydrophobic interactions while green dash depicts hydrogen bonding. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)
Fig. 5 in Cytotoxicity of methanolic extract of Swertia petiolata against gastric cancer cell line SNU-5 is via induction of apoptosis ⁎
Fig. 5. Mitochondrial membrane potential (MMP %) of SNU-5 cells treated with indicated concentration of methanolic extract of plant for 48 h. The S. petiolata methanolic extracttreated SNU5 cells showed induction of MMP loss in a dose dependent manner.
Fig. 2. SNU5 in Cytotoxicity of methanolic extract of Swertia petiolata against gastric cancer cell line SNU-5 is via induction of apoptosis ⁎
Fig. 2. SNU5 cells were treated with indicated concentrations of methanolic extract of S. petiolata for 48 h and cells were prolonged for clonogenic assay for 21 days. The extract-treated SNU5 cells showed decrease in the percentage of colony formation as compared to control in a dose-dependent manner.
Fig. 1. SNU5 in Cytotoxicity of methanolic extract of Swertia petiolata against gastric cancer cell line SNU-5 is via induction of apoptosis ⁎
Fig. 1. SNU5 cells were treated with indicated concentration of methanolic extract of S. petiolata for 48 h. The extract-treated SNU5 cells showed nuclear condensation and marked fragmented in a dose dependent manner.
Fig. 6 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 6. Binding interactions of active site residues of CDK2 (PDB ID: 2DUV) with lupeol (black colour). Purple dash depicts hydrophobic interactions while green dash depicts hydrogen bonding. (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)
Fig. 3 in Cytotoxicity of methanolic extract of Swertia petiolata against gastric cancer cell line SNU-5 is via induction of apoptosis ⁎
Fig. 3. Cells were treated with different concentrations (a) control (b) 50 (c) 100 and (d) 150 μg/ml of extract for 48 h wherein apoptosis enhanced in a dose dependent manner as shown by flow cytometer using Annexin V/PI.
Fig. 4 in Cytotoxicity of methanolic extract of Swertia petiolata against gastric cancer cell line SNU-5 is via induction of apoptosis ⁎
Fig. 4. Production of ROS (%) SNU-5 cells treated with indicated concentration of methanolic extract of plant for 48 h. The S. petiolata methanolic extract-treated SNU5 cells showed the production of reactive oxygen species generated in a dose dependent manner.
Fig. 4 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 4. Effect of V. glaberrima compounds on the viability of A375 cell lines as determined by MTS assay.
Fig. 1 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 1. Photograph of Vernonia glaberrima (personally taken by Alhassan AM in Nasarawa LGA, Nasarawa State, Nigeria on 12, May, 2015).
Fig. 2 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 2. Structures of isolated compounds from Vernonia glaberrima leaves viz., nonacosanoic acid (1), lupeol (2), 5-methylcoumarin-4-β-glucoside (3) and 4-hydroxy-5-methylcoumarin (4).
Fig. 3 in Phytoconstituents from Vernonia glaberrima Welw. Ex O. Hoffm. leaves and their cytotoxic activities on a panel of human cancer cell lines
Fig. 3. Effect of Vernonia glaberrima leaves crude methanolic extract on cancer cells viability as determined by MTS assay.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.