Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

35 results for “whole slide imaging”

Learn how ShareScore rates datasets ↗
zenodo48/100

Testing whole slide image for OpenPhi - Open Pathology Interface

<p>An anonymous whole slide image in Philips iSyntax format for running software tests on OpenPhi - Open Pathology Interface (https://zenodo.org/record/4680748#.YNnBxDqxXJU). See the repository (https://gitlab.com/BioimageInformaticsGroup/openphi/) for up to date information.</p>

openmit-licenseJun 2021View details →
zenodo48/100

DICOM converted whole slide hematoxylin and eosin images of rhabdomyosarcoma from Children's Oncology Group trials

<p>Rhabdomyosarcoma (RMS) is an aggressive soft-tissue sarcoma, which primarily occurs in children and young adults. This dataset contains manifests referring to the hematoxylin and eosin (H&amp;E) stained images in Digital Imaging and Communications in Medicine (DICOM) format available from National Cancer Institute Imaging Data Commons (IDC) [1] (also see IDC Portal at&nbsp;<a href="https://imaging.datacommons.cancer.gov">https://imaging.datacommons.cancer.gov</a>) as of data release v16. The original images in vendor-specific format were collected on IRB-approved clinical trials or tissue banking studies from Children&rsquo;s Oncology Group (COG) patients enrolled on ARST0331, ARST0431, D9602, D9803, and D9902 trials, as described in [2]. Those images, augmented with the metadata describing their content, were provided to the IDC team for the purposes of archival, and were converted into DICOM Whole Slide Microscopy (SM) representation [3], [4] using custom open source scripts and tools available and described here [5]. The resulting converted images were released in IDC in the RMS-Mutation-Prediction collection with the data release v16.</p> <p>To conveniently explore the data available for this dataset, please use this dashboard: <a href="https://lookerstudio.google.com/reporting/7f267400-8774-42e1-b5d1-ca11863c52a9">https://lookerstudio.google.com/reporting/7f267400-8774-42e1-b5d1-ca11863c52a9</a>.</p> <p>Notebooks demonstrating how to use this data are available here: <a href="https://github.com/ImagingDataCommons/IDC-Tutorials/tree/master/notebooks/collections_demos/rms_mutation_prediction">https://github.com/ImagingDataCommons/IDC-Tutorials/tree/master/notebooks/collections_demos/rms_mutation_prediction</a>.</p> <p>Clinical data accompanying the images is available via SQL interface in IDC BigQuery tables, see details on accessing IDC clinical data in the respective tutorial (<a href="https://github.com/ImagingDataCommons/IDC-Tutorials/blob/master/notebooks/clinical_data_intro.ipynb">https://github.com/ImagingDataCommons/IDC-Tutorials/blob/master/notebooks/clinical_data_intro.ipynb</a>).</p> <p>The images referred to by the accompanying manifests can be explored and visualized using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/">https://portal.imaging.datacommons.cancer.gov/explore/</a>. Direct link to open the collection is <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=rms_mutation_prediction">https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=rms_mutation_prediction</a>.</p> <p>The GCP and AWS manifests provided with this dataset record can be used to download the corresponding files from the IDC Google Cloud Storage (GCS) or Amazon S3 (AWS) buckets free of charge following the instructions available in IDC documentation here: <a href="https://learn.canceridc.dev/data/downloading-data">https://learn.canceridc.dev/data/downloading-data</a>. Specifically, you will need to install the s5cmd command line tool on your computer (see instructions at <a href="https://github.com/peak/s5cmd#installation">https://github.com/peak/s5cmd#installation</a>), and follow the manifest-specific download instructions accompanying the file list below.</p> <p>If you use the files referenced in the attached manifests, we ask you to please cite this dataset, as well as the publication describing the original dataset [2] and the publication acknowledging IDC [1].</p> <p>Specific files included in the record are:</p> <ol> <li> <p><strong><code>rms_mutation_prediction_gcs.s5cmd</code></strong>: GCS-based manifest (to download the files described in the manifest, execute this command: <code>s5cmd --no-sign-request --endpoint-url https://storage.googleapis.com run rms_mutation_prediction_gcs.s5cmd</code>)</p> </li> <li> <p><strong><code>rms_mutation_prediction_aws.s5cmd</code></strong>: AWS-based manifest (to download the files described in the manifest, execute this command: <code>s5cmd --no-sign-request --endpoint-url https://s3.amazonaws.com run rms_mutation_prediction_aws.s5cmd</code>)</p> </li> <li> <p><strong><code>rms_mutation_prediction_dcf.csv</code></strong>: Gen3-based manifest (see details in <a href="https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>).</p> </li> </ol> <p><strong>References</strong></p> <p>[1] A. Fedorov et al., "NCI Imaging Data Commons," Cancer Res., vol. 81, no. 16, pp. 4188&ndash;4193, Aug. 2021, doi: <a href="https://dx.doi.org/10.1158/0008-5472.CAN-21-0950">10.1158/0008-5472.CAN-21-0950</a>.&nbsp;</p> <p>[2] D. Milewski et al., "Predicting molecular subtype and survival of rhabdomyosarcoma patients using deep learning of H&amp;E images: A report from the Children's Oncology Group," Clin. Cancer Res., vol. 29, no. 2, pp. 364&ndash;378, Jan. 2023, doi: <a href="https://dx.doi.org/10.1158/1078-0432.CCR-22-1663">10.1158/1078-0432.CCR-22-1663</a>.</p> <p>[3] National Electrical Manufacturers Association (NEMA), "DICOM PS3.3 - Information Object Definitions: A.32.8 VL Whole Slide Microscopy Image IOD." Accessed: Aug. 11, 2023. [Online]. Available: <a href="https://dicom.nema.org/medical/dicom/current/output/html/part03.html#sect_A.32.8">https://dicom.nema.org/medical/dicom/current/output/html/part03.html#sect_A.32.8</a></p> <p>[4] M. D. Herrmann et al., "Implementing the DICOM standard for digital pathology," J. Pathol. Inform., vol. 9, no. 1, p. 37, Jan. 2018, doi: <a href="https://dx.doi.org/10.4103/jpi.jpi_42_18">10.4103/jpi.jpi_42_18</a>.&nbsp;</p> <p>[5] D. Clunie, A. Fedorov, and M. D. Herrmann, ImagingDataCommons/idc-wsi-conversion: Initial release. Zenodo, 2023. doi: <a href="https://dx.doi.org/10.5281/zenodo.8240154">10.5281/zenodo.8240154</a>.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Artefact segmentation in digital pathology whole-slide images

<p>Dataset with examples of Artefacts in Digital Pathology.</p> <p>The dataset contains 22 Whole-Slide Images, with H&amp;E or IHC staining, showing various types and levels of defect to the slides. Annotations were made by a biomedical engineer based on examples given by an expert.</p> <p>The dataset is split in different folders:</p> <ul> <li>train <ul> <li>18 whole-slide images (extracted at 1.25x &amp; 2.5x magnification)</li> <li>All from the same Block (colorectal cancer tissue)</li> <li>1/2 with H&amp;E &amp; 1/2 with anti-pan-cytokeratin IHC staining.</li> </ul> </li> <li>validation <ul> <li>3 whole-slide images (1.25x + 2.5x mag)</li> <li>2 from the same Block as the training set (1 IHC, 1 H&amp;E)</li> <li>1 from another Block (IHC anti-pan-cytokerating, gastroesophageal junction lesion)</li> </ul> </li> <li>validation_tiles <ul> <li>patches of varying sizes taken from the 3 validation whole-slide images @1.25x magnification.</li> <li>7 patches from each slide.</li> </ul> </li> <li>test <ul> <li>1 whole-slide image (1.25x + 2.5x mag)</li> <li>From another block: IHC staining (anti-NR2F2), mouth cancer</li> </ul> </li> </ul> <p>For the train, validation and test whole-slide images, each slide has:<br> - The RGB images @1.25x &amp; 2.5x mag<br> - The corresponding background/tissue masks<br> - The corresponding annotation masks containing examples of artefacts (note that a majority of artefacts are not annotated. In total, 918 artefacts are in the train set)</p> <p>For the validation tiles, the following table gives the &quot;patch-level&quot; supervision:</p> <p>tile#&nbsp;&nbsp; Artefact(s)<br> 00&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 01&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold<br> 02&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Ink<br> 03&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 04&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 05&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold<br> 06&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold + Blur<br> 07&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Knife damage<br> 08&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Knife damage<br> 09&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Ink<br> 10&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 11&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold<br> 12&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold<br> 13&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 14&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 15&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Knife damage<br> 16&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tear&amp;Fold<br> 17&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 18&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; None/Few<br> 19&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Blur<br> 20&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Knife damage</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Representative Sample Dataset for Resolution-Agnostic Tissue Segmentation in Whole-Slide Histopathology Images

<p>This is a representative sample from the dataset that was used to develop resolution-agnostic convolutional neural networks for tissue segmentation1 in whole-slide histopathology images.</p> <p>The dataset is composed of two parts: <strong>development set</strong> and <strong>dissimilar set</strong>.</p> <p>Sample images from the development set:</p> <ul> <li>breast_hne_00.tif</li> <li>breast_lymph_node_hne_00.tif</li> <li>tongue_ae1ae3_00.tif</li> <li>tongue_hne_00.tif</li> <li>tongue_ki67_00.tif</li> </ul> <p>Sample images from the dissimilar set:</p> <ul> <li>brain_alcianblue_00.tif</li> <li>cornea_grocott_00.tif</li> <li>kidney_cab_00.tif</li> <li>skin_perls_00.tif</li> <li>uterus_vonkossa_00.tif</li> </ul>

opencc-by-4.0Aug 2019View details →
zenodo44/100

PESO: Prostate Epithelium Segmentation on H&E-stained prostatectomy whole slide images

<p>Large set of whole-slide-images (WSI) of prostatectomy specimens with various grades of prostate cancer (PCa). More information can be found in the corresponding paper:&nbsp;<a href="https://doi.org/10.1038/s41598-018-37257-4">https://doi.org/10.1038/s41598-018-37257-4</a></p> <p>The WSIs in this dataset can be viewed using the open-source software <a href="https://github.com/computationalpathologygroup/ASAP">ASAP</a>&nbsp;or <a href="https://openslide.org/">Open Slide</a>.</p> <p>Due to the large size of the complete dataset, the data has been split up in to multiple archives.</p> <p>The data from the training set:</p> <ul> <li><strong>peso_training_masks.zip:&nbsp;</strong>Training masks (N=62)&nbsp;that have been used to train the main network of our paper. These masks are generated by a trained U-Net on the corresponding IHC slides.</li> <li><strong>peso_training_masks_corrected.zip:&nbsp;</strong>A subset of the color deconvolution masks (N=25)&nbsp;on which manual annotations have been made. Within these regions, stain and other artifacts have been removed.</li> <li><strong>peso_training_colordeconvolution.zip:&nbsp;</strong>Mask files (N=62)&nbsp;containing the P63&amp;CK8/18 channel&nbsp;of the color deconvolution operation. These masks mark all regions that are stained by either P63 or CK8/18 in the IHC version of the slides.</li> <li><strong>peso_training_wsi_{1-6}.zip:&nbsp;</strong>Zip files containing the whole slide images of the training set (N=62). Each archive contains 10 slides, excluding the last which contains 12.&nbsp;These images are exported at a pixel resolution of 0.48mu/pixels.&nbsp;</li> </ul> <p>The data from the test set:</p> <ul> <li><strong>peso_testset_regions.zip:&nbsp;</strong>Collection of annotation XML files with outlines of the test regions. These can be used to view the test regions in more detail using ASAP.</li> <li><strong>peso_testset_png.zip:&nbsp;</strong>Export of the test set regions in PNG format (2500x2500 pixels per region).</li> <li><strong>peso_testset_png_padded.zip:&nbsp;</strong>Export of the test regions in PNG format padded with a 500 pixel wide border (3500x3500 pixels per region). Useful for segmenting pixels at the border of the regions.</li> <li><strong>peso_testset_mapping.csv:&nbsp;</strong>A csv file mapping files from the test set (numbered 1-160) to regions in the xml files. The csv file also contains the label (benign or cancer) for each region.</li> <li><strong>peso_testset_groundtruth_masks.zip: </strong>The ground truth (pixel) masks (N=40) of all regions in the test set. For each pixel in the test set regions, these masks contain the ground truth: 0 for unlabelled, 1 for background and 2 for epithelial tissue.</li> <li><strong>peso_testset_wsi_{1-4}.zip:&nbsp;</strong>Zip files containing the whole slide images of the test set (N=40). Each archive contains 10 slides of the test set. These images are exported at a pixel resolution of 0.48mu/pixels.&nbsp;</li> </ul> <p>This study was financed by a grant from the Dutch Cancer Society (KWF), grant number KUN 2015-7970.</p> <p><strong>If you make use of this dataset please cite both the dataset itself and the corresponding paper:&nbsp;</strong><a href="https://doi.org/10.1038/s41598-018-37257-4">https://doi.org/10.1038/s41598-018-37257-4</a></p> <p><strong>Update July 2021: </strong>We have added the ground truth masks for the test set.</p>

opencc-by-nc-sa-4.0Nov 2018View details →
zenodo40/100

DICOM WG-26 Whole Slide Imaging Annotations Connectathon: Imaging Data Commons entry

<p>This data descriptor contains DICOM Slide Microscopy (SM modality) images and DICOM 2D point and polygon Bulk Annotations (ANN modality) submitted by the Imaging Data Commons team as part of the participation in the 2024 DICOM Working Group 26 Annotations Connectathon (<a href="https://dicom-wg26-connectathons.github.io/2024-annotations/" target="_blank" rel="noopener">https://dicom-wg26-connectathons.github.io/2024-annotations/</a>).</p> <p>Detailed content of the descriptor is as follows:</p> <ol> <li>Images correspond to 3 series selected from the DICOM-converted slides incuded in the TCGA-READ collection (originally shared in vendor-specific format in <a href="https://portal.gdc.cancer.gov/projects/TCGA-READ" target="_blank" rel="noopener">https://portal.gdc.cancer.gov/projects/TCGA-READ</a>) and available from NCI Imaging Data Commons [1] at <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=tcga_read" target="_blank" rel="noopener">https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=tcga_read</a>. Specifically, the following series are included, as defined by their `PatientID`, `StudyInstanceUID` and `SeriesInstanceUID` values. Each individual series contains multiple files corresponding to different resolution layers, shared as zip files.&nbsp; <ol> <li>`TCGA-AF-2687.zip`: series `1.3.6.1.4.1.5962.99.1.2251401802.152239158.1638633974346.2.0` <a href="https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.213661963103110408605329613498871186883/series/1.3.6.1.4.1.5962.99.1.2251401802.152239158.1638633974346.2.0" target="_blank" rel="noopener">https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.213661963103110408605329613498871186883/series/1.3.6.1.4.1.5962.99.1.2251401802.152239158.1638633974346.2.0</a></li> <li>`TCGA-AF-2689.zip`: series `1.3.6.1.4.1.5962.99.1.2259090539.712983657.1638641663083.2.0` <a href="https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.312916405820155829215771528638931942827/series/1.3.6.1.4.1.5962.99.1.2259090539.712983657.1638641663083.2.0" target="_blank" rel="noopener">https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.312916405820155829215771528638931942827/series/1.3.6.1.4.1.5962.99.1.2259090539.712983657.1638641663083.2.0</a></li> <li>`TCGA-AF-2690.zip`: series `1.3.6.1.4.1.5962.99.1.2247972296.1080138101.1638630544840.2.0` <a href="https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.158926540358526295486644564130790202309/series/1.3.6.1.4.1.5962.99.1.2247972296.1080138101.1638630544840.2.0" target="_blank" rel="noopener">https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.158926540358526295486644564130790202309/series/1.3.6.1.4.1.5962.99.1.2247972296.1080138101.1638630544840.2.0</a></li> </ol> </li> <li>`polygon_annotations.zip`: 2D polygon annotations of the boundary of the cell nuclei. Conversion into DICOM ANN was performed from the original content shared in [2].</li> <li>`point_annotations.zip`: 2D point annotations corresponding to the centroids of the cell nuclei defined in the prior polygon annotations.</li> <li>`rectangle_annotations.zip`: 2D rectangle annotations corresponding to the axis-aligned bounding box of the cell nuclei defined in the prior polygon annotations.</li> <li>`ellipse_annotations.zip`: 2D ellipse annotations corresponding to the ellipse fit to the cell nuclei defined in the prior polygon annotations.</li> <li>`screenshots.pdf`: screenshots demonstrating visualization of the included annotations in the open source Slim viewer (https://github.com/ImagingDataCommons/slim) (note these visualizations were produced using a development branch of the software)</li> <li>`run_verify.sh`: script that was used to run `dciodvfy` validator and save validation results.</li> </ol> <p>Zip files with the annotations also include the output of the `dciodvfy` validator (https://dclunie.com/dicom3tools/dciodvfy.html) version `20240227102104` for each of the files.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Whole slide images of mouse liver serial sections - Test registration dataset

<p>15 H&amp;E serial section of mouse liver and small intestine.</p> <p>Sampled prepared in the <a href="https://www.epfl.ch/research/facilities/histology-core-facility/">EPFL histology core facility</a> by Nathalie M&uuml;ller, Gian-Filippo Mancini, and Agn&egrave;s Hautier.</p> <p>All slides where imaged with a VS200 Evident slide scanner from the<a href="http://biop.epfl.ch/"> EPFL BIOP imaging facility</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Complete results of Procedure section in manuscript "PathML: A unified framework for whole-slide image analysis with deep learning"

<p>The results from a complete run of the Procedure section of the paper &quot;PathML: A unified framework for whole-slide image analysis with deep learning&quot;.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Dual-modality imaging of immunofluorescence and imaging mass cytometry for high-resolution whole slide imaging with accurate single-cell segmentation

<p>Imaging mass cytometry (IMC) is a powerful multiplexed tissue imaging technology that allows simultaneous detection of more than 30 makers on a single slide. It has been increasingly used for single-cell based spatial phenotyping in a wide range of samples. However, it only acquires a small, rectangle field of view (FOV) with a low image resolution that hinders downstream analysis. Here, we reported a highly practical dual-modality imaging method that combines high-resolution immunofluorescence (IF) and high-dementional IMC on the same tissue slide. Our computational pipeline uses the whole slide image (WSI) of IF as spatial reference, &nbsp;integrates small FOV IMC into a WSI of IMC. The high-resolution IF images enable accurate single-cell segmentation to extract robust high-dimensional IMC features for downstream analysis. We applied this method in esophageal adenocarcinoma of different stages, identified the single-cell pathology landscape via reconstruction of WSI IMC images and demonstrated the advantage of the dual-modality imaging strategy.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

DeepHisto: Dataset for glioma subtype classification from Whole Slide Images

<p>DeepHisto dataset contains tiles (patches) of hematoxylin and eosin stained Whole Slide Images (WSI) of 28 adult-type diffuse glioma cases collected at the National Center of Pathology (NCP), Luxembourg National Health Laboratory (Laboratoire national de sant&eacute; - LNS) from 2017 to 2021. WSIs were acquired with an IntelliSite Ultra Fast digital slide scanner from Philips containing a 20x/0.75 NA Plan Apo objective with an average slide resolution of 0.25um/pixel.</p> <p>Three primary diffuse glioma subtypes are classified into IDH-mutant, 1p/19q codeleted oligodendroglioma, IDH-mutant astrocytoma, and IDH-wildtype glioblastoma according to the 5th edition of the WHO classification of central nervous system tumors. The brain WSIs of a non-cancer patients were used as normal brain (white and gray matter) controls.</p> <p>Region annotation of WSIs was done by a board-certified pathologist, and the regions of interest are further divided into square 512&times;512 tiles, each of them associated with a particular class denoting the respective tumor entity, normal brain tissue or necrosis.<br> Tiles are further divided into training and test subsets patient-wise.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Image tiles of TCGA-CRC-DX histological whole slide images, non-normalized, tumor only

<p>These are image tiles of tumor tissue of N=604 colorectal cancer (CRC) histological whole slide images in the TCGA database. The tumor tissue was manually outlined and cut into non-overlapping tiles of 512x512 px at 0.5 &micro;m/px. No further preprocessing was applied. No color normalization was applied. Tiles are compressed with JPEG. Each ZIP file corresponds to one whole slide image (original format: SVS).</p> <p>Please cite our previous publication related to this dataset: https://www.nature.com/articles/s41591-019-0462-y</p> <p>Please observe the original TCGA licenses if you use the data https://portal.gdc.cancer.gov/</p> <p>Original image credit to TCGA: https://portal.gdc.cancer.gov/</p> <p>Image tiles were created with QuPath v0.1.2 https://qupath.github.io</p> <p>Genetic data matching these images are available at https://cbioportal.org</p> <p>More information on the procedures: https://zenodo.org/record/3694994</p>

opencc-by-4.0May 2020View details →
zenodo36/100

High resolution images for 'Identification of "BRAF-positive" cases based on whole-slide image analysis'

<p>This archive contains high resolution images to accompany the article 'Identification of “BRAF-positive” cases based on whole-slide image analysis' by V. Popovici, A. Krenek and E. Budinska.</p>

opencc-by-nc-4.0Mar 2017View details →
zenodo36/100

Datasets for "Automated Detection of Portal Fields and Central Veins in Whole-Slide Images of Liver Tissue"

<p>Datasets and results for the manuscript &ldquo;Automated Detection of Portal Fields and Central Veins in Whole-Slide<br> Images of Liver Tissue&rdquo; (Journal of Pathology Informatics 13 (2022) 100001, https://doi.org/10.1016/j.jpi.2022.100001)</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Feature Re-calibration based Multiple Instance Learning for Whole Slide Image Classification - Features

<p>A release of the sources employed in the FRMIL work - &quot;Feature Re-calibration based Multiple Instance Learning for Whole Slide Image Classification&quot;, presented at MICCAI 2022.</p> <p>We release the Camelyon16 Whole Slide Image (WSI) extracted patch-level features for reproducibility.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Source data for VALIS: Virtual Alignment of pathoLogy Image Series for multi-gigapixel whole slide images publication

<p>Source data used to create figures in&nbsp;<em>VALIS: Virtual Alignment of pathoLogy Image Series for multi-gigapixel whole slide images</em>&nbsp;(Nature Communications, 2023)</p>

opencc-by-4.0May 2023View details →
ClinicalTrials.gov36/100

Development and Validation of a Deep Learning Model to Predict Distant Metastases in Nasopharyngeal Carcinoma Using Whole Slide Imaging and MRI

ClinicalTrials.gov study NCT06831357. IPD Sharing: NO. Countries: 1. Publications: 8.

closedIPD-NOFeb 2026View details →
dryad36/100

Generation of synthetic whole-slide image tiles of tumours from RNA-sequencing data via cascaded diffusion models

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo32/100

Analysis of multiplexed whole slide images with QuPath and Cytomap

<p>This image was acquired during MIFOBIO 2023 and was used for workshop entitled "Analysis of multiplexed whole slide images with QuPath and Cytomap".</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

HAPPY: a deep learning pipeline for mapping cell-to-tissue graphs across placenta histology whole slide images

<p>These two zipped folders contain all data necessary to train, validate and reproduce results from the paper.</p> <p>Unzipping the files will create 6 folders. Data from folders with the same name across both zips should be combined into one folder. The 'annotations' folder contains all ground truth annotations for training all three deep learning models. The 'datasets' folder contains images for training the nuclei localisation and cell classification models. The 'embeddings' folder contains cell embedding vectors and nuclei coordinates from two slides used to create nodes to train the graph tissue classification model. The 'graph_splits' folder contains regions defining the validation and test splits for the graph model. The 'slides' folder contains a sample region of a whole slide image as a .tiff file for running the inference demo. The 'trained_models' folder contains trained weights for each of the three models.</p> <p>Further instructions for dataset use and creation of custom datasets are available in the GitHub readme: https://github.com/Nellaker-group/happy.</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Histology images from uniform tumor regions in TCGA Whole Slide Images (TCGA-UT)

<h1><strong>TCGA-UT Dataset Documentation&nbsp;</strong></h1> <h2>Quick Links</h2> <ul> <li><strong>Dataset on Hugging Face</strong>: For users interested in benchmarking foundation models or feature extractors, please visit <a href="https://huggingface.co/datasets/dakomura/tcga-ut">TCGA-UT on Hugging Face</a></li> <li><strong>Original Paper</strong>: <a href="https://doi.org/10.1016/j.celrep.2022.110424">Universal encoding of pan-cancer histology by deep texture representations</a></li> </ul> <p>&nbsp;</p> <h2>Dataset Overview</h2> <p>The TCGA-UT dataset is a large-scale collection of histopathological image patches from human cancer tissues. It contains 1,608,060 image patches extracted from hematoxylin &amp; eosin (H&amp;E) stained histological samples across 32 different types of solid cancers.</p> <h3>Key Features</h3> <ul> <li><strong>Size</strong>: Over 1.6 million image patches</li> <li><strong>Resolution</strong>: All patches are standardized to 256 x 256 pixels</li> <li><strong>Source</strong>: Derived from The Cancer Genome Atlas (TCGA) dataset</li> <li><strong>Quality</strong>: Curated by trained pathologists</li> <li><strong>Coverage</strong>: 32 different cancer types</li> <li><strong>Patient Base</strong>: 7,175 patients from 8,736 diagnostic slides</li> </ul> <h2>Data Collection Process</h2> <ol> <li><strong>Image Source</strong>: Whole Slide Images (WSI) were downloaded from the GDC legacy database between December 2016 and June 2017</li> <li><strong>Expert Annotation</strong>: Two trained pathologists selected at least three representative tumor regions per slide</li> <li><strong>Quality Control</strong>: 926 slides were removed due to various quality issues (poor staining, low resolution, focus problems, etc.)</li> <li><strong>Patch Extraction</strong>: 10 patches were randomly cropped at 6 different magnification levels from each annotated region</li> </ol> <h2>File Structure</h2> <p>Files are organized using the following format:</p> <div> <div>&nbsp;</div> <div> <div><span>Copy</span></div> </div> <div> <div><code>[cancer_type]/[resolution]/[TCGA Barcode]/[region]-[number]-[pixel resolution].jpg</code></div> </div> </div> <h3>Resolution Key</h3> <ul> <li>0: 0.5 &mu;m/pixel</li> <li>1: 0.6 &mu;m/pixel</li> <li>2: 0.7 &mu;m/pixel</li> <li>3: 0.8 &mu;m/pixel</li> <li>4: 0.9 &mu;m/pixel</li> <li>5: 1.0 &mu;m/pixel</li> </ul> <h2>License</h2> <ul> <li><strong>Non-Commercial Use</strong>: CC-BY-NC-SA 4.0</li> <li><strong>Commercial Use</strong>: Please contact <a href="mailto:ishum-prm@m.u-tokyo.ac.jp">ishum-prm@m.u-tokyo.ac.jp</a> for licensing</li> </ul> <h2>Citation</h2> <p>If you use this dataset in your research, please cite:</p> <div> <div>&nbsp;</div> <div> <div><span>Copy</span></div> </div> <div> <div><code>Komura, D., et al. (2022). Universal encoding of pan-cancer histology by deep texture representations. Cell Reports 38, 110424. https://doi.org/10.1016/j.celrep.2022.110424</code></div> </div> </div> <h2>For Model Benchmarking</h2> <p>If you're interested in using this dataset for benchmarking foundation models or feature extractors, we recommend accessing the dataset through the Hugging Face Hub at <a href="https://huggingface.co/datasets/dakomura/tcga-ut">dakomura/tcga-ut</a>. The Hugging Face version provides:</p> <ul> <li>Predefined train/validation/test splits (both internal and external facility-based splits)</li> <li>Ready-to-use benchmarking framework for foundation models</li> <li>WebDataset format support for efficient data loading</li> <li>Example implementations for state-of-the-art model evaluation</li> </ul> <p>&nbsp;</p>

openother-ncDec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record