Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

35 results for “whole slide imaging”

Learn how ShareScore rates datasets ↗
zenodo32/100

WebMicroscope's Deep Learning AI platform automates image analyses with an approach that is faster and able to understand tissue context, which reduces steps needed for accurate results. Researchers can gain access to digitized samples, such as this image of breast-cancer tissue (left), and analyze results through the cloud platform anywhere, anytime. This is a whole slide image of a tissue section of an adrenal gland (right). Fimmic's WebMicroscope cloud platform allows researchers to manage, share, and view digital gigapixel images with any modern browser. Researchers can rapidly pan, zoom, and analyze a digital sample. Photographs: Courtesy of Fimmic Oy. in Deep learning brings speed, accuracy to the life sciences.

WebMicroscope's Deep Learning AI platform automates image analyses with an approach that is faster and able to understand tissue context, which reduces steps needed for accurate results. Researchers can gain access to digitized samples, such as this image of breast-cancer tissue (left), and analyze results through the cloud platform anywhere, anytime. This is a whole slide image of a tissue section of an adrenal gland (right). Fimmic's WebMicroscope cloud platform allows researchers to manage, share, and view digital gigapixel images with any modern browser. Researchers can rapidly pan, zoom, and analyze a digital sample. Photographs: Courtesy of Fimmic Oy.

opennotspecifiedJan 2018View details →
zenodo32/100

GTEx: DICOM converted whole slide hematoxylin and eosin stained images from the Genotype-Tissue Expression (GTEx) Project

<p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute&nbsp;<a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=gtex" target="_blank" rel="noopener">GTEx</a>. You can use the manifests included in this Zenodo record to download the content of the collection following the&nbsp;<strong>Download instructions</strong>&nbsp;below.</p> <h3>Collection description</h3> <p>The<a href="https://commonfund.nih.gov/GTEx"> Genotype-Tissue Expression (GTEx) Project</a> established a data resource and tissue bank to study the relationship between genetic variants and gene expression in multiple human tissues and across individuals. The project included contributions from numerous groups with diverse expertise in biospecimen collection and processing, pathology review, molecular analysis, and data management. The contributors are collectively called the GTEx Consortium.</p> <p>GTEx collected a total of 26,468 unique tissue samples from 50+ different tissue types, from 956 healthy postmortem donors. The standardized biospecimen collection and analysis practices applied during the study served to minimize preanalytical variability associated with specimen-related factors and their potential impact on analytic endpoints. Each GTEx tissue was divided into two tissue blocks, one for histology and one for molecular analysis; both tissue blocks were preserved in PAXgene Tissue Fixative (Qiagen) solution for 6 to 24 hours, followed by PAXgene Tissue Stabilizer (Qiagen) as specified in the project-specific<a href="https://biospecimens.cancer.gov/resources/sops/library.asp"> standard operating procedures</a>. Tissue blocks were processed and embedded in paraffin at the GTEx central repository at the Van Andel Institute (MI) and hematoxylin and eosin&ndash;stained slides were generated from all GTEx donors. Digitally scanned whole slide images of PAXgene-fixed/stabilized, paraffin-embedded tissue sections were created using Aperio Scanscope software (Leica Biosystems). The digital images were then reviewed and annotated by one of four board-certified pathologists assigned to the GTEx study. There are a total of 25,503 digital histology images in the GTEx collection.</p> <p>GTEx was supported by the NIH Common Fund (2010 &ndash; 2019).&nbsp; Additional resources include the<a href="https://gtexportal.org/home/biobank"> GTEx Biobank</a>, the<a href="https://gtexportal.org/home/"> GTEx Portal</a>, and the full dataset at<a href="https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000424.v9.p2"> dbGaP</a> (accession number phs000424).</p> <p>Please refer to the listed GTEx publications below for more details [2-7].&nbsp;</p> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example,&nbsp;<code>collection_id-idc_v8-aws.s5cmd</code>&nbsp;corresponds to the contents of the&nbsp;<code>collection_id</code>&nbsp;collection introduced in IDC data release v8. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <ol> <li><code>gtex-idc_v19-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>gtex-idc_v19-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>gtex-idc_v19-dcf.dcf</code>: Gen3 manifest (for details see&nbsp;<a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in&nbsp;<code>-aws.s5cmd</code>&nbsp;reference files stored in Amazon Web Services (AWS) buckets, while&nbsp;<code>-gcs.s5cmd</code> reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using&nbsp;<code>.s5cmd</code>&nbsp;manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the&nbsp;<code>.s5cmd</code>&nbsp;manifest file:&nbsp;<code>idc download&nbsp;manifest.s5cmd</code></li> </ol> <p>To download the files using&nbsp;<code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <div>Please acknowledge the GTEx Consortium in any published work that includes the images.&nbsp;A sample statement for the acknowledgment of the Genotype-Tissue Expression (GTEx) Project dataset(s) follows.</div> <p>The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health (<a href="http://commonfund.nih.gov/GTEx" target="_blank" rel="noopener">commonfund.nih.gov/GTEx</a>). Additional funds were provided by the NCI, NHGRI, NHLBI, NIDA, NIMH, and NINDS. Donors were enrolled at Biospecimen Source Sites funded by NCI/Leidos Biomedical Research, Inc. subcontracts to the National Disease Research Interchange (10XS170), Roswell Park Cancer Institute (10XS171), and Science Care, Inc. (X10S172). The Laboratory, Data Analysis, and Coordinating Center (LDACC) was funded through a contract (HHSN268201000029C) to the Broad Institute of MIT and Harvard. Biorepository operations were funded through a Leidos Biomedical Research, Inc. subcontract to Van Andel Research Institute (10ST1035). Additional data repository and project management were provided by Leidos Biomedical Research, Inc. (HHSN261200800001E). The Brain Bank was supported with supplements to University of Miami grant DA006227. Statistical Methods development grants were made to the University of Geneva (MH090941&amp; MH101814), the University of Chicago (MH090951, MH090937, MH101825, &amp; MH101820), the University of North Carolina - Chapel Hill (MH090936), North Carolina State University (MH101819), Harvard University (MH090948), Stanford University (MH101782), Washington University (MH101810), and to the University of Pennsylvania (MH101822).</p> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <h3>References</h3> <p>[1] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. &amp; Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence.&nbsp;<em>Radiographics</em>&nbsp;<strong>43,</strong> (2023).</p> <p>[2] Sobin, L., Barcus, M., Branton, P. A., Engel, K. B., Keen, J., Tabor, D., Ardlie, K. G., Greytak, S. R., Roche, N., Luke, B., Vaught, J., Guan, P. &amp; Moore, H. M. Histologic and quality assessment of genotype-Tissue Expression (GTEx) research samples: A large postmortem tissue collection. Arch. Pathol. Lab. Med. (2024). doi:<a href="http://dx.doi.org/10.5858/arpa.2023-0467-OA">10.5858/arpa.2023-0467-OA</a></p> <p>[3] GTEx Consortium. The Genotype-Tissue Expression (GTEx) project. Nat. Genet. 45, 580&ndash;585 (2013).</p> <p>[4] GTEx Consortium. Human genomics. The Genotype-Tissue Expression (GTEx) pilot analysis: multitissue gene regulation in humans. Science 348, 648&ndash;660 (2015).</p> <p>[5] GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318&ndash;1330 (2020).</p> <p>[6] Carithers, L. J., Ardlie, K., Barcus, M., Branton, P. A., Britton, A., Buia, S. A., Compton, C. C., DeLuca, D. S., Peter-Demchok, J., Gelfand, E. T., Guan, P., Korzeniewski, G. E., Lockhart, N. C., Rabiner, C. A., Rao, A. K., Robinson, K. L., Roche, N. V., Sawyer, S. J., Segr&egrave;, A. V., Shive, C. E., Smith, A. M., Sobin, L. H., Undale, A. H., Valentino, K. M., Vaught, J., Young, T. R., Moore, H. M. &amp; GTEx Consortium. A novel approach to high-quality postmortem tissue procurement: The GTEx project. Biopreserv. Biobank. 13, 311&ndash;319 (2015).</p> <p>[7] Branton, P. A., Sobin, L., Barcus, M., Engel, K. B., Greytak, S. R., Guan, P., Vaught, J. &amp; Moore, H. M. Notable histologic findings in a &lsquo;normal&rsquo; cohort: The National Institutes of Health Genotype-Tissue Expression (GTEx) project. Arch. Pathol. Lab. Med. (2024). doi:<a href="http://dx.doi.org/10.5858/arpa.2023-0468-OA">10.5858/arpa.2023-0468-OA</a></p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

GrandQC: QC Masks for TCGA whole-slide images

<p>These masks for WSIs of 32 TCGA cohorts were generated by GrandQC tool at MPP of 1.5 &micro;m/px.</p> <p><em>Note: for some cohorts tar file with '_2' mean part 2</em>).</p> <p>The masks contain the following 7 classes:</p> <ol> <li>Tissue</li> <li>Fold</li> <li>Dark Spot and Foreign Object</li> <li>Pen Marker</li> <li>Edge and Air Bubble</li> <li>Out of Focus</li> <li>Background</li> </ol> <p>For using these masks, the original publication should be cited:</p> <p>Weng Z. et al. "<strong>GrandQC: </strong><strong>a</strong><strong> </strong><strong>comprehensive</strong><strong> </strong><strong>s</strong><strong>olution to </strong><strong>q</strong><strong>uality </strong><strong>c</strong><strong>ontrol </strong><strong>p</strong><strong>roblem in </strong><strong>d</strong><strong>igital </strong><strong>p</strong><strong>athology</strong>"</p> <p>Nature Communications 2024</p> <p>&nbsp;</p> <p>These masks are for NON-COMMERCIAL use only.&nbsp;&nbsp;</p>

opencc-by-nc-sa-4.0Nov 2024View details →
dryad32/100

Orbit Image Analysis: An open-source whole slide image analysis tool

<p>This is a whole slide image (WSI) dataset for glomeruli segmentation on kidney tissue, in total 88 images.</p> <p>The train-set (58 images) and test-set (32 images) has been used in the publication "Orbit Image Analysis: An open-source whole slide image analysis tool" to train and test the<br> glomeruli segmentation model.</p>

opencc-zeroJan 2020View details →
zenodo32/100

Trained CNN for analysis of melanoma whole slide images to automatically assess the infiltration of TILs

<p>The file is the trained convolutional neural network (CNN) developed in &quot;Ugolini F et al.Tumor infiltrating lymphocytes recognition in primary melanoma by deep learning convolutional neural network&quot; submitted to&nbsp;American Journal of Pathology (2023). The CNN is&nbsp;based on a a pre-trained Inception-ResNet-v2&nbsp;to&nbsp;automatically&nbsp;to recognize areas in the tumor region containing TILs and area without TILs in&nbsp;histopathological digitalized slides of primary melanoma. The file is in a Matlab format (.mat).</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Test Dataset for Whole Slide Image Registration

<p>Mouse duodenum fixed in 4% PFA overnight at 4&deg;C, processed for paraffin infiltration using a standard histology procedure and cut at 4 microns were dewaxed, rehydrated, permeabilized with 0.5% Triton X-100 in PBS 1x and stained with Azide - Alexa Fluor 555 (Thermo Fisher) to detect EdU and DAPI for nuclei. The images were taken using a Leica DM5500 microscope with a 40X N.A.1 objective (black&amp;white camera: DFC350FXR2, pixel dimension: 0.161 microns). Next, the slide was unmounted and stained using the fully automated Ventana Discovery xT autostainer (Roche Diagnostics, Rotkreuz, Switzerland). All steps were performed on automate with Ventana solutions. Sections were pretreated with heat using the CC1 solution under mild conditions. The primary rat anti BrDU (clone: BU1/75 (ICR1), Serotec, diluted 1:300) was incubated 1 hour at 37&deg;C. After incubation with a donkey anti rat biotin diluted 1:200 (Jackson ImmunoResearch Laboratories), chromogenic revelation was performed with DabMap kit. The section was counterstained with Harris hematoxylin (J.T. Baker) before a second round of imaging on DM5500 PL Fluotar 40X N.A.1.0 oil (color camera: DFC 320 R2, pixel dimension: 0.1725 microns). Before acquisition, a white-balance as well as a shading correction is performed according to Leica LAS software wizard. The fluorescence and DAB images were converted in ome.tiff multiresolution file with the <a href="https://github.com/BIOP/ijp-kheops">kheops Fiji Plugin</a>.</p> <p>Sampled prepared in the <a href="https://www.epfl.ch/research/facilities/histology-core-facility/">EPFL histology core facility</a> by Nathalie M&uuml;ller and Gian-Filippo Mancini.</p> <p>Associated documents:</p> <ul> <li><a href="https://c4science.ch/w/bioimaging_and_optics_platform_biop/teaching/dab-intensity/">https://c4science.ch/w/bioimaging_and_optics_platform_biop/teaching/dab-intensity/</a></li> <li>https://imagej.net/plugins/bdv/warpy/warpy</li> </ul> <p>This document contains a full QuPath project with an example of registered image.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2021View details →
dryad32/100

Orbit Image Analysis: an open-source whole slide image analysis tool

Open the record for dataset details and reuse information.

publicFeb 2020View details →
dryad28/100

Data from: High-throughput adaptive sampling for whole-slide histopathology image analysis (HASHI) via convolutional neural networks: application to invasive breast cancer detection

Precise detection of invasive cancer on whole-slide images (WSI) is a critical first step in digital pathology tasks of diagnosis and grading. Convolutional neural network (CNN) is the most popular representation learning method for computer vision tasks, which have been successfully applied in digital pathology, including tumor and mitosis detection. However, CNNs are typically only tenable with relatively small image sizes (200x200 pixels). Only recently, Fully convolutional networks (FCN) are able to deal with larger image sizes (500x500 pixels) for semantic segmentation. Hence, the direct application of CNNs to WSI is not computationally feasible because for a WSI, a CNN would require billions or trillions of parameters. To alleviate this issue, this paper presents a novel method, High-throughput Adaptive Sampling for whole-slide Histopathology Image analysis (HASHI), which involves: i) a new efficient adaptive sampling method based on probability gradient and quasi-Monte Carlo sampling, and, ii) a powerful representation learning classifier based on CNNs. We applied HASHI to automated detection of invasive breast cancer on WSI. HASHI was trained and validated using three different data cohorts involving near 500 cases and then independently tested on 195 studies from The Cancer Genome Atlas. The results show that (1) the adaptive sampling method is an effective strategy to deal with WSI without compromising prediction accuracy by obtaining comparative results of a dense sampling (~6 million of samples in 24 hours) with far fewer samples (~2,000 samples in 1 minute), and (2) on an independent test dataset, HASHI is effective and robust to data from multiple sites, scanners, and platforms, achieving an average Dice coefficient of 76%.

opencc-zeroDec 2017View details →
zenodo28/100

Contrastive learning-based histopathological feature infers molecular subtypes and clinical outcomes of breast cancer from unannotated whole slide images

<p>The breast cancer cohort&nbsp;came from the Changzhou Second&nbsp;People&#39;s Hospital&nbsp;(CZSPH)&nbsp;in&nbsp;Jiangsu, China.&nbsp;This cohort&nbsp;collected 91 FFPE WSIs from 90 breast cancer&nbsp;patients, including 15&nbsp;recurrence cases within 5 years.</p>

opencc-by-4.0May 2023View details →
ClinicalTrials.gov28/100

Whole-slide Image and CT Radiomics Based Deep Learning System for Prognostication Prediction in Bladder Cancer

ClinicalTrials.gov study NCT06389019. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov28/100

Whole-slide Image and CT Radiomics Based Deep Learning System for Prognostication Prediction in Upper Tract Urothelial Carcinoma

ClinicalTrials.gov study NCT06993779. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
dryad28/100

Data from: High-throughput adaptive sampling for whole-slide histopathology image analysis (HASHI) via convolutional neural networks: application to invasive breast cancer detection

Open the record for dataset details and reuse information.

publicJun 2018View details →
geo24/100

Self-supervised learning for predicting transcriptomic groups on whole slides images in intrahepatic cholangiocarcinoma

GEO Series GSE244807. Homo sapiens. 246 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2023View details →
zenodo20/100

UTOM dataset: whole-slide pathological images of various kinds of cancers

<p>UTOM dataset: whole-slide pathological images of various kinds of cancers</p>

opencc-by-4.0Sep 2023View details →
zenodo12/100

Generating highly accurate pathology reports from gigapixel whole slide images with HistoGPT

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record