Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
240
datasets available to search
ShareScore release 0.9.0
Dataset results
240 results for “Tissue imaging”
Data from: Three-dimensional single-cell transcriptome imaging of thick tissues
Open the record for dataset details and reuse information.
Data for: PIEZO1-HaloTag hiPSCs: bridging molecular, cellular and tissue imaging
Open the record for dataset details and reuse information.
Data from: Photoacoustic imaging of rat kidney tissue oxygenation using NIR-II wavelengths
Open the record for dataset details and reuse information.
Magnetic resonance imaging reveals human brown adipose tissue is rapidly activated in response to cold
<p class="MsoNoSpacing"><b>Context.</b> In rodents, cold exposure induces the activation of brown adipose tissue (BAT) and the induction of intracellular triacylglycerol (TAG) lipolysis. However, in humans, the kinetics of supraclavicular (SCV) BAT activation and the potential importance of TAG stores remain poorly defined.</p> <p class="MsoNoSpacing"><b>Objective.</b> To determine the time course of BAT activation and changes in intracellular TAG using magnetic resonance imaging (MRI) assessment of the SCV (i.e. BAT depot) and fat in the posterior neck region (i.e. non BAT).</p> <p class="MsoNoSpacing"><b>Design.</b> Cross-sectional.</p> <p class="MsoNoSpacing"><b>Setting.</b> Clinical research centre.</p> <p class="MsoNoSpacing"><b>Patients or Other Participants.</b> Twelve healthy male volunteers ages 18-29 years [BMI=24.7±2.8kg/m<sup>2</sup> and body fat percentage = 25.0±7.4% (both mean±SD)].</p> <p class="MsoNoSpacing"><b>Intervention(s).</b> Standardized whole-body cold exposure (180 minutes at 18<span>°</span>C) and immediate re-warming (30 minutes at 32°C).</p> <p class="MsoNoSpacing"><b>Main Outcome Measure(s).</b> Proton density fat fraction (PDFF) and T2* of the SCV and posterior neck fat pads. Acquisitions occurred at 5-15 minute intervals during cooling and subsequent warming.</p> <p class="MsoNoSpacing"><b>Results.</b> SCV PDFF declined significantly after only 10 minutes of cold exposure [-1.6% (standard error (SE) 0.44%), <i>p</i>=0.007) and continued to decline until 35 minutes after which time it remained stable until 180 minutes. A similar time course was also observed for SCV T2*. In the posterior neck fat (non-BAT) there were no cold-induced changes in PDFF or T2*. Re-warming did not result in a change in SCV PDFF or T2*.</p> <p class="MsoNoSpacing"><b>Conclusions.</b> The rapid cold-induced decline in SCV PDFF suggests that in humans, BAT is activated quickly in response to cold and that TAG is a primary substrate.</p>
Expert and AI-generated annotations of the tissue types for the RMS-Mutation-Prediction microscopy images
<div> <p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute <a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.. You can use the manifests included in this Zenodo record to download the content of the collection following the <strong>Download instructions</strong> below.</p> <h3>Collection description</h3> </div> <div> <div> <p>This dataset contains 2 components:</p> <ol> <li>Annotations of multiple regions of interest performed by an expert pathologist with eight years of experience for a subset of hematoxylin and eosin (H&E) stained images from the RMS-Mutation-Prediction image collection [1,2]. Annotations were generated manually, using the Aperio ImageScope tool, to delineate regions of alveolar rhabdomyosarcoma (ARMS), embryonal rhabdomyosarcoma (ERMS), stroma, and necrosis [3]. The resulting planar contour annotations were originally stored in ImageScope-specific XML format, and subsequently converted into Digital Imaging and Communications in Medicine (DICOM) Structured Report (SR) representation using the open source conversion tool [4].</li> <li>AI-generated annotations stored as probabilistic segmentations.</li> </ol> <p><strong>WARNING</strong>: After the release of IDC v20 (v2 of this data record), it was discovered that a mistake had been made during data conversion that affected the newly-released segmentations accompanying the "RMS-Mutation-Prediction" collection. Segmentations released in v20 for this collection have the segment labels for alveolar rhabdomyosarcoma (ARMS) and embryonal rhabdomyosarcoma (ERMS) switched in the metadata relative to the correct labels. Thus segment 3 in the released files is labelled in the metadata (the SegmentSequence) as ARMS but should correctly be interpreted as ERMS, and conversely segment 4 in the released files is labelled as ERMS but should be correctly interpreted as ARMS. This mistake was fixed in the version v3 of this record (IDC data release v21).</p> <p>Many pixels from the whole slide images annotated by this dataset are not contained inside any annotation contours and are considered to belong to the background class. Other pixels are contained inside only one annotation contour and are assigned to a single class. However, cases also exist in this dataset where annotation contours overlap. In these cases, the pixels contained in multiple contours could be assigned membership in multiple classes. One example is a necrotic tissue contour overlapping an internal subregion of an area designated by a larger ARMS or ERMS annotation. The ordering of annotations in this DICOM dataset preserves the order in the original XML generated using ImageScope. These annotations were converted, in sequence, into segmentation masks and used in the training of several machine learning models. Details on the training methods and model results are presented in [1]. In the case of overlapping contours, the order in which annotations are processed may affect the generated segmentation mask if prior contours are overwritten by later contours in the sequence. It is up to the application consuming this data to decide how to interpret tissues regions annotated with multiple classes. The annotations included in this dataset are available for visualization and exploration from the National Cancer Institute Imaging Data Commons (IDC) [5] (also see IDC Portal at <a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov</a>) as of data release v18. Direct link to open the collection in IDC Portal: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.</p> </div> <div> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example, <code>pan_cancer_nuclei_seg_dicom-collection_id-idc_v19-aws.s5cmd</code> corresponds to the annotations for th eimages in the <code>collection_id</code> collection introduced in IDC data release v19. DICOM Binary segmentations were introduced in IDC v20. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <p>For each of the collections, the following manifest files are provided:</p> <ol> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-dcf.dcf</code>: Gen3 manifest (for details see <a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in <code>-aws.s5cmd</code> reference files stored in Amazon Web Services (AWS) buckets, while <code>-gcs.s5cmd</code> reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using <code>.s5cmd</code> manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the <code>.s5cmd</code> manifest file: <code>idc download manifest.s5cmd</code></li> </ol> <p>To download the files using <code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <p>If you use the files referenced in the attached manifests, we ask you to cite this dataset, as well as the publication describing the original dataset <a href="https://paperpile.com/c/NHiBXI/njdR">[2]</a> and publication acknowledging IDC <a href="https://paperpile.com/c/NHiBXI/uJJZ">[5]</a>.</p> <h3>References</h3> </div> </div> <div> <p>[1] D. Milewski et al., "Predicting molecular subtype and survival of rhabdomyosarcoma patients using deep learning of H&E images: A report from the Children's Oncology Group," Clin. Cancer Res., vol. 29, no. 2, pp. 364–378, Jan. 2023, doi: 10.1158/1078-0432.CCR-22-1663.</p> <p>[2] Clunie, D., Khan, J., Milewski, D., Jung, H., Bowen, J., Lisle, C., Brown, T., Liu, Y., Collins, J., Linardic, C. M., Hawkins, D. S., Venkatramani, R., Clifford, W., Pot, D., Wagner, U., Farahani, K., Kim, E., & Fedorov, A. (2023). DICOM converted whole slide hematoxylin and eosin images of rhabdomyosarcoma from Children's Oncology Group trials [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8225132" rel="noopener">https://doi.org/10.5281/zenodo.8225132</a></p> <p>[3] Agaram NP. Evolving classification of rhabdomyosarcoma. Histopathology. 2022 Jan;80(1):98-108. doi: 10.1111/his.14449. PMID: 34958505; PMCID: PMC9425116,https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9425116/</p> <p>[4] Chris Bridge. (2024). ImagingDataCommons/idc-sm-annotations-conversion: v1.0.0 (v1.0.0). Zenodo. <a href="https://doi.org/10.5281/zenodo.10632182" rel="noopener">https://doi.org/10.5281/zenodo.10632182</a></p> <p>[5] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. & Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence. Radiographics 43, (2023).</p> </div>
Data for MAPS: Pathologist-level Cell Type Annotation from Tissue Images through Machine Learning
<p>Extracted datasets and processed images used to generate figures in manuscript <i>"MAPS: Pathologist-level Cell Type Annotation from Tissue Images through Machine Learning".</i></p>
HAPPY: a deep learning pipeline for mapping cell-to-tissue graphs across placenta histology whole slide images
<p>These two zipped folders contain all data necessary to train, validate and reproduce results from the paper.</p> <p>Unzipping the files will create 6 folders. Data from folders with the same name across both zips should be combined into one folder. The 'annotations' folder contains all ground truth annotations for training all three deep learning models. The 'datasets' folder contains images for training the nuclei localisation and cell classification models. The 'embeddings' folder contains cell embedding vectors and nuclei coordinates from two slides used to create nodes to train the graph tissue classification model. The 'graph_splits' folder contains regions defining the validation and test splits for the graph model. The 'slides' folder contains a sample region of a whole slide image as a .tiff file for running the inference demo. The 'trained_models' folder contains trained weights for each of the three models.</p> <p>Further instructions for dataset use and creation of custom datasets are available in the GitHub readme: https://github.com/Nellaker-group/happy.</p>
Plant tissue images
<p>All raw imaging data generated during my postdoctoral fellowship with Jacques Dumais (2007-2011) and used for https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3076879/</p>
Reproducible, high-dimensional imaging in archival human tissue by Multiplexed Ion Beam Imaging by Time-of-Flight (MIBI-TOF)
<p>1. SingleChannelMIBI.zip: Single-channel MIBI-TOF images</p> <p>All folders are labeled as Slide[Number]Stain[Number]_Point[Number]_[TMACoreIndex], where the slide number and stain number correspond to the slide and day of staining, the point number corresponds to the order in which the images were collected for each slide, and the TMA core index corresponds to the ID of the tissue microarray core. Each folder contains single-channel TIFFs for each marker. See paper for details.</p> <p>2. SegmentationOutput.zip: Segmentation output of MIBI-TOF images</p> <p>Cell segmentation was performed using Mesmer (Greenwald NF, Nature Biotechnology 2021, https://www.deepcell.org/predict). Output of Mesmer that delineates the single cells in each of the images is included here. Naming convention is the same as above.</p> <p>3. DataTables.zip: Data tables that are needed to run mpi_ppp_ihc_regression.ipynb</p> <p>Contains MIBI-TOF data (ionpath_processed_data.csv), MIBI-TOF calibration data (calibration_data.csv), IHC data (ihc_data.csv), and a map of each sample to its tissue type (tissue_data.csv). Also includes cell table output from Mesmer with the cell clusters appended to the table (cell_table_size_normalized_clusters.csv).</p>
CODEX multiplexed imaging cell datasets used for using STELLAR to transfer cell type annotations to other tissues and donors
<p>We performed CODEX (co-detection by indexing) multiplexed imaging on 24 sections of the human intestine from 3 donors (B004, B005, B006) using a panel of 47 oligonucleotide-barcoded antibodies. We also performed CODEX imaging on both human tonsil and Barrett's esophagus (BE) using a panel of 57 oligonucleotide-barcoded antibodies. Subsequently images underwent standard CODEX image processing (tile stitching, drift compensation, cycle concatenation, background subtraction, deconvolution, and determination of best focal plane), single cell segmentation, and column marker z-normalization by tissue. Output of this process were dataframes of 870,000 cells and 220,000 cells respectively with fluorescence values quantified from each marker.</p>
Colorectal Cancer Histology Image Tiles for Tissue Multi-class Classification
<p><strong>Content</strong></p> <p>The present dataset is linked to a research aimed at discovering the best normalization pipeline and classification model for colorectal cancer multi-class tissue classification.<br> The 15,856 histological image tiles are completely anomized and are extracted from 10 formalin-fized paraffine-embedded samples of patients affected by colorectal cancer.</p> <p>The materials are inside the following zip file:</p> <p>“CRC_Tiles_IRCCS_ISTITUTO_TUMORI_BARI.zip”: a zipped folder containing tiles (n=15,856) annotated by a pathologist, grouped in 6 subdirectories, each of them representing a class. Tiles are of size 224 x 224 px, taken at a resolution of 0.5 μm/px.<br> <br> <strong>Ethical Statement</strong></p> <p>The study has been funded by “Tecnopolo per la Medicina di Precisione (CUP B84I18000540002)”. The institutional Ethic Committee approved the study (Prot n. 780/CE).</p> <p><br> <strong>Related Datasets and Works</strong><br> <br> For further details concerning the aforementioned dataset, refer to the papers below. <br> Please cite the following articles if you need this dataset for your research.</p> <p>Altini N. et al. (2021) Multi-class Tissue Classification in Colorectal Cancer with Handcrafted and Deep Features. In: Huang DS., Jo KH., Li J., Gribova V., Bevilacqua V. (eds) Intelligent Computing Theories and Application. ICIC 2021. Lecture Notes in Computer Science, vol 12836. Springer, Cham. <br> https://doi.org/10.1007/978-3-030-84522-3_42</p> <p>Altini, N., Marvulli, T. M., Zito, F. A., Caputo, M., Tommasi, S., Azzariti, A., ... & Bevilacqua, V. (2023). The Role of Unpaired Image-to-Image Translation for Stain Color Normalization in Colorectal Cancer Histology Classification. <em>Computer Methods and Programs in Biomedicine</em>, 107511. <br> <a href="https://doi.org/10.1016/j.cmpb.2023.107511">https://doi.org/10.1016/j.cmpb.2023.107511</a></p> <p>Please also consider the dataset offered in our previous work:</p> <p>Altini N. et al. (2021). Pathologist's Annotated Image Tiles for Multi-Class Tissue Classification in Colorectal Cancer (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4785131</p>
WebMicroscope's Deep Learning AI platform automates image analyses with an approach that is faster and able to understand tissue context, which reduces steps needed for accurate results. Researchers can gain access to digitized samples, such as this image of breast-cancer tissue (left), and analyze results through the cloud platform anywhere, anytime. This is a whole slide image of a tissue section of an adrenal gland (right). Fimmic's WebMicroscope cloud platform allows researchers to manage, share, and view digital gigapixel images with any modern browser. Researchers can rapidly pan, zoom, and analyze a digital sample. Photographs: Courtesy of Fimmic Oy. in Deep learning brings speed, accuracy to the life sciences.
WebMicroscope's Deep Learning AI platform automates image analyses with an approach that is faster and able to understand tissue context, which reduces steps needed for accurate results. Researchers can gain access to digitized samples, such as this image of breast-cancer tissue (left), and analyze results through the cloud platform anywhere, anytime. This is a whole slide image of a tissue section of an adrenal gland (right). Fimmic's WebMicroscope cloud platform allows researchers to manage, share, and view digital gigapixel images with any modern browser. Researchers can rapidly pan, zoom, and analyze a digital sample. Photographs: Courtesy of Fimmic Oy.
Multi-antigen imaging data of skin tissue samples from cutaneous T cell lymphoma, atopic dermatitis, and psoriasis patients
<h1>Data</h1> <p>Multi-antigen imaging data for 69 skin tissue samples (21 cutaneous T cell lymphoma (CTCL), 23 atopic dermatitis (AD), and 25 pseudolymphoma (PSO)), obtained from 27 patients (8 CTCL, 7 AD, 12 PSO) treated at the University Hospital Erlangen. Each sample contains at least 36 protein channels, each of resolution 512 X 512 pixels. The data were generated using multi-epitope ligand cartography (MELC) (https://doi.org/10.1038/nbt1250, https://doi.org/10.1007/3-540-36459-5_8).</p> <h1>Metadata</h1> <p>For each sample, we information on sex, age, and condition of the corresponding patient. Samples and patients are numbered as S01 through S69 and P01 through P27, respectively.</p> <h1>Structure of the repository</h1> <ul> <li>The file <strong>metadata.csv</strong> contains the metadata.</li> <li>The archive <strong>data.zip</strong> contains one directory for each sample, named with the sample numbers S01 through S69.</li> <li>The directories for the individual samples contain images named<strong> <marker>-<flourescent protein>.tif</strong>, where <marker> is the name of the quantified protein marker and <flourescent protein> is the name of the flourescent protein that was used for image acquisition.</li> </ul> <p> </p> <p> </p>
Image Datasets for "Virtual tissue microstructure reconstruction across species using generative deep learning"
<p>Training and velautaion image datasets used in the manuscript "Virtual tissue microstructure reconstruction across species using generative deep learning"</p>
Accelerating Whole-Sample Polarization-Resolved Second Harmonic Generation imaging in Mammary Gland Tissue via Generative Adversarial Networks
<p>Authors:</p> <p>Arash Aghigh, Jysiane Cardot, Melika Saadat Mohammadi, Gaëtan Jargot, Heide Ibrahim, Isabelle Plante, François Légaré</p> <p>Affiliations:</p> <p> 1. Centre Énergie Matériaux Télécommunications, Institut National de la Recherche Scientifique, Varennes, Québec, Canada.<br> 2. Centre Armand-Frappier Santé Biotechnologie, Institut National de la Recherche Scientifique, Laval, Québec, Canada.</p> <p>Corresponding Author:</p> <p>Arash Aghigh, arash.aghigh@inrs.ca</p> <p>Description:</p> <p>This dataset accompanies the research on improving whole-sample Polarization-Resolved Second Harmonic Generation (P-SHG) imaging in mammary gland tissue using Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN). The novel approach significantly reduces imaging time while maintaining high image quality and analytical accuracy, demonstrating a reduction in imaging time by more than 95%. This method also minimizes laser-induced photodamage, lowers costs of optical components, and increases the accessibility and applicability of P-SHG imaging in various fields.</p> <p>Keywords:</p> <p>Polarization-Resolved Second Harmonic Generation, P-SHG, Generative Adversarial Networks, GAN, ESRGAN, Mammary Gland Imaging, Super-Resolution, Image Upscaling, Deep Learning, Biomedical Imaging</p> <p>Funding Information:</p> <p> • Canada Foundation for Innovation<br> • Fonds de recherche du Québec–Nature et technologies<br> • Natural Sciences and Engineering Research Council of Canada<br> • New Frontiers Research Fund<br> • NSERC CREATE program (scholarship for Arash Aghigh)</p> <p>Related Identifiers:</p> <p> • GitHub repository for ChaiNNer program: https://github.com/chaiNNer-org/chaiNNer<br> • Download links for models used: https://openmodeldb.info</p> <p>Additional Information:</p> <p>Animal studies were conducted according to the procedures provided by the Canadian Council on Animal Care. The protocol (2005-02) was reviewed and approved by the Institutional Committee for Animal Protection of the Laboratoire National de Biologie Expérimentale (LNBE), the animal facilities based at the Institut National de Recherche Scientifique (INRS).</p>
The features of Tissues and Patches for "Predicting microsatellite instabilitiy from histology images with a three-level hierarchical graph fusion model"
<p>This repository contains features and corresponding coordinates of patches and tissues extracted from 430 and 326 histologic images from patients with colorectal and gastric cancers from the TCGA cohort (original whole section SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this library are from formalin-fixed paraffin-embedded (FFPE) diagnostic sections (“DX” on the GDC Data Portal). This blog explains this in detail: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing.</strong></p> <p>All SVS slices were pre-processed as follows.</p> <p>According to “Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer” these histology images were categorized into The histology images were classified as “MSS” (microsatellite stable) or “MSIMUT” (microsatellite unstable or highly mutated) according to “Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer”, which corresponds to the division of the training and test sets in the article.<br><br></p> <p>Patches were extracted at 40x objective magnification and 20x objective magnification, respectively, and the corresponding features were extracted by pre-training resnet48, respectively</p> <p>The features of Tissues are thumbnails obtained at 2.5x objective magnification and further extracted by MedSAM after extracting the masks of the tissues.</p>
GTEx: DICOM converted whole slide hematoxylin and eosin stained images from the Genotype-Tissue Expression (GTEx) Project
<p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute <a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=gtex" target="_blank" rel="noopener">GTEx</a>. You can use the manifests included in this Zenodo record to download the content of the collection following the <strong>Download instructions</strong> below.</p> <h3>Collection description</h3> <p>The<a href="https://commonfund.nih.gov/GTEx"> Genotype-Tissue Expression (GTEx) Project</a> established a data resource and tissue bank to study the relationship between genetic variants and gene expression in multiple human tissues and across individuals. The project included contributions from numerous groups with diverse expertise in biospecimen collection and processing, pathology review, molecular analysis, and data management. The contributors are collectively called the GTEx Consortium.</p> <p>GTEx collected a total of 26,468 unique tissue samples from 50+ different tissue types, from 956 healthy postmortem donors. The standardized biospecimen collection and analysis practices applied during the study served to minimize preanalytical variability associated with specimen-related factors and their potential impact on analytic endpoints. Each GTEx tissue was divided into two tissue blocks, one for histology and one for molecular analysis; both tissue blocks were preserved in PAXgene Tissue Fixative (Qiagen) solution for 6 to 24 hours, followed by PAXgene Tissue Stabilizer (Qiagen) as specified in the project-specific<a href="https://biospecimens.cancer.gov/resources/sops/library.asp"> standard operating procedures</a>. Tissue blocks were processed and embedded in paraffin at the GTEx central repository at the Van Andel Institute (MI) and hematoxylin and eosin–stained slides were generated from all GTEx donors. Digitally scanned whole slide images of PAXgene-fixed/stabilized, paraffin-embedded tissue sections were created using Aperio Scanscope software (Leica Biosystems). The digital images were then reviewed and annotated by one of four board-certified pathologists assigned to the GTEx study. There are a total of 25,503 digital histology images in the GTEx collection.</p> <p>GTEx was supported by the NIH Common Fund (2010 – 2019). Additional resources include the<a href="https://gtexportal.org/home/biobank"> GTEx Biobank</a>, the<a href="https://gtexportal.org/home/"> GTEx Portal</a>, and the full dataset at<a href="https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000424.v9.p2"> dbGaP</a> (accession number phs000424).</p> <p>Please refer to the listed GTEx publications below for more details [2-7]. </p> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example, <code>collection_id-idc_v8-aws.s5cmd</code> corresponds to the contents of the <code>collection_id</code> collection introduced in IDC data release v8. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <ol> <li><code>gtex-idc_v19-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>gtex-idc_v19-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>gtex-idc_v19-dcf.dcf</code>: Gen3 manifest (for details see <a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in <code>-aws.s5cmd</code> reference files stored in Amazon Web Services (AWS) buckets, while <code>-gcs.s5cmd</code> reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using <code>.s5cmd</code> manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the <code>.s5cmd</code> manifest file: <code>idc download manifest.s5cmd</code></li> </ol> <p>To download the files using <code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <div>Please acknowledge the GTEx Consortium in any published work that includes the images. A sample statement for the acknowledgment of the Genotype-Tissue Expression (GTEx) Project dataset(s) follows.</div> <p>The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health (<a href="http://commonfund.nih.gov/GTEx" target="_blank" rel="noopener">commonfund.nih.gov/GTEx</a>). Additional funds were provided by the NCI, NHGRI, NHLBI, NIDA, NIMH, and NINDS. Donors were enrolled at Biospecimen Source Sites funded by NCI/Leidos Biomedical Research, Inc. subcontracts to the National Disease Research Interchange (10XS170), Roswell Park Cancer Institute (10XS171), and Science Care, Inc. (X10S172). The Laboratory, Data Analysis, and Coordinating Center (LDACC) was funded through a contract (HHSN268201000029C) to the Broad Institute of MIT and Harvard. Biorepository operations were funded through a Leidos Biomedical Research, Inc. subcontract to Van Andel Research Institute (10ST1035). Additional data repository and project management were provided by Leidos Biomedical Research, Inc. (HHSN261200800001E). The Brain Bank was supported with supplements to University of Miami grant DA006227. Statistical Methods development grants were made to the University of Geneva (MH090941& MH101814), the University of Chicago (MH090951, MH090937, MH101825, & MH101820), the University of North Carolina - Chapel Hill (MH090936), North Carolina State University (MH101819), Harvard University (MH090948), Stanford University (MH101782), Washington University (MH101810), and to the University of Pennsylvania (MH101822).</p> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <h3>References</h3> <p>[1] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. & Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence. <em>Radiographics</em> <strong>43,</strong> (2023).</p> <p>[2] Sobin, L., Barcus, M., Branton, P. A., Engel, K. B., Keen, J., Tabor, D., Ardlie, K. G., Greytak, S. R., Roche, N., Luke, B., Vaught, J., Guan, P. & Moore, H. M. Histologic and quality assessment of genotype-Tissue Expression (GTEx) research samples: A large postmortem tissue collection. Arch. Pathol. Lab. Med. (2024). doi:<a href="http://dx.doi.org/10.5858/arpa.2023-0467-OA">10.5858/arpa.2023-0467-OA</a></p> <p>[3] GTEx Consortium. The Genotype-Tissue Expression (GTEx) project. Nat. Genet. 45, 580–585 (2013).</p> <p>[4] GTEx Consortium. Human genomics. The Genotype-Tissue Expression (GTEx) pilot analysis: multitissue gene regulation in humans. Science 348, 648–660 (2015).</p> <p>[5] GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318–1330 (2020).</p> <p>[6] Carithers, L. J., Ardlie, K., Barcus, M., Branton, P. A., Britton, A., Buia, S. A., Compton, C. C., DeLuca, D. S., Peter-Demchok, J., Gelfand, E. T., Guan, P., Korzeniewski, G. E., Lockhart, N. C., Rabiner, C. A., Rao, A. K., Robinson, K. L., Roche, N. V., Sawyer, S. J., Segrè, A. V., Shive, C. E., Smith, A. M., Sobin, L. H., Undale, A. H., Valentino, K. M., Vaught, J., Young, T. R., Moore, H. M. & GTEx Consortium. A novel approach to high-quality postmortem tissue procurement: The GTEx project. Biopreserv. Biobank. 13, 311–319 (2015).</p> <p>[7] Branton, P. A., Sobin, L., Barcus, M., Engel, K. B., Greytak, S. R., Guan, P., Vaught, J. & Moore, H. M. Notable histologic findings in a ‘normal’ cohort: The National Institutes of Health Genotype-Tissue Expression (GTEx) project. Arch. Pathol. Lab. Med. (2024). doi:<a href="http://dx.doi.org/10.5858/arpa.2023-0468-OA">10.5858/arpa.2023-0468-OA</a></p>
Intravital imaging data of resident tissue macrophages in the peritoneum
<p>Imaging data for the manuscript titled 'Cellular morphodynamics as quantifiers for functional states of resident tissue macrophages <em>in vivo</em>', to be submitted to PLOS Computational Biology</p>
Arctique - ARtificial Colon Tissue Images for Qualitative Uncertainty Evaluation
<p>This dataset was introduced and published in the NeurIPS 2024 paper, <em>"Arctique: An Artificial Histopathological Dataset Unifying Realism and Controllability for Uncertainty Quantification"</em>. It includes various versions, in particular a version containing 50,000 training images and 1,000 test images, each paired with corresponding instance and semantic masks, as well as 400 additional variations across 50 selected test images to support research on uncertainty quantification.</p> <p><strong>Versions:</strong></p> <ul> <li><strong>Version v3</strong>: The version used for experimental results presented in the associated research paper. This dataset features improved realism and includes the core images and labels as in <strong>v2</strong>, supplemented with noise-augmented variations specifically used to evaluate algorithmic performance under challenging conditions. It comprises 1,500 synthetic images<br>without noise, 1,500 with added noise, and 1,500 depth mask images, along with a range of noisy variations.</li> <li><strong>Version v2</strong>: The full dataset, consisting of 50,000 training images and 1,000 test images, along with their associated instance and semantic masks. Additionally, this version includes 400 augmented variations for 50 selected test images.</li> <li><strong>Version v1</strong>: A small example subset of the dataset provided for review purposes, containing a limited number of images and annotations to allow preliminary exploration.</li> </ul> <p>Each version of the Arctique dataset is split into training and test sets and variations, each containing the following directories:</p> <p>The <strong>images </strong>directory contains all synthetically generated images stored as PNG files. Each image has a resolution of 512x512 pixels with RGB channels and is named "img_<ID>", where <ID> is a unique integer identifier for each image.</p> <p>The <strong>masks </strong>directory includes subdirectories containing various masks related to the images:</p> <ul> <li><em>cytoplasm</em>: Contains 2D semantic masks for the cell cytoplasm. Each mask corresponds to an image named "<ID>.tif", where "<ID>" is the identifier for that image. The mask file is named using the same identifier.</li> <li><em>instance_3d</em>: Contains a directory for each image, named "<ID>. Inside each directory, there is a 3D stack numpy file representing the instance IDs in a 3D volumetric array. Additionally, it includes a sequence of 2D instance segmentation masks, named "slice_<ID>_<slice_count>.png", each representing equidistant slices through the 3D volume along the depth axis.</li> <li><em>instance</em>: Contains 2D instance masks for the cell nuclei. Each mask corresponds to an image named "<ID>.tif", and the mask file is named with the same identifier.</li> <li><em>semantic</em>: Contains 2D semantic masks for the cell nuclei. Similar to the instance masks, each mask corresponds to an image named "<ID>.tif", with the mask file named using the same identifier.</li> </ul> <p>Note that all semantic masks appear as black images when viewed with a standard image viewer. This is because the cell type IDs, ranging from 1 to 5, are used as greyscale values, which appear dark in the images. The modelled cell types are:</p> <p>Cell Types</p> <table> <tbody> <tr> <td>1</td> <td>Epithelial Cells</td> <td>'EPI'</td> </tr> <tr> <td>2</td> <td>Plasma Cells</td> <td>'PLA'</td> </tr> <tr> <td>3</td> <td>Lymphocytes</td> <td>'LYM'</td> </tr> <tr> <td>4</td> <td>Eosinophils</td> <td>'EOS'</td> </tr> <tr> <td>5</td> <td>Fibroblasts</td> <td>'FIB'</td> </tr> </tbody> </table> <p> </p> <p>The <strong>metadata </strong>directory contains JSON metadata files named "metadata_<ID>" for each image. Each JSON file includes a list of Python dictionaries, one for each cell object visible in the image. Consider submission appendix F for a detailed explanation of each dictionary.</p> <p>The <strong>parameters </strong>directory contains JSON files named "parameters_<ID>", which detail the parameters used to generate each image. Each JSON file is a Python dictionary with all the parameter values necessary to reproduce the scene. Consider submission appendix F for a detailed explanation of each parameter.</p>
Source Data for "Imaging biological tissue with high-throughput single-pixel compressive holography"
<p>This file contains five subfolders, which are archived with relevant data that are necessary for reconstructing the holographic images of biological samples and resolution targets, respectively. <br> Here we introduce in order:<br> 1. 'dataset 1' is prepared for holographic reconstruction of stained tissue from mouse tails;<br> 2. 'dataset 2' is provided for holographic reconstruction of 80-um unstained tissue from mouse brains.<br> 3. 'dataset 3' is provided for verification of amplitude resolution in large-FOV mode;<br> 4. 'dataset 4' is provided for verification of amplitude resolution in high-resolution mode;<br> 5. 'dataset 5' is provided for verification of phase resolution in high-resolution mode;<br> 6. 'additional dataset 1' is prepared for additional holographic reconstruction of another stained tissue from mouse tails;<br> 7. 'additional dataset 2' is prepared for additional holographic reconstruction of 100-um unstained tissue from mouse brains;<br> 8. 'additional dataset 3' is prepared for additional holographic reconstruction of 120-um unstained tissue from mouse brains;<br> 9. 'additional dataset 4' is prepared for additional holographic reconstruction of 10-um unstained tissue from mouse brains;</p> <p>Both subfolders have the same structures, including the MATLAB data and raw data collected from the data acquisition card, which are necessary for holographic imaging reconstruction.<br> Here we introduce in order:<br> *) biological_sample.mat: The raw data of imaging biological sample. The format of the data has been converted from .tdms to .mat file.</p> <p>*) target_sample.mat: The raw data of imaging resolution target. The format of the data has been converted from .tdms to .mat file.</p> <p>*) background_curvature.mat: The raw data used to correct for phase contaminations from system aberrations. The format of the data has been converted from .tdms to .mat file.</p> <p>*) biological_sample_rawdata.tdms: The raw data of imaging biological sample. The data was collected through DAC and was in the format of TDMS.</p> <p>*) target_sample_rawdata.tdms: The raw data of imaging resolution target. The data was collected through DAC and was in the format of TDMS.</p> <p>*) background_curvature_rawdata.tdms: The raw data used to correct for phase contaminations from system aberrations. The data was collected through DAC and was in the format of TDMS.<br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.