Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

979

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

979 results for “image dataset”

Learn how ShareScore rates datasets ↗
zenodo36/100

Test Dataset for 3D semantic image segmentation of the various organs from CT and MR scans

<p>These test cases are for the <a href="https://github.com/MIC-DKFZ/nnUNet/releases/tag/v1.7.1">nnUnet v1</a> models trained on the following datasets:<br><br></p> <table> <tbody> <tr> <td>Dataset&nbsp;</td> <td>Task</td> <td>Model Details on Zenodo</td> </tr> <tr> <td>&nbsp;<a href="../record/6802614">TotalSegmentator</a>&nbsp;and&nbsp;<a href="../record/5903672">FLARE21</a> datasets</td> <td>Segment Liver from CT scans</td> <td>https://zenodo.org/record/8274976</td> </tr> <tr> <td><a href="https://kits-challenge.org/kits23/">KiTS23</a> datasets and a subset of the<a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=5800386#5800386566e265abf95408aa64c4917f0cbe5d9">&nbsp;TCGA-KIRC&nbsp;</a>dataset</td> <td>Segment Kidney, Cyst, and Tumors from CT Scans</td> <td>https://zenodo.org/records/8277846</td> </tr> <tr> <td><a href="http://ji%20yuanfeng.%20(2022).%20amos%20a%20large-scale%20abdominal%20multi-organ%20benchmark%20for%20versatile%20medical%20image%20segmentation%20[data%20set].%20zenodo.%20https">AMOS</a>&nbsp;and&nbsp;<a href="http://macdonald,%20jacob%20a.,%20zhu,%20zhe,%20konkel,%20brandon,%20mazurowski,%20maciej,%20wiggins,%20walter,%20&amp;%20bashir,%20mustafa.%20(2020).%20duke%20liver%20dataset%20(mri)%20v2%20(2.0.0)%20[data%20set].%20zenodo.%20https//doi.org/10.5281/zenodo.7774566">DUKE Liver</a> datasets</td> <td>Segment Liver from the MR scans</td> <td>https://zenodo.org/record/8290124</td> </tr> <tr> <td>Data from m&nbsp;<a href="../record/6624726">pi-cai</a></td> <td>Segment Prostate region from MR scans</td> <td>https://zenodo.org/record/8290093</td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

A dataset of 1600 images extracted from 5 cm RGB orthophotos for the classification of 12 classes of roofing materials

<p>This dataset contains a collection of 1601 image tiles of 64x64 pixels (3.2x3.2m&sup2;) annotated for 12 roofing materials. These tiles were extracted from 5 cm RGB orthophotos acquired by the city of Namur (Belgium) in 2017. The additional data used to create this dataset are (a) a Namur roof section mask, and (b) a set of 1601 material samples acquired using stratified random sampling. The tiles were obtained as follows: the centroid of each roof section containing a sample is used to extract tiles. A size of 64x64 pixels has been chosen so that a tile contains information for only one roof section, in order to learn only the colour and texture of the roof materials. This also avoids adding information outside the given roof section. The tiles are thus extracted for each orthophoto spectral band and labelled with the identifier of the class of roofing materials to which they belong. Here are the 12 material classes considered, preceded by their labels:</p> <p>0- Solar panels<br>1- Brown tiles<br>2- Orange tiles<br>3- Black tiles<br>4- Dark membranes<br>5- White membranes<br>6- Slates containing asbestos<br>7- Slates without asbestos<br>8- Corrugated asbestos-cement sheets<br>9- Gravel<br>10- Vegetation<br>12- Metals</p> <p>There are approximately 140 tiles by material class except for the vegetated roof sections (class 10) which contains only 47 samples due to its rarety.</p> <p>The dataset contains 1 folder for each spectral band. Each folder contains 1601 thumbnails in tif format named as follows:</p> <p><strong>img[tile id]_[class label].tif</strong></p> <p>It is suggested to apply pre-processing to these images as done by<a href="https://doi.org/10.1109/jurse57346.2023.10144142"> Wyard et al. (2023).</a></p>

opencc-by-nc-sa-4.0Dec 2023View details →
zenodo36/100

dataset- Tuberculosis detection using Squid Game Optimization with Deep Learning Model on Chest X-Ray Images

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo36/100

Neural Image Segmentation for Redacted Text Detection Dataset

<p>This dataset was created for the "<span>Redacted Text Detection Using Neural Image </span><span>Segmentation Methods</span>" project, and contains roughly 1000 pages with manually annotated redactions in Dutch documents released under the WOO, with the trained model files and the model outputs also included in the dataset. More details on the usage of the dataset and models can be found on Github: https://github.com/RubenvanHeusden/NeuralRedactedTextDetection/</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset for: AI-enabled Lorentz microscopy for quantitative imaging of nanoscale magnetic spin textures

<p>Data used in the publication: AI-enabled Lorentz microscopy for quantitative imaging of nanoscale magnetic spin textures</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset of "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders"

<p>Dataset of the paper "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders", published in Experimental Thermal and Fluid Science (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.expthermflusci.2024.111195" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.expthermflusci.2024.111195</a>)</p> <p>The project received funding from: the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (grant agreement No 949085); the National Natural Science Foundation of China (NSFC No 12227803 and No 12372276).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

DeviantArt Image and Metric Dataset: results by topic

<p>Dataset of images stored by deviantart.com website, as a result of the search for 20 topics and the images included in the first 15 pages of each topic,&nbsp;downloaded on 2024-04-14.&nbsp;The topics searched are:&nbsp;</p> <p>"Fantasy art", "Science fiction art", "Anime and manga art", "Fan art (for specific fandoms)", "Digital paintings", "Traditional drawings", "Character designs", "Creature concepts", "Landscape art", "Abstract art", "Surrealism", "Steampunk art", "Cyberpunk art", "Gothic art", "Horror art", "Cosplay photography", "Pixel art", "Concept art", "Comics and graphic novels", "Street art and graffiti".</p> <p>The dataset contains 7067 image records related to these topics.</p> <p>The fields included in the dataset are:<br>- <strong>search_topic</strong>: the image is a search result for this topic<br>- <strong>page_num</strong>: the image appears on this search page number<br>- <strong>image_page</strong>: link to the page with the image information<br>- <strong>image_url</strong>: link to the image<br>- <strong>image_title</strong>: image title<br>- <strong>image_author</strong>: author of the image<br>- <strong>image_favs</strong>: number of times the image has been &ldquo;liked&rdquo;<br>- <strong>image_com</strong>: number of comments that the image has<br>- <strong>image_views</strong>: number of views to the image<br>- <strong>private_collections</strong>: number of times it has been included in a private collection<br>- <strong>tags</strong>: tags that have been assigned to the image to facilitate its discovery<br>- <strong>location</strong>: country or geographical location, if the author wants to identify it<br>- <strong>description</strong>: open text field created by the author. It may include technical details or links to the author's social networks.<br>- <strong>image_px</strong>: dimensions of the image, in pixels.<br>- <strong>image_size</strong>: image weight in MB.<br>- <strong>published_date</strong>: image publication date.<br>- <strong>last_comment</strong>: the last comment added to the image.<br>- <strong>license</strong>: image license.</p> <p>Individual image licenses must be respected by users of this dataset.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Image datasets associated with Gut Analysis Toolbox

<div>The images are sample image datasets associated with the software: <a href="https://gut-analysis-toolbox.gitbook.io/docs/">Gut Analysis Toolbox (GAT)</a>.</div> <div>The dataset contains immunofluorescence images of enteric neurons and glia labeled with different markers.&nbsp; The data is mostly from mouse and human colon or small intestine.&nbsp;</div> <div>Channels corresponding to Hu labelling can be used for segmenting enteric neurons in GAT.&nbsp;</div> <div>Channels corresponding to GFAP (enteric glia) or neurons with markers labelling the cell body and processes (ChAT, Calbindin, Calretinin) can be used as a ganglia marker for segmenting the ganglia</div> <div>The data is two-dimensional (2D) with some images having multiple channels. The data is mostly in tif format, except for one dataset that is czi (Fiji using bioformats or aicspylibczi in Python). Calcium imaging data is 2D+Time.</div> <div>&nbsp;</div> <div>Data curated by:&nbsp;<a href="https://www.linkedin.com/in/rajapradeep/">Pradeep Rajasekhar, Walter and Eliza Hall Institute of Medical Research</a>, Australia</div> <h2><strong>Data from<a href="https://www.monash.edu/mips/themes/drug-discovery-biology/labs/inm"> INM lab, Monash University</a> (mouse images)</strong></h2> <p><strong>Immunofluorescence images</strong></p> <ul> <li>&nbsp;181107_ms_distal_colon_GFAP_Hu_40X.tif <ul> <li>Channel 1: GFAP</li> <li>Channel 2: Hu</li> </ul> </li> </ul> <div> <ul> <li>181107_ms_distal_colon_nNOS_GFAP_Hu_40X.tif (Same as above, but got an extra channel)</li> </ul> </div> <ul> <li> <ul> <li>Channel 1: nNOS</li> <li>Channel 2: GFAP</li> <li>Channel 3: Hu</li> </ul> </li> </ul> <div>In both images above, GFAP can be used as ganglia segmentation channel in GAT using DeepImageJ.</div> <div> <ul> <li>ms_distal_colon_Hu_20X.tif</li> </ul> </div> <div> <ul> <li> <ul> <li>Hu</li> </ul> </li> </ul> </div> <div> <ul> <li>ms_distal_colon_Hu_40X_1.tif</li> </ul> </div> <div> <ul> <li> <ul> <li>Hu</li> </ul> </li> </ul> </div> <div> <ul> <li>Tilescan_GAT_ms_distal_colon_MP_hu.tif</li> </ul> </div> <div> <ul> <li> <ul> <li>Hu</li> </ul> </li> </ul> </div> <h3><strong>Calcium imaging data (video)</strong></h3> <div>&bull; calcium_imaging_mouse_distal_colon_25X.tif</div> <div>Tissue was incubated with calcium dye Fluo8-AM calcium dye (Myenteric wholemount from distal colon of mouse)</div> <div>142 frames in total acquired at 1.162 frames per second. At frame 52, 100 uM of ATP is added which causes enteric glia, neurons and blood vessels to respond. This causes slow drifting in the field of view.</div> <div>&nbsp;</div> <div>&bull; mouse_GCamp_calcium_movement.tif</div> <div>&nbsp;</div> <div>Wnt1-GCaMP3 mouse where GCamP3 is a genetically encoded calcium sensor expressed by enteric neurons and enteric glia</div> <div>Imaging performed at Monash University. 745 frames in total acquired at 1.162 frames per second with drifting over time</div> <div>Tissue source: <a href="https://biomedicalsciences.unimelb.edu.au/sbs-research-groups/anatomy-and-physiology-research/neuroscience/development-of-the-enteric-nervous-system">Stamp &amp; Hao laboratory, University of Melbourne.</a></div> <div>&nbsp;</div> <h2><strong>Images from McQuade Lab, University of Melbourne</strong></h2> <ul> <li>ms_28_wk_colon_DAPI_nNOS_Hu_10X.tif (mouse colon)</li> <li>ms_28_wk_colon_DAPI_nNOS_Hu_10X.tif (mouse ileum) <ul> <li>Channel 1: DAPI</li> <li>Channel 2: nNOS&nbsp;</li> <li>Channel 3: Hu</li> </ul> </li> </ul> <div> <ul> <li>ms_distal_colon_nNOS_Hu_10X.czi&nbsp; &nbsp; (This is a .czi file which can be opened in Fiji using bioformats or aicspylibczi in Python)</li> </ul> </div> <ul> <li> <ul> <li>Channel 1: DAPI</li> <li>Channel 2: nNOS&nbsp;</li> <li>Channel 3: Hu</li> </ul> </li> </ul> <h2><strong>Images from public repository (SPARC)</strong>:</h2> <ul> <li>DYM_22_7_Pr_Chat_BYFP_DIN_GFP-g_nNOS-m_VIP-r_Hu-b.tif is a crop from File 100 05-07-2019 DYM 22 7 Pr Chat%3BYFP DIN GFP-g nNOS-m VIP-r Hu-b (Mouse Proximal Colon)</li> <li>DYM_22_7_Pr_Hu_crop.tif is a crop from above. <ul> <li>Channel 1: Choline acetyltransferase</li> <li>Channel 2: nNOS</li> <li>Channel 3: Calretinin</li> <li>Channel 4: Hu (pan-neuronal marker)</li> </ul> </li> </ul> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Channel 1 and 3 can be used as ganglia segmentation channels in GAT using DeepImageJ.</div> <div>&nbsp;</div> <div> <ul> <li>146_02_14_20DYM8_6_mouse_Mid_Chat-g CalB-r CalR-b_max.tif is from File 146 02-14-20 DYM 8 6 Mid Chat-g CalB-r CalR-b (Mouse mid colon)</li> </ul> </div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; No Hu staining, any channel could be used as ganglia segmentation channel in GAT using DeepImageJ.</div> <div> <ul> <li> <ul> <li>Channel 1: ChAT</li> <li>Channel 2: Calbindin</li> <li>Channel 3: Calretinin</li> </ul> </li> </ul> </div> <h3><strong>Reference</strong>:</h3> <div>Thanks goes to Marthe Howard for depositing the data in the SPARC repository.</div> <div>Howard, M. (2021). 3D imaging of enteric neurons in mouse (Version 1) [Data set]. SPARC Consortium.<a href="https://doi.org/10.26275/9FFG-482D"> https://doi.org/10.26275/9FFG-482D</a></div> <div>**************</div> <h2><strong>Multiplex data (Flinders University)</strong></h2> <div><strong>Multiplexing_H2202Desc_Layer 1_Ganglia1_Hu.zip&nbsp;</strong>is from:</div> <div><a href="https://pubmed.ncbi.nlm.nih.gov/37355216/">Chen, B. N., Humenick, A., Yew, W. P., Peterson, R. A., Wiklendt, L., Dinning, P. G., Spencer, N. J., Wattchow, D. A., Costa, M., &amp; Brookes, S. J. H. (2023). Types of Neurons in the Human Colonic Myenteric Plexus Identified by Multilayer Immunohistochemical Coding. Cellular and molecular gastroenterology and hepatology, 16(4), 573&ndash;605.</a></div> <div>This data is a myenteric wholemount from the descending colon of a Human. It has 14 different markers, 6 different rounds of staining. Every round has pan-neuronal marker Hu as a reference marker.There are 19 images. The filenames follow the convention:</div> <div>&nbsp;</div> <div>H2202Desc_<em>layer num</em>_<em>ganglia num</em>_<em>markername</em>.tif&nbsp;</div> <div><em>H2202&nbsp;</em>is the sample name, <em>Desc </em>means descending colon</div> <div>Here <em>layer num</em> corresponds to the round of staining, so Layer3, means its the 3rd round of staining.</div> <div><em>ganglia num</em> is specified as multiple ganglia can be imaged from same tissue.&nbsp;</div> <div><em>markername</em> corresponds to the marker used.&nbsp;</div> <div>&nbsp;</div> <div>Markers used are: Hu, 5HT, ChAT, NOS, CGRP, Enk, SP, Somat, VACht, NPY, Calbindin, Calretinin, NF, VIP</div> <div>&nbsp;</div> <div><strong>Abbreviations:</strong></div> <div>&nbsp;</div> <ul> <li>Hu: Pan-neuronal marker</li> <li>5HT: Serotonin (5-Hydroxytryptamine)</li> <li>ChAT: Choline acetyltransferase</li> <li>nNOS: neuronal Nitric Oxide Synthase (NOS in these images are actually nNOS)</li> <li>CGRP: Calcitonin Gene-Related Peptide</li> <li>Enk: Enkephalin</li> <li>SP: Substance P</li> <li>Somat: Somatostatin</li> <li>VACht: Vasoactive Intestinal Peptide (VIP)&nbsp;</li> <li>NPY: Neuropeptide Y</li> <li>NF: neurofilament 200&nbsp;</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo36/100

ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models

<p>We introduce a new image-text dataset, providing high-quality natural language descriptions for global-scale satellite data. Specifically, we utilize Sentinel-2 data for its global coverage as the foundational image source, employing semantic segmentation labels from the European Space Agency's WorldCover project to enrich the descriptions of land covers. By conducting in-depth semantic analysis, we formulate detailed prompts to elicit rich descriptions from ChatGPT. We then include a manual verification process to enhance the dataset's quality further. This step involves manual inspection and correction to refine the dataset. Finally, we offer the community ChatEarthNet, a large-scale image-text dataset characterized by global coverage, high quality, wide-ranging diversity, and detailed descriptions. ChatEarthNet consists of 163,488 image-text pairs with captions generated by ChatGPT-3.5 and an additional 10,000 image-text pairs with captions generated by ChatGPT-4V(ision). This dataset has significant potential for both training and evaluating vision-language geo-foundation models for remote sensing.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

[Dataset for] Whole-brain meso-vein imaging in living humans using fast 7 T MRI

<p>This dataset is associated with:</p> <ul> <li>Gulban, Stirnberg, Tse, Pizzuti, Koiso, Archila-Melendez, Huber, Bollmann, Goebel, Kay, Ivanov, 2025. Whole-brain meso-vein imaging in living humans using fast 7 T MRI (Preprint).</li> </ul> <p>This dataset is also used in:</p> <ul> <li>Pizzuti, Bazin, Ivanov, Dresbach, Peter, Goebel, Gulban, 2024.&nbsp; Multimodal laminar characterization of visual areas along the cortical hierarchy (Preprint).</li> </ul> <p>More data are going to be be added as we progress with our manuscripts though their publications or upon request (please contact Omer Faruk Gulban).</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Toxic Sentence Classification Dataset with labels of categories such as religion, mental health, race, sex, body image, disability, physical abuse, and politics

<p>The dataset has a collection of various toxic sentences belonging to different categories. It was collected from various sources. It indicates which category each sentence belongs to. The values of the category columns are binary 1 or 0 indicating whether the sentence belongs to that particular category or not. Each sentence belongs to only 1 category.&nbsp;</p> <p>&nbsp;</p> <p>Columns:<br>1.comment_text: Contains toxic sentences that are insensitive and offensive, focusing on various categories.<br>2.mental_health: Binary value 1 indicates that the sentence focuses on mental health.<br>3.Race:Binary value 1 indicates that the sentence is racist.<br>4.sex:Binary value 1 indicates that the sentence focuses on sexuality.<br>5.body_image:Binary value 1 indicates that the sentence focuses on body image.<br>6.disability:Binary value 1 indicates that the sentence focuses on physical disability and related issues.<br>7.religion:Binary value 1 indicates that the sentence can be triggering to people who are extremely religious.<br>8.physical_abuse:Binary value 1 indicates that the sentence focuses on physical abuse issues.<br>9.politics:Binary value 1 indicates that the sentence focuses on political issues.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Pan-Cancer-Nuclei-Seg-DICOM: DICOM converted Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images

<div> <p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute&nbsp;<a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=Pan-Cancer-Nuclei-Seg-DICOM" target="_blank" rel="noopener">Pan-Cancer-Nuclei-Seg-DICOM</a>. You can use the manifests included in this Zenodo record to download the content of the collection following the <strong>Download instructions</strong>&nbsp;below.</p> <h3>Collection description</h3> </div> <div> <div>This collection contains automatic nucleus segmentation data of 5,060 whole slide tissue images of 10 cancer types earlier published in [2] (<a href="https://doi.org/10.7937/TCIA.2019.4A4DKP9U">https://doi.org/10.7937/TCIA.2019.4A4DKP9U</a>) stored in DICOM Bulk Annotation and DICOM Segmentation formats.</div> <div>&nbsp;</div> <div>DICOM Bulk Annotation nuclei annotations are stored as closed polygons along with the area of each nuclei. DICOM Segmentation version contains binary segmentations obtained by rasterizing the polygon contours.&nbsp;</div> <div>&nbsp;</div> <div>The annotations correspond to digital pathology images from the TCGA-BLCA,TCGA-BRCA,TCGA-CESC,TCGA-COAD,TCGA-GBM,TCGA-LUAD,TCGA-LUSC,TCGA-PAAD,TCGA-PRAD,TCGA-READ,TCGA-SKCM,TCGA-STAD,TCGA-UCEC,TCGA-UVM collections available in NCI Imaging Data Commons.</div> <div>&nbsp;</div> <div>To learn how these files are organized and how to access the content programmatically, see this documentation page: <a href="https://highdicom.readthedocs.io/en/latest/ann.html">https://highdicom.readthedocs.io/en/latest/ann.html</a>.</div> <div>&nbsp;</div> <div>Conversion of the nuclei segmentations from the original format into DICOM ANN and SEG representations was done using the code available in <a href="https://doi.org/10.5281/zenodo.13871765">10.5281/zenodo.10632181</a>.</div> <div>&nbsp;</div> <div>Annotations corresponding to this container ID in the source failed to convert due to the pixel matrix being too large to store:&nbsp; <code>TCGA-OL-A66K-01Z-00-DX1</code></div> <div>&nbsp;</div> <div>The following container IDs from the source annotations have failed due to inability to find the annotated images using the container IDs:</div> <div> <pre><code>TCGA-CU-A3QU-01Z-00-DX1 TCGA-A2-A0D1-01Z-00-DX1 TCGA-AQ-A1H2-01Z-00-DX1 TCGA-AQ-A1H2-01Z-00-DX1 TCGA-AQ-A1H3-01Z-00-DX1 TCGA-AQ-A1H3-01Z-00-DX1 TCGA-BH-A0B2-01Z-00-DX1 TCGA-E2-A15E-01Z-00-DX1 TCGA-E2-A1IP-01Z-00-DX1 TCGA-F4-6857-01Z-00-DX1 TCGA-12-0773-01Z-00-DX4 TCGA-35-3621-01Z-00-DX1 TCGA-49-4486-01Z-00-DX1 TCGA-33-4587-01Z-00-DX1 TCGA-D9-A1X3-01Z-00-DX1 TCGA-D9-A1X3-01Z-00-DX2 TCGA-D9-A4Z6-01Z-00-DX1 TCGA-EE-A17Y-01Z-00-DX1 TCGA-EE-A29R-01Z-00-DX1 TCGA-EE-A2A0-01Z-00-DX1 TCGA-EE-A2MS-01Z-00-DX1 TCGA-ER-A199-01Z-00-DX1 TCGA-ER-A1A1-01Z-00-DX1 TCGA-ER-A2NC-01Z-00-DX1 TCGA-FS-A1Z7-06Z-00-DX10 TCGA-FS-A1Z7-06Z-00-DX11 TCGA-FS-A1Z7-06Z-00-DX12 TCGA-FS-A1Z7-06Z-00-DX13 TCGA-FS-A1ZN-01Z-00-DX10 TCGA-FS-A1ZN-01Z-00-DX11 TCGA-FS-A1ZW-06Z-00-DX10 TCGA-FS-A1ZW-06Z-00-DX11 TCGA-GN-A261-01Z-00-DX1 TCGA-GN-A266-01Z-00-DX1 TCGA-GN-A268-01Z-00-DX1 TCGA-GN-A26A-01Z-00-DX1 TCGA-XV-AB01-01Z-00-DX1 TCGA-AJ-A23O-01Z-00-DX1 TCGA-AP-A056-01Z-00-DX1 TCGA-BK-A139-01Z-00-DX1 TCGA-E6-A1M0-01Z-00-DX1</code></pre> </div> <div> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example,&nbsp;<code>pan_cancer_nuclei_seg_dicom-collection_id-idc_v19-aws.s5cmd</code> corresponds to the annotations for th eimages in the <code>collection_id</code> collection introduced in IDC data release v19. DICOM Binary segmentations were introduced in IDC v20. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <p>For each of the collections, the following manifest files are provided:</p> <ol> <li><code>pan_cancer_nuclei_seg_dicom-&lt;collection_id&gt;-idc_v20-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>pan_cancer_nuclei_seg_dicom-&lt;collection_id&gt;-idc_v20-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>pan_cancer_nuclei_seg_dicom-&lt;collection_id&gt;-idc_v20-dcf.dcf</code>: Gen3 manifest (for details see&nbsp;<a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in&nbsp;<code>-aws.s5cmd</code>&nbsp;reference files stored in Amazon Web Services (AWS) buckets, while&nbsp;<code>-gcs.s5cmd</code> reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using&nbsp;<code>.s5cmd</code>&nbsp;manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the&nbsp;<code>.s5cmd</code>&nbsp;manifest file:&nbsp;<code>idc download&nbsp;manifest.s5cmd</code></li> </ol> <p>To download the files using&nbsp;<code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <h3>References</h3> </div> </div> <div>[1] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. &amp; Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence. Radiographics 43, (2023).</div> <div>&nbsp;</div> <div>[2] Hou, L., Gupta, R., Van Arnam, J. S., Zhang, Y., Sivalenka, K., Samaras, D., Kurc, T., &amp; Saltz, J. H. (2019). Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images of 10 Cancer Types [Data set]. The Cancer Imaging Archive. https://doi.org/10.7937/TCIA.2019.4A4DKP9U</div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Regulatory T cell therapy is associated with distinct immune regulatory lymphocytic infiltrates in kidney transplants: Spatial transcriptomic dataset and images

<p>The outputs of the NanoString GeoMx DSP platform were concatenated into three xlsx files, each illustrating a separate experiment along with their sample annotations. This technique analyzes protein or RNA abundance within regions of interest (ROIs) or specific cell segments selected based on histological features and immunofluorescence. In this repository, the concatenated GeoMx output files are presented, along with PowerPoint presentations for each biopsy that show immunofluorescence images of the selected ROIs and/or cell segments.</p> <ul> <li><strong>Protein_Full ROI:</strong> This experiment measured the abundance of 41 proteins in discrete regions of interest (ROIs) within transplant kidney biopsies.</li> <li><strong>Protein_Rare cell:</strong> This experiment measured the abundance of 40 proteins in specific cell segments, such as CD4+FoxP3- cells vs. CD4+FoxP3+ cells, within transplant kidney biopsies.</li> <li><strong>RNA:</strong> This experiment measured the abundance of 90 genes in discrete ROIs within transplant kidney biopsies.</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Optimizing Deep Learning Models for Aflatoxin Detection: A Case of Artificial Intelligence-Driven Classified Groundnut Image Datasets for Postharvest Management

<p><strong>DATASET DESCRIPTION&nbsp;</strong><br>This dataset comprises a curated collection of classified groundnut images, specifically designed for deep learning applications in aflatoxin detection. The dataset is organized into four distinct categories: Healthy, Moldy, Insect-Infested, and Physiological Disorder, making it a vital resource for training AI and machine learning models aimed at advancing agricultural research. These classifications are crucial for the development of AI-driven solutions addressing aflatoxin contamination, enhancing crop quality assessments, and improving postharvest management practices.<br>The dataset has been developed to support research in agricultural Artificial Intelligence (AI), machine learning (ML), and food safety, with a focus on aiding resource-constrained regions in combating postharvest losses due to contamination. By leveraging this dataset, researchers can contribute to safeguarding public health, promoting food security, and supporting smallholder farmers.</p> <p><strong>POTENTIAL APPLICATIONS</strong><br>This dataset provides numerous opportunities for innovation in agriculture through AI and deep learning technologies. Its key applications include:<br><strong>Early Aflatoxin Detection</strong>: Facilitates the development of AI-powered models for prompt identification of aflatoxins in groundnuts, helping mitigate associated health risks.<br><strong>Postharvest Management Improvement</strong>: Enables the creation of innovative solutions to enhance storage, handling, and processing, reducing contamination and losses.<br><strong>Food Safety and Quality Assurance</strong>: Strengthens agricultural value chains by supporting the production of safe and high-quality food products.</p> <p><strong>BROADER IMPACT</strong><br>This resource is invaluable for fostering AI innovation in agriculture, particularly in resource-limited environments. It addresses critical challenges such as postharvest losses and food contamination while contributing to global efforts in sustainable agricultural development. By utilizing this dataset, researchers can improve food security, support smallholder farmers, and drive advancements in agricultural practices that benefit both local and global communities.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

"WLRI-HRC" - A Dataset of Infrared Images for Human-Robot Collaboration in Manufacturing Environment

<p>This repository contains all needed data sets for the contribution in&nbsp; Journal of Sensors and Sensor Systems&nbsp; "Enhancing human&ndash;robot collaboration with thermal images and deep neural networks: the unique thermal industrial dataset WLRI-HRC and evaluation of convolutional neural networks". You may use this data for scientific, non-commercial purposes, provided that you give credit to the owners when publishing any work based on this data.</p> <p><strong>DOI: 10.5194/jsss-14-37-2025</strong></p> <p>&nbsp;</p> <p><strong>or as BibTex:</strong></p> <div> <div>@article{sume_enhancing_2025,</div> <div>&nbsp; &nbsp; title = {Enhancing human&ndash;robot collaboration with thermal images and deep neural networks: the unique thermal industrial dataset {WLRI}-{HRC} and evaluation of convolutional neural networks},</div> <div>&nbsp; &nbsp; volume = {14},</div> <div>&nbsp; &nbsp; issn = {2194-8771},</div> <div>&nbsp; &nbsp; shorttitle = {Enhancing human&ndash;robot collaboration with thermal images and deep neural networks},</div> <div>&nbsp; &nbsp; url = {https://jsss.copernicus.org/articles/14/37/2025/},</div> <div>&nbsp; &nbsp; doi = {10.5194/jsss-14-37-2025},</div> <div>&nbsp; &nbsp; abstract = {This contribution introduces the use of convolutional neural networks to detect humans and collaborative robots (cobots) in human&ndash;robot collaboration (HRC) workspaces based on their thermal radiation fingerprint. The unique data acquisition includes an infrared camera, two cobots, and up to two persons walking and interacting with the cobots in real industrial settings. The dataset also includes different thermal distortions from other heat sources. In contrast to data from the public environment, this data collection addresses the challenges of indoor manufacturing, such as heat distortions from the environment, and allows for it to be applicable in indoor manufacturing. The Work-Life Robotics Institute HRC (WLRI-HRC) dataset contains 6485 images with over 20 000 instances to detect. In this research, the dataset is evaluated for implementation by different convolutional neural networks: first, one-stage methods, i.e., You Only Look Once (YOLO v5, v8, v9 and v10) in different model sizes and, secondly, two-stage methods with Faster R-CNN with three variants of backbone structures (ResNet18, ResNet50 and VGG16). The results indicate promising results with the best mean average precision at an intersection over union (IoU) of 50 (mAP50) value achieved by YOLOv9s (99.4 \%), the best mAP50-95 value achieved by YOLOv9s and YOLOv8m (90.2 \%), and the fastest prediction time of 2.2 ms achieved by the YOLOv10n model. Further differences in detection precision and time between the one-stage and multi-stage methods are discussed. Finally, this paper examines the possibility of the Clever Hans phenomenon to verify the validity of the training data and the models&rsquo; prediction capabilities.},</div> <div>&nbsp; &nbsp; language = {English},</div> <div>&nbsp; &nbsp; number = {1},</div> <div>&nbsp; &nbsp; journal = {Journal of Sensors and Sensor Systems},</div> <div>&nbsp; &nbsp; author = {S&uuml;me, Sinan and Ponomarjova, Katrin-Misel and Wendt, Thomas M. and Rupitsch, Stefan J.},</div> <div>&nbsp; &nbsp; month = feb,</div> <div>&nbsp; &nbsp; year = {2025},</div> <div>&nbsp; &nbsp; note = {Publisher: Copernicus GmbH},</div> <div>&nbsp; &nbsp; pages = {37--46},</div> <div>}</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo36/100

(12)-Pereyra2024A-DS0001--0009 – Nine Tribolium castaneum long-term live imaging datasets of embryonic development acquired with light sheet fluorescence microscopy

<p>(12)-Pereyra2024A-DS0001--0009 &ndash; Nine <em>Tribolium castaneum</em> long-term live imaging datasets of embryonic development acquired with light sheet fluorescence microscopy</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Microscopy images and datasets of Sphingomonas and Methylobacterium on Arabidopsis leaves

<p>This repository contains the supplemental material and raw data to analyse the population density at the CFU-level and single cell-resolution, and spatial distribution of bacterial communities composed of&nbsp;<em>Methylobacterium</em> and/or <em>Sphingomonas</em> species on <em>Arabidopsis thaliana</em>.</p> <p>Identities of each community can be found in metadata.csv and comm_id.csv. Data analysis can be found in the GitHub repository associated to the manuscript: <a href="https://github.com/relab-fuberlin/schlechter_phyllosphere_spatial_distribution" target="_blank" rel="noopener">https://github.com/relab-fuberlin/schlechter_phyllosphere_spatial_distribution</a>.</p> <p>File bacimg.tar.gz contains the pre-processed images of near-isogenic controls, two- and three-species communities (C, S2 and S3).</p> <p>Abbreviations:</p> <p>Fluorescence channels:</p> <ul> <li>C0: Channel 0 (Red fluorescence)</li> <li>C1: Channel 1 (Yellow/Cyan fluorescence)</li> <li>C2: Channel 2 (Cyan fluorescence)</li> </ul> <p>Independent experiments:</p> <ul> <li>e1: Experiment #1</li> <li>e2: Experiment #2</li> </ul> <p>Days post-inoculation:</p> <ul> <li>7d: 7 days-post-inoculation</li> <li>14d: 14 days-post-inoculation</li> </ul> <p>SynCom ID (SynID):</p> <ul> <li>c: near-isogenic controls (C)</li> <li>syn2: two-species communities (S2)</li> <li>syn3: three-species communities (S3)</li> </ul> <p>Community ID (ComID):</p> <ul> <li>com01-com15: Strain composition of each community (see comm_id.csv)</li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo36/100

A dataset for semantic segmentation of typical oceanic and atmospheric phenomena from Sentinel-1 images

<p>We have constructed a SAR (Synthetic Aperture Radar) image semantic segmentation dataset that includes 12 oceanic and atmospheric phenomena: Atmospheric Front (AF), Oceanic Front (OF), Rainfall (RF), Iceberg (IC), Sea Ice (SI), Pure Ocean Wave (POW), Wind Streak (WS), Low Wind Area (LWA), Biological Slick (BS), Micro Convective Cells (MCC), Internal Wave (IW), and Eddy.</p> <p>This dataset is built using Sentinel-1 IW and WV mode images. For WV mode data, we referenced TenGeoP-SARwv and SAR_WV_SemanticSegmentation and selected 2,383 images for semantic segmentation and annotation. For IW mode images, we incorporated some images from Tao et al.'s internal wave detection dataset. We selected 484 Sentinel-1 IW mode images obtained from 2015 to 2022 and divided them into 2,628 sub-images.</p> <p>The dataset contains a total of 5,011 image slices, with approximately 400 images for each phenomenon. All images are 16-bit .tiff files with a resolution of 100m and a size of 256x256 pixels. The images were manually annotated using the Labelme software, generating corresponding JSON files, which were then used to create the related annotation .png files.</p> <p>The updated version(V2) provides geographic information for each image.</p> <p>Thank you for your interest in our dataset. Here are the meanings of each label:</p> <p>1. BG: The unlabelled parts in JSON files are "BG" (Background)<br>2. AF: Atmospheric Front<br>3. BS: Biological Slick<br>4. I: &ldquo;I&rdquo; is equivalent to &ldquo;IB&rdquo;, representing icebergs<br>5. LWA: Low Wind Area<br>6. MCC: Micro Convective Cells<br>7. OF: Oceanic Front<br>8. POW: Pure Ocean Wave<br>9. RC: &ldquo;RC&rdquo; (Rain Cells) is equivalent to &ldquo;RF&rdquo; (Rainfall), both representing the&nbsp; rainfall phenomenon in the SAR image.&nbsp;<br>10. SI: Sea Ice<br>11. WS: Wind Streak<br>12. Eddy<br>13. IW: Internal Wave<br><em>14. HM: Represents the artificial objects appearing in the image, such as ships, aquaculture floating rafts, wind power facilities, etc.</em><br><em>15. OS: Unlike &ldquo;BS&rdquo;,&ldquo;OS&rdquo; represents mineral oil spills appearing in the SAR image (currently, there is insufficient data available for training, which will be supplemented in the future).</em></p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

TotalSegmentator MRI dataset: 616 MRI images with segmentations for 50 anatomical regions

<p>In 616 MR images we segmented 50 anatomical structures covering a majority of relevant classes for most use cases. The MR images were randomly sampled from clinical routine, thus representing a real world dataset which generalizes to clinical application. The dataset contains a wide range of different pathologies, scanners, sequences and institutions. Moreover, it contains some images from <a href="https://portal.imaging.datacommons.cancer.gov/" target="_blank" rel="noopener">IDC</a> for further data diversity (see column "source" in meta.csv).</p> <p>Link to a copy of this dataset on Dropbox for much quicker download: <a href="https://www.dropbox.com/scl/fi/gskhmz2n9mmlt7rcn77wo/TotalsegmentatorMRI_dataset_v200.zip?rlkey=kf1eocb6xjmqwzafuohhv19by&amp;st=ka3m6wae&amp;dl=0" target="_blank" rel="noopener">Dropbox Link</a></p> <p>You can find a segmentation model trained on this dataset <a href="https://github.com/wasserth/TotalSegmentator">here</a>.<br><br>More details about the dataset can be found in the corresponding <a href="https://arxiv.org/abs/2405.19492">paper</a>. Please cite this paper if you use the dataset. The CT images described in the paper can be found <a href="../doi/10.5281/zenodo.6802613" target="_blank" rel="noopener">here</a>.</p> <p>This dataset contains all 50 structures from the TotalSegmentator "total" task. It does not contain the structures of other TotalSegmentator MRI subtasks.</p> <p>This dataset was created by the department of&nbsp;<a href="https://www.unispital-basel.ch/en/radiologie-nuklearmedizin/forschung-radiologie-nuklearmedizin">Research and Analysis at University Hospital Basel</a>.</p> <p><strong>UPDATE</strong>: on 2025-01-21 we uploaded version 2.0.0 which increases the number of images from 298 to 616. It also contains slightly different structures.</p>

opencc-by-nc-sa-2.0May 2024View details →
zenodo36/100

Glacial Lake Image Dataset for "Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset"

<p>Glacial lake dataset for the paper <strong>"Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.jhydrol.2025.133072" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.jhydrol.2025.133072</a>)</strong></p> <p>The GLID dataset contains a total of 18,367 samples, and the size of each sample is 512*512. Each sample consists of an image and a corresponding label. Four glacial lake types including supraglacial lake, proglacial lake, ice-marginal lake, and unconnected glacial lake are involved. The pixel value of the glacier lake in annotation map is labeled 255 and background is labeled 0.</p> <p>GLID.rar contains the training dataset (16,000 samples),&nbsp; and val.zip contains the validation dataset (2,367 samples).</p> <p>GLID_annotation.zip contains the annotated shapefile of GLID with a CRS of WGS 84.</p> <p>Optical_images_source.xlsx contains the optical images ID/names and acquisition time of each platform (e.g., WV2, LC08, S2B, and GF02) used in GLID.</p> <p>Transferability validation.zip contains the images, labels, and predictions for transferability validation, which is independent of GLID (not used for model training or validation). The file structure is shown below:</p> <p>Transferability validation.zip</p> <ul> <li>images <ul> <li>AS.tif</li> <li>GL.tif</li> <li>NA.tif</li> <li>SA.tif</li> </ul> </li> <li>labels <ul> <li>AS_gt.tif</li> <li>GL_gt.tif</li> <li>NA_gt.tif</li> <li>SA_gt.tif</li> </ul> </li> <li>predictions <ul> <li>AS_pred.tif</li> <li>GL_pred.tif</li> <li>NA_pred.tif</li> <li>SA_pred.tif</li> </ul> </li> </ul> <p>AS, GL, NA, and SA represent Asia, Greenland, North America, and South America, respectively. Four high-quality Landsat-8/9 images (each cloud cover less than 6%) were used for testing, and we manually annotated the glacial lakes in each image as labels. The pixel value of the glacier lake in annotation map is labeled 255 and background is labeled 0.&nbsp; Files in transferability validation.zip have a same CRS of WGS 84.</p> <p>If you find this dataset is helpful in your research, please consider <strong><em>cite </em></strong>this paper:</p> <blockquote> <p><em>Ma D, Li J, Jiang L. 2025. Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset. Journal of Hydrology 657: 133072.</em></p> </blockquote>

opencc-by-4.0Dec 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record