Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “image dataset”
DSB (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>DSB n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
DSB (noise 0) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>DSB n0 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Mouse (noise 0) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Mouse n0 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Flywing (noise 20) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Flywing n20 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Mouse (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Mouse n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Mouse (noise 20) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Mouse n20 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
CNN for image-based sediment detection applied to a large terrestrial and airborne dataset
<p>This repository contains data sets and model used in the manuscript "CNN for image-based sediment detection applied to a large terrestrial and airborne dataset" submitted to Earth Surface Dynamics (2021) by Xingyu Chen, Marwan A. Hassan, and Xudong Fu.</p>
Dataset: lhp1 FLC-Venus time course imaging – Hybrid protein assembly-histone modification mechanism for PRC2-based epigenetic switching and memory
<p>The histone modification H3K27me3 plays a central role in Polycomb-mediated epigenetic silencing. H3K27me3 recruits and allosterically activates Polycomb Repressive Complex 2 (PRC2), which adds this modification to nearby histones, providing a read/write mechanism for inheritance through DNA replication. However, for some PRC2 targets, a purely histone-based system for epigenetic inheritance may be insufficient. We address this issue at the Polycomb target Flowering Locus C (FLC) in Arabidopsis thaliana, as a narrow nucleation region of only ~three nucleosomes within FLC mediates epigenetic state switching and subsequent memory over many cell cycles. To explain the memory's unexpected persistence, we introduce a mathematical model incorporating extra protein memory storage elements with positive feedback that persist at the locus through DNA replication, in addition to histone modifications. Our hybrid model explains many features of epigenetic switching/memory at FLC and encapsulates generic mechanisms that may be widely applicable.</p>
Performance of chemical structure string representations for chemical image recognition using transformers dataset
<p>The datasets contain string representations used for DECIMER short communication paper.</p> <p><strong>ChEMBL dataset:</strong></p> <p>Train and test datasets downloaded from ChEMBL and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.</p> <p>String representations contain SMILES, DeepSMILES, SELFIES and InChIs.</p> <ul> <li>Train dataset: 1.5 Mio molecules</li> <li>Test dataset: ~100K molecules</li> </ul> <p><strong>Pubchem dataset:</strong></p> <p>Train and test datasets downloaded from PubChem and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.</p> <p>String representations contain SMILES, DeepSMILES and SELFIES.</p> <ul> <li>Train dataset: 3 Mio molecules</li> <li>Test dataset: 250K molecules</li> </ul>
Re-identification of Individuals in Genomic Datasets Using Public Face Images
<p>Image-genome pairs in these synthetic datasets were created by combining a subset of the publicly available face image dataset, CelebA, and genotypes from OpenSNP. The genome in a given pair does not correspond to the individual in the image (taken from CelebA), but comes instead from an individual with the same set of phenotypes (taken from OpenSNP). Artificial genotypes were created for each image (genotype refers only to the small subset of SNPs we are interested in) using all available data from OpenSNP where self-reported phenotypes are present.</p> <p>In the Synthetic-Ideal dataset, to each image, we assigned a genotype from OpenSNP that corresponds to an individual with the same phenotypes, such that the probability of the selected phenotypes is maximized, given the genotype. In other words, we picked the genotype from the OpenSNP data that is most representative of an individual with a given set of phenotypes.</p> <p>In the Synthetic-Realistic dataset, to each image, we assigned a genotype from OpenSNP that corresponds to an individual with the same phenotypes, but at random according to the empirical distribution of phenotypes for particular SNPs in our data.</p> <p>Since CelebA does not have labels for all considered phenotypes, 1000 images from this dataset were manually labeled by one of the authors. After cleaning and removing ambiguous cases, the resulting datasets consist of 456 records.</p>
Multiple Image Splicing Dataset (MISD): A Dataset for Multiple Splicing
<p>Multiple Image Splicing dataset is the first publicly available Multiple Image Splicing Dataset. It consists of 918 images classified as Authentic or Spliced.618 Authenticate images of type JPG and 300 Multiple Spliced JPG images of size 384 ×<em> </em>256.</p>
test images dataset for ear tracking example
<p>Test images for ear tracking example, including parameters about these images</p>
The dataset of Sentinel-1 SAR images for sea ice classification
<p>The dataset implementation of the paper "A Multi-scale Dual Attention Network for Automatic Polar Sea Ice Classification Based on Sentinel-1 SAR Images".</p> <p>There are 7381 images as the training set, 1210 images as the validation set, and 3630 images as the test set. </p> <p>The file contains original images and processed images.</p>
Model 4 dataset for the manuscript "Improving trajectory calculations by FLEXPART 10.4+ using deep learning inspired single image superresolution"
<p>Model 4 dataset for the manuscript "Improving trajectory calculations by FLEXPART 10.4+ using deep learning inspired single image superresolution"</p>
Model 2 dataset for the manuscript "Improving trajectory calculations by FLEXPART 10.4+ using deep learning inspired single image superresolution"
<p>Dataset produced by the model 2 neural network for the manuscript "Improving trajectory calculations by FLEXPART 10.4+ using deep learning inspired single image superresolution"</p>
Large-scale annotation dataset for cell/tissue segmentation in H&E-stained images : anti-CD3/CD20 (lymphocytes)
<p><strong>LICENSE</strong></p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (<strong>CC-BY-NC-SA 4.0</strong>)</p> <p>For non-commercial use, please use the dataset under CC-BY-NC-SA.<br> If you would like to use the dataset for commercial purposes, please contact us (ishum-prm@m.u-tokyo.ac.jp).</p> <p>A Tar.gz file contains the following files:</p> <p>- HE image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_HE.png</p> <p>- Mask image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_mask.png</p> <p>Each image file is 984x984 px.</p> <p>posX and posY are the leftmost position in WSI coordinate.</p> <p>Mask files store binary segmentation mask (background : 0, target : 1)</p> <p> </p> <p>A csv file contains the following information:</p> <p>antigen : Antibodies for this antigen were used to create the segmentation mask.</p> <p>filename: filename of image or mask file.</p> <p>train_val_test : train, validation, or test sample in the paper.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for your research, please cite our paper.</p> <p>Daisuke Komura, Takumi Onoyama, Koki Shinbo, Hiroto Odaka, Minako Hayakawa, Mieko Ochi, Ranny Rahaningrum Herdiantoputri, Haruya Endo, Hiroto Katoh, Tohru Ikeda, Tetsuo Ushiku, Shumpei Ishikawa,<br> Restaining-based annotation for cancer histology segmentation to overcome annotation-related limitations among pathologists, Patterns, Volume 4, Issue 2, 2023, 100688, https://doi.org/10.1016/j.patter.2023.100688.</p>
Large-scale annotation dataset for cell/tissue segmentation in H&E-stained images : anti-ERG (endothelial cells)
<p><strong>LICENSE</strong></p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (<strong>CC-BY-NC-SA 4.0</strong>)</p> <p>For non-commercial use, please use the dataset under CC-BY-NC-SA.<br> If you would like to use the dataset for commercial purposes, please contact us (ishum-prm@m.u-tokyo.ac.jp).</p> <p>A Tar.gz file contains the following files:</p> <p>- HE image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_HE.png</p> <p>- Mask image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_mask.png</p> <p>Each image file is 984x984 px.</p> <p>posX and posY are the leftmost position in WSI coordinate.</p> <p>Mask files store binary segmentation mask (background : 0, target : 1)</p> <p> </p> <p>A csv file contains the following information:</p> <p>antigen : Antibodies for this antigen were used to create the segmentation mask.</p> <p>filename: filename of image or mask file.</p> <p>train_val_test : train, validation, or test sample in the paper.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for your research, please cite our paper.</p> <p>Daisuke Komura, Takumi Onoyama, Koki Shinbo, Hiroto Odaka, Minako Hayakawa, Mieko Ochi, Ranny Rahaningrum Herdiantoputri, Haruya Endo, Hiroto Katoh, Tohru Ikeda, Tetsuo Ushiku, Shumpei Ishikawa,<br> Restaining-based annotation for cancer histology segmentation to overcome annotation-related limitations among pathologists, Patterns, Volume 4, Issue 2, 2023, 100688, https://doi.org/10.1016/j.patter.2023.100688.</p>
Large-scale annotation dataset for cell/tissue segmentation in H&E-stained images : anti-MNDA (myeloid cells)
<p><strong>LICENSE</strong></p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (<strong>CC-BY-NC-SA 4.0</strong>)</p> <p>For non-commercial use, please use the dataset under CC-BY-NC-SA.<br> If you would like to use the dataset for commercial purposes, please contact us (ishum-prm@m.u-tokyo.ac.jp).</p> <p>A Tar.gz file contains the following files:</p> <p>- HE image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_HE.png</p> <p>- Mask image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_mask.png</p> <p>Each image file is 984x984 px.</p> <p>posX and posY are the leftmost position in WSI coordinate.</p> <p>Mask files store binary segmentation mask (background : 0, target : 1)</p> <p> </p> <p>A csv file contains the following information:</p> <p>antigen : Antibodies for this antigen were used to create the segmentation mask.</p> <p>filename: filename of image or mask file.</p> <p>train_val_test : train, validation, or test sample in the paper.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for your research, please cite our paper.</p> <p>Daisuke Komura, Takumi Onoyama, Koki Shinbo, Hiroto Odaka, Minako Hayakawa, Mieko Ochi, Ranny Rahaningrum Herdiantoputri, Haruya Endo, Hiroto Katoh, Tohru Ikeda, Tetsuo Ushiku, Shumpei Ishikawa,<br> Restaining-based annotation for cancer histology segmentation to overcome annotation-related limitations among pathologists, Patterns, Volume 4, Issue 2, 2023, 100688, https://doi.org/10.1016/j.patter.2023.100688.</p>
Large-scale annotation dataset for cell/tissue segmentation in H&E-stained images : anti-CD235a (red blood cells)
<p><strong>LICENSE</strong></p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (<strong>CC-BY-NC-SA 4.0</strong>)</p> <p>For non-commercial use, please use the dataset under CC-BY-NC-SA.<br> If you would like to use the dataset for commercial purposes, please contact us (ishum-prm@m.u-tokyo.ac.jp).</p> <p>A Tar.gz file contains the following files:</p> <p>- HE image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_HE.png</p> <p>- Mask image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_mask.png</p> <p>Each image file is 984x984 px.</p> <p>posX and posY are the leftmost position in WSI coordinate.</p> <p>Mask files store binary segmentation mask (background : 0, target : 1)</p> <p> </p> <p>A csv file contains the following information:</p> <p>antigen : Antibodies for this antigen were used to create the segmentation mask.</p> <p>filename: filename of image or mask file.</p> <p>train_val_test : train, validation, or test sample in the paper.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for your research, please cite our paper.</p> <p>Daisuke Komura, Takumi Onoyama, Koki Shinbo, Hiroto Odaka, Minako Hayakawa, Mieko Ochi, Ranny Rahaningrum Herdiantoputri, Haruya Endo, Hiroto Katoh, Tohru Ikeda, Tetsuo Ushiku, Shumpei Ishikawa,<br> Restaining-based annotation for cancer histology segmentation to overcome annotation-related limitations among pathologists, Patterns, Volume 4, Issue 2, 2023, 100688, https://doi.org/10.1016/j.patter.2023.100688.</p>
Large-scale annotation dataset for cell/tissue segmentation in H&E-stained images : anti-CD45RB (leukocytes)
<p><strong>LICENSE</strong></p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (<strong>CC-BY-NC-SA 4.0</strong>)</p> <p>For non-commercial use, please use the dataset under CC-BY-NC-SA.<br> If you would like to use the dataset for commercial purposes, please contact us (ishum-prm@m.u-tokyo.ac.jp).</p> <p>A Tar.gz file contains the following files:</p> <p>- HE image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_HE.png</p> <p>- Mask image file: {antigen}_{celltype}_{slideID}_{posx}_{posy}_mask.png</p> <p>Each image file is 984x984 px.</p> <p>posX and posY are the leftmost position in WSI coordinate.</p> <p>Mask files store binary segmentation mask (background : 0, target : 1)</p> <p> </p> <p>A csv file contains the following information:</p> <p>antigen : Antibodies for this antigen were used to create the segmentation mask.</p> <p>filename: filename of image or mask file.</p> <p>train_val_test : train, validation, or test sample in the paper.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for your research, please cite our paper.</p> <p>Daisuke Komura, Takumi Onoyama, Koki Shinbo, Hiroto Odaka, Minako Hayakawa, Mieko Ochi, Ranny Rahaningrum Herdiantoputri, Haruya Endo, Hiroto Katoh, Tohru Ikeda, Tetsuo Ushiku, Shumpei Ishikawa,<br> Restaining-based annotation for cancer histology segmentation to overcome annotation-related limitations among pathologists, Patterns, Volume 4, Issue 2, 2023, 100688, https://doi.org/10.1016/j.patter.2023.100688.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.