Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
98
datasets available to search
ShareScore release 0.9.0
Dataset results
98 results for “image generation”
Survey2Survey: A deep learning generative model approach for cross-survey image mapping
<p>During the last decade, there has been an explosive growth in survey data and deep learning techniques, both of which have enabled great advances for astronomy. The amount of data from various surveys from multiple epochs with a wide range of wavelengths, albeit with varying brightness and quality, is overwhelming, and leveraging information from overlapping observations from different surveys has limitless potential in understanding galaxy formation and evolution. Synthetic galaxy image generation using physical models has been an important tool for survey data analysis, while deep learning generative models show great promise. In this paper, we present a novel approach for robustly expanding and improving survey data through cross survey feature translation. We trained two types of neural networks to map images from the Sloan Digital Sky Survey (SDSS) to corresponding images from the Dark Energy Survey (DES). This map was used to generate false DES representations of SDSS images, increasing the brightness and S/N while retaining important morphological information. We substantiate the robustness of our method by generating DES representations of SDSS images from outside the overlapping region, showing that the brightness and quality are improved even when the source images are of lower quality than the training images. Finally, we highlight several images in which the reconstruction process appears to have removed large artifacts from SDSS images. While only an initial application, our method shows promise as a method for robustly expanding and improving the quality of optical survey data and provides a potential avenue for cross-band reconstruction.</p><p>------------</p><p>This repository contains the image files from Survey2Survey: a deep learning generative model approach for cross-survey image mapping. Please cite https://arxiv.org/abs/2011.07124 if you use this data in a publication. For more information, contact Brandon Buncher at buncher2(at)illinois.edu</p><p><strong>--- Directory structure ---</strong></p><p>tutorial.ipynb demonstrates how to load the image files (uploaded here as tarballs). Images were obtained from the SDSS DR16 cutout server (https://skyserver.sdss.org/dr16/en/help/docs/api.aspx) and DES DR1 cutout server (https://des.ncsa.illinois.edu/desaccess/</p><ul><li>./sdss_train/ and ./des_train/ contain the original SDSS and DES images used to train the neural network (Stripe82)</li><li>./sdss_test/ and ./des_test/ contain the original SDSS and DES images used for the validation dataset (Stripe82)</li><li>./sdss_ext/ contain images from the external SDSS dataset (SDSS images without a DES counterpart, outside Stripe82)</li><li>./cae and ./cyclegan contain images generated by the CAE and CycleGAN, respectively. train_decoded/ and test_decoded/ contain the reconstructions of the images from the training dataset and test dataset, respectively. external_decoded/ contain the DES-like image reconstructions of SDSS objects from the external dataset (outside Stripe82).</li></ul>
European Registry of Next Generation Imaging in Advanced Prostate Cancer
ClinicalTrials.gov study NCT06866782. IPD Sharing: NO. Countries: 8. Publications: 1.
Chest X-Ray Image Diagnosis and Report Generation Dedicated Model Based on Deepseek
ClinicalTrials.gov study NCT06874647. IPD Sharing: NO. Countries: 1. Publications: 1.
Aminoacyl-tRNA synthetase gene alignments from multiple Sileneae species generated from full-length transcripts using Iso-Seq and raw microscopy image files
Open the record for dataset details and reuse information.
Data from: Second harmonic generation imaging reveals entanglement of collagen fibers in the elephant trunk skin dermis
Open the record for dataset details and reuse information.
Pot-1::mCherry transgene images from: Gametes deficient for Pot1 telomere binding proteins alter levels of telomeric foci for multiple generations
Open the record for dataset details and reuse information.
Synthetic histology images of colorectal cancer, generated by conditional generative adversarial networks
<p>These are generated (synthetic) histology images of colorectal cancer. These images were generated by conditional GANs and are in two classes: MSIH (microsatellite instable high) and nonMSIH. There are two sets: one set with 10K images per class and another one with 75K images per class. All images are RGB, 512x512 px at a resolution of 0.5 micrometers per pixel. For more information, please stay tuned for our upcoming manuscript on www.kather.ai. </p>
A new dataset of rain cell generated from observations of the Tropical Rainfall Measuring Mission (TRMM) precipitation radar and visible and infrared scanner and microwave imager
<p>This new dataset (M.TRMM-1B01-1B11-2A25-PMD-Rain) contains orbit-level data with 5 km spatial resolution and 0.25 km vertical resolution. It is produced by merging TRMM PR, VIRS and TMI measurements at PR pixel resolution combined with rain cell identification. The near-surface rain rate, profiles of rain rate and precipitation reflectivity factor, visible and infrared signals and microwave signals can be obtained in the dataset. The dataset provides new important data for in-depth research on the structural characteristics of rain cells and supports the study of precipitation mechanisms.</p>
GAN Generated Images for Facial Expression Recognition systems
<p>Most facial expression recognition (FER) systems rely on machine learning approaches that require large databases (DBs) for effective training. As these are not easily available, a good solution is to augment the DBs with appropriate techniques, which are typically based on either geometric transformation or deep learning based technologies (e.g., Generative Adversarial Networks (GANs)). Whereas the first category of techniques has been fairly adopted in the past, studies that use GAN-based techniques are limited for FER systems. To advance in this respect, we evaluate the impact of the GAN techniques by creating a new DB containing the generated synthetic images. </p> <p>The face images contained in the KDEF DB serve as the basis for creating novel synthetic images by combining the facial features of two images (i.e., Candie Kung and Cristina Saralegui) selected from the YouTube-Faces DB. The novel images differ from each other, in particular concerning the eyes, the nose, and the mouth, whose characteristics are taken from the Candie and Cristina images.</p> <p>The total number of novel synthetic images generated with the GAN is 980 (70 individuals from KDEF DB x 7 emotions x 2 subjects from YouTube-Faces DB).</p> <p>The zip file "GAN_KDEF_Candie" contains the 490 images generated by combining the KDEF images with the Candie Kung image. The zip file "GAN_KDEF_Cristina" contains the 490 images generated by combining the KDEF images with the Cristina Saralegui image. The used image IDs are the same used for the KDEF DB. The synthetic generated images have a resolution of 562x762 pixels.</p> <p> </p> <p><strong>If you make use of this dataset, please consider citing the following publication:</strong></p> <p>Porcu, S., Floris, A., & Atzori, L. (2020). Evaluation of Data Augmentation Techniques for Facial Expression Recognition Systems. Electronics, 9, 1892, doi: 10.3390/electronics9111892, url: https://www.mdpi.com/2079-9292/9/11/1892.</p> <p>BibTex format:</p> <p>@article{porcu2020evaluation, title={Evaluation of Data Augmentation Techniques for Facial Expression Recognition Systems}, author={Porcu, Simone and Floris, Alessandro and Atzori, Luigi}, journal={Electronics}, volume={9}, pages={108781}, year={2020}, number = {11}, article-number = {1892}, publisher={MDPI}, doi={10.3390/electronics9111892} }</p> <p> </p>
CelebA-HQ Dataset Generated Images with StyleSwin Method
Open the record for dataset details and reuse information.
Machine Learning for Analyzing Atomic Force Microscopy (AFM) Images Generated from Polymer Blends
Open the record for dataset details and reuse information.
Image Synthesis with a Convolutional Capsule Generative Adversarial Network- Prepared Data
<p>A set of prepared datasets for running experiments to replicate paper (see below).</p> <p>List of data:</p> <p>Training capspix2pix:</p> <ul> <li>crops256.zip - folder containing 256x256 crops from the original dataset for training capspix2pix. Images are in the "train/original" folder, and labels are in the "train/mask" folder.</li> <li>syn256_x_data_val.npy + syn256_y_data_val.npy + syn256_y_points_data_val.npy (images + labels + centrelines) - validation synthetic dataset, used while training capspix2pix for plotting</li> </ul> <p>Training u-net:</p> <ul> <li>capspix2pix_AR_data_train.npy + capspix2pix_AR_mask_train.npy (images + labels) - data generated from a capspix2pix model from real labels</li> <li>capspix2pix_SSM_data_train.npy + capspix2pix_AR_mask_train.npy (images + labels) - data generated from a capspix2pix model from synthetic labels</li> <li>PBAM_SSM_data_train.npy + PBAM_SSM_mask_train.npy (images + labels) - data generated from PBAM (Physics-based model) for training u-net</li> <li>pix2pix_AR_data_train.npy + pix2pix_AR_mask_train.npy (images + labels) - data generated from a pix2pix model from real labels for training u-net</li> <li>pix2pix_SSM_data_train.npy + pix2pix_SSM_mask_train.npy (images + labels) - data generated from a pix2pix model from synthetic labels for training u-net</li> <li>real_data_data_train.npy + real_data_mask_train.npy (images + labels) - augmented real dataset for training u-net</li> </ul> <p>Testing u-net:</p> <ul> <li>org64_data_test.npy + org64_mask_test.npy (images + labels) - crops from original test dataset for testing u-net</li> </ul> <p>Interpolation:</p> <ul> <li>crops256_inter_data_train.npy + crops256_inter_mask_train.npy (images + labels) - example data for interpolation</li> </ul> <p><strong>Please cite the following paper when using this dataset:</strong></p> <p>Bass, C., Dai, T., Billot, B., Arulkumaran, K., Creswell, A., Clopath, C., De Paola, V., and Bharath, A. A., 2019. “Image synthesis with a convolutional capsule generative adversarial network,” <em>Medial Imaging with Deep Learning.</em></p> <p><strong>See Github page for further instructions:</strong></p> <p>https://github.com/CherBass/CapsPix2Pix</p> <p> </p>
"water droplet on a leaf" - epistemic insight and an AI generated image
Open the record for dataset details and reuse information.
Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models
<h2>[Update] Looking for a larger unsafe image dataset? We publish a new dataset named UnsafeBench on Hugging Face. Take a look at <a href="https://huggingface.co/datasets/yiting/UnsafeBench">here</a>!</h2> <p>This dataset used in the paper <a href="https://arxiv.org/pdf/2305.13873.pdf">https://arxiv.org/pdf/2305.13873.pdf</a> contains four prompt sets and one image set.</p> <p>The four prompt sets were used to query Text-to-Image models and generate images for safety assessment. These sets include three harmful prompt sets and one harmless prompt set. The harmful prompts originate from different sources and contain various unsafe concepts, such as sexually explicit, violent, disturbing, hateful, and political content.</p> <p><strong>Prompt Sets</strong>:</p> <ul> <li>4chan Prompts: Harmful</li> <li>Lexica Prompts: Harmful</li> <li>Template Prompts: Harmful</li> <li>COCO Prompts: Harmless</li> </ul> <p><strong>Image Dataset</strong>:</p> <p>This dataset consists of 800 images, which were randomly selected from all the generated images from Text-to-Image models.</p> <ul> <li>Safe: 580 images</li> <li>Sexually Explicit: 48 images</li> <li>Violent: 45 images</li> <li>Disturbing: 68 images</li> <li>Hateful: 35 images</li> <li>Political: 50 images</li> </ul>
Realistic in Generation of HEp-2 Cell Images Using Latent Diffusion Models: a Multi-center Visual Turing Test
ClinicalTrials.gov study NCT06542783. IPD Sharing: NO. Countries: 0. Publications: 6.
The Examination of Woman Generations' on Exercise Preferences, Body Image, Physical Activity and Social Media Use
ClinicalTrials.gov study NCT05921708. IPD Sharing: NO. Countries: 0. Publications: 4.
GRIP METEOSAT SECOND GENERATION (MSG) IMAGE DATA V1
The GRIP Meteosat Second Generation (MSG) Image Data was collected during the Genesis and Rapid Intensification Processes (GRIP) experiment from August 15, 2010 to September 30, 2010. The major goal was to better understand how tropical storms form and develop into major hurricanes. Infrared and visible radiances, and water vapor were measured. Meteosat Second Generation (MSG) consists of a series of four geostationary meteorological satellites, along with ground-based infrastructure, that will operate consecutively until 2020. The MSG system is established under cooperation between The European Organization for the Exploitation of Meteorological Satellites (EUMETSAT) and the European Space Agency (ESA) to ensure the continuity of meteorological observations from geostationary orbit. The MSG satellites carry an impressive pair of instruments, the Spinning Enhanced Visible and InfraRed Imager (SEVIRI), which has the capacity to observe the Earth in 12 spectral channels and provide image data which is core to operational forecasting needs, and the Geostationary Earth Radiation Budget (GERB) instrument supporting climate studies.
NOAA GHRSST Level 2P Atlantic Ocean Regional Skin Sea Surface Temperature v1.0 from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) on the Meteosat Second Generation-4 (MSG-4) satellite
The GHRSST L2P MSG04 SST v1.0 dataset is produced by the US National Oceanic and Atmospheric Administration (NOAA) National Environmental Satellite, Data, and Information Service (NESDIS) from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) onboard the Meteosat-11 (MSG4) satellite. It provides the full disk SEVIRI imagery covering the Atlantic Ocean region from its position at 0.0°E longitude. The L2P SST is produced at approximately 3 km resolution with a 15 minute duty cycle. On Feb. 2, 2018 the Meteosat-11 (MSG4) took over the Meteosat-10 (MSG3) (MSG03-OSPO-L2P-v1.0) and produced the L2P SST data from Sept 10. 2018 to March 24, 2023. In March 2023, Meteosat-10 and Meteosat-11 were swapped roles and orbital positions. The MSG03 has started to produce the L2P SST data again over the Atlantic Ocean region. Be aware that the granules before Dec. 1, 2022 contain some uncorrected metadata errors. <br><br>The SST measurements from SEVIRI are parameters in study of the weather, atmosphere, climate and ocean environments. Meteosat satellites have been providing crucial data for weather forecasting since 1977. <br><br>This L2P SST product which includes Single Sensor Error Statistics (i.e., uncertainty statistics) follows the GHRSST Data Processing Specification (GDS) version 2.0 format guidelines. Please refer to the user guide for more information.
NOAA GHRSST Level 2P Indian Ocean Regional Skin Sea Surface Temperature v1.0 from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) on the Meteosat Second Generation-1 (MSG-1) satellite
The GHRSST L2P MSG01 SST v1.0 dataset is produced by the US National Oceanic and Atmospheric Administration (NOAA) National Environmental Satellite, Data, and Information Service (NESDIS) from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) onboard the Meteosat-8 (MSG1) satellite. It provides the full disk SEVIRI imagery covering the Indian Ocean region from its position at 45.5°E longitude. The L2P SST is produced at approximately 3 km resolution with a 15 minute duty cycle. The full data records stretch from Sept. 18, 2018 to June 1, 2022. After June 1, 2022, the Meteosat-9 (MSG2) took over as the prime geostationary satellite for the Indian Ocean region (MSG02-OSPO-L2P-v1.0). Be aware that the granules before Dec. 1, 2022 contain some uncorrected metadata errors. <br><br>The SST measurements from SEVIRI are key parameters in study of the weather, atmosphere, climate and ocean environments. Meteosat satellites have been providing crucial data for weather forecasting since 1977. <br><br>This L2P SST product which includes Single Sensor Error Statistics (i.e., uncertainty statistics) follows the GHRSST Data Processing Specification (GDS) version 2.0 format guidelines. Please refer to the user guide for more information.
NOAA GHRSST Level 2P Indian Ocean Regional Skin Sea Surface Temperature v1.0 from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) on the Meteosat Second Generation-2 (MSG-2) satellite
The GHRSST L2P MSG02 SST v1.0 dataset is produced by the US National Oceanic and Atmospheric Administration (NOAA) National Environmental Satellite, Data, and Information Service (NESDIS) from the Spinning Enhanced Visible and InfraRed Imager (SEVIRI) onboard the Meteosat-9 (MSG2) satellite. It provides the full disk SEVIRI imagery covering the Indian Ocean region from its position at 45.5°E longitude. The L2P SST is produced at approximately 3 km resolution with a 15 minute duty cycle. On June 1, 2022, the Meteosat-9 (MSG2) replaced the Meteosat-8 (MSG1) (MSG01-OSPO-L2P-v1.0) and produced the L2P SST data from June 11. 2022 to the present. This dataset will be updated every 15 minutes as a forward data stream with 3-24 hours nominal latency. Be aware that the granules before Dec. 1, 2022 contain some uncorrected metadata errors.<br><br>The SST measurements from SEVIRI are key parameters in study of the weather, atmosphere, climate and ocean environments. Meteosat satellites have been providing crucial data for weather forecasting since 1977. <br><br>This L2P SST product which includes Single Sensor Error Statistics (i.e., uncertainty statistics) follows the GHRSST Data Processing Specification (GDS) version 2.0 format guidelines. Please refer to the user guide for more information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.