Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
98
datasets available to search
ShareScore release 0.9.0
Dataset results
98 results for “image generation”
Subjective human thresholds over computer generated images
<p>Realistic image computation mimics the natural process of acquiring pictures by simulating the physical interactions of light between all the objects, lights and cameras lying within a modelled 3D scene. This process is known as global illumination and was formalised by Kajiya with the following rendering Equation:<br> <span class="math-tex">\(\begin{equation} \label{eq:rendering_equation} L_o(x, \omega_o) = {L_e(x, \omega_o)} + \int_{\Omega}^{} {L_i(x, \omega_i)} \cdot f_r(x, \omega_i \rightarrow \omega_o) \cdot \cos \theta_i d\omega_i \end{equation}\)</span></p> <p>where:</p> <ul> <li> <span class="math-tex">\(L_o(x, \omega_o)\)</span> is the luminance traveling from point <span class="math-tex">\(x\)</span> in direction <span class="math-tex">\(\omega_o\)</span>;</li> <li><span class="math-tex">\(L_e(x, \omega_o)\)</span> is point <span class="math-tex">\(x\)</span> emitted luminance (it is null if point x does not lie on a ligth source surface);</li> <li>the integral represents the set of luminances <span class="math-tex">\(L_i\)</span>incident in <span class="math-tex">\(x \)</span> from the hemisphere of the directions <span class="math-tex">\(\Omega\)</span> and reflected in the direction <span class="math-tex">\(\omega_o\)</span>. The reflected luminances are weighted by the materials reflecting properties (bidirectionnal reflectance function <span class="math-tex">\(f_r(x, \omega_i \rightarrow \omega_o)\)</span>) and the cosinus of the incident angle.</li> </ul> <p>This equation cannot be analytically solved and Monte Carlo approaches are generally used to estimate the value of the pixels of the final image.</p> <p>This proposed dataset is composed of 80 points of view of photo realistics images with different level of samples (following the Monte Carlo approach) for each. Each image is 800 x 800 pixels in size. The most noisy image is of 20 samples and the reference one (the most converged image obtained) is of 10000 samples. The <a href="https://www.pbrt.org/index.html">pbrt</a> rendering engine (version 3) was used to generate these images.</p> <p>By exploiting these levels of samples obtained and therefore of noise perceptible in the images, average subjective human thresholds were collected. For this purpose, the images were divided into 16 areas of 200 x 200 pixels in size for each point of view.</p> <p>The proposed image database is composed of the following files:</p> <ul> <li><strong>human-thresholds.csv</strong> : the set of human subjective thresholds obtained on 40 points of view. A line is composed of the name of the point of view followed by all the thresholds obtained for each of the 16 zones;</li> <li><strong>SIN3D_dataset.tar.gz</strong> : is an archive containing all the images from 20 to 10000 samples in steps of 20 samples for each point of view (i.e. 500 images per point of view). Each folder in the archive corresponds to a point of view.</li> </ul> <p><em>This image database has been exploited in order to propose an objective model for noise detection in photo-realistic computer-generated images (article referenced to this image database).</em></p> <p><strong>Note:</strong> Some of the proposed scenes come from:</p> <ul> <li><a href="https://pbrt.org/scenes-v3">https://pbrt.org/scenes-v3</a></li> <li><a href="https://benedikt-bitterli.me/resources/">https://benedikt-bitterli.me/resources/</a></li> </ul> <p><strong>Funding:</strong> This research was funded by ANR support: project ANR-17-CE38-0009.</p> <p> </p>
Pythia Generated Jet Images with Alternative Rotation Scheme for Location Aware Generative Adversarial Network Training
<p>Dataset containing 300k jet images that can be used to train Location Aware Generative Adversarial Networks (LAGAN) for High Energy Physics, such as the one in [arXiv:1701.05927].</p> <p><strong>Format</strong>:</p> <p>HDF5 file with the following fields:</p> <ul> <li>'image' : array of dim (300000, 25, 25), contains the pixel intensities of each 25x25 image</li> <li>'signal' : binary array to identify signal (1, i.e. W boson) vs background (0, i.e. QCD)</li> <li>'jet_eta': eta coordinate per jet</li> <li>'jet_phi': phi coordinate per jet</li> <li>'jet_mass': mass per jet</li> <li>'jet_pt': transverse momentum per jet</li> <li>'jet_delta_R': distance between leading and subleading subjets if 2 subjets present, else 0</li> <li>'tau_1', 'tau_2', 'tau_3': substructure variables per jet (a.k.a. n-subjettiness, where n=1, 2, 3)</li> <li>'tau_21': tau<sub>2</sub>/tau<sub>1</sub> per jet</li> <li>'tau_32': tau<sub>3</sub>/tau<sub>2</sub> per jet</li> </ul> <p><strong>Details</strong>:</p> <ul> <li>Simulated using Pythia 8.219 at √ s = 14 TeV</li> <li>Image pre-processing using method from in L. de Oliveira et al., <em>Jet-Images -- Deep Learning Edition </em>[arXiv:1511.05190]</li> <li>scikit-image==0.10.0 implementation of cubic spline rotation with fewer low energy artifacts than scikit-image>=0.12.0</li> <li>Finite calorimeter granularity simulated with 0.1×0.1 grid in η and φ, with η × φ ∈ [−1.25, 1.25] × [−1.25, 1.25]</li> <li>Jet clustering with anti-k<sub>t</sub> algorithm with a radius R = 1.0 using FastJet 3.2.1; constituent re-clustering into R = 0.3 k<sub>t</sub> subjets</li> <li>Intensity of pixel = p<sub>T</sub> of cell</li> <li>60 GeV < m<sup>jet</sup> < 100 GeV</li> <li>250 GeV < p<sub>T</sub><sup>jet</sup> < 300 GeV</li> <li>Sparse images (~10% NNZ)</li> </ul> <p>Full dataset description in [arXiv:1701.05927].</p>
Third harmonic generation images of the lacuno-canalicular network in bone femoral diaphysis of mice from the BionM1 project (space flight)
<p>Data set for 11 samples in 3 groups of Control, Space Flight and Synchro (ground control with space flight housing and feeding conditions). Contains THG images in tif format of 2D mosaic of selected samples and 3D stacks in selected anatomical regions of interest. See readme file for more information.</p>
COSMOS isolated galaxy images (parametric models generated with GalSim)
<p>Isolated galaxy images generated with GalSim from parametric models extracted from the Hubble COSMOS catalog.</p> <p>These files contain images and data for 10 000 images of isolated galaxy:</p> <ul> <li><strong>galaxies_isolated_10000_images.npy: </strong>numpy array of shape (10 000, 10, 64, 64), 10 000 images of size 64x64 pixels, in 10 filters (4 Euclid filters and 6 <em>ugrizy</em> LSST filters, in that order). Images contain Poissonian noise.</li> <li><strong>galaxies_isolated_10000_data.csv: </strong>corresponding parameters: <ul> <li>SNR: signal-to-noise ratio</li> <li>redshift: redshift of the galaxy</li> <li>e1: e1 parameter of ellipticity (e = e1 + i.e2)</li> <li>e2: e2 parameter of ellipticity (e = e1 + i.e2)</li> <li>mag: magnitude</li> </ul> </li> </ul>
A collection of AI generated images visualising various RDM aspects
<p>This publication contains images visualising various RDM aspects. These images were generated by the <a href="https://www.forschungsdaten.uni-bonn.de/en" target="_blank" rel="noopener">Research Data Service Center</a> team at the University of Bonn and are used in the workshop "Research Data Management: A Crash Course" conducted since 2021 by the Research Data Service Center. The slide deck is available as a related publication (see the related works section below for details).</p> <p>The images were generated with the help of <a href="https://help.openai.com/en/articles/8932459-creating-images-in-chatgpt">ChatGPT</a>. </p> <p>In this version, due to legal reasons, we changed the images.</p>
Synthbuster: Towards Detection of Diffusion Model Generated Images
<p>Dataset described in the paper "Synthbuster: Towards Detection of Diffusion Model Generated Images" (Quentin Bammey, 2023, <i>Open Journal of Signal Processing</i>)</p><p>This dataset contains synthetic, AI-generated images from 9 different models:</p><ul><li>DALL·E 2</li><li>DALL·E 3</li><li>Adobe Firefly</li><li>Midjourney v5</li><li>Stable Diffusion 1.3</li><li>Stable Diffusion 1.4</li><li>Stable Diffusion 2</li><li>Stable Diffusion XL</li><li>Glide</li></ul><p> </p><p>1000 images were generated per model. The images are loosely based on raise-1k images (Dang-Nguyen, Duc-Tien, et al. "Raise: A raw images dataset for digital image forensics." Proceedings of the 6th ACM multimedia systems conference. 2015.). For each image of the raise-1k dataset, a description was generated using the Midjourney /describe function and CLIP interrogator (https://github.com/pharmapsychotic/clip-interrogator/). Each of these prompts was manually edited to produce results as photorealistic as possible and remove living persons and artists names.</p><p> </p><p>In addition to this, parameters were randomly selected within reasonable values for methods requiring so.</p><p>The prompts and parameters used for each method can be found in the `prompts.csv` file.</p><p> </p><p>This dataset can be used to evaluate AI-generated image detection methods. We recommend matching the generated images with the real Raise-1k images, to evaluate whether the methods can distinguish the two of them. Raise-1k images are not included in the dataset, they can be downloaded separately at (http://loki.disi.unitn.it/RAISE/download.html).</p><p> </p><p>None of the images suffered degradations such as JPEG compression or resampling, which leaves room to add your own degradations to test robustness to various transformation in a controlled manner.</p><p> </p>
An urban traffic dataset composed of visible images and their semantic segmentation generated by the CARLA simulator
<p><strong>If you use this dataset please cite this paper: Rosende, S.B.; Gavilán, D.S.J.; Fernández-Andrés, J.; Sánchez-Soriano, J. An Urban Traffic Dataset Composed of Visible Images and Their Semantic Segmentation Generated by the CARLA Simulator. <em>Data</em> 2024, <em>9</em>, 4. <a href="https://doi.org/10.3390/data9010004">https://doi.org/10.3390/data9010004</a></strong></p> <p>A dataset of aerial urban traffic images and their semantic segmentation is presented to be used to train computer vision algorithms, among which those based on convolutional neural networks stand out. The images have been generated using the CARLA simulator (but would be like those that could be obtained with fixed aerial cameras or by using AUVs) in the field of intelligent transportation management. The presented dataset is available and accessible to improve the performance of vision and road traffic management systems, especially for the detection of incorrect or dangerous maneuvers.</p>
Touché25-Image-Retrieval-and-Generation-for-Arguments
<p>Data for the <a href="https://touche.webis.de/clef25/touche25-web/image-retrieval-for-arguments.html">Image Retrieval/Generation for Arguments</a> task at Touché 2025.</p> <p> </p> <p>Only the main.zip and nodes.zip are uploaded here due to space restrictions. Find the web page screenshots and web archives here: <a href="https://files.webis.de/corpora/corpora-webis/corpus-touche-image-search-25/version-2025-04-02/">https://files.webis.de/corpora/corpora-webis/corpus-touche-image-search-25/version-2025-04-02/</a></p>
deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques
<p>In this paper, we present datasets that can be utilised for synthetic near infrared (NIR) image and bounding box level fruit detection system. It is undeniable fact that high-caliber machine learning software frameworks such as Tensorflow or Pytorch and large scale dataset such as ImageNet and COCO, and accelerated GPU hardware support have pushed the limit of machine learning for more than decades.</p> <p>Among these breakthroughs quality dataset is one of important key building blocks that can lead to success in model generalisation and deployment for data-driven deep neural networks. Particularly, synthetic data generation such as generative adversarial networks often requires relatively larger scale data than other supervised approaches. In addition, posing constrains such as geometrical facial constrains in fake face generation or consistent and radiometrically calibrated reflectances from satellite imagery commonly yield better results. We share NIR+RGB dataset that are re-processed from other two public datasets (nirscene and SEN12MS) and our own novel sweetpepper dataset to be able to timely adopt to other following studies.</p> <p>We oversampled from original nirscene dataset at 10, 100, 200, and 400 ratios and total of 127k pair of images. For SEN12MS satellite multispectral dataset, we selected one largest subset; Summer (45k) and All seasons (180k). Our sweetpeppr dataset consists of 1,615 pairs of NIR+RGB images. We demonstrate these NIR+RGB datasets are sufficient to be used for synthetic NIR generation quantitatively and qualitatively. We achieved Frechet Inception Distance (FID) of 11.36, 26.53, and 40.15 for nirscene1, SEN12MS, and sweetpepper dataset respectively.</p> <p>We also release <em>11</em> fruits' bounding box annotations that can be exported as various formats using cloud service. 4 newly added fruits [blueberry, cherry, kiwi, and wheat] compounds 11 novel bounding box dastaset together with our previous work in deepFruits project [apple, avocado, capsicum, mango, orange, rockmelon, strawberry]. The total number of bounding box instances is 162k and all bounding box dataset is ready for use from cloud service. For evaluation of these dataset, Yolov5 single stage detector is exploited and reported impressive mean-average-precision, mAP[0.5:0.95] results of [min:0.49, max:0.812]. We hope these dataset is useful and serves as one of baseline for the following up studies.</p>
Carbon Nanotube Uptake in Cyanobacteria for Near-infrared Imaging and Enhancing Bioelectricity Generation in Living Photovoltaics
<p>Dataset of the work entitled "Carbon Nanotube Uptake in Cyanobacteria for Near-infrared Imaging and Enhancing Bioelectricity Generation in Living Photovoltaics".</p>
Figure 3: The microscopy images fo neuronal cells generated by SWCNT (a) and MWCNT (b)-COMPARATIVE STUDY OF SINGLE- AND MULTI-WALL CARBON NANOTUBES WITH APPLICATION IN CEREBRAL ANEURYSM
<p>Carbon nanotubes (CNTs) are nanometer-scale cylindrical graphitic struc-<br> tures that exhibit extraordinary physical properties as determined by their<br> structure [6]. Developing neural implants and the process of neuron regener-<br> ation are extremely di±cult. Nerve cells require the right environment and<br> the right growth factors at the right time to grow and proliferate. The elec-<br> trical conductive properties of these nanotubes o®er the possibility of using<br> it as a replacement to transmit and receive signals. The resulting 'hair like'<br> conductive wires that incorporate the properties of electrodes, permeable mi-<br> cro°uidic conduits and the porosity of the CNTs was found to promote cell<br> growth, migration and proliferation. The bridging consists either of an axon<br> or bundles of axons and dendrites. In some cases the bridge is covered with<br> clusters of cells [7]. These bridges form very e±ciently over quartz surfaces<br> which are apparently very poor surfaces for cell attachment. Fig. 2 shows the<br> evolution of a network generated by SWCNT and MWCNT. The data show<br> that cells ¯rst aggregate at the NT islands. As they complete this step axons<br> and dendrites begin to form and to build connections.<br> Also, has been observed for MWCNT higher connections than for SWCNT,<br> Figure 3.</p>
Figure 2: The microscopy images fo neuronal cells control (a) generated by MWCNT (b) and SWCNT (c)-COMPARATIVE STUDY OF SINGLE- AND MULTI-WALL CARBON NANOTUBES WITH APPLICATION IN CEREBRAL ANEURYSM
<p>Fig. 2 shows the evolution of a network generated by SWCNT and MWCNT. The data show<br> that cells ¯rst aggregate at the NT islands. As they complete this step axons and dendrites begin to form and to build connections.</p>
Touché24-Image-Retrieval-and-Generation-for-Arguments
<div> <p>Data for the <a href="https://touche.webis.de/clef24/touche24-web/image-retrieval-for-arguments.html">Image Retrieval/Generation for Arguments</a> task at Touché 2024.</p> <p>Only the main.zip and nodes.zip are uploaded here due to space restrictions. Find the web page screenshots and web archives here: <a href="https://files.webis.de/corpora/corpora-webis/corpus-touche-image-retrieval-and-generation-24/">https://files.webis.de/corpora/corpora-webis/corpus-touche-image-retrieval-and-generation-24/</a></p> </div>
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter
<p><strong>Update May 2024: Fixed a data type issue with "id" column that prevented twitter ids from rendering correctly.</strong></p> <p>Recent progress in generative artificial intelligence (gen-AI) has enabled the generation of photo-realistic and artistically-inspiring photos at a single click, catering to millions of users online. To explore how people use gen-AI models such as DALLE and StableDiffusion, it is critical to understand the themes, contents, and variations present in the AI-generated photos. In this work, we introduce TWIGMA (TWItter Generative-ai images with MetadatA), a comprehensive dataset encompassing 800,000 gen-AI images collected from Jan 2021 to March 2023 on Twitter, with associated metadata (e.g., tweet text, creation date, number of likes).</p> <p>Through a comparative analysis of TWIGMA with natural images and human artwork, we find that gen-AI images possess distinctive characteristics and exhibit, on average, lower variability when compared to their non-gen-AI counterparts. Additionally, we find that the similarity between a gen-AI image and human images (i) is correlated with the number of likes; and (ii) can be used to identify human images that served as inspiration for the gen-AI creations. Finally, we observe a longitudinal shift in the themes of AI-generated images on Twitter, with users increasingly sharing artistically sophisticated content such as intricate human portraits, whereas their interest in simple subjects such as natural scenes and animals has decreased. Our analyses and findings underscore the significance of TWIGMA as a unique data resource for studying AI-generated images.</p> <p>Note that in accordance with the privacy and control policy of Twitter, <strong>NO raw content from Twitter is included</strong> in this dataset and users could and need to retrieve the original Twitter content used for analysis using the Twitter id. In addition, users who want to access Twitter data should consult and follow rules and regulations closely at the official Twitter developer policy at https://developer.twitter.com/en/developer-terms/policy. </p> <p> </p> <p> </p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-In Figure 15 we were able to generate the symbol H without any help from a small image we used only for row and column data
<p>In Figure 15 we were able to generate the symbol H without any help from a small image we used only for row and column data.</p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-Figure 16. Failed Generations
<p>For some images there is a chance to get stuck where fitness function maxed, but we are not near to the original image like in figure 16, both row and column fitness matched. But image lost a key portion from original image, in these cases we should increase the weight of fitness function Fx(X), which will solve the issue.</p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-Figure 12. After few generation
<p>In figure 11 it is the initial population showed and figure 12 the population started to change and figure 13 we reached a convergence.</p>
Image Synthesis with a Convolutional Capsule Generative Adversarial Network- Dataset
<p><strong>Dataset of 152 two-photon images (512x512) of axons with segmentation labels. </strong></p> <p>We combined data from two published sources (Bass et al., 2017; Canty et al., 2018) to get 152 (512×512) 2D images (produced from 3D image stacks), and manually produced the corresponding labels. These images were collected using in-vivo two-photon microscopy from the mouse somatosensory cortex. To generate the 2D images, we used a max projection over the 3D stack. The labels are binary segmentation maps of the axons.</p> <p>This dataset is split into a train (132 images) and test (20 images) sets. The raw 2D images of axons are in /original folder, and the segmentation labels are in /mask folder.</p> <p><strong>Please cite the following paper when using this dataset:</strong></p> <p>Bass, C., Dai, T., Billot, B., Arulkumaran, K., Creswell, A., Clopath, C., De Paola, V., and Bharath, A. A., 2019. “Image synthesis with a convolutional capsule generative adversarial network,” <em>Medial Imaging with Deep Learning.</em></p> <p><strong>This dataset was complied from the following publications:</strong><br> Bass, C., Helkkula, P., De Paola, V., Clopath, C. and Bharath, A.A., 2017. Detection of axonal synapses in 3D two-photon images. PloS one, 12(9), p.e0183309.<br> Canty, A.J., Jackson, J.S., Huang, L., Trabalza, A., Bass, C., Little, G. and De Paola, V., 2018. Single-axon-resolution intravital imaging reveals a rapid onset form of Wallerian degeneration in the adult neocortex. <em>bioRxiv</em>, p.391425.</p>
Images of the data brushes generated for the ApPEARS deliverable D5.1 Database of vector based brush strokes and sample prints that demonstrate the range of printed materials
<p>These images are appendices of ApPEARS deliverable D5.1 Database of vector based brush strokes and sample prints that demonstrate the range of printed materials. They show the generated data brushes.</p>
Using CycleGANs to Generate Realistic STEM Images for Machine Learning
<p>This data set contains part of the images that were used in the manuscript "Using CycleGANs to Generate Realistic STEM<br> Images for Machine Learning", including experimental, simulated, and CycleGAN-processed monolayer WSe<sub>2</sub> images. The acquisition and simulated parameters are publicly available in the manuscript (arXiv:2301.07743).</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.