Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “image dataset”
A geotagged image dataset with compass directions for studying the drivers of farmland abandonment
<p>In this work, we present a dataset containing a collection of pictures taken during the fieldwork of a farmland abandonment study. Data was taken in 2010 with a compact camera which incorporates GPS and a digital compass sensor. The photographs are taken as a part of a GIS database. Using their Exif metadata we created a layer of Geographic Fields Of View (GeoFOVs) that can be used to perform very specific spatial queries. The dataset contains 2,235 pictures and GIS layers of GeoFOVs contextualizing the agricultural plots being photographed.</p>
Dataset for "Unsupervised Learning of Lagrangian Dynamics from Images for Prediction and Control"
<p>Dataset to reproduce results of "Unsupervised Learning of Lagrangian Dynamics from Images for Prediction and Control"</p> <p> </p> <p><strong>Abstract</strong>: Recent approaches for modelling dynamics of physical systems with neural networks enforce Lagrangian or Hamiltonian structure to improve prediction and generalization. However, when coordinates are embedded in high-dimensional data such as images, these approaches either lose interpretability or can only be applied to one particular example. We introduce a new unsupervised neural network model that learns Lagrangian dynamics from images, with interpretability that benefits prediction and control. The model infers Lagrangian dynamics on generalized coordinates that are simultaneously learned with a coordinate-aware variational autoencoder (VAE). The VAE is designed to account for the geometry of physical systems composed of multiple rigid bodies in the plane. By inferring interpretable Lagrangian dynamics, the model learns physical system properties, such as kinetic and potential energy, which enables long-term prediction of dynamics in the image space and synthesis of energy-based controllers.</p>
Scanning electron microscopy image dataset -- Abundances and morphotypes of the coccolithophore Emiliania huxleyi in southern Patagonian fjords and channels
<p>Data set 1: E.huxleyi_morphotypes_Patagonia.zip</p> <p>Scanning electron microscopy images of <em>Emiliania huxleyi</em> cells found inhabit the southern Patagonia fjords during the late-spring 2015 and early-spring 2017.</p> <p> </p> <p>Data set 2: E.huxleyi_abundances_Patagonia.zip</p> <p>Scanning electron microscope images of filters of plankton samples taken in 2015 and 2017 throughout southern Patagonia fjords.</p> <p>The "m" in sample name refer to the depth from which the sample was obtained. </p> <p>Tables are provided to associate <em>Emiliania huxleyi</em> morphotypes' counts and taxonomic identifications to environmental variables from the samples for which data was used in statistical analysis.</p>
Malaria Blood Smear Image Dataset Creation
<p><strong>Dataset Creation</strong></p> <p>The dataset was collected from Tanzania. We sought out ethical clearance that gave permission to collect samples of patients that tested positive for malaria and as well as negative. The blood samples were stained and images were taken using the iPhone 6s mounted on top of an Olympus microscope. Afterward, the images were labeled by the three Lab technologists by drawing bounding boxes around the malaria parasites and white blood cells. </p> <p> </p> <p>Ethical Statement</p> <p>The nationally recognized ethics committee at The University of Dodoma and Benjamin Mkapa Hospital Research Center approved this research. It granted permission to isolate the samples of the positive cases so as to capture images for the purpose of this research. The data was collected from patients with suspected cases of malaria who willingly went to the hospital for diagnosis and treatment. We took images of what the lab technician was examining under the microscope. To avoid privacy violations of patients, no details about the patient identity were taken for this research rather than the images of their stained blood samples, age, gender, and location.</p> <p> </p> <p>Sample Collected</p> <p>The samples were collected from patients that had been requested, by a doctor, to receive a malaria test. A total of 40 cases were included in the present study,<em> </em>20 patients had positive confirmation of having malaria.<em> </em>Their mean age was 23.7 years (SD: 17.9 years) and 44.3% of cases were males and 53.7% were females. The mean ages of positive confirmed cases and negative confirmed cases were 23.30±17.7 and 25.89±18.7 years, respectively. All cases were residents of the Morogoro Region in Tanzania which have a higher rate of malaria patients.<em> </em></p> <p> </p> <p>Reagent Preparation</p> <p>Before subjecting a blood sample to a microscope for observation and image capturing, it had to be stained using a reagent. For that case, a buffer solution using 1 liter of distilled water and 1 buffer tablet were prepared with the aim of making a 7.2 PH solution. Thereafter, a Giemsa working solution was prepared by taking 2.5 ml of Giemsa stain stock into 25 ml of water making a 10% concentration. The working solution was then filtered using a circle filter paper. After filtration, the samples were placed horizontally and stained for 10 minutes. The stained samples were washed by using tap-water and placed vertically using a staining rack for the water to run off. At this stage, the dried stained blood samples were ready for observation under a 100 magnification of Olympus microscope. </p> <p> </p> <p>Image Collection</p> <p>This phase involved using a smartphone (iPhone 6s+) to capture images of stained blood smear that were observed under a microscope. A small portion of immersion oil was applied to the stained thick blood smear to enhance visibility. The slide was then placed under an Olympus CX 21 microscope for observation. The lens used had 100x magnification as recommended by the WHO (D Payne 1988). The microscope was continuously adjusted by a lab technician to ensure proper focus. At the same time, the iPhone 6s+ mobile phone was mounted to the microscope using the Labcam Microscope Adapter as shown in figure 1, and pictures were taken.</p> <p>The standard malaria diagnosis involves a lab technologist examining not less than 100 fields for a single slide under observation (D Payne, 1988). Therefore, approximately 100 images were captured for every blood smear slide placed under observation. For the 100 patients, we had a total of 100,000 images captured with 5000 images from positive patients. All 5000 images from 50 positive (infected) patients required annotation (labeling of the parasites and white blood cells). On the other hand, the 5000 images from uninfected patients did not require any annotation. The images captured were in JPG format, with a resolution of 4302 X 3204 pixels and a size of approximately 1 MB. The images were stored in a folder labeled with a date the slide was taken followed by a sample number for identification of the image.</p> <p> </p> <p>Image Annotation</p> <p>A team of three experts from the College of Health Science of the University of Dodoma and Benjamin Mkapa Hospital performed the annotation of the 2000 images altogether. The images were annotated using the LabelImg annotation tool. The annotation involved creating bounding boxes for the plasmodium and white blood cell classes. Annotators were instructed to label a target class by drawing the smallest possible box that contains all the visible parts of the plasmodium and the white blood cells. The output of the annotation was a Pascal VOC XML file with specific details on where the image is stored, the size of the image, filename, and coordinates of bounding boxes of all objects present in the image (Plasmodium and white blood cells). The time taken to annotate a single image with a fewer number of parasites, this means less than 20 parasites, took approximately 2 minutes while for a case with a higher number of parasites approximately more than 100 parasites, took around 15 to 20 minutes for a single image. In general for one patient, it took about 8 hours to annotate an image of the stained blood sample. The table below shows a summary of the dataset that was created in this stage.</p> <p> </p>
Dataset for interactive course on Deep Learning for Imaging
<p>This is a companion dataset for the course available at <a href="https://github.com/guiwitz/DLImaging">https://github.com/guiwitz/DLImaging</a> on Deep Learning for Imaging. It should be downloaded, decompressed, and placed at the same level as the "notebooks" folder of the course.</p> <p>The list of the original location of the data as well as their licenses can be found in the LICENSE file.</p>
Human Interaction Image (HII) dataset
<p>The Human Interaction Image (HII) dataset is a new dataset containing Web images from Commercial Search Engines (Google, Bing and Flickr). We use keyword search to collect images corresponding to four types of interactions: handshake, highfive, hug, kiss. Then we manually filter the irrelevant images. The dataset contains 2410 images with at least 550 images per interaction.</p> <p>The dataset can be applied, but not limited to the following research areas:</p> <ul> <li>interaction recognition/prediction</li> <li>action recognition</li> <li>video analysis</li> <li>transfer learning</li> </ul> <p>Please cite the following paper if you use the HII dataset in your work (papers, articles, reports, books, software, etc):</p> <ul> <li>J. Li, Y. Wong, Q.Zhao, M. Kankanhalli<br> <strong>Attention Transfer from Web Images for Video Recognition</strong><br> <em>ACM Multimedia</em>, 2017.<br> http://doi.org/10.1145/3123266.3123432</li> </ul> <ul> </ul>
A new dataset of rain cell generated from observations of the Tropical Rainfall Measuring Mission (TRMM) precipitation radar and visible and infrared scanner and microwave imager
<p>This new dataset (M.TRMM-1B01-1B11-2A25-PMD-Rain) contains orbit-level data with 5 km spatial resolution and 0.25 km vertical resolution. It is produced by merging TRMM PR, VIRS and TMI measurements at PR pixel resolution combined with rain cell identification. The near-surface rain rate, profiles of rain rate and precipitation reflectivity factor, visible and infrared signals and microwave signals can be obtained in the dataset. The dataset provides new important data for in-depth research on the structural characteristics of rain cells and supports the study of precipitation mechanisms.</p>
Truck Image Dataset
<p>Collection of annotated truck images, from a side point view, used to extract information about truck axles, collected on a highway in the State of São Paulo, Brazil. This is still a work in progress dataset and will be updated regularly, as new images are acquired. More info can be found on: <a href="https://www.researchgate.net/lab/Andre-Luiz-Cunha-Lab">Researchgate Lab Page</a>, OrcID Profiles, or <a href="https://github.com/labITS-stt-eesc">ITS Lab page on Github</a>.</p> <p>The dataset includes 1053 cropped images of trucks, with mixed real world trucks and synthetic trucks from Euro Truck Simulator 2.</p> <p>727 images were taken with three different cameras, on five different locations.</p> <ul> <li>727 images</li> <li>Format: JPG</li> <li>Resolution: 1920xVarious, 96dpi, 24bits</li> <li>Naming pattern: <video_name>_<color|gray>-<Region_of_Interest_ID>-<truck_ID>.jpg</li> </ul> <p>326 images were collected from <a href="https://truckersmp.com/" target="_blank" rel="noopener">Trucker's MP website</a>.</p> <ul> <li>326 images</li> <li>Format: JPG</li> <li>Resolution: 1920xVarious, 96dpi, 24bits</li> <li>Naming Pattern: <HEXID>.jpg</li> </ul> <p>All annotated objects were created with <a href="https://github.com/wkentaro/labelme">LabelMe</a>, and saved in JSON files for each image. For more information about the annotation format, please refer to the LabelMe documentation.</p> <p>Annotated objects are all related to truck axles, in 4 categories, Truck, Axle, Tandem, Tridem. Tandem is a double axle composition, and tridem is a triple axle composition. The number of objects in each category is as follows: </p> <ul> <li>Truck: 1053 </li> <li>Axle: 3927</li> <li>Tandem: 1172</li> <li>Tridem: 188</li> </ul> <p>If this dataset helps in any way your research, please feel free to contact the authors. We really enjoy knowing about other researcher's projects and how everybody is making use of the images on this dataset. We are also open for collaborations and to answer any questions. We also have a paper that uses this dataset, so if you want to officially cite us in your research, please do so! We appreciate it!</p> <p>Marcomini, Leandro Arab, and André Luiz Cunha. "Truck Axle Detection with Convolutional Neural Networks." <em>arXiv preprint <a href="https://arxiv.org/abs/2204.01868">arXiv:2204.01868</a></em> (2022).</p> <p> </p>
Remote sensing image classification dataset
Open the record for dataset details and reuse information.
Training dataset for land type detection on Sentinel-2 images (annotations: build-up, rural, forest, water)
Open the record for dataset details and reuse information.
Sofubi image and text description pair dataset
Open the record for dataset details and reuse information.
The image files for the WJocondeMM dataset
<p>The image files for the WJocondeMM dataset, contain 608 images of artworks.</p>
Craters in Historical Aerial Images (CHAI) Dataset
<p>This dataset contains 99 aerial images from Austria and Germany from 1943 - 1945. There are three versions of the dataset: <strong>CHAI-raw</strong>, <strong>CHAI-full</strong>, and <strong>CHAI-light</strong>. The CHAI-raw contains the original 99 historical aerial images with a Region of Interest (ROI) mask as well as the crater annotations. CHAI-full and CHAI-light are the derived datasets that were used for the evaluation of the paper <strong>"CHAI: Craters in Historical Aerial Images"</strong> presented at <a href="https://wacv2024.thecvf.com/">WACV2024</a>. For both datasets we extracted 960×960 patches with an overlap of 20%, both come with the images in .png format and the train, val, and test as .json files in the COCO style. The difference between the CHAI-full and CHAI-light datasets is that for the light version, all patches without any annotations have been removed, which results in the same amount of annotations, but fewer patches.</p><h2>Technical Details</h2><p>CHAI-raw contains all images unnormalized, additionally, the -info.csv contains the GSD in m, the -mask.png contains the region of interest (only in which craters were annotated), as well as the craters.csv and craters-manual-adapted.csv, which contain the craters and the manually refined craters, details can be found in the original publication. For ease of use, please consider using the derived datasets, which were used for the evaluation of our paper:<br><br>Will be linked once published.</p><p>Please cite the WACV paper when publishing results on these datasets.</p><h2>Access</h2><p>If you would like to request access to these files, please fill out the form below.</p><p>You need to satisfy these conditions in order for this request to be accepted:</p><p>The dataset is freely available for non-commercial research use. In order to get access to the dataset you have to fill in and sign the <a href="https://cvl.tuwien.ac.at/wp-content/uploads/2023/12/Form-Agreement-for-Usage-of-CHAI-Dataset.pdf">usage agreement form</a> and send it to <a href="https://cvl.tuwien.ac.at/staff/marvin-burges/">Marvin Burges</a><a href="mailto:sebastian.zambanini@tuwien.ac.at">.</a></p>
TINKER_WP3_TU49 2D and 3D image dataset_071123
<p>Contains images taken during and after the pick-and-place assembly of the RADAR use case for project TINKER. </p><p>Machine assembly, before curing: 1) empty cavity snapshot, 2) Epoxy bright exposure, 3) Epoxy dark exposure, 4) Die placement </p><p>After curing: 5) Die after curing, 6) 3D topology data after curing</p>
OrganoIDNetData: A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models
<p>OrganoIDNetData encompasses 180 images with 34113 organoids of human and murine Pancreatic Ductal Adenocarcinoma co-cultured with immune cells.</p> <p><strong>For further information and to cite this dataset, please refer to:</strong></p> <p>Kulkarni, A., Ferreira, N., Scodellaro, R. <em>et al.</em> A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models. <em>Sci Data</em> <strong>11</strong>, 820 (2024). https://doi.org/10.1038/s41597-024-03631-3</p>
Coastal Wetland Crab Image Dataset and Image Processing Model Weights
<p>This data was collected from the distribution range of mangroves along the coast of China and has been randomly selected and manually corrected to be labeled as a crab image analysis dataset. The weights of deep learning models trained on YOLOv5/v8 and EfficientNet are also uploaded simultaneously. Additionally, it includes the necessary test datasets and some test results. Please cite when using this dataset, and contact the administrator if you need help. The specific directories are as follows:</p> <p>- R-crab: Contains code needed to test model performance, with the subfolder data containing the test dataset required. It includes (1) crab-man as human-marked references, crab-ref as model detection results. (2) luoyuan-burrow for the detection results of crab burrows in the Luoyuan area case study, and luoyuan-crab for crab detection results, including crab classification, localization, and carapace width information. (3) method-test for testing different methods, i.e., whether to use a two-stage detection model. (4) size-conf-test records the model detection results under different image input sizes and confidence threshold levels. (1), (3), and (4) are completed in Out-of-sample data, while (2) is completed in Luoyuan. fig_attr.csv records the test results of image attributes on detection accuracy. label_results.csv records the comparison results between traits measured manually using ImageJ software and our designed model for detecting crab carapace width. luoyuan_list.csv records the numbering information of Luoyuan sampling plots.<br>- Aiweights: Contains model weights trained based on Object detection data, with cpm-model under v8n-seg-crab.pt for YOLOv8 trained to extract crab carapace width. Sfc-model under adam20.pth is a two-stage detection model trained based on EfficientNet. Trained_weights under burrow-baseline is a crab burrow detection model trained based on our improved YOLOv5 (improvements stored at https://github.com/GuuX29/crab-yolo-add, same below); frame-s6-2560.pt is a plot frame detection model for obtaining standard 50*50cm plot images; s-simam-20.pt and x-simam-20.pt are both for crab detection and classification models, with x having higher accuracy.<br>- Luoyuan: crop stores standardized processed Luoyuan image data, numbered as above, test stores detection results, same as R-crab.<br>- Object detection: Contains cropped 640 pixels crab field sampling images, with images in the images folder and bounding box labels in the labels folder, divided into training and testing at a 9:1 ratio, train.txt and val.txt record the allocation information.<br>- Out-of-sample: Records 100 independent test images not used to train the model, stored in images, crab-man, and crab-ref are the same as in R-crab.<br>- Segmentation: Stores image data used to test model detection of crab carapace width, with results also stored in R-crab.</p> <p>We hope this data and method will benefit the progress of research in this field. For any suggestions for improvement and help with usage, please contact the administrator. guuxuan1994@gmail.com </p>
Animal image dataset
Open the record for dataset details and reuse information.
k-space dataset for Robust multishot diffusion-weighted imaging of the abdomen with region-based shot rejection
<p>k-space dataset for https://onlinelibrary.wiley.com/doi/epdf/10.1002/mrm.30102</p>
NLSTseg: A Pixel-level Lung Cancer Dataset Based on NLST LDCT Images
Open the record for dataset details and reuse information.
Thin cloud removal dataset for Sentinel-2 images
<p>This is a thin cloud removal dataset for Sentinel-2A images in CR-GAN-PM[1] article. This dataset contains 20 paired thin cloud and clear Sentinel-2 images.</p> <p>If you use this dataset for your research, please cite us accordingly:</p> <p>#Reference: </p> <p>[1] J. Li, Z. W, Z. Hu, J. Z, M. Li, L. Mo and M. Molinier, “Thin cloud removal in optical remote sensing images based on generative adversarial networks and physical model of cloud distortion,” ISPRS J. Photogramm. Remote Sens., vol. 166, pp. 373–389, Aug. 2020. <a href="http://doi.org/10.1016/j.isprsjprs.2020.06.021">http://doi.org/10.1016/j.isprsjprs.2020.06.021</a>.</p> <p><br> [2] J. Li, Z. Wu, Z. Hu, Z. Li, Y. Wang, and M. Molinier, “Deep learning based thin cloud removal fusing vegetation red edge and short wave infrared spectral information for Sentinel-2A imagery,” Remote Sens., vol. 13, no. 1, p. 157, Jan. 2021. <a href="http://doi.org/10.3390/rs13010157">http://doi.org/10.3390/rs13010157</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.