Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38,240
datasets available to search
ShareScore release 0.7.1
Dataset results
38,240 results for “image”
Long-term seasonally and annually aggregated climatic variables for the greater Phoenix, Arizona, USA, metropolitan area and the surrounding Sonoran desert, derived from single-day NASA Daymet images, 2000 to 2022
This data package consists of multiple decades of bioclimatic raster data across the Central Arizona-Phoenix Long-Term Ecological Research (CAP LTER) study area within metropolitan Phoenix, Arizona, USA, temporally aggregated by year and by four meteorological seasons (winter, spring, summer, fall). We sourced each bioclimatic variable from 1-km resolution gridded estimates of daily climatic data from NASA Daymet V4, including daily mean (ppt) and total precipitation (ppt_sum), daily maximum air temperature (temp_max), daily minimum air temperature (temp_min), incident shortwave radiation flux density (srad), and daily average partial pressure of water vapor (vp). For each of these six variables, we created temporally aggregated raster images by calculating mean pixel-values of each for each season and year, as well as producing a seventh variable of seasonally and annually summed precipitation (ppt_sum). Finally, we exported images as individual GeoTIFF raster files, each with five bands corresponding values summarized annually (band 1) and seasonally (bands 2-5). All imagery retrieval and data processing were completed with Google Earth Engine (Gorelick et al. 2017) and program R. A complete description of data processing methods, including the aggregation of imagery by year and season, can be found in the data package metadata (see 'Methods and Protocols') and accompanying Javascript code. ### citations - Gorelick N, Hancher M, Dixon M, et al. (2017) Google Earth Engine: Planetary-scale geospatial analysis for everyone. Remote Sensing of Environment 202:18–27. https://doi.org/10.1016/j.rse.2017.06.031
PhenoCam Images and Canopy Phenology at the Harvard Forest EMS Tower since 2008
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (http://phenocam.sr.unh.edu) for more information and additional images.
R Code and Images for Developing a Spatial Concordance Coefficient at Harvard Forest 2010
Concordance correlation coefficients have been developed in a variety of different contexts. This problem has been widely addressed in a non-spatial context, but here we consider a coefficient that for a fixed spatial lag allows the comparison of two spatial sequences (e.g., images). We define a spatial concordance coefficient for second-order stationary processes.
PhenoCam Images and Canopy Phenology at the Harvard Forest LPH Tower 2010-2021
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (https://phenocam.nau.edu/webcam/) for more information and additional images.
PhenoCam Images and Canopy Phenology at the Harvard Forest Barn Tower since 2011
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (https://phenocam.nau.edu/webcam/) for more information and additional images.
PhenoCam Images and Canopy Phenology at the Harvard Forest Farm since 2015
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (https://phenocam.nau.edu/webcam/) for more information and additional images.
PhenoCam Images and Canopy Phenology at the Harvard Forest Hemlock Tower since 2010
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (https://phenocam.nau.edu/webcam/) for more information and additional images.
PhenoCam Images and Canopy Phenology at the Harvard Forest Witness Tree since 2014
The PhenoCam Network uses imagery from digital cameras to track vegetation phenology and seasonal changes in vegetation activity in diverse ecosystems across North America and around the world. Imagery is uploaded to the PhenoCam server at the University of New Hampshire, where it is made publicly available in near-real time, every 30 minutes from sunrise to sunset, 365 days a year. The data are processed using simple image analysis tools to yield a measure of canopy greenness, from which phenological metrics are extracted, characterizing the start and end of the growing season. These transition dates have been shown to align well with on-the-ground observations of tree phenology at Harvard Forest (HF003). Long-term PhenoCam data can be used to track the impact of climate variability and change on the rhythm of the seasons. This dataset contains one mid-day image for each camera. Please see the PhenoCam Network website (https://phenocam.nau.edu/webcam/) for more information and additional images.
Mueller matrix imaging combining optical parameters of mice non-melanoma skin cancer tissue
<p>The dataset consists of the Mueller matrix elements and optical parameters acquired from the backscattered light using a CCD camera and Mueller matrix imaging technique.</p><p>This dataset contains 90 samples including 20 feature vectors for SCC, 33 feature vectors for normal and 37 feature vectors for papilloma.</p>
Damage Localisation in Fresh Cement Mortar Observed via In Situ (Timelapse) X-ray uCT imaging.
<p>This is dataset to paper: Damage Localisation in Fresh Cement Mortar Observed via In Situ (Timelapse) X-ray uCT imaging.</p>
Disentangling the origins of confidence in speeded perceptual judgments through multimodal imaging
Open the record for dataset details and reuse information.
MASiVar: Multisite, Multiscanner, and Multisubject Acquisitions for Studying Variability in Diffusion Weighted Magnetic Resonance Imaging
Open the record for dataset details and reuse information.
Sharpening of Hierarchical Visual Feature Representations of Blurred Images
Open the record for dataset details and reuse information.
ColoPola: A dataset of colorectal cancer polarimetric images (Mueller matrix elements) for colorectal cancer detection
<p><strong>ColoPola</strong> dataset is <strong>Colo</strong>rectal cancer <strong>Pola</strong>rimetric images dataset</p> <p>The dataset consists of 572 slices (specimens) with 20,592 images, 284 slices of which were designated as cancer samples and 288 as normal samples.</p> <p>Each sample has 36 polarimetric images (i.e., HH, HV, HP, HM, HR, HL, VH, VV, VP, VM, VR, VL, PH, PV, PP, PM, PR, PL, MH, MV, MP, MM, MR, ML, RH, RV, RP, RM, RR, RL, LH, LV, LP, LM, LR, and LL).</p> <p>Each folder in the <strong>ColoPola</strong> dataset consists of 36 polarimetric images. Each image is 1280x1024 pixels in size and was created in the TIF file format (HH.tif, HV.tif, ..., LL.tif). </p>
DNA Origami Raw AFM Data - NanoLocz: Image analysis platform for AFM, high-speed AFM and localization AFM
<p>The data file is in the original ARIS data format as captured on a Cypher VRS1250 AFM (Oxford Instruments)<br><br><br></p>
Cebulka (Polish dark web cryptomarket and image board) messages data
<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Cebulka (Polish dark web cryptomarket and image board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Patrycja Cheba (Jagiellonian University); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish dark web cryptomarket and image board called Cebulka (<a href="http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php">http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php</a>). </p> <p>5. <strong>Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project focuses on studying internet behavior concerning disruptive actions, particularly emphasizing the online narcotics market in Poland. The research seeks to (1) investigate how the open internet, including social media, is used in the drug trade; (2) outline the significance of darknet platforms in the distribution of drugs; and (3) explore the complex exchange of content related to the drug trade between the surface web and the darknet, along with understanding meanings constructed within the drug subculture.</p> <p>Within this context, Cebulka is identified as a critical digital venue in Poland’s dark web illicit substances scene. Besides serving as a marketplace, it plays a crucial role in shaping the narratives and discussions prevalent in the drug subculture. The dataset has proved to be a valuable tool for performing the analyses needed to achieve the project’s objectives.</p> <h3><strong>Data Content</strong></h3> <p>6. <strong>Data Description</strong></p> <p>The data was collected in three periods, i.e., in January 2023, June 2023, and January 2024.</p> <p>The dataset comprises a sample of messages posted on Cebulka from its inception until January 2024 (including all the messages with drug advertisements). These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories. The “cebulka_adverts” directory contains posts related to drug advertisements (both advertisements and comments). In contrast, the “cebulka_community” directory holds a sample of posts from other parts of the cryptomarket, i.e., those not related directly to trading drugs but rather focusing on discussing illicit substances. The dataset consists of 16,842 posts.</p> <p>7. <strong>Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>8. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (“cebulka_adverts.zip” and “cebulka_community.zip”) containing all messages. These files are organized into individual directories that mirror the folder structure found on Cebulka.</li> <li>Two .csv files that list all the messages, including file names and the content of each post. The first .csv lists messages from “cebulka_adverts.zip,” and the second .csv lists messages from “cebulka_community.zip.”</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>9. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Pre-processed (in Detectron2 and YOLO format) planetary images and boulder labels collected during the BOULDERING Marie Skłodowska-Curie Global fellowship
<p>This database contains 4976 planetary images of boulder fields located on Earth, Mars and Moon. The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024. The data was already splitted into train, validation and test datasets, but feel free to re-organize the labels at your convenience. </p> <p>For each image, all of the boulder outlines within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript <a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). In addition, the boulder outlines were also pre-processed so that it can be ingested directly in YOLOv8.</p> <p>A description of what is what is given in the README.txt file (in addition in how to load the custom datasets in Detectron2 and YOLO). Most of the other files are mostly self-explanatory. Please see previous dataset or manuscript for more information. If you want to have more information about specific lunar and martian planetary images, the IDs of the images are still available in the name of the file. Use this ID to find more information (e.g., M121118602_00875_image.png, ID M121118602 ca be used on https://pilot.wr.usgs.gov/). I will also upload the raw data from which this pre-processed dataset was generated (see <a href="https://zenodo.org/records/14250970">https://zenodo.org/records/14250970</a>).</p> <p>Thanks to this database, you can easily train a Detectron2 Mask R-CNN or YOLO instance segmentation models to automatically detect boulders. </p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── boulder2024/ ├── jupyter-notebooks/ │ └── REGISTERING_BOULDER_DATASET_IN_DETECTRON2.ipynb ├── test/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── train/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── validation/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── detectron2_inst_seg_boulder_dataset.json ├── README.txt ├── yolo_inst_seg_boulder_dataset.yaml</code></pre> <p> </p> <pre><code>detectron2_inst_seg_boulder_dataset.json</code></pre> <p>is a json file containing the masks as expected by Detectron2 (see <a href="https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a> for more information on the format). In order to use this custom dataset, you need to register the dataset before using it in the training. There is an example how to do that in the jupyter-notebooks folder. You need to have detectron2, and all of its depedencies installed. </p> <pre><code>yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>can be used as it is, however you need to update the paths in the .yaml file, to the test, train and validation folders. More information about the YOLO format can be found here (<a href="https://docs.ultralytics.com/datasets/segment/">https://docs.ultralytics.com/datasets/segment/</a>).</p>
Product Images for Life Cycle Assessment Dataset For Peritoneal Dialysis and Haemodialysis in Modena
<p>The database contains a collection of images showcasing the individual components of peritoneal dialysis (PD) products, along with their corresponding weights. These images serve as a visual record for life cycle assessment (LCA) purposes, focusing on the material composition and environmental impact of each product.</p> <ol> <li> <p><strong>Patient Education Materials</strong>: Photographs of educational materials provided to patients, with accompanying data on the weight of the paper and packaging.</p> </li> <li> <p><strong>Catheters and Surgical Kits</strong>: Images display the disassembled components of PD catheters and surgical kits, including tubing, connectors, and packaging. Each image is annotated with the precise weight of the individual components.</p> </li> <li> <p><strong>Dialysis Solution Bags</strong>: The database includes images of both CAPD and APD solution bags, separated into their constituent parts (e.g., plastic bag, solution, and protective wrapping), with weights noted for each component.</p> </li> <li> <p><strong>Connection Devices and Consumables</strong>: Detailed images of connection devices, clamps, and other consumable items, with individual component weights clearly labeled.</p> </li> <li> <p><strong>Packaging and Transport Materials</strong>: Photographs of transport packaging, such as cardboard boxes and plastic wraps, alongside recorded weights for each element.</p> </li> <li> <p><strong>Maintenance Items</strong>: Visuals of terminal catheter sets, cleaning agents, and related products, each accompanied by their respective weight data.</p> </li> <li> <p><strong>Disposal Components</strong>: Images of used solution bags, syringes, and other single-use items, separated into recyclable and non-recyclable components, with weights specified for each.</p> </li> </ol> <p>This image-based database provides a clear and comprehensive reference for the material breakdown and weight distribution of PD product components, essential for conducting a thorough LCA and identifying areas for environmental improvement.</p>
Dataset of "Neutron imaging and molecular simulation of systems from methane and p‑xylene"
<p>The dataset contains parameterizations, and input files for molecular dynamics simulations used in the study of methane dissolution in p-xylene. For selected conditions, full simulation data, i.e., trajectories and energetics are provided. All used simulation results data are provided in the table, along with the measured experimental data.</p>
Example Microscopy Metadata JSON files produced using Micro-Meta App to document the acquisition of example images using a custom-built TIRF Epifluorescence Structured Illumination Microscope
<p><strong>Example Microscopy Metadata JSON files produced using the <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">Micro-Meta App</a> documenting an example raw-image file acquired using the custom-built TIRF Epifluorescence Structured Illumination Microscope.</strong></p> <p>For this use case, which is presented in Figure 5 of <a href="http://doi: https://doi.org/10.1101/2021.05.31.446382">Rigano et al., 2021</a>, Micro-Meta App was utilized to document:</p> <p>1) The <strong>Hardware Specifications</strong> of the custom build TIRF Epifluorescence Structured light Microscope (TESM; <a href="https://www.pnas.org/content/109/8/E471.long">Navaroli et al., 2010</a>) developed, built on the basis of the based on Olympus IX71 microscope stand, and owned by the Biomedical Imaging Group (http://big.umassmed.edu/) at the Program in Molecular Medicine of the University of Massachusetts Medical School. Because TESM was custom-built the most appropriate documentation level is <strong>Tier 3</strong> (<em>Manufacturing/Technical Development/Full Documentation</em>) as specified by the <a href="https://doi.org/10.5281/zenodo.4710731">4DN-BINA-OME</a> Microscopy Metadata model (<a href="https://doi.org/10.1101/2021.04.25.441198">Hammer et al., 2021</a>).</p> <p>The TESM Hardware Specifications are stored in: <strong>Rigano et al._Figure 5_UseCase_Biomedical Imaging Group_TESM.JSON</strong></p> <p>2) The <strong>Image Acquisition Settings</strong> that were applied to the TESM microscope for the acquisition of an example image (FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif) obtained by Nicholas Vecchietti and Caterina Strambio-De-Castillia. For this image, TZM-bl human cells were infected with HIV-1 retroviral three-part vector (FSWT+PAX2+pMD2.G). Six hours post-infection cells were fixed for 10 min with 1% formaldehyde in PBS, and permeabilized. Cells were stained with mouse anti-p24 primary antibody followed by DyLight488-anti-Mouse secondary antibody, to detect HIV-1 viral Capsid. In addition, cells were counterstained using rabbit anti-Lamin B1 primary antibody followed by DyLight649-anti-Rabbit secondary antibody, to visualize the nuclear envelope and with DAPI to visualize the nuclear chromosomal DNA.</p> <p>The Image Acquisition Settings used to acquire the FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif image are stored in: <strong>Rigano et al._Figure 5_UseCase_AS_fswt-6hvirus-10minfix-stk_4-epi.tif.JSON</strong></p> <p><em><strong>Instructional video tutorials on how to use these example data files:</strong></em><br> Use these videos to get started with using Micro-Meta App after downloading the example data files available here.</p> <ul> <li><a href="https://vimeo.com/562022222">Part 1/2</a></li> <li><a href="https://vimeo.com/562022281">Part 2/2</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.