Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
145
datasets available to search
ShareScore release 0.7.1
Dataset results
145 results for “Image processing”
Advances in Leukemia Detection and Classification: A Systematic Review of AI and Image Processing Techniques
Open the record for dataset details and reuse information.
Multi-modality medical image dataset for medical image processing in Python lesson
<p>This dataset contains a collection of medical imaging files for use in the <a href="https://github.com/esciencecenter-digital-skills/medical-image-processing">"Medical Image Processing with Python" lesson</a>, developed by the <a href="https://www.esciencecenter.nl/">Netherlands eScience Center</a>. </p> <p>The dataset includes:</p> <ol> <li>SimpleITK compatible files: MRI T1 and CT scans (<em>training_001_mr_T1.mha, training_001_ct.mha</em>), digital X-ray (<em>digital_xray.dcm</em> in DICOM format), neuroimaging data (<em>A1_grayT1.nrrd, A1_grayT2.nrrd</em>). Data have been downloaded from <a href="https://insightsoftwareconsortium.github.io/SimpleITK-Notebooks/Python_html/00_Setup.html">here</a>. </li> <li>MRI data: a T2-weighted image (<em>OBJECT_phantom_T2W_TSE_Cor_14_1.nii</em> in NIfTI-1 format). Data have been downloaded from <a href="../records/6467772">here</a>. </li> <li>Example images for the machine learning lesson: chest X-rays (<em>rotatechest.png, other_op.png</em>), cardiomegaly example (<em>cardiomegaly_cc0.png</em>).</li> <li>Array data: Array data for the Intro to Medical Imaging lesson. Numpy arrays were created by processing and manipulation of publicly available data i.e. from <a href="https://doi.org/10.1109/TNS.1974.6499235">the Schepp Logan phantom</a> and from the <a href="https://fastmri.med.nyu.edu/">NYU FastMRI dataset</a></li> <li>Data for the anonymization exercises: ultrasound (<em>identifiable_us.jpg</em>) dowloaded from <a href="https://www.flickr.com/photos/jcarter/2461223727">here</a>, and DICOM data (<em>our_sample_dicom.dcm</em>) shared for this course specifically by a colleague</li> <li>Histopathology data: histopathology slide images from <a href="https://openslide.org/">openslide</a> library samples in the freely distributable test data </li> </ol> <p>These files represent various medical imaging modalities and formats commonly used in clinical research and practice. They are intended for educational purposes, allowing students to practice image processing techniques, machine learning applications, and statistical analysis of medical images using Python libraries such as scikit-image, pydicom, and SimpleITK.</p>
Polarimetric Radar Signatures of Lightning Initiation Processes Imaged with VHF Broadband Interferometer
<p>This spreadsheet contains the details of the lightning events analyzed in this study, including the date, time, INTF mapping information, NLDN-reported information, and the positions and location errors referred from the combined analysis of the INTF and NLDN reports. </p> <p> </p>
Incorporating the image formation process into deep learning improves network performance
<p>These are representative source data for the main figures (Figs, 1d, 2a, 2d, 3c, 4b, 5a) in paper "Incorporating the image formation process into deep learning improves network performance".</p>
Synergistic olfactory processing for social plasticity in desert locusts (widefield imaging sub-repo-04)
<p>Main Zenodo repository: https://doi.org/10.5281/zenodo.10931394</p> <p>Data deposited here: Physiology\Cal520_Widefield\COL_HEX3_MOL_N2_OC3L2_ZHAE2\solitarious</p>
Synergistic olfactory processing for social plasticity in desert locusts (widefield imaging sub-repo-03)
<p>Main Zenodo repository: https://doi.org/10.5281/zenodo.10931394</p> <p>Data deposited here: Physiology\Cal520_Widefield\COL_HEX3_MOL_N2_OC3L2_ZHAE2\gregarious</p>
Synergistic olfactory processing for social plasticity in desert locusts (widefield imaging sub-repo-02)
<p>Main Zenodo repository: https://doi.org/10.5281/zenodo.10931394</p> <p>Data deposited here: Physiology\Cal520_Widefield\BERRYLEAF_COL_MOL_N2_ZHAE2\solitarious</p>
Synergistic olfactory processing for social plasticity in desert locusts (widefield imaging sub-repo-01)
<p>Main Zenodo repository: https://doi.org/10.5281/zenodo.10931394</p> <p>Data deposited here: Physiology\Cal520_Widefield\BERRYLEAF_COL_MOL_N2_ZHAE2\gregarious</p>
Supporting dataset for "Magnetotelluric images of Paleoproterozoic accretion and Mesoproterozoic to Neoproterozoic reworking processes in the northern Sao Francisco Craton, central-eastern Brazil"
<p>This dataset presents .edi files from 38 magnetotelluric stations across the northern São Francisco Craton in central-eastern Brazil. Data collection was financially supported by CNPq Project 573713/2008-1.</p>
Data of Sedimentation process of ashfall during a Vulcanian eruption as revealed by high-temporal-resolution grain size analysis and high-speed camera imaging
<p>Data of ash falling rate, grain size analysis of ash deposit, and high-speed camera imaging of airborne ash particles used in "Sedimentation process of ashfall during a Vulcanian eruption as revealed by high-temporal-resolution grain size analysis and high-speed camera imaging". </p>
MATLAB Code for "Joint Image Processing with Learning-Driven Data Representation and Model Behavior for Non-Intrusive Anemia Diagnosis in Pediatric Patients"
<p>This MATLAB code is part of the study titled <em>"Joint Image Processing with Learning-Driven Data Representation and Model Behavior for Non-Intrusive Anemia Diagnosis in Pediatric Patients"</em>, which has been accepted for publication in the <em>Journal of Imaging (MDPI)</em>. The code supports image processing, feature extraction, and deep learning model training (including LSTM and RexNet) to classify pediatric patients as anemic or non-anemic based on palm, conjunctival, and fingernail images. Full study details are available in this paper:</p> <p>Berghout T. Joint Image Processing with Learning-Driven Data Representation and Model Behavior for Non-Intrusive Anemia Diagnosis in Pediatric Patients. <em>Journal of Imaging</em>. 2024; 10(10):245. <a href="https://doi.org/10.3390/jimaging10100245">https://doi.org/10.3390/jimaging10100245 </a></p> <p>The datsets use in this work are:</p> <p>Asare, J. W., Appiahene, P. & Donkoh, E. (2022). Anemia Detection using Palpable Palm Image Datasets from Ghana. Mendeley Data. https://doi.org/10.17632/ccr8cm22vz.1<br>Asare, J. W., Appiahene, P. & Donkoh, E. (2023). CP-AnemiC (A Conjunctival Pallor) Dataset from Ghana. Mendeley Data. https://doi.org/10.17632/m53vz6b7fx.1<br>Asare, J. W., Appiahene, P. & Donkoh, E. (2020). Detection of Anemia using Colour of the Fingernails Image Datasets from Ghana. Mendeley Data. https://doi.org/10.17632/2xx4j3kjg2.1</p>
Processing steps to generate a Digital Surface Model based on SPOT-7 tri-stereo images published in the study "An assessment of the effects of DEM quality and spatial resolution on a model for mapping lahar inundation areas at volcan Copahue (Argentina & Chile)" in the Journal of South American Earth Sciences https://doi.org/10.1016/j.jsames.2022.104138
<p>The Digital Surface Model (DSM) was created from SPOT-7 tri-stereo images for the Copahue volcano between the border of Argentina and Chile. Two versions of the DSM are provided: an unfiltered product and a final, filtered product. The final product has a spatial resolution of 5-m and was used for lahar inundation modeling for the Copahue volcano (Viotto, Toyos, and Bookhagen 2022, <a href="https://doi.org/10.1016/j.jsames.2022.104138">https://doi.org/10.1016/j.jsames.2022.104138</a> : An assessment of the effects of DEM quality and spatial resolution on a model for mapping lahar hazard inundation at Volcán Copahue (Argentina & Chile). <em>Journal of South American Earth Sciences</em> ). The dataset provided should be cited together with the article. </p> <p><strong>DSM processing </strong></p> <p>The source images were given by a SPOT-7 snow- and cloud-free triplet (Nadir, Backward and Forward) of 1.5 m spatial resolution from 19 April 2018 (SPOT Image, Airbus Defence and Space GmbH, distributed by CONAE; Dataset ID: <em>SEN_SPOT7_20180419_142955500_000</em>, delivered by CONAE as <em>DS_SPOT7_20180419</em>).</p> <p>The data were processed with the suite of digital photogrammetry tools AMES Stereo Pipeline ASP (Beyer et al., 2018). The procedure for the generation of the DSM is summarized by following steps: </p> <ol> <li> <p>The orbital parameters (RCP models) were adjusted using the bundle adjustment tool with no ground control points, since they were unavailable.</p> </li> <li> <p>The scenes were map-projected onto the NASADEM (spatial resolution of 30 m) elevation dataset, assisted by the results of the orbital adjustment in Step 1.</p> </li> <li>The stereo correlation of the map-projected scenes including the results of the adjusted orbital parameters, was performed three times, using as first scene (i.e., primary image) the nadir (N), backward (B), and forward (F) images . In each run, the order of images to perform the stereo correlation was: N-F-B, F-N-B, and B-N-F. Thus, three point clouds were generated. Specific ASP correlator settings (other than defaults parameters; for details see the provided stereo-default file) were set in the following way: <em>Correlation Kernel</em>: 15 x 15 pixels; <em>Sub-pixel Refinement Kernel</em>: 21 x 21 pixels; <em>Subpixel Refinement Mode</em>: 2 (Weighted Affine Adaptive Window Correlator EM)</li> <li> <p>The three point clouds were merged into one point cloud with a regular grid of 5 m (unfiltered product, known as <em>DSM_Copahue_UTM19S_WGS84_5m_raw.tif</em>).</p> </li> </ol> <p>The quality of the final point cloud was assessed by comparing the unfiltered DSM with a spatial resolution of 12-m against the WorldDEM<sup>TM</sup> elevation dataset (Collins et al., 2015). The WorldDEM was provided by Airbus Defence and Space GmbH under license for the scope of the Viotto et al., 2022 study. The comparison of the pixel-to-pixel heights above the ellipsoid (WGS84) between the two datasets resulted in a mean difference of 0.67 m and a standard deviation of +/- 4.82 m. </p> <p>Comprehensive details on the methodologies evaluated to create the dataset with ASP, can be found in the corresponding master's thesis “Topografía digital y modelado de lahares en el Volcán Copahue, Argentina-Chile” from S. Viotto (link: https://rdu.unc.edu.ar/handle/11086/15384). Recommended literature about processing DEMs from SPOT imagery is given by Mueting et al., 2021 (<a href="https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2021JF006330">https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2021JF006330</a>). </p> <p><strong>Creation of the Final, Filtered DSM product</strong></p> <p>The corrections and improvements applied to the unfiltered product to create the final, filtered DSM (named DSM_Copahue_UTM19S_WGS84_5m_VoidFilled.tif) are summarized by following steps. </p> <p> </p> <ol> <li> <p><em>Water Bodies Delineation</em></p> </li> </ol> <p>The delineation of the water bodies was based on a mask created from the free access water bodies datasets provided by the Instituto Geográfico Nacional of Argentina (<a href="https://www.ign.gob.ar/NuestrasActividades/InformacionGeoespacial/CapasSIG">https://www.ign.gob.ar/ NuestrasActividades/InformacionGeoespacia l/CapasSIG</a>) and by the Ministerio de Bienes Nacionales in Chile ( <a href="https://www.ide.cl/index.php/aguas-continentales/item/1508-catastro-de-lagos">https://www.ide.cl/index.php /aguas-continentales/item/1508-catastro-de-lagos</a>). A total of 45 lakes within the area of interest were considered. Lakes with areas below or equal to 25 m2 were smoothed with a median filter in the last step. Lakes with areas above this threshold were filled in with a constant value and their borders were smoothed with a median filter to provide smooth shorelines.</p> <p><em>2 . Void Filling</em></p> <p>Voids (other than water bodies) were filled with the tool “Close Gaps” from Saga GIS software. </p> <p><em>3. Smoothing</em></p> <p>Finally, the elevation dataset was smoothed with a median filter using a 3 x 3 pixel window, excluding water bodies filled in the step 1. </p> <p><strong>Final Remarks and Suggestion</strong></p> <p>The quality assessment of the final version by visual inspection of the hillshades suggested an improvement of the signal to noise ratio. However, the void filling process may be improved.</p> <p><br> </p> <p><strong>Dataset Description</strong></p> <table align="center"> <caption> </caption> <tbody> <tr> <td>Digital Surface Models</td> <td> <p>No Data Value = -9999</p> <p>Format = float 32 bit</p> <p>File Format = GeoTiff</p> <p>Vertical Datum: WGS84</p> <p>Projection information: EPSG 32719 (UTM19S)</p> <p>Spatial Resolution: 5m (subfix: <em>_5m</em>) </p> <p>Versions: </p> <ul> <li> <p>Unfiltered product: without corrections <em>DSM_Copahue_UTM19S_WGS84_5m_raw.tif</em></p> </li> <li> <p>Final, filtered product: smoothed and void filled <em>DSM_Copahue_UTM19S_WGS84_5m_VoidFilled.tif</em></p> </li> </ul> </td> </tr> <tr> <td>Water Bodies Mask</td> <td> <p>No Lake Value = 0</p> <p>Lakes Values = 1 to 45</p> <p>File Format= GeoTiff</p> <p>Spatial Resolution: 5m (subfix: <em>_5m</em>)</p> <p>Projection information : EPSG 32719 (UTM19S)</p> <p><em>WB_mask_5m_UTM19S.tif</em></p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p><strong>Repository structure</strong></p> <p>|__ 01_Scripts</p> <p> |+ run21_CopahueDSM_AMES_sviotto.sh</p> <p> |+ stereo.default</p> <p>|__ 02_DSMs</p> <p> |+ DSM_Copahue_UTM19S_WGS84_5m_raw.tif</p> <p> |+ DSM_Copahue_UTM19S_WGS84_5m_VoidFilled.tif</p> <p> |+ WB_mask_5m_UTM19S.tif</p> <p><strong>References</strong></p> <p>Beyer, R. A., Alexandrov, O., & McMichael, S. (2018). The Ames Stereo Pipeline: NASA's open source software for deriving and processing terrain data. <em>Earth and Space Science</em>, 5, 537– 548. <a href="https://doi.org/10.1029/2018EA000409">https://doi.org/10.1029/2018EA000409</a></p> <p>Collins, J., Riegler, G., Schrader, H., Tinz, M., 2015. Applying terrain and hydrological editing to TanDEM-X data to create a consumer-ready worlddem product. Int. Arch. Photogram. Rem. Sens. Spatial Inf. Sci. 40 (7), 1149. https://doi.org/10.5194/isprsarchives-XL-7-W3-1149-2015.</p> <p>Mueting, A., Bookhagen, B., & Strecker, M. R. (2021). Identification of debris-flow channels using high-resolution topographic data: A case study in the Quebrada del Toro, NW Argentina. <em>Journal of Geophysical Research: Earth Surface</em>, 126, e2021JF006330. <a href="https://doi.org/10.1029/2021JF006330">https://doi.org/10.1029/2021JF006330</a></p> <p>Viotto, S., Toyos, G., & Bookhagen, B. (2022). An assessment of the effects of DEM quality and spatial resolution on a model for mapping lahar hazard inundation at volcán copahue (Argentina & Chile). Journal of South American Earth Sciences, 104138. https://doi.org/10.1016/j.jsames.2022.104138</p> <p> </p> <p> </p>
Our processed LoveDA dataset for "LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images"
<p>Our processed LoveDA dataset is used for the paper "<a href="https://doi.org/10.7717/peerj-cs.1467">LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images</a>"</p>
Our processed CITY_OSM dataset for "LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images"
<p>Our processed CITY_OSM dataset is used for the paper "<a href="https://doi.org/10.7717/peerj-cs.1467">LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images</a>".</p>
2023 IEEE SPS Video and Image Processing (VIP) Cup: Ophthalmic Biomarker Detection
<p>Ophthalmic clinical trials that study treatment efficacy of eye diseases are performed with a specific purpose and a set of procedures that are predetermined before trial initiation. Hence, they result in a controlled data collection process with gradual changes in the state of a diseased eye. In general, these data include 1D clinical measurements and 3D optical coherence tomography (OCT) imagery. Physicians interpret structural biomarkers for every patient using the 3D OCT images and clinical measurements to make personalized decisions for every patient.</p> <p>Two main challenges in medical image processing has been <em>generalization</em> and <em>personalization</em>.</p> <p>Generalization aims to develop algorithms that work well across diverse patients and scenarios, providing standardized and widely applicable solutions. Personalization, in contrast, tailors algorithms to individual patients based on their unique characteristics, optimizing diagnosis and treatment planning. Generalization offers broad applicability but may overlook individual variations. Personalization provides tailored solutions but requires patient-specific data. While deep learning has shown an affinity towards generalization, it is lacking in personalization.</p> <p>The presence and absence of biomarkers is a personalization challenge rather than a generalization challenge. The variation within OCT scans of patients between visits can be minimal while the difference in manifestation of the same disease across patients may be substantial. The domain difference between OCT scans can arise due to pathology manifestation across patients, clinical labels, and the visit along the treatment process when the scan is taken. Morphological, texture, statistical and fuzzy image processing techniques through adaptive thresholds and preprocessing may prove substantial to overcome these fine-grained challenges. This challenge provides the data and application to address personalization.</p>
2023 IEEE SPS Video and Image Processing (VIP) Cup: Ophthalmic Biomarker Detection Phase 2
<p>Ophthalmic clinical trials that study treatment efficacy of eye diseases are performed with a specific purpose and a set of procedures that are predetermined before trial initiation. Hence, they result in a controlled data collection process with gradual changes in the state of a diseased eye. In general, these data include 1D clinical measurements and 3D optical coherence tomography (OCT) imagery. Physicians interpret structural biomarkers for every patient using the 3D OCT images and clinical measurements to make personalized decisions for every patient.</p> <p>Two main challenges in medical image processing has been <em>generalization</em> and <em>personalization</em>.</p> <p>Generalization aims to develop algorithms that work well across diverse patients and scenarios, providing standardized and widely applicable solutions. Personalization, in contrast, tailors algorithms to individual patients based on their unique characteristics, optimizing diagnosis and treatment planning. Generalization offers broad applicability but may overlook individual variations. Personalization provides tailored solutions but requires patient-specific data. While deep learning has shown an affinity towards generalization, it is lacking in personalization.</p> <p>The presence and absence of biomarkers is a personalization challenge rather than a generalization challenge. The variation within OCT scans of patients between visits can be minimal while the difference in manifestation of the same disease across patients may be substantial. The domain difference between OCT scans can arise due to pathology manifestation across patients, clinical labels, and the visit along the treatment process when the scan is taken. Morphological, texture, statistical and fuzzy image processing techniques through adaptive thresholds and preprocessing may prove substantial to overcome these fine-grained challenges. This challenge provides the data and application to address personalization.</p> <p> </p> <p>These files constitute the second phase of the VIP CUP 2023 Challenge at ICIP 2023. This test set has a more general patient base than the first one and as such is a better indicator of the performance of models. This test set was created by taking a subset of the data from a publicly available OCT dataset and then asking our medical partners to provide fine-grained biomarker labels for the competition. We provide the citation for the source of these images below: </p> <p>Kermany D, Goldbaum M, Cai W et al. Identifying Medical Diagnoses and Treatable Diseases by Image-Based Deep Learning. Cell. 2018; 172(5):1122-1131. doi:10.1016/j.cell.2018.02.010.</p> <p> </p> <p>This zenodo repository contains the images and submission template file needed for the second phase of the competition.</p>
Dispersion Engineered Metasurfaces for Broadband, High-NA, High-Efficiency, Dual-Polarization Analog Image Processing
<p>This dataset contains tabulated versions of the experimental data shown in the pictures of the publication "Dispersion Engineered Metasurfaces for Broadband, High-NA, High-Efficiency, Dual-Polarization Analog Image Processing"</p><p><strong>Figure 2</strong></p><ul><li>The file Fig_2_tpp.xlsx contains the raw data of Fig. 2g. The first sheet contains a 31 x 121 matrix. The second sheet contains a column vector with 121 entries, corresponding to the wavelength axis. The third sheet contains a column vector with 31 entries, corresponding to the angle theta axis.</li><li>The file Fig_2_tss.xlsx contains the raw data of Fig. 2i. The first sheet contains a 31 x 121 matrix. The second sheet contains a column vector with 121 entries, corresponding to the wavelength axis. The third sheet contains a column vector with 31 entries, corresponding to the angle theta axis.</li><li>The data of the plots in Fig. 2h and 2j can be extracted from the corresponding row/columns in the files Fig_2_tpp.xlsx and Fig_2_tss.xlsx</li></ul><p><strong>Figure 4</strong></p><ul><li>The file Fig_4b_MSoff.xlsx contains the raw data of the first panel ('NO Metasurface') of Fig. 4b</li><li>The files Fig_4b_MSon_X.xlsx, with X = [1435, 1445, 1450, 1452, 1455:5:1480, 1483, 1485, 1490, 1495], contain the raw data of the other panels of Fig. 4b.</li><li>For all panels of Fig. 4b, each pixel corresponds to 0.2864 microns.</li><li>The data of the plots in Fig. 4c can be extracted from the corresponding row/columns of the plots in Fig. 4b</li></ul><p><strong>Figure 5</strong></p><ul><li>The file Fig_5a.xlsx contains the spectra shown in Fig. 5a</li><li>The files Fig_5b_MSoff.xlsx, Fig_5b_MSon_xpol.xlsx, Fig_5b_MSon_ypol.xlsx, Fig_5b_MSon_+45.xlsx, Fig_5b_MSon_-45.xlsx, Fig_5b_MSon_Lpol.xlsx, Fig_5b_MSon_Rpol.xlsx contains the raw data of the panels b, c, d, e, f, g, h, respectively</li><li>For all 2D images in Fig. 5, each pixel corresponds to 0.2864 microns.</li><li>The data of the plots in Fig. 5i can be extracted from the corresponding row/columns of the plots in Figs. 5c-5h</li></ul>
The Effectiveness of 4D Image Acquisition and Post-processing With Vios Works
ClinicalTrials.gov study NCT03128268. IPD Sharing: NO. Countries: 1. Publications: 7.
The Process of Blood Collection With a Vascular Imaging Device
ClinicalTrials.gov study NCT05678504. IPD Sharing: NO. Countries: 1. Publications: 10.
Brain Imaging, Attention, and Auditory Processing in Schizophrenia
ClinicalTrials.gov study NCT03068806. IPD Sharing: NO. Countries: 1. Publications: 41.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.