Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “image dataset”
Geometric-Phase Microscopy for Quantitative Phase Imaging of Isotropic, Birefringent and SpaceVariant Polarization Samples_experimental dataset
<p>This dataset shares the data presented in the paper "Geometric-Phase Microscopy for Quantitative Phase Imaging of Isotropic, Birefringent and SpaceVariant Polarization Samples" available in open access under http://doi.org/10.1038/s41598-019-40441-9. </p>
Datasets used for Automatic Acquisition of Non-Saturated Hyperspectral Images
<p>data-sets acquired for studying correlation between automatic exposure times and hyper-spectral images, with the aim of devising procedures for automatic acquisition of non-saturated hyper-spectral images</p>
Image Dataset for FIQA
<p>Parte do banco de dados necessário para o treinamento do algoritmo de avaliação da qualidade das imagens de rostos. São <strong>392 </strong>imagens com resolução <strong>896*592</strong>, <strong>192 </strong>dpi e formato <strong>jpg</strong>. As imagens foram adquiridas do Caltech Face Dataset.</p>
Dataset for image analysis of LCDV susceptibility in sea bream
<p>This dataset comprises the photographs used to carry out the genetic estimates of LCDV susceptibility in sea bream. The numbers 1-499 correspond to batch1 and 500-100 to batch2</p>
AIDERv2 (Aerial Image Dataset for Emergency Response Applications)
<div>SUMMARY OF DATASET</div> <div> </div> <div>• This dataset consist of 167723 aerial images divided into 4 classes.</div> <div> </div> <div>• The dataset contains three commonly occurring natural disasters</div> <div>earthquake/collapsed buildings, flood, wildfire/fire, and a normal class; do not reflect any disaster</div> <div> </div> <div>• The images can be loaded as numpy arrays using Python programming language and then used to train a Convolutional Neural Network to detect natural disasters from aerial images.</div> <div> </div> <div>• The images are resized to 224x224x3 (heighty,width,channel number) when loaded as numpy arrays.</div> <div> </div> <div>• The dataset is an extension of the AIDER dataset (Aerial Image Dataset for Emergency Response Applications). </div> <div> </div> <div>• Additional images were collected by open source databases and extracted images as frames of videos downloaded from YouTube. </div> <div> </div> <div> </div> <div>The table below shows the number of images in each set.</div> <div> </div> <div> Train Validation Test Total</div> <div> Earthquakes 1927 239 239 2405</div> <div> Floods 4063 505 502 5070</div> <div> Fire 3509 439 436 4384</div> <div> Normal 3900 487 477 4864</div> <div> Total 13399 1670 1654 16723</div> <div> </div> <div> </div> <div>If you use this dataset please cite the following publications:</div> <div> </div> <div>[1] Shianios, D., Kyrkou, C., Kolios, P.S. (2023). A Benchmark and Investigation of Deep-Learning-Based Techniques for Detecting Natural Disasters in Aerial Images. In: Tsapatsoulis, N., <em>et al.</em> Computer Analysis of Images and Patterns. CAIP 2023. Lecture Notes in Computer Science, vol 14185. Springer, Cham. https://doi.org/10.1007/978-3-031-44240-7_24</div> <div>Link: https://link.springer.com/chapter/10.1007/978-3-031-44240-7_24</div> <div> </div> <div>[2] D. Shianios, P. Kolios, C. Kyrkou, "DiRecNetV2: A Transformer-Enhanced Network for Aerial Disaster Recognition", SN Computer Science, 2024 (Accepted to Appear)</div> <div> </div> <div> </div> <div> </div> <div> </div> <div>DATASET FOLDERS FORMAT</div> <div> </div> <div>└───data</div> <div>│ │</div> <div>│ └───Dataset_Images</div> <div>│ │ └───Train</div> <div>│ │ │ | └───Earthquake</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Flood</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Normal</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Wildfire</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ └───Val</div> <div>│ │ │ | └───Earthquake</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Flood</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Normal</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Wildfire</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ └───Test</div> <div>│ │ │ | └───Earthquake</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Flood</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Normal</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div>│ │ │ | └───Wildfire</div> <div>│ │ │ | img (1).jpg</div> <div>│ │ │ | img (2).jpg</div> <div>│ │ │ | .....</div> <div> </div> <div> </div> <div> </div> <div> </div> <div> </div> <div>DATA SOURCES AND DATA COLLECTION</div> <div> </div> <div>OPEN SOURCE DATABASES</div> <div> </div> <div>└───AIDER </div> <div>https://zenodo.org/record/3888300#.Yuu11nZBxD-</div> <div>Kyrkou, C. and Theocharides, T., 2020. EmergencyNet: Efficient aerial image classification for drone-based emergency monitoring using atrous convolutional feature fusion. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13, pp.1687-1699.</div> <div> </div> <div>└───ERA </div> <div>https://lcmou.github.io/ERA_Dataset/</div> <div>Mou, L., Hua, Y., Jin, P. and Zhu, X.X., 2020. Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets]. IEEE Geoscience and Remote Sensing Magazine, 8(4), pp.125-133.</div> <div>@article{eradataset,</div> <div> title = {{ERA: A dataset and deep learning benchmark for event recognition in aerial videos}},</div> <div> author = {Mou, L. and Hua, Y. and Jin, P. and Zhu, X. X.},</div> <div> journal = {IEEE Geoscience and Remote Sensing Magazine},</div> <div> year = {in press}</div> <div>}</div> <div> </div> <div> </div> <div> </div> <div>└───ISBDA</div> <div>https://drive.google.com/file/d/1kEKJ8kr1aScXz_1El7Mn-Yi0ANducQIW/view</div> <div>Zhu, X., Liang, J. and Hauptmann, A., 2021. Msnet: A multilevel instance segmentation network for natural disaster damage assessment in aerial videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 2023-2032).</div> <div>@misc{zhu2020msnet,</div> <div> title={MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos},</div> <div> author={Xiaoyu Zhu and Junwei Liang and Alexander Hauptmann},</div> <div> year={2020},</div> <div> eprint={2006.16479},</div> <div> archivePrefix={arXiv},</div> <div> primaryClass={cs.CV}</div> <div>}</div> <div> </div> <div> </div> <div>└───Floods 2013</div> <div>https://github.com/cvjena/eu-flood-dataset</div> <div>Barz, B., Schröter, K., Münch, M., Yang, B., Unger, A., Dransch, D. and Denzler, J., 2019. Enhancing flood impact analysis using interactive retrieval of social media images. arXiv preprint arXiv:1908.03361.</div> <div>@article{barz2019enhancing,</div> <div> title={Enhancing flood impact analysis using interactive retrieval of social media images},</div> <div> author={Barz, Bj{\"o}rn and Schr{\"o}ter, Kai and M{\"u}nch, Moritz and Yang, Bin and Unger, Andrea and Dransch, Doris and Denzler, Joachim},</div> <div> journal={arXiv preprint arXiv:1908.03361},</div> <div> year={2019}</div> <div>}</div> <div> </div> <div>└───Wildfire Research</div> <div>http://wildfire.fesb.hr/index.php?option=com_content&view=article&id=58&Itemid=54</div> <div> </div> <div> </div> <div>└───PyImages</div> <div>https://drive.google.com/file/d/1NvTyhUsrFbL91E10EPm38IjoCg6E2c6q/view</div> <div>The dataset was curated by PyImageSearch reader, Gautam Kumar.</div> <div> </div> <div> </div> <div> </div> <div> </div> <div>YOUTUBE VIDEOS</div> <div> </div> <div>└───Collapsed Buildings/Earthquakes</div> <div>• https://www.youtube.com/watch?v=TMow3WPcZrQ&t=133s&ab_channel=GORKHALYFOUNDATION</div> <div>• https://www.youtube.com/watch?v=_HT0tYKKjBI&t=47s&ab_channel=Effect.org</div> <div>• https://www.youtube.com/watch?v=rkb3y6K3waU</div> <div>• https://www.youtube.com/watch?v=yir6ArRZY4o&t=109s&ab_channel=UnicefUK</div> <div>• https://www.youtube.com/watch?v=CM9APmIR9Fk&ab_channel=ToonsZilla</div> <div>• https://www.youtube.com/watch?v=tmx2w6drAeU&ab_channel=AssociatedPress</div> <div>• https://www.youtube.com/watch?v=kuSEe8Emwrk&ab_channel=BloombergQuicktake%3ANow</div> <div>• https://www.youtube.com/watch?v=qoFHA3-m5ag&ab_channel=NBCNews</div> <div>• https://www.youtube.com/watch?v=MM3PToqEPhQ&ab_channel=GuardianNews</div> <div>• https://www.youtube.com/watch?v=zB_-TRnGuZE&ab_channel=DISASTERNEWS</div> <div>• https://www.youtube.com/watch?v=TqAMQQOEsBs&ab_channel=WHAS11</div> <div>• https://www.youtube.com/watch?v=0ixjTt-jmok&ab_channel=EveningStandard</div> <div>• https://www.youtube.com/watch?v=bNGA8Ms3d70&ab_channel=CatersClips</div> <div>• https://www.youtube.com/watch?v=wJ-2d5t23Lg&ab_channel=DailyDose</div> <div>• https://www.youtube.com/watch?v=ewUcI7I6Gf4&ab_channel=NBCNews</div> <div>• https://www.youtube.com/watch?v=Wx1cjOdlMZ4&ab_channel=ABCNews%28Australia%29</div> <div>• https://www.youtube.com/watch?v=jiMK_sVmbXk&t=12s&ab_channel=NewChinaTV </div> <div>• https://www.youtube.com/watch?v=M9au_9A2YRo&ab_channel=GuardianNews</div> <div>• https://www.youtube.com/watch?v=i6Lh8IXPjso&ab_channel=TheSun</div> <div>• https://www.youtube.com/watch?v=CKwxEr3I4Y8&ab_channel=GuardianNews</div> <div>• https://www.youtube.com/watch?v=hxqzcajBCNg&list=RDCMUCD3KREyo3IqCLBC-4khGgIw&index=3&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=2GEeTDuf9mI&list=RDCMUCD3KREyo3IqCLBC-4khGgIw&index=6&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=bDOuZWxIyNQ&list=RDCMUCD3KREyo3IqCLBC-4khGgIw&index=9&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=vzoSADijLCQ&list=RDCMUCD3KREyo3IqCLBC-4khGgIw&index=15&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=ZaL1fldTEAk&list=RDCMUCD3KREyo3IqCLBC-4khGgIw&index=17&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=QSV81FdilZE&ab_channel=GlobalNews</div> <div>• https://www.youtube.com/watch?v=KgOk9otW1Bg&ab_channel=EricFeijten</div> <div> </div> <div> </div> <div> </div> <div>└───Floods</div> <div>• https://www.youtube.com/watch?v=DJqgv8Sa5bA&t=317s&ab_channel=7NEWSAustralia</div> <div>• https://www.youtube.com/watch?v=w5FintiCLJU&t=9s&ab_channel=GuardianNews</div> <div>• https://www.youtube.com/watch?v=HjMymNN6Ajc&t=143s&ab_channel=BioLogicTreeServices</div> <div>• https://www.youtube.com/watch?v=Tmba18C94C8&ab_channel=AL.com</div> <div>• https://www.youtube.com/watch?v=Dqvpv4Vg4lk&t=63s&ab_channel=ElevenEleven</div> <div>• https://www.youtube.com/watch?v=N7QGicNtN2A&ab_channel=PKSVideoProductions</div> <div>• https://www.youtube.com/watch?v=8CHagyQG16Q&ab_channel=Stolly-Sven</div> <div>• https://www.youtube.com/watch?v=vjH3zFqdzcE&ab_channel=BenChilders</div> <div>• https://www.youtube.com/watch?v=GFw89UB4fE8&ab_channel=BenChilders</div> <div>• https://www.youtube.com/watch?v=heP3LEJ_NkE&ab_channel=7NEWSAustralia</div> <div> </div> <div> </div> <div> </div> <div>└───Fires</div> <div>• https://www.youtube.com/watch?v=gbM_NPx2GPc&t=201s&ab_channel=WXChasing</div> <div>• https://www.youtube.com/watch?v=M97sJdyeEM4&t=72s&ab_channel=Sanuck176</div> <div>• https://www.youtube.com/watch?v=1Z2K6lDt76M&t=557s&ab_channel=TheRelaxationChannel</div> <div> </div> <div>└───Normal</div> <div>• https://www.youtube.com/watch?v=SyxjsuNHWhM&t=328s&ab_channel=OneManWolfPack</div> <div>• https://www.youtube.com/watch?v=f1PTWsBtrtc&ab_channel=ChernobylPug</div> <div> </div> <div> </div> <div> </div> <div> </div> <div> </div>
Images dataset for Chemical Images Classifier model
<h1>Original paper</h1> <p>The manually curated images dataset is a part of the Supplementary Materials of the paper: <code>A. Krasnov, S. Barnabas, T. Böhme, S. Boyer, L. Weber, Comparing software tools for optical chemical structure recognition, Digital Discovery (2024). <a href="https://doi.org/10.1039/D3DD00228D">https://doi.org/10.1039/D3DD00228D</a></code><br><br></p> <h1>Images dataset description</h1> <p>The dataset was used to generate the image classifier model. The dataset consists of <strong>16,000</strong> images that were collected from different sources:</p> <p>1) Chemical data images extracted from EP, US, and WO patents by OntoChem GmbH.</p> <p>2) Images from the MolScribe datasets <code><a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480">https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480</a></code></p> <p>3) DECIMER–hand-drawn molecule images dataset <code>H.O. Brinkhaus, A. Zielesny, C. Steinbeck, K. Rajan, “DECIMER - hand-drawn molecule images dataset”, 2022, Journal of Cheminformatics, 14, 36. <a href="https://doi.org/10.1186/s13321-022-00620-9">https://doi.org/10.1186/s13321-022-00620-9</a></code></p> <p>4) Images from the Rxnscribe training set <code>Y. Qian, J. Guo, Z. Tu, C.W. Coley, R. Barzilay, “RxnScribe: A Sequence Generation Model for Reaction Diagram Parsing”, 2023, arXiv:2305.11845v1, <a href="https://doi.org/10.48550/arXiv.2305.11845">https://doi.org/10.48550/arXiv.2305.11845</a> </code></p> <p>5) Formulas images from the im2latex-100k dataset <code>A prebuilt dataset for OpenAI's task for image-2-latex system, <a href="../record/56198#.YJjuCGZKgox">https://zenodo.org/record/56198#.YJjuCGZKgox</a> (accessed 16 Januar 2024)</code></p> <h1>Structure of dataset</h1> <p>The dataset consists of two directories:</p> <p>The "<strong>classified</strong>" directory contains manually labeled images. These images are divided into four distinct categories, with each category including 4000 images:</p> <p>● one_molecule</p> <p>● several_molecules</p> <p>● reactions</p> <p>● other</p> <p>In the “<strong>for_model</strong>” folder, we have split the images for training, validation, and testing in order to create a Chemical Image Classifier model:</p> <p>● training: 12,804 images</p> <p>● test: 1,604 images</p> <p>● validation: 1,604 images.</p> <p> </p>
Image dataset for disease quantification in barley and wheat using the Macrobot system and BluVision Macro software
<p>This dataset contains multi-spectral high-resolution macroscopic images acquired using the Macrobot system, tailored for the automated analysis of plant disease phenotyping. The dataset includes images focused on three specific diseases: </p> <ul> <li>Bipolaris sorokiniana on barley, with images captured 8 days after inoculation </li> <li>Blumeria graminis f. sp. hordei resp. tritici (wheat powdery mildew) on wheat with images captured 8 days after inoculation</li> <li>Puccinia striiformis (yellow rust) on wheat, with images taken 15 days after inoculation</li> </ul> <p>For each sample, the dataset includes individual R, G, B, UV, and backlight illumination images, along with a combined RGB preview image.</p> <p>The BluVision Macro software is compatible with this dataset, enabling precise quantification and analysis of disease severity on barley and wheat leaves. </p>
INSA soil image dataset
<h2>Camera description</h2> <h3>1 RGB Camera</h3> <p><br>Images of each mixture are captured at a distance of ∼ 40cm from a Luxonis OAK-D Pro RGB camera under artificial (ambient) lighting. The camera hasthe following specifications:<br>• Resolution: 12MP (4056x3040)<br>• Frame rate: 60 FPS<br>• Lens Size: 1/2.3 inch<br>• 3 color channels Red+Green+Blue (RGB)</p> <h3>2 InfraRed Camera</h3> <p>To generate the image data under artificial (ambiant) lighting, an Aerial Map-ping Camera (MAPIR) is used which is a low-cost, high-resolution Near-InfraRed (NIR) camera capable of capturing multiple non-distorted images. Some specifications of the camera include:<br>• Resolution: 4000x3000, JPG (24bit)<br>• Shutter Speed: automatic<br>• Lens: 19mm 87° HFOV, F2.8 Fixed Aperture (Survey3W)<br>• Color channels NIR+Green+Blue (NGB)<br>Distance of cameras from soil: ∼40cm</p> <h2>Soil description</h2> <p>The dataset contains:<br>• 3 coarse soils: gravel, Hostun sand and Synthetic coarse sand. Different mixtures are realised with gravel/Hostun sand and gravel/Synthetic sand. Images are taken with the RGB camera.<br>• 2 fine soils: Kaolin (white clay) and silt of Dieppe. The soils were mixed with water at different water content values, and images are taken with the RGB and the InfraRed camera.</p> <p> </p> <h2>Acknowledgment</h2> <p><br>The authors thank the Department of Civil Engineering and Urban Planning of INSA Lyon (GCU INSA Lyon) for the donation and preparation of soils.</p> <p>This work was supported by ANR within the joint French-German project RE- MATCH ANR-21-FAI2-0009. We would like to express our sincere appreciation to them for their support.</p> <p> </p>
MyoD1 localization at the nuclear periphery is mediated by association of WFS1 with active enhancers - Image Dataset
<p>The set of raw images used in this study published on Nature Communications:</p> <p><a href="https://www.nature.com/articles/s41467-025-57758-x" target="_blank" rel="noopener">https://www.nature.com/articles/s41467-025-57758-x</a></p> <p>are available as 4 subsets below:</p> <p><a href="https://osf.io/w5n43/" target="_blank" rel="noopener">https://osf.io/w5n43/</a></p> <p><a href="https://osf.io/bfxdw/" target="_blank" rel="noopener">https://osf.io/bfxdw/</a></p> <div><a href="https://osf.io/rkguq/" target="_blank" rel="noopener">https://osf.io/rkguq/</a></div> <div> </div> <div><a href="https://osf.io/q6b4j/" target="_blank" rel="noopener">https://osf.io/q6b4j/</a> <div> </div> <div>The reason of this Zenodo entry is to combine these subsets into a single DOI link.</div> <div>The total size adds up to 200GB and it is not possible to host such a large dataset on Zenodo, and it is only possible as maximum of 50GB pieces on Open Science Framework.</div> <div> </div> <div>Summary of the study:</div> <div>Spatial organization of the mammalian genome influences gene expression and cell identity. While association of genes with the nuclear periphery is commonly linked to transcriptional repression, also active, expressed genes can localize at the nuclear periphery. The transcriptionally active MyoD1 gene, a master regulator of myogenesis, exhibits peripheral localization in proliferating myoblasts, yet the underlying mechanisms remain elusive. Using a newly generated reporter cell line, we demonstrate here that peripheral association of the MyoD1 locus is independent of mechanisms involved in heterochromatin anchoring. We identify a set of nuclear envelope transmembrane proteins, particularly WFS1, that actively tether MyoD1 to the nuclear periphery. WFS1 primarily associates with active distal enhancer elements upstream of MyoD1, and with a subset of enhancers enriched in active histone marks genome-wide, which are linked to expressed myogenic genes. Overall, our data identify a novel mechanism involved in tethering active genes to the nuclear periphery.</div> <div> </div> <div>This research was funded in whole or in part by the Austrian Science Fund (FWF) [P29713-B28, P32512-B and P36503-B] to Roland Foisner and a doctorate program funded by the Austrian Science Fund (FWF) [W1261-B28].</div> <div> </div> </div>
Test Dataset for 3D semantic image segmentation of the Breast, Fibrograndular Tissue, and Breast Carcinoma
Open the record for dataset details and reuse information.
Dataset: Combining video telemetry and wearable MEG for naturalistic imaging
<p>OPM-MEG and Openpose keypoint data from the study "Combining video telemetry and wearable MEG for naturalistic imaging".</p> <p><strong>Changelog</strong></p> <p><strong>v1.10</strong></p> <ul> <li>Subject 004 from v1.01 has been renamed 005 (to reflect addition of new subject recorded prior to 005 during acquisition).</li> <li><strong>NEW </strong>sub-004</li> <li>Subjects 003-004 have a proof-of-principle motor paradigm added.</li> <li>README changes</li> </ul> <p><strong>v1.01</strong></p> <ul> <li>Corrected sub-004 *_channel.json files to include bad channel identifiers</li> <li>Telemetry data zipped prior to uploading to zenondo</li> <li>Updates to README</li> </ul>
Datasets corresponding to "Direct evaluation of antiplatelet therapy in coronary artery disease by comprehensive image-based profiling of circulating platelets"
<p><strong>Datasets corresponding to "Direct evaluation of antiplatelet therapy in coronary artery disease by comprehensive image-based profiling of circulating platelets"</strong></p> <p><strong>02_CNN_PhenotypeClassif.7z</strong></p> <p>CNN Phenotype classification. Model was trained using AIDeveloper. using manually labelled data. Labelled Data is contained in folder "03_GatedData". The AIDeveloper session file in "02_Model\M10_Nitta6l_32pix_8class_meta.xlsx" shows, which files correspond to which subpopulation. The final model "M10_Nitta6l_32pix_8class_448.model" and corresponding .pb files are also located in that folder.</p> <p><strong>codeforclassification.zip</strong></p> <p>Code for applying model on unlabelled data. Test data is contained in 'sampledata.zip'</p> <p> </p> <p> </p> <p> </p>
Mass spectrometry Imaging dataset for a study on fungicide application to tomato leaves - II
<p>The dataset uploaded here is in association to a manuscript in press by Ajith et al., titled, "Visualizing active fungicide formulation mobility in tomato leaves with Desorption Electrospray Ionisation Mass Spectrometry Imaging". This dataset contains .imzML files along with zipped .ibd files of mass spectrometry imaging data and .mzML format LC-MS data for a fungicide applicaion study on tomato leaves. DESI Imprint imaging data here is for tomato leaves applied with Azoxystrobin standard after 24 hours, 56 hours and a week after formulation application along with a blank leaf data with no Azoxystrobin application.</p> <table> <tbody> <tr> <td><strong>File name</strong></td> <td><strong>Time point</strong></td> </tr> <tr> <td>DTIM_Std_24h</td> <td>24h</td> </tr> <tr> <td>DTIM_Std_56h</td> <td>56h</td> </tr> <tr> <td>DTIM_Std_1week</td> <td>1 week</td> </tr> <tr> <td>DTIM_Blank</td> <td> No application</td> </tr> </tbody> </table> <p>This data upload contains a zipped LC-MS data folder (LC_mzML.7z). The details of the files, the time point of sampling and part of leaf is in the excel sheet uploaded along with the data (LC-MS_datanames.xls).</p> <p> </p>
SeagrassFinder: An Underwater Eelgrass Image Classification Dataset
<p><span>This dataset is published as part of the publishing of the paper “SeagrassFinder: Deep Learning for Eelgrass Detection and Coverage Estimation in the Wild” in the Journal Ecological Informatics. The dataset is created as a machine learning dataset for training computer vision models to classify the presence of eelgrass. </span></p> <p><span>This dataset was created by the main author Jannik Els</span>äßer as part of his bachelor's thesis. The original video transect data in this dataset comes from DHI A/S work providing By og Havn a “Summer Status” report on the maritime environmental impacts of the Lynetteholm project. More information on the project and the report is available here: <span><a href="https://byoghavn.dk/mediebibliotek/lynetteholm-sommerstatus-2023/">https://byoghavn.dk/mediebibliotek/lynetteholm-sommerstatus-2023/</a></span><br><br>The dataset consists of underwater images taken on a sled, dragged through the water by a survey vessel. The camera used is a Subsea HD-Camera made by LH-Camera. Images were created by taking 5 video frames each second, and then randomly sampling. Each image is labeled True or False for eelgrass presence. In total, the dataset consists of 8500 images from 6 different transects, with 4482 images containing eelgrass, and 4042 images not containing eelgrass. All images have been annotated by a both domain-experts, and non-domain experts. Images were annotated using a uniform sampling process. In the occurrence of any disagreement between annotators, images have been removed from the dataset. For more information on the dataset creation, please refer to the corresponding paper.</p> <p>We recommend using one transect as a test dataset, and not using a random split of all images to create the test dataset. When using a random split of all images, a form of data leakage occurs, since some images can be very similar to other images.</p> <p>An unfortunate limitation, we believe caused by the compression of the videos in the camera system, is some frames contain an echo or form of motion trail. This can lead to ghost like eelgrass features in some frames. This should be taken into consideration when applying the dataset in future locations.</p>
Large Scale Medical Image Dataset
<p>Multiple dataset from different sources has been aggregated to create a large-scale medical image benchmark dataset in order to measure its performance. As each of the dataset’s images are of different sizes, the images are resized to 3 × 224 × 224 before the training process. This dataset contains total of 35 diseases of 4 different modality and is divided into train, validation, and test with a ratio of 7 : 1 : 2.</p>
Dataset of histopathological image crops from GTEx project
<p>This is a dataset of histological slides from the GTEx project that has been balanced for 3 major factors (organ, sex, and age bracket) that may be useful to train models in supervised or self-supervised modes.</p> <p>Four datasets are avaialble:</p> <ul> <li><code>gtex_histology_balanced_3_slides_200_tiles.tar.gz</code>: Conditioned on the 3 factors, 3 slides were selected per group, and 200 tiles in tissue segmented areas selected randomly per slide.</li> <li><code>gtex_histology_balanced_3_slides_2000_tiles.tar.gz</code>: Conditioned on the 3 factors, 3 slides were selected per group, and 2000 tiles in tissue segmented areas selected randomly per slide.</li> <li><code>gtex_histology_balanced_10_slides_100_tiles.tar.gz</code>: Conditioned on the 3 factors, 10 slides were selected per group (when possible), and 100 tiles in tissue segmented areas selected randomly per slide. This dataset matches closely the "gtex_histology_balanced_3_slides_200_tiles.tar.gz" dataset in total number of tiles.</li> <li><code>gtex_histology_balanced_10_slides_800_tiles.tar.gz</code>: Conditioned on the 3 factors, 10 slides were selected per group (when possible), and 800 tiles in tissue segmented areas selected randomly per slide. This dataset matches closely the "gtex_histology_balanced_3_slides_200_tiles.tar.gz" dataset in total number of tiles.</li> </ul> <p>Each archive file contains the following:</p> <ul> <li><code>slide_annotation.csv</code>: a slide-level annotation of the slides (see below)</li> <li><code>train</code>: a directory with image tiles to be used to train a model</li> <li><code>valid</code>: a directory with image tiles to be used to validate a model</li> </ul> <p>The slide_annotation file contains publicly available information on the slides in addition to 3 columns:</p> <ul> <li>"Tissue_simple": the organ of the slide</li> <li>"split": whether the slide was assign the 'train' or 'valid' split for training. The validation split slides have 1/10th of the tiles from training.</li> <li>"n_tiles": the number of image tiles in the dataset for each slide</li> </ul> <p>Example:</p> <table> <tbody> <tr> <td>Tissue Sample ID</td> <td>Tissue</td> <td>Subject ID</td> <td>Sex</td> <td>Age Bracket</td> <td>Hardy Scale</td> <td>Pathology Categories</td> <td>Pathology Notes</td> <td>Tissue_simple</td> <td>split</td> <td>n_tiles</td> </tr> <tr> <td>GTEX-1128S-1426</td> <td>Esophagus - Mucosa</td> <td>GTEX-1128S</td> <td>female</td> <td>60-69</td> <td>Fast death - natural causes</td> <td> </td> <td>6 pieces, near- total autolysis/mucosa completely sloughed</td> <td>Esophagus</td> <td>train</td> <td>200</td> </tr> <tr> <td>GTEX-113JC-1226</td> <td>Stomach</td> <td>GTEX-113JC</td> <td>female</td> <td>50-59</td> <td>Fast death - natural causes</td> <td> </td> <td>6 pieces, well dissected mucosa; some areas are severely autolyzed</td> <td>Stomach</td> <td>valid</td> <td>20</td> </tr> <tr> <td>GTEX-1192W-2526</td> <td>Muscle - Skeletal</td> <td>GTEX-1192W</td> <td>male</td> <td>60-69</td> <td>Fast death - natural causes</td> <td> </td> <td>2 pieces, ~10-20% interstitial fat, rep foci delineated</td> <td>Muscle</td> <td>train</td> <td>200</td> </tr> <tr> <td>GTEX-1192X-0426</td> <td>Muscle - Skeletal</td> <td>GTEX-1192X</td> <td>male</td> <td>50-59</td> <td>Slow death</td> <td> </td> <td>2 pieces, 5-10% interstitial fat, rep. foci delineated</td> <td>Muscle</td> <td>valid</td> <td>20</td> </tr> <tr> <td>GTEX-11DXX-1326</td> <td>Stomach</td> <td>GTEX-11DXX</td> <td>female</td> <td>60-69</td> <td>Ventilator case</td> <td>gastritis</td> <td>6 pieces, mild chronic active gastritis</td> <td>Stomach</td> <td>train</td> <td>200</td> </tr> </tbody> </table> <p>Inside <code>train</code> and <code>valid</code> and JPEG files named with the following convention: <code><Tissue Sample ID>.<Tissue_simple>.<Sex>.<Age Bracket>.<Y position>.<X position>.jpg</code> such that the origin of the crops can be traced and the file name serve as a direct class label if desired.</p> <p>Examples: "GTEX-ZYT6-1326.Pancreas.male.30-39.47492.16064.jpg", "GTEX-WWYW-2726.Ovary.female.50-59.5024.15008.jpg.</p>
Brazilian highways dataset - satellite images
<p>The <strong>Brazilian Highways</strong> dataset was developed to meet the growing demand for data that represents the reality of Brazilian highways, particularly in vehicle detection in satellite images. This dataset contains images from eight Brazilian highways (SP 160, SP 150, BR 101, BR 116, SP 021, SP 041, SP 070, SP 280) high-quality RGB images. Each image includes detailed annotations in two main classes—light vehicles and heavy vehicles—with a ground sampling distance (GSD) of 0.074 m/pixel and 0.3 m/pixel, providing essential data for training deep learning models in computer vision. This dataset is available and offers an unprecedented and valuable resource for developing technologies adapted to the characteristics of the Brazilian fleet and highways.</p> <p>The dataset development was conducted under the guidance of <strong>Prof. Dr. André Luiz Cunha</strong>, with the active participation of researchers <strong>Luan André Contel</strong> and <strong>Crhistian Emilio Ribeiro</strong>. The project stands out for its relevance in the field of vehicle detection by satellite images through computer vision and deep learning techniques and its potential applications in urban planning and traffic management.</p> <p>Funded by the <strong>University of São Paulo (USP)</strong>, through the PUB scholarship program, the work represents a significant advance in the use of technologies for vehicle detection for innovative solutions in urban mobility. The initiative not only contributes to the academic development of those involved but also to technological progress in the area of intelligent transport systems.</p>
Dataset of High-Resolution Micro-CT Imaging of Tumor Invasion and Metastasis in a Murine Esophageal Cancer PDX Model
<p>This dataset features high-resolution micro-CT imaging data capturing the progression of tumor invasion and metastasis in an orthotopic patient-derived xenograft (PDX) model of esophageal cancer. Using contrast-enhanced micro-CT, we visualized detailed patterns of tumor invasion, including budding, multicellular streaming, and expansive growth, across multiple abdominal organs such as the stomach, pancreas, liver, and spleen. The dataset includes two specimens, highlighting both the primary tumor site and extensive metastases throughout the abdominal cavity. Our imaging preserved the native tissue architecture, providing a unique three-dimensional view of tumor-host interactions. This collection offers valuable insights for researchers studying the dynamics of esophageal cancer invasion and metastasis. Detailed descriptions of the micro-CT scanning parameters, image analysis, and sample preparation are provided within the dataset archive.</p>
SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection
<h2>Abstract </h2> <blockquote> <p><em>Developing robust drone detection systems is often constrained by the limited availability of large-scale annotated training data and the high costs associated with real-world data collection. However, synthetic data presents a promising and cost-effective solution to overcome this issue. Therefore, we present SynDroneVision, a synthetic dataset specifically designed for RGB-based drone detection in surveillance applications. Featuring diverse backgrounds, lighting conditions, and drone models, SynDroneVision offers a comprehensive training foundation for deep learning algorithms. To evaluate the dataset's effectiveness, we perform a comparative analysis across a selection of recent YOLO detection models. Our findings demonstrated that SynDroneVision is a valuable resource for real-world data enrichment, achieving notable enhancements in model performance and robustness, while significantly reducing the time and costs of real-world data acquisition. </em> </p> </blockquote> <h2><strong>Paper</strong></h2> <h3><strong>Published in the Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV2025)!</strong></h3> <p>SynDroneVision is presented in the paper <strong>SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection</strong> by Tamara R. Lenhard, Andreas Weinmann, Kai Franke, and Tobias Koch. This work is published in the Proceedings of the <strong>2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV2025)</strong>.</p> <p>The preprint is currently available on ArXiv: <a href="https://arxiv.org/abs/2411.05633v1">here</a></p> <p>The final version is now published in the proceedings of WACV 2025: <a href="https://ieeexplore.ieee.org/document/10943801" target="_blank" rel="noopener">here</a> </p> <h2><strong>Dataset Details</strong></h2> <p>SynDroneVision comprises a total of<strong> 140,038</strong> <strong>annotaed RGB images</strong> (131,238 for training, 8,800 for validation, and 4,000 for testing), featuring a resolution of <strong>2560x1489</strong> pixels. All images are recorded in a sequential manner using <a href="https://www.unrealengine.com/en-US/">Unreal Engine 5.0</a> in combination with <a href="https://github.com/CodexLabsLLC/Colosseum">Colosseum</a>. Apart from drone images, SynDroneVision also includes ~7% of background images (i.e., imag frames without drone instances).</p> <p><strong>Annotation Format:</strong> Annotations (bounding boxes) are provided via text files according to the <strong>YOLO standard format</strong></p> <pre><code><object-class> <x> <y> <width> <height></code></pre> <p>Here, <code><x></code> and <code><y></code> represent the normalized coordinates of the bounding box center, while <code><width></code> and <code><height></code> denote the normalized bounding box wisth and height. In SynDroneVision, <code><object-class></code> is always 0, indicating the drone class.</p> <h2><strong>Download</strong></h2> <p>The SynDroneVision dataset offers around 900 GB of data dedicated to image-based drone detection. To facilitate the download process, we have partitioned the dataset into smaller sections. Specifically, we have divided the training data into 10 segments, organized by sequences.</p> <p>Annotations are available below, with image data accessible via the following links:</p> <table> <tbody> <tr> <td><strong>Dataset Split<br></strong></td> <td><strong>Sequences<br></strong></td> <td><strong>File Name<br></strong></td> <td><strong>Link</strong></td> <td><strong>Size (GB)<br></strong></td> </tr> <tr> <td>Training Set</td> <td>Seq. 001 - 009</td> <td>images_train_seq001-009.zip</td> <td><a href="https://datastore.dlr-pi.de/s/ted96QK36RmXgLs" target="_blank" rel="noopener">Training images PART 1</a></td> <td>57</td> </tr> <tr> <td> </td> <td>Seq. 010 - 018</td> <td>images_train_seq010-018.zip</td> <td><a href="https://datastore.dlr-pi.de/s/28GtS3qeNGksWk7" target="_blank" rel="noopener">Trainng images PART 2</a></td> <td>95.4</td> </tr> <tr> <td> </td> <td>Seq. 019 - 027</td> <td>images_train_seq019-027.zip</td> <td><a href="https://datastore.dlr-pi.de/s/DLCwRBRnsf53Rg7" target="_blank" rel="noopener">Training images PART 3</a></td> <td>96.2</td> </tr> <tr> <td> </td> <td>Seq. 028 - 035</td> <td>images_train_seq028-035.zip</td> <td><a href="https://datastore.dlr-pi.de/s/oF9PmFCexHbB2bp" target="_blank" rel="noopener">Training images PART 4</a></td> <td>83.9</td> </tr> <tr> <td> </td> <td>Seq. 036 - 043</td> <td>images_train_seq036-043.zip</td> <td><a href="https://datastore.dlr-pi.de/s/eepMQrixWXNXNYS" target="_blank" rel="noopener">Training images PART 5</a></td> <td>77.1</td> </tr> <tr> <td> </td> <td>Seq. 044 - 050</td> <td>images_train_seq044-050.zip</td> <td><a href="https://datastore.dlr-pi.de/s/FPnyEAcmpjomY9q" target="_blank" rel="noopener">Training images PART 6</a></td> <td>84.7</td> </tr> <tr> <td> </td> <td>Seq. 051 - 056</td> <td>images_train_seq051-056.zip</td> <td><a href="https://datastore.dlr-pi.de/s/awsHAJqKe85AqiG" target="_blank" rel="noopener">Training images PART 7</a></td> <td>86.8</td> </tr> <tr> <td> </td> <td>Seq. 057 - 065</td> <td>images_train_seq057-065.zip</td> <td><a href="https://datastore.dlr-pi.de/s/2LH8TmAjG94r26P" target="_blank" rel="noopener">Training images PART 8</a></td> <td>86.2</td> </tr> <tr> <td> </td> <td>Seq. 066 - 070</td> <td>images_train_seq066-070.zip</td> <td><a href="https://datastore.dlr-pi.de/s/BGkLANPT3mAwEw8" target="_blank" rel="noopener">Training images PART 9</a></td> <td>75.7</td> </tr> <tr> <td> </td> <td>Seq. 071 - 073</td> <td>images_train_seq071-073.zip</td> <td><a href="https://datastore.dlr-pi.de/s/WCEQPCioNxdqJr9" target="_blank" rel="noopener">Training images PART 10</a></td> <td>38.5</td> </tr> <tr> <td>Validation Set</td> <td>Seq. 001 - 073</td> <td>images_val.zip</td> <td><a href="https://datastore.dlr-pi.de/s/9nYjwxGXJe5ws7g" target="_blank" rel="noopener">Validation images</a></td> <td>55.2</td> </tr> <tr> <td>Test Set</td> <td>Seq. 001 - 073</td> <td>images_test.zip</td> <td><a href="https://datastore.dlr-pi.de/s/5YgqR75oBByEz6R" target="_blank" rel="noopener">Test images</a></td> <td>26.5</td> </tr> </tbody> </table> <h2><strong>Citation</strong></h2> <p>If you find SynDroneVision helpful in your research, we kindly ask that you cite the associated paper. Below is the citation in BibTeX format for your convenience:</p> <p><strong>BibTeX:</strong></p> <pre>@INPROCEEDINGS{10943801, author={Lenhard, Tamara R. and Weinmann, Andreas and Franke, Kai and Koch, Tobias}, booktitle={2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)}, title={SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection}, year={2025}, volume={}, number={}, pages={7637-7647}, doi={10.1109/WACV61041.2025.00742}} <br><br></pre> <p><em>SynDroneVision uses Unreal® Engine. Unreal® is a trademark or registered trademark of Epic Games, Inc. in the United States of America and elsewhere.</em></p>
Imaging thermocline microstructure in 2D with swaths traced by wave-pumped χpods: dataset and code
<p><a href="https://doi.org/10.1029/2024JC022134">Associated paper</a> published in <em>J. Geophys. Res. Oceans</em></p> <h2>Dataset summary</h2> <p>Location: 0°N, 140°W<br>Depth: 120 m<br>Period: 16-Sep-2014 to 19-Oct-2015</p> <p>This data archive contains two types of data files:<br>data_yymmdd.mat<br>grid_yymmdd.mat<br>Each file contains 24 hours of data. There are 394 of each type.</p> <p>Arrays in data_yymmdd.mat are single precision (except the 'time' array) to keep file sizes small.</p> <h2>Contents of the data files</h2> <p>Each Matlab file contains a single struct. These structs include readmes, which are reproduced in the full PDF readme (chipod_swaths_readme.pdf).</p> <p>Files are grouped into months and zipped (yymm.zip) to ease downloading.</p> <h2>Reading the data file with Python</h2> <p>Example code to read the files into Python as dictionaries is given in the full PDF readme (chipod_swaths_readme.pdf).</p> <h2>Matlab code to produce the processed data</h2> <p>The code to read in raw chipod data and process them is primarily contained in the file 'swaths_paper_data_preparation.m'. This file calls three other files ('raw_load_chipod.m', 'deglitch.m', and 'bin.m'). All of these files are provided for completeness, but we are only archiving the processed outputs (not the raw voltage signals). Please email if more information is needed.</p> <h2>Matlab code for the convolutional neural network</h2> <p>See 'chipod_swaths_convolutional_neural_network.m'.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.