Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

979

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

979 results for “Image Dataset”

Learn how ShareScore rates datasets ↗
zenodo52/100

ColoPola: A dataset of colorectal cancer polarimetric images (Mueller matrix elements) for colorectal cancer detection

<p><strong>ColoPola</strong> dataset is <strong>Colo</strong>rectal cancer <strong>Pola</strong>rimetric images dataset</p> <p>The dataset consists of 572 slices (specimens) with 20,592 images, 284 slices of which were designated as cancer samples and 288 as normal samples.</p> <p>Each sample has 36 polarimetric images (i.e., HH, HV, HP, HM, HR, HL, VH, VV, VP, VM, VR, VL, PH, PV, PP, PM, PR, PL, MH, MV, MP, MM, MR, ML, RH, RV, RP, RM, RR, RL, LH, LV, LP, LM, LR, and LL).</p> <p>Each folder in the <strong>ColoPola</strong> dataset consists of 36 polarimetric images. Each image is 1280x1024 pixels in size and was created in the TIF file format (HH.tif, HV.tif, ..., LL.tif).&nbsp;</p>

opencc-zeroNov 2023View details →
zenodo52/100

Product Images for Life Cycle Assessment Dataset For Peritoneal Dialysis and Haemodialysis in Modena

<p>The database contains a collection of images showcasing the individual components of peritoneal dialysis (PD) products, along with their corresponding weights. These images serve as a visual record for life cycle assessment (LCA) purposes, focusing on the material composition and environmental impact of each product.</p> <ol> <li> <p><strong>Patient Education Materials</strong>: Photographs of educational materials provided to patients, with accompanying data on the weight of the paper and packaging.</p> </li> <li> <p><strong>Catheters and Surgical Kits</strong>: Images display the disassembled components of PD catheters and surgical kits, including tubing, connectors, and packaging. Each image is annotated with the precise weight of the individual components.</p> </li> <li> <p><strong>Dialysis Solution Bags</strong>: The database includes images of both CAPD and APD solution bags, separated into their constituent parts (e.g., plastic bag, solution, and protective wrapping), with weights noted for each component.</p> </li> <li> <p><strong>Connection Devices and Consumables</strong>: Detailed images of connection devices, clamps, and other consumable items, with individual component weights clearly labeled.</p> </li> <li> <p><strong>Packaging and Transport Materials</strong>: Photographs of transport packaging, such as cardboard boxes and plastic wraps, alongside recorded weights for each element.</p> </li> <li> <p><strong>Maintenance Items</strong>: Visuals of terminal catheter sets, cleaning agents, and related products, each accompanied by their respective weight data.</p> </li> <li> <p><strong>Disposal Components</strong>: Images of used solution bags, syringes, and other single-use items, separated into recyclable and non-recyclable components, with weights specified for each.</p> </li> </ol> <p>This image-based database provides a clear and comprehensive reference for the material breakdown and weight distribution of PD product components, essential for conducting a thorough LCA and identifying areas for environmental improvement.</p>

opencc-by-4.0Dec 2024View details →
zenodo52/100

Dataset of "Neutron imaging and molecular simulation of systems from methane and p‑xylene"

<p>The dataset contains parameterizations, and input files for molecular dynamics simulations used in the study of methane dissolution in p-xylene. For selected conditions, full simulation data, i.e., trajectories and energetics are provided. All used simulation results data are provided in the table, along with the measured experimental data.</p>

opencc-by-4.0Dec 2024View details →
zenodo52/100

Dataset for the manuscript "Event-triggered STED imaging"

<p>Dataset that supports the implementation of the event-triggered STED method and support the findings in the manuscript: &quot;Event-triggered STED imaging&quot; (Jonatan Alvelid, Martina Damenti, Chiara Sgattoni, Ilaria Testa, preprint: https://doi.org/10.1101/2021.10.26.465907). The data files are organized according to the various experiments performed for characterization or application of the method. References to the specific figures in the manuscript that uses the different data is provided in the info file.</p> <p>The scripts provided at https://github.com/jonatanalvelid/etSTEDanalysis (https://doi.org/10.5281/zenodo.6469723) have been used for image and data handling, and image analysis of the here provided dataset.</p>

opencc-by-4.0Oct 2021View details →
zenodo52/100

RCS-Image Dataset

<p>Note: There is a single zip file of the project available under the name &quot;<a href="https://zenodo.org/api/files/30cdde03-6e96-4019-9a1a-bef54f3d2737/RCS%20Image%20Dataset.zip">RCS Image Dataset.zip&nbsp;</a>&quot;</p> <p>A synthetic dataset for object detection, semantic segmentation and depth recognition was created using images from a virtual environment in Unreal Engine. This dataset that could be used to retrain neural networks on virtual aerial images, for object detection, segmentation and depth planning.</p> <p>There is ground truth for full pixel level semantic segmentation and object detction by way of bounding box coordinates in .exif files. The main&nbsp; labels present in the dataset are train truck and gas cylinders&nbsp; The size of the data is currently more than 10GB, with 1000 aerial images, from 100 different waypoints at 10 different heights. For each raw image, there is also an associated depth image, a segmented image, object ground truths with bounding box, and the metadata .exif file.</p> <p>If you found this dataset useful for your reserach, please cite as,</p> <pre>@article{smyth2018virtual, title={A Virtual Environment with Multi-Robot Navigation, Analytics, and Decision Support for Critical Incident Investigation}, author={Smyth, David L and Fennell, James and Abinesh, Sai and Karimi, Nazli B and Glavin, Frank G and Ullah, Ihsan and Drury, Brett and Madden, Michael G}, journal={arXiv preprint arXiv:1806.04497}, year={2018} }</pre> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Evaluation Framework for Multiband Image Enhancement and Blending Algorithms in Enhanced Flight Vision Systems - Image Dataset

<p>This dataset contains data used in the research published by MLabs Optronics in the paper:</p> <p>Medina Heierle, Victor, Mar&iacute;a Tejada Casado, Alberto Briasco Gonz&aacute;lez, Hugo Jestes Zoilo, Jes&uacute;s Mart&iacute;n Tapia, Adeodato Altamirano Aguilar, and Javier Mu&ntilde;oz De Luna Clemente. Evaluation Framework for Multiband Image Enhancement and Blending Algorithms in Enhanced Flight Vision Systems. Proceedings of the 14th International Conference on Signal-Image Technology &amp; Internet-Based Systems (SITIS), pp. 274-280. IEEE, 2018.</p> <p><br> The dataset is classified into 3 folders:</p> <p>- IR_VIS: Contains 28 pairs of images in the IR (some images may be in the NIR spectrum instead) and Visual spectrum, taken from different public repositories off the internet, which are typically used in multispectral fusion research.<br> &nbsp;<br> - Fusion: Contains 8 sets with the results of applying each of the 4 fusion algorithms described in the paper on some of the images in folder &quot;IR_VIS&quot;.</p> <p>- VIS haze filtering: Contains 24 images taken with a CCD camera of a contrast target inside a fog simulation cabin in a laboratory. For comparison purposes, all images have been taken with a similar amount of fog, which is as much as was possible while still being able to see the target with the camera through the fog. Each image has been taken with a different type of filter (filter information is provided in another image inside the folder).</p> <p>&nbsp;</p> <p>Mlabs Optronics<br> PTA<br> Calle Pierre Laffitte, 8<br> 29590 M&aacute;laga (Spain)</p> <p>www.mlabsoptronics.com<br> info@mlabsoptronics.com</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Dataset of cracks on DIC images

<p>This dataset contains crack images and corresponding annotated ground truth masks. This data was used to train, validate, and test a deep convolutional neural network to detect crack pixels on images taken as input for the digital image correlation (DIC) method.&nbsp;</p> <p>For more information about the trained network, please refer to our publication at this <a href="https://www.sciencedirect.com/science/article/pii/S095006182032479X?via%3Dihub">link</a>.&nbsp;&nbsp;</p> <p>The source codes to reproduce the results are shared at this <a href="https://github.com/amirrezaie1415/Deep-DIC-Crack">link</a>.</p> <p>Please cite the following articles:</p> <blockquote> <p>[1] Rezaie, A., Achanta, R., Godio, M., &amp; Beyer, K. (2020). Comparison of crack segmentation using digital image correlation measurements and deep learning. Construction and Building Materials, 261, 120474.&nbsp;doi:https://doi.org/10.1016/j.conbuildmat.2020.120474</p> <p>[2]&nbsp;Rezaie, A., Godio, M., &amp; Beyer, K. (2021). Investigating the cracking of plastered stone masonry walls under shear&ndash;compression loading.&nbsp;Construction and Building Materials,&nbsp;306, 124831.<br> doi:https://doi.org/10.1016/j.conbuildmat.2021.124831</p> </blockquote>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Diffraction images of crystals of the first and second spectrin repeats (mutant C420A/C435A) of human plectin (PDB code 2ODV): 2-wavelength SeMet MAD dataset

<p>Diffraction images of SeMet labeled crystals of a fragment of human plectin that includes the first and second spectrin repeats (SR1-SR2) of the plakin domain. The two Cys in the wild type sequence were replaced by Ala.</p> <p>This Se-Met MAD dataset was used for the <em>de novo</em> phasing of the pdb entry 2ODV (http://www.rcsb.org/pdb/explore/explore.do?structureId=2ODV).</p> <p>&nbsp;</p> <p>Data was collected at the BM14 beamline of the European Synchrotron Radiation Facility (ESRF, Grenoble, France) using a Mar CCD detector. Data from the same crystal were collected at two wavelengths :</p> <ul> <li>Remote wavelength (0.9185 &Aring;): 180 images (1 degree oscillation per image).</li> <li>Peak wavelength ( 0.9785 &Aring;): 360 images (1 degree oscillation per image).</li> </ul>

opencc-by-sa-4.0Mar 2016View details →
zenodo48/100

Hyperspectral Imaging Dataset for Laser Thermal Ablation Monitoring in Vital Organs

<p><strong>Objectives:</strong> The objective of the research was to use hyperspectral imaging (HSI) to detect thermal damage induced in vital organs (such as the liver, pancreas, and stomach) during laser thermal therapy. The experimental study was conducted during thermal ablation procedures on live pigs.</p> <p><strong>Ethical Approval:</strong> The experiments were performed at the Institute for Image Guided Surgery in Strasbourg, France. This experimental study was approved by the local Ethical Committee on Animal Experimentation (ICOMETH No. 38.2015.01.069) and by the French Ministry of Higher Education and Research (protocol №APAFiS-19543-2019030112087889, approved on March 14, 2019). All animals were treated in accordance with the ARRIVE guidelines, the French legislation on the use and care of animals, and the guidelines of the Council of the European Union (2010/63/EU).</p> <p><strong>Description:</strong> During our experimental study, we used a TIVITA hyperspectral camera to acquire hypercubes of size 640x480x100 voxels, indicating 640x480 pixels for 100 bands, and regular RGB images at each acquisition step. These bands were acquired directly from the hyperspectral camera without additional pre-processing. The hypercube was acquired in approximately 6 seconds and synchronized with the absence of breathing motion using a protocol implemented for animal anesthesia. Polyurethane markers were placed around the target area to serve as references for superimposing the hyperspectral images, which were acquired using target areas selected according to the hyperspectral camera manufacturer's guidelines.</p> <p>As part of our investigation, we included hyperspectral cubes from 20 experiments conducted under identical conditions in our study. The hyperspectral cubes were collected in three distinct stages. In the first stage, the cubes were gathered before laparotomy at a temperature of 37&deg;C. In the second stage, we obtained the cubes as the temperature gradually increased from 60&deg;C to 110&deg;C at 10&deg;C intervals. Finally, in the last stage, the cubes were collected after turning off the laser during the post-ablation phase. Thus, we obtained a total of 233 hyperspectral cubes, each consisting of 100 wavelengths, resulting in a dataset of 23,300 two-dimensional images. The temperature changes were recorded, and the &ldquo;<em>Temperature profile during laser ablation</em>&rdquo; image illustrates the corresponding profile, highlighting the specific time intervals during which the hyperspectral camera and laser were activated and deactivated. To provide a visual representation of the collected data, we have included several examples of images captured from different organs in the &ldquo;<em>Examples of ablation areas</em>&rdquo; figure.</p> <p>The raw dataset, comprising 233 hyperspectral cubes of 100 wavelengths each, was transformed into 699 single-channel images using PCA and t-SNE decompositions. These images were then divided into training and test subsets and prepared in the COCO object detection format. This COCO dataset can be used for training and testing different neural networks.</p> <p><strong>Access to the Study:</strong> Further information about this study, including curated source code, dataset details, and trained models, can be accessed through the following repositories:</p> <ul> <li><strong>Source code:</strong> <a href="https://github.com/ViacheslavDanilov/hsi_analysis" target="_blank" rel="noopener">https://github.com/ViacheslavDanilov/hsi_analysis</a></li> <li><strong>Dataset:</strong> <a href="https://doi.org/10.5281/zenodo.10444212" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10444212</a></li> <li><strong>Models:</strong> <a href="https://doi.org/10.5281/zenodo.10444269" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10444269</a></li> </ul>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Dataset of 400 pomegranate tree (Punica granatum L. 'Wonderful') images

<p>Dataset of 400 pomegranate tree (Punica granatum L. &lsquo;Wonderful&rsquo;) images, with the corresponding fruit masks.</p> <p>The dataset is designed for training artificial intelligence models for instance segmentation.</p> <p>The pictures were collected by means of mobile devices (smartphones), in random trees, from different distances, orientations and in varying lighting conditions. The resolution of the images and masks is 640x480 pixels. The dataset is divided into training (70%), validation (15%) and test set (15%). Stratification was performed in 3 periods of the season to ensure that all fruit ripening stages were present in each subset.&nbsp;Masks consist of a very detailed manual annotation of the visible part for each of the fruits in the images.</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

R-CAUSTIC: Rippling CAUSTICs underwater Image dataset

<p><strong>Description</strong></p><p>Rippling caustics seem to be the main factor degrading the underwater RGB image quality and affecting the image- based 3D reconstruction process in very shallow waters. These effects are adversely affecting image matching algorithms by throwing off most of them, leading to less&nbsp;accurate matches&nbsp;and causing issues in the Simultaneous Localization and Mapping (SLAM) based navigation of the Remotely Operated Vehicles (ROV) and Autonomous Underwater Vehicles (AUV) on shallow waters. Also, they are the main cause for dissimilarities in the generated textures and orthoimages. In order to fill the&nbsp;gap in the literature regading underwater rippling caustics imagery with real ground truth and reference images, the first real-world underwater caustics benchmark dataset which contains 1465 underwater images is presented. Together with the RGB imagery, the corresponding generated ground truth images are delivered for facilitating the training and testing of machine learning and deep learning methods for image classification. R-CAUSTIC&nbsp;dataset also provides the necessary data to evaluate, at least to some extent, the performance of 3D reconstruction approaches. Data were acquired using a GoPro Hero 4 Black action camera with image dimensions of 4000 x 3000 pixels, focal length of 2.77mm and pixel size of 1.55μm and a tripod. Action cameras are widely used for underwater image acquisition. The dataset was captured in near-shore underwater sites at depths varying from 0.5 to 2m. No artificial light sources were used. Due to the wind, the turbulent surface of the water created dynamic rippling caustics on the seabed. In total 1465 RGB images were collected, separated in 7 different datasets; five of them containing stereo images, one of them tri-stereo images and one consists of multi-stereo imagery acquired in 7 different camera poses.</p><p>&nbsp;</p><p><strong>Publication</strong></p><p>The paper is availbale in Open Access here: https://ieeexplore.ieee.org/document/10172291</p><p><strong>If you use this dataset please cite it as R-CAUSTIC</strong> [Reference].<br>[Reference]: <strong>P. Agrafiotis, K. Karantzalos and A. Georgopoulos, "Seafloor-Invariant Caustics Removal From Underwater Imagery," in </strong><i><strong>IEEE Journal of Oceanic Engineering</strong></i><strong>, vol. 48, no. 4, pp. 1300-1321, Oct. 2023, doi: 10.1109/JOE.2023.3277168.</strong></p><p>BibTeX:</p><p>@ARTICLE{10172291, &nbsp;author={Agrafiotis, Panagiotis and Karantzalos, Konstantinos and Georgopoulos, Andreas}, &nbsp;journal={IEEE Journal of Oceanic Engineering}, &nbsp;title={Seafloor-Invariant Caustics Removal From Underwater Imagery}, &nbsp;year={2023}, &nbsp;volume={48}, &nbsp;number={4}, &nbsp;pages={1300-1321}, &nbsp;doi={10.1109/JOE.2023.3277168}}</p><p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

Dataset of pomegranate tree (Punica granatum L. 'Wonderful') image times series

<p>Dataset of pomegranate tree (Punica granatum L. &lsquo;Wonderful&rsquo;) image times series.&nbsp;The pictures were collected by means of Raspberry Pi cameras with OV5647 sensor (5 MP, f2.9). Sensors were installed on fixed platforms for continuous measurement with zenithal orientation at a distance of approximately 1 metre from the canopy. Images were captured daily at 9 a.m. (GMT+2) from July to mid-October in 2021 and 2022.&nbsp;The resolution of the images is 640x480 pixels.</p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Videos of the processed microscope images and time series of the petrophysical parameters from image processing and geochemical simulation and of the measured induced polarisation [Video][Dataset]

<p>Supporting Information for the manuscript&nbsp;<em>Microfluidics and&nbsp;spectral induced polarization for direct observation and petrophysical modeling of calcite dissolution</em> published in Geophysical Research Letters</p> <ul> <li><strong>Data Set S1.</strong> Porosity, water saturation, and calcite sample perimeter from image<br>processing.</li> <li><strong>Data Set S2.</strong> Porosity, water conductivity, and pH from geochemical simulation.</li> <li><strong>Data Set S3.</strong> Real and imaginary components of the complex electrical conductivity at<br>2.5 Hz and CEC from petrophysical modeling.</li> <li><strong>Movie S1.</strong> Dissolution of the calcite sample with the detected contour superimposed in<br>white on the grayscale images. Time, length scale, and flow direction are indicated. In<br>case of problems launching the file, we recommend using VLC Media Player software.</li> <li><strong>Movie S2.</strong> Segmented images of the CO2 bubbles produced by the calcite dissolution.<br>Time, length scale, and flow direction are indicated. In case of problems launching the<br>file, we recommend using VLC Media Player software.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Training and test dataset of STED images of microtubules in fixed cells

<p>Training and test dataset of microtubule used in the manuscript "Denoising diffusion models for high-resolution microscopy image restoration".</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Live-cell STED dataset of mitochondria containing ground truth and corresponding low intensity noisy images

<p>The dataset was acquired as part of the manuscript "Denoising diffusion models for high-resolution microscopy image restoration". The dataset contains ground truth and low intensity STED images of mitochondria acquired in live U2-OS cells stably expressing TOM20 coupled to the dead mutant of HaloTag7 which was made fluorescent by using the exchangeable ligand Hy4 bound to the fluorophore SiR.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

A dataset of colorectal cancer histopathological images

<p>The dataset contains the histopathological images of the ColoPola dataset (https://doi.org/10.5281/zenodo.10068018).</p> <p>CLCXYYZZNN_Hx</p> <p>CLC: colorectal (cancer) tissue</p> <p>NLC: normal tissue</p> <p>X - Times<br>YY - Sample number<br>ZZ - Serial number<br>NN - Image number<br>H - Magnification</p>

opencc-zeroNov 2024View details →
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Pl@ntNet-300K image dataset

<p>This paper presents a novel image dataset with high intrinsic ambiguity and a long-tailed distribution built from the database of Pl@ntNet citizen observatory. It consists of 306146&nbsp;plant images covering 1081&nbsp;species. We highlight two particular features of the dataset, inherent to the way the images are acquired and to the intrinsic diversity of plants morphology:</p> <p>&nbsp; &nbsp; (i) the dataset has a strong class imbalance, i.e. a few species account for most of the images, and,</p> <p>&nbsp; &nbsp; (ii) many species are visually similar, rendering identification difficult even for the expert eye.</p> <p>&nbsp; &nbsp; These two characteristics make the present dataset well suited for the evaluation of set-valued classification methods and algorithms. Therefore, we recommend two set-valued evaluation metrics associated with the dataset (macro-average top-k accuracy&nbsp;and macro-average average-k accuracy) and we provide baseline results established by training deep neural networks using the cross-entropy loss.</p> <p>A full description of the dataset as well as baseline experiments can be found in the following&nbsp;publication:</p> <p>&quot;<a href="https://openreview.net/forum?id=eLYinD0TtIt">Pl@ntNet-300K: a plant image dataset with high label ambiguity and a long-tailed distribution</a>&quot;, Camille Garcin, Alexis Joly, Pierre Bonnet, Antoine Affouard, Jean-Christophe Lombardo, Mathias Chouet, Maximilien Servajean,&nbsp;Titouan Lorieul and Joseph Salmon, in Proc. of Thirty-fifth Conference on Neural Information Processing Systems, Datasets and Benchmarks Track, 2021.</p> <p>&nbsp;Please cite the above&nbsp;reference for any publication using the dataset.</p> <p>Utilities to load the data and train models with pytorch can be found here: <a href="https://github.com/plantnet/PlantNet-300K/">https://github.com/plantnet/PlantNet-300K/</a></p>

opencc-by-4.0Apr 2021View details →
zenodo48/100

DeepOrchidSeries: A Sentinel-2 Dataset to inform convolutional SDMs with twelve-month Sentinel-2 image time-series, Orchid family

<p><strong>Deep Species Distribution Modelling from Sentinel-2 Image Time-series: a Global Scale Analysis on the Orchid Family</strong>&nbsp;</p> <ul> <li><strong><em>DeepOrchidSeries</em></strong> dataset gathers Sentinel-2 image time-series around geolocated orchid occurrences. Seasonal evolutions of the habitats are captured in the twelve-month RGB/IR time-series with 640x640m spatial resolution. It allows novel Species Distribution Models (SDMs) coupled with convolutional networks to take advantage of both spatial and temporal information.</li> <li>Our <strong>associated article</strong> is describing the modeling choices made to shape this ambitious dataset. It is submitted to <a href="https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence">https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence</a>. We believe such global data, methods and scripts are valuable to the conservation ecology community and especially deep-SDMs users. To our knowledge, no similar ready-to-use dataset is available. In the article, the dataset&#39;s temporal dimension is proven to significantly improve SDMs performances.</li> <li><strong><em>sen2patch</em></strong> is the gitlab project gathering the code to create such dataset. It is available at <a href="https://gitlab.inria.fr/jestopin/sen2patch">https://gitlab.inria.fr/jestopin/sen2patch</a>.</li> <li><strong><em>DeepOrchidSeries.csv</em></strong> contains all occurrences-level information. <ul> <li>We advice to load it with: <pre><code class="language-python">import pandas as pd df = pd.read_csv("path/to/DeepOrchidSeries.csv", sep=';') df.columns ['gbifid', 'canonical_name', 'decimallatitude', 'decimallongitude', 'speciesKey', 'cell_index', 'bot_country', 'bot_code', 'lvl2_code', 'continent_code']</code></pre> <ul> <li>&#39;gbifid&#39; is the occurrences GBIF ID</li> <li>&#39;canonical_name&#39;, is the species canonical name</li> <li>&#39;decimallatitude&#39;, &#39;decimallongitude&#39; are the species coordinates in decimal degrees</li> <li>&#39;speciesKey&#39; is the species GBIF unique identifier</li> <li>&#39;cell_index&#39; is&nbsp;a unique cell ID in a 0.0025&deg; lon/lat grid partitioning the Earth (used to stratify train/val/test set by geographic blocks)</li> <li>&#39;bot_country&#39;, &#39;bot_code&#39;, &#39;lvl2_code&#39;, &#39;continent_code&#39; are geographic subdivisions defined in <a href="https://github.com/tdwg/wgsrpd">https://github.com/tdwg/wgsrpd</a> (code and string for WGSRPD level 1, the botanical countries)</li> </ul> </li> </ul> </li> <li> <p>Initial <a href="https://www.gbif.org/">GBIF</a> query DOI is <a href="http://https://doi.org/10.15468/dl.4bijtu">https://doi.org/10.15468/dl.4bijtu</a> (26 August 2019).</p> </li> <li><strong><em>DeepOrchidSeries.tar</em></strong> file contains the satellite image time-series and is available at <a href="https://lab.plantnet.org/deeporchidseries/">https://lab.plantnet.org/deeporchidseries/</a> <ul> <li><em>.tar</em> archive measure 286 GB and extends to 432 GB once decompressed.</li> <li>Image time-series relative tree paths are constructed from the occurrences unique GBIF IDs.</li> <li>For a given occurence <em>gbifid</em>, matching patches are located in: <em>final_dataset_by_gbifid/gbifid[-2:]/gbifid[-4:-2]</em>, <em>i.e.</em> in a first folder named with the <em>gbifid</em> last two numbers and a subfolder with the previous two ones. Example: the time-series files matching occurrence 2236837714 are located at <em>final_dataset_by_gbifid/14/77/</em>.&nbsp;</li> <li>Image time-series are composed of twelve 16 bits RGB <em>.png</em>&nbsp; and twelve 16 bits IR <em>.png</em> files containing data identical to the original L1C products, no lossy compression was made. There are one RGB and one IR .png file per month.</li> <li>Patches from month MM/YYYY of occurrence <em>gbifid</em> are named<em> </em><em>RGB_YYYY_MM_gbifid_.png</em> and <em>IR0_YYYY_MM_gbifid_.png</em>.</li> </ul> </li> <li><em><strong>models.zip</strong></em> is the archive containing the four PyTorch models weights described in our article and<strong><em> </em></strong><em><strong>inception_env.py</strong></em> the used Inception V3 architecture. <em><strong>index.json</strong></em> contains the dictionnary linking the models class indexes from 0 to 14128 with our labels <em>speciesKey</em>: {&quot;class_index&quot;:speciesKey}.</li> </ul> <p>&nbsp;</p> <ul> <li><strong>ACKNOWLEDGMENTS</strong>: We warmly thank Alexander Zizka et al. for providing us the geographically and taxonomically curated set of Orchids occurrences. This dataset contains modified Copernicus Sentinel data and Copernicus Service information (2018). Sentinel-2 MSI data used were available at no cost from ESA Sentinels Scientific Data Hub.</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Twitter dataset of flood-related images for September 2021, Thailand and June/July 2021, Nepal floods

<p>Twitter dataset related to flood events onsets in Thailand and Nepal, focused on&nbsp;September 26/27, 2022, June 16/17 2021 and July 01/02 2021. The dataset has been&nbsp;processed with a VisualCit pipeline in order to automatically filter a relevant subset of posts through automated image analysis, using deep learning techniques. The posts were then geolocated using the CIME algorithm. Additional information about the data collection and data processing are described in <a href="http://arxiv.org/abs/2202.12014">http://arxiv.org/abs/2202.12014</a></p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record