Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “Image dataset”
ROCOv2: Radiology Objects in COntext Version 2, An Updated Multimodal Image Dataset
<p>Recent advances in deep learning techniques have enabled the development of systems for automatic analysis of medical images. These systems often require large amounts of training data with high quality labels, which is difficult and time consuming to generate.</p> <p>Here, we introduce Radiology Object in COntext Version 2 (ROCOv2), a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PubMed Open Access subset. Concepts for clinical modality, anatomy (X-ray), and directionality (X-ray) were manually curated and additionally evaluated by a radiologist. Unlike MIMIC-CXR, ROCOv2 includes seven different clinical modalities.</p> <p>It is an updated version of the ROCO dataset published in 2018, and includes 35,705 new images added to PubMed since 2018, as well as manually curated medical concepts for modality, body region (X-ray) and directionality (X-ray). The dataset consists of 79,789 images and has been used, with minor modifications, in the concept detection and caption prediction tasks of ImageCLEFmedical 2023. The participants had access to the training and validation sets after signing a user agreement.</p> <p>The dataset is suitable for training image annotation models based on image-caption pairs, or for multi-label image classification using the UMLS concepts provided with each image, e.g., to build systems to support structured medical reporting.</p> <p>Additional possible use cases for the ROCOv2 dataset include the pre-training of models for the medical domain, and the evaluation evaluation of deep learning models for multi-task learning.</p>
Using the traditional microscope for mineral grain orientation determination: A prototype image analysis pipeline for optic-axis mapping (POAM). Original dataset.
<p>The data repository contains data obtained with the microscope Nikon Eclipse LV100ND that was stitched with <a href="https://imagej.net/plugins/trakem2/">TrakEM2 software</a>. The files allow reproducing the results obtained and plot in <a href="https://doi.org/10.1111/jmi.13284">Acevedo et al. (2024)</a> <strong>"Using the traditional microscope for mineral grain orientation determination: A prototype image analysis pipeline for optic-axis mapping (POAM)."</strong> by Acevedo Zamora, M. A., Schrank, C. E., & Kamber, B. S.</p> <p>The prototype uses MatLab scripts (<a href="https://github.com/marcoaaz/AcevedoEtAl._2024a_POAM">AcevedoEtAl._2024a_POAM</a>) that were documented in the paper Supplementary Material 1. The metadata can be found in Supplementary Material 3 and follows the structure of this data repository. The user needs downloading and changing the paths to run the same scripts and reproduce the results.</p> <p>Note: After download, unzip and merge (copy-paste) the folders (parts 1, 2 and 3). Before merging, the containing folder should be re-named to 'paper 2_datasets' to match exactly the MatLab scripts and reproduce our work.</p> <p>The remaining questions should be addressed to Marco Acevedo (maaz.geologia@gmail.com ; marco.acevedozamora@qut.edu.au)</p> <p>Thanks.</p>
Improving the accuracy of automated labeling of specimen images datasets via a confidence-based process - Datasets
<p>This dataset contains supporting data for a research project aimed at analysing herbarium samples from the New England area at a large scale with deep learning techniques. Details on the methodology are shared in the acompanying paper (to be published).</p> <p>Content:</p> <ul> <li>dataset600k_withAI.csv : A dataset of over 600.000 herbarium samples with its record metadata and a corresponding AI phenological annotations with matching confidence scores. The entirety of the record headers are provided, extracted directly from the NEVP portal. In addition, the AI labels are defined by the following headers. These 8 columns represent 4 binary classifiers with the Presence/Absence of each 4 traits and corresponding confidence (as a percentage - presence/absence percentages sum to 1).<br> <ul> <li> <table> <tbody> <tr> <td>Flowering</td> <td>Not Flowering</td> <td>Budding</td> <td>Not Budding</td> <td>Fruiting</td> <td>Not Fruiting</td> <td>Reproductive</td> <td>Not Reproductive</td> </tr> </tbody> </table> </li> </ul> </li> </ul> <ul> <li>data_species_with_statuses.csv: A processed dataset summarizing flowering period shift at a species level. Two types of headers are provided. <ul> <li>First metadata concerning the flowering shift and the data used to compute that value: <ul> <li> <table> <tbody> <tr> <td>genus</td> <td>genus_species</td> <td>slope</td> <td>nb_specimens</td> <td>p_value_significance</td> <td>trend_category</td> </tr> <tr> <td>Genus of the species</td> <td>Binomial name of the species</td> <td>Regression slope defining the flowering shift as a slope</td> <td>Number of herbarium specimens used to compute the shift</td> <td>P-value significance of the slope being non-zero. ('Non Significant'/'Significant')</td> <td>Summary of the shift as a binary characteristic ('Earlier'/'Later')</td> </tr> </tbody> </table> </li> </ul> </li> <li>Second, metadata summarizing various traits associated to each species: <ul> <li> <table> <tbody> <tr> <td>lifeform_status</td> <td>native_introduced_status</td> <td>wetland_status</td> <td>seasonality_average</td> <td>seasonality_spread</td> </tr> <tr> <td>Growth form from the USDA PLANTS Database. 'Forb_Herb', 'Shrub_Tree' or 'Vine'</td> <td>'Native'/'Introduced' status from the USDA PLANTS Database.</td> <td> <p>National Wetland Plant List (NWPL) Wetland Indicator Status within the Northcentral and Northeast Region</p> <p>'OBL'/'FACW'/'FAC'/'FACU'/'UPL'</p> </td> <td>A characteristic of the flowering season of the species based on the mean Day of Year of the analysed specimens: if <=180: 'Early', else 'Late'</td> <td>A characteristic of the flowering season of the species based on the spread of the flowering season. Less than 28 days: 'Narrow', larger: 'Large'.</td> </tr> </tbody> </table> <p> </p> </li> </ul> </li> </ul> </li> <li>phylogenetic_tree.tre: The raw data used to generate the visualization of the flowering seasonality character and the detected flowering shift foreach species on a phylogenetic tree.</li> <li>phylogenetic_processed_dataset.csv: The processed dataset resuting from the phylogenetic signal analysis. For each trait, an associated significance binary value is provided.</li> </ul>
Dataset for marine vessel detection from Sentinel 2 images in the Finnish coast
<p>This dataset contains annotated marine vessels from 15 different Sentinel-2 product, used for training object detection models for marine vessel detection. The vessels are annotated as bounding boxes, covering also some amount of the wake, if present.</p> <h2>Source data</h2> <div> <div>Individual products used to generate annotations are shown in the following table:</div> </div> <div> </div> <div> <table style="width: 58.034%; height: 411.47px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"><strong>Location</strong></td> <td style="width: 79.3617%; height: 19.5938px;"><strong>Product name</strong></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Archipelago sea</td> <td style="width: 79.3617%; height: 39.1875px;">S2A_MSIL1C_20220515T100031_N0510_R122_T34VEM_20240617T162344.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220619T100029_N0510_R122_T34VEM_20240627T204751.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220721T095041_N0510_R079_T34VEM_20240712T224506.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220813T095601_N0510_R122_T34VEM_20240717T115958.SAFE</td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Gulf of Finland</td> <td style="width: 79.3617%; height: 39.1875px;">S2B_MSIL1C_20220606T095029_N0510_R079_T35VLG_20240619T111429.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220626T095039_N0510_R079_T35VLG_20240620T013500.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220703T094039_N0510_R036_T35VLG_20240702T075354.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220721T095041_N0510_R079_T35VLG_20240712T224506.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Bay</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220627T100611_N0510_R022_T34WFT_20240628T041908.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220712T100559_N0510_R022_T34WFT_20240718T033027.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220828T095549_N0510_R122_T34WFT_20240708T035231.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Sea</td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20210714T100029_N0500_R122_T34VEN_20230224T120043.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220619T100029_N0510_R122_T34VEN_20240627T204751.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220624T100041_N0510_R122_T34VEN_20240714T110124.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220813T095601_N0510_R122_T34VEN_20240717T115958.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Kvarken</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220617T100611_N0510_R022_T34VER_20240627T094433.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220712T100559_N0510_R022_T34VER_20240718T033027.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220826T100611_N0510_R022_T34VER_20240705T062429.SAFE</td> </tr> </tbody> </table> </div> <div> <div> </div> <div>Even though the reference data IDs are for L1C products, L2A products from the same acquisition dates can be used along with the annotations. However, Sen2Cor has been known to produce incorrect reflectance values for water bodies.</div> <div> </div> <div>The corresponding L2A product identifiers are:</div> </div> <div> </div> <div> <table style="width: 58.034%; height: 411.47px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"><strong>Location</strong></td> <td style="width: 79.3617%; height: 19.5938px;"><strong>Product name</strong></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Archipelago sea</td> <td style="width: 79.3617%; height: 39.1875px;">S2A_MSIL2A_20220515T100031_N0400_R122_T34VEM_20220515T141508.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220619T100029_N0510_R122_T34VEM_20240628T011619.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220721T095041_N0510_R079_T34VEM_20240713T035445.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220813T095601_N0510_R122_T34VEM_20240717T165127.SAFE</td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Gulf of Finland</td> <td style="width: 79.3617%; height: 39.1875px;">S2B_MSIL2A_20220606T095029_N0510_R079_T35VLG_20240619T162121.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220626T095039_N0510_R079_T35VLG_20240620T063951.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220703T094039_N0510_R036_T35VLG_20240702T130032.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220721T095041_N0510_R079_T35VLG_20240713T035445.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Bay</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220627T100611_N0510_R022_T34WFT_20240628T095704.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220712T100559_N0510_R022_T34WFT_20240718T063657.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220828T095549_N0510_R122_T34WFT_20240708T091048.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Sea</td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20210714T100029_N0500_R122_T34VEN_20230224T182455.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220619T100029_N0510_R122_T34VEN_20240628T011619.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220624T100041_N0510_R122_T34VEN_20240714T162313.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220813T095601_N0510_R122_T34VEN_20240717T165127.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Kvarken</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220617T100611_N0510_R022_T34VER_20240627T130404.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220712T100559_N0510_R022_T34VER_20240718T063657.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220826T100611_N0510_R022_T34VER_20240705T120522.SAFE</td> </tr> </tbody> </table> </div> <div><br> <div>The raw products can be acquired from <a href="https://dataspace.copernicus.eu" target="_blank" rel="noopener">Copernicus Data Space Ecosystem.</a> The products listed above can be unavailable due to e.g. processing level updates and old versions being deleted. In those cases, try searching with the tile identifier and acquisition date in order to get the correct product ID.</div> <br> <h2>Annotations</h2> <br> <div>The annotations are bounding boxes drawn around marine vessels so that some amount of their wakes, if present, are also contained within the boxes. The data are distributed as geopackage files, so that one geopackage corresponds to a single Sentinel-2 tile, and each package has separate layers for individual products as shown below:</div> <br> <blockquote> <div>T34VEM</div> <div>|-20220515</div> <div>|-20220619</div> <div>|-20220721</div> <div>|-20220813</div> </blockquote> <br> <div>All layers have a column <strong>id</strong>, which has the value <strong>b</strong><strong>oat</strong> for all annotations.</div> <br> <div>CRS is EPSG:32634 for all products except for the Gulf of Finland (35VLG), which is in EPSG:32635. This is done in order to have the bounding boxes to be aligned with the pixels in the imagery.</div> <br> <div>As tiles 34VEM and 34VEN have an overlap of 9.5x100 km, 34VEN is not annotated from the overlapping part to prevent data leakage between splits.</div> <br> <h3>Annotation process</h3> The minimum size for an object to be considered as a potential marine vessel was set to 2x2 pixels. Three separate acquisitions for each location were used to detect smallest objects, so that if an object was located at the same place in all images, then it was left unannotated. The data were annotated by two experts. <div> </div> <table style="width: 63.327%; height: 391.876px;"> <tbody> <tr style="height: 39.1875px;"> <td style="width: 72.7285%; height: 39.1875px;"><strong>Product name</strong></td> <td style="width: 23.0224%; height: 39.1875px;"><strong>Number of annotations</strong></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220515T100031_N0510_R122_T34VEM_20240617T162344.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">183</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220619T100029_N0510_R122_T34VEM_20240627T204751.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">519</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220721T095041_N0510_R079_T34VEM_20240712T224506.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1518</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220813T095601_N0510_R122_T34VEM_20240717T115958.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1371</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220606T095029_N0510_R079_T35VLG_20240619T111429.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">277</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220626T095039_N0510_R079_T35VLG_20240620T013500.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1205</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220703T094039_N0510_R036_T35VLG_20240702T075354.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">746</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220721T095041_N0510_R079_T35VLG_20240712T224506.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">971</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220627T100611_N0510_R022_T34WFT_20240628T041908.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">122</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220712T100559_N0510_R022_T34WFT_20240718T033027.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">162</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220828T095549_N0510_R122_T34WFT_20240708T035231.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">98</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20210714T100029_N0500_R122_T34VEN_20230224T120043.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">450</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220619T100029_N0510_R122_T34VEN_20240627T204751.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">66</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220624T100041_N0510_R122_T34VEN_20240714T110124.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">424</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220813T095601_N0510_R122_T34VEN_20240717T115958.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">399</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220617T100611_N0510_R022_T34VER_20240627T094433.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">83</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220712T100559_N0510_R022_T34VER_20240718T033027.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">184</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;">S2A_MSIL1C_20220826T100611_N0510_R022_T34VER_20240705T062429.SAFE</td> <td style="width: 23.0224%; height: 19.5938px;">88</td> </tr> </tbody> </table> <br><br> <h3>Annotation statistics</h3> <br>Sentinel-2 images have spatial resolution of 10 m, so below statistics can be converted to pixel sizes by dividing them by 10 (diameter) or 100 (area).</div> <div> <table> <tbody> <tr> <td> </td> <td><strong>mean</strong></td> <td><strong>min</strong></td> <td><strong>25%</strong></td> <td><strong>50%</strong></td> <td><strong>75%</strong></td> <td><strong>max</strong></td> </tr> <tr> <td><strong>Area (m²)</strong></td> <td>5305.7</td> <td>567.9</td> <td>1629.9</td> <td>2328.2</td> <td>5176.3</td> <td>414795.7</td> </tr> <tr> <td><strong>Diameter (m)</strong></td> <td>92.5</td> <td>33.9</td> <td>57.9</td> <td>69.4</td> <td>108.3</td> <td>913.9</td> </tr> </tbody> </table> <br><br> <div>As most of the annotations cover also most of the wake of the marine vessel, the bounding boxes are significantly larger than a typical boat. There are a few annotations larger than 100 000 m², which are either cruise or cargo ships that are travelling along ordinal directions instead of cardinal directions, instead of e.g. smaller leisure boats.</div> <br> <div>Annotations typically have diameter less than 100 meters, and the largest diameters correspond to similar instances than the largest bounding box areas.</div> <br> <h3>Train-test-split</h3> <br> <div>We used tiles 34VEN and 34VER as the test dataset. For validation, we split the other three tile areas into 5x5 equal sized grid, and used 20 % of the area (i.e 5 cells) for the validation. The same split also makes it possible to do cross-validation.</div> <div> </div> <div> </div> <div> </div> </div> <div> <h3>Post-processing</h3> </div> <div><br> <div>Before evaluating, the predictions for the test set are cleaned using the following steps:</div> <br> <div>1. All prediction whose centroid points are not located on water are discarded. The water mask used contains layers `jarvi` (Lakes), `meri` (Sea) and `virtavesialue` (Rivers as polygon geometry) from the Topographical database by the National Land Survey of Finland. Unfortunately this also discards all points not within the Finnish borders.</div> <div>2. All predictions whose centroid points are located on water rock areas are discarded. The mask is the layer `vesikivikko` (Water rock areas) from the Topographical database.</div> <div>3. All predictions that contain an above water rock within the bounding box are discarded. The mask contains classes `38511`, `38512`, `38513` from the layer `vesikivi` in the Topographical database.</div> <div>4. All predictions that contain a lighthouse or a sector light within the bounding box are discarded. Lighthouses and sector lights come from Väylävirasto data, `ty_njr` class ids are 1, 2, 3, 4, 5, 8</div> <div>5. All predictions that are wind turbines, found in Topographical database layer `tuulivoimalat`</div> <div>6. All predictions that are obviously too large are discarded. The prediction is defined to be "too large" if either of its edges is longer than 750 meters.</div> </div> <div> </div> <div>Model checkpoint for the best performing model is available on Hugging Face platform: <a href="https://huggingface.co/mayrajeo/marine-vessel-detection-yolov8">https://huggingface.co/mayrajeo/marine-vessel-detection-yolo</a><br> <h2>Usage</h2> The simplest way to chip the rasters into suitable format and convert the data to COCO or YOLO formats is to use <a href="https://github.com/mayrajeo/geo2ml">geo2ml</a>. First download the raw mosaics and convert them into GeoTiff files and then use the following to generate the datasets. <div> </div> To generate COCO format dataset run</div> <div> </div> <div> <pre><code>from geo2ml.scripts.data import create_coco_dataset raster_path = '<path_to_raster>' outpath = '<path_to_save_the_dataset>' poly_path = '<path_to_gpkg>' layer = '<date_of_raster>' create_coco_dataset(raster_path=raster_path, polygon_path=poly_path, target_column='id', gpkg_layer=layer, outpath=outpath, save_grid=False, dataset_name='<name_of_dataset>', gridsize_x=320, gridsize_y=320, ann_format='box', min_bbox_area=0)</code></pre> </div> <div><br> <div>To generate YOLO format dataset run</div> <div> <pre><code>from geo2ml.scripts.data import create_yolo_dataset raster_path = '<path_to_raster>' outpath = '<path_to_save_the_dataset>' poly_path = '<path_to_gpkg>' layer = '<date_of_raster>' create_yolo_dataset(raster_path=raster_path, polygon_path=poly_path, target_column='id', gpkg_layer=layer, outpath=outpath, save_grid=False, gridsize_x=320, gridsize_y=320, ann_format='box', min_bbox_area=0)</code></pre> </div> </div>
Events Dataset for Image Sanitization.
<p><strong>Overview</strong></p> <p>Datasets introduced in "Semi-Supervised Feature Embedding for Data Sanitization in Real-World Events"</p> <p>It includes the links for the images that consists the five datasets.</p> <p> </p> <p><strong>NotreDame</strong>: Remarkable partial destruction of the Parisian Cathedral by a fire in 2019.</p> <p><strong>Grenfell</strong>: 2017 tragic fire incident in the Grenfell Tower in London.</p> <p><strong>NationalMuseum</strong>: Total destruction of the National Museum in Brazil by flames in 2018.</p> <p><strong>BangladeshFire</strong>: Fast-moving fire in a district in Dakha, that took place in 2019.</p> <p><strong>BostonMarathon</strong>: 2013 terrorist attack on the traditional Bostonian event.</p> <p> </p> <p><strong>Observations:</strong></p> <p>- We did not publish the links for the positive images for Grenfell dataset due to copyrights reasons.</p> <p>- BostonMarathon and Grenfell negative sets are pictures from Flickr100k dataset</p> <p>- Since BostonMarathon positive samples contains frames from Youtube videos, we published the Youtube video URL and the number of the frame that was extracted.</p> <p>- BostonMarathon positive sample contains augmentation, that are crops which size is half of the original image. In this sense, columns <em>i</em> and <em>j</em> indicates the upper-left pixel of the crop. For instance, if the image is a <em>x </em>b, the crop will be the square between the points: j <em>x</em> i; (j + a/2) <em>x</em> i; j <em>x</em> (i + b/2); and (j + a/2) <em>x</em> (i + b/2).</p> <p> </p> <p><strong>Media Content</strong></p> <p>Due to the terms of use from the social networks, we do not make publicly available the texts, images and videos that were collected (only the links). However, we can provide some extra piece of media content related to one (or more) events by contacting the authors.</p> <p> </p> <p><strong>Funding</strong></p> <p>DéjàVu thematic project, São Paulo Research Foundation (2017/12646-3, 2018/05668-3, and 2020/02241-9)</p> <p> </p>
Image dataset for the evaluation of a low-cost high-throughput plant phenotyping system
<p>This dataset contains the raw and processed images from a low-cost high-throughput plant phenotyping (HTP) system, as well as the raw and processed images that were manually acquired for comparison. The HTP images were automatically and wirelessly acquired for entire benches of plants with a system composed of a Raspberry Pi and eight GoPro cameras. The entire file system of each GoPro camera was copied directly into a subfolder of finalGoProImages (numbered by camera). The raw HTP images were processed by correcting for lens distortion, computing the "greenness index" for each individual pixel, and filtering out extreme high and low values. These processed HTP images were then saved in the "greenness" subfolder of finalGoProImages. The manually acquired images in the finalDSLR folder each represent an individual plant from one of five time points during the same greenhouse experiment. The raw manually acquired images were processed in the same manner as the raw HTP images by computing the greenness index for each individual pixel and filtering out extreme high and low values. The two tab-delimited text files include the number of green pixels and mean greenness index for each HTP (greennessGoProTable2.txt) and manually acquired (greennessDSLRTable2.txt) image.</p>
Four angle fused dataset for Ascidian embryo imaged via light sheet
<p>Original dataset imaged and published here: </p> <pre>DOI: 10.6084/m9.figshare.8235473.v1</pre> <p>This dataset is provided as Raw dataset for training deep neural networks for segmentation tasks. The binary masks and the integer labels are provided separately.</p>
AUTH-OpenDR Mixed Image Annotated Dataset for Human-centric Perception Tasks
<p>The dataset was generated through a mixed (real and synthetic) image data generation method which utilizes real background images and DL-generated human models. It contains 50000 real images depicting urban scenes, populated by synthetic human models in various positions and poses and is suitable for training/evaluating (a) pose estimation, (b) person detection, (c) identity recognition methods. Annotations for 2D bounding boxes of the depicted humans, their IDs and 2D keypoints etc are provided. The 133 3D human models, required by the method, were generated using the Pixel-aligned Implicit Function (PIFu) and full-body images of people from the Clothing Co-Parsing (CCP) dataset. As background images, a subset of the Cityscapes dataset was used. The Cityscapes license prohibits the distribution of any modified versions of itself. Thus, we provide code that can re-generate the exact same dataset, given that the Cityscapes dataset is downloaded by the website of its authors.</p> <p>Code and instructions for re-generating the dataset are provided <a href="https://github.com/opendr-eu/opendr/tree/master/projects/python/simulation/human_dataset_generation">here</a>.</p> <p>The dataset was developed by Aristotle University of Thessaloniki (AUTH) within the H2020 OpenDR Project.</p>
Dataset for Automated Image Analysis for Single-Atom Detection in Catalytic Materials by Transmission Electron Microscopy
<p>Raw and processed image data resulting from the paper "Automated Image Analysis for Single-Atom Detection in Catalytic Materials by Transmission Electron Microscopy", by S. Mitchell, F. Parés, D. Faust Akl, S. M. Collins, D. M. Kepaptsoglou, Q. M. Ramasse, D. Garcia-Gasulla, J. Pérez-Ramírez, and N. López (JACS, 2021). </p> <p>The corresponding code can be found under: <a href="https://github.com/HPAI-BSC/AtomDetection_ACSTEM">GitHub - HPAI-BSC/AtomDetection_ACSTEM</a></p>
A benchmark dataset of herbarium specimen images with label data: Summary
<p>This landing page contains a CSV file compiling all data associated with herbarium specimens that are part of this dataset, as they could be found on GBIF, JACQ or FinBIF. A CSV file with and without Darwin Core extension data is available, as some CSV readers have trouble with the JSON format that is used for those extensions.</p> <p>In addition, DOI's of the individual specimens uploaded to Zenodo and direct links to the different files (JPEG, TIFF, JSON, PNG) are also included. Index of these added variables:</p> <p>- persistentID: Persistent Identifier of the collection specimen. Data uploaded as part of this dataset will not be kept in sync with changes at the collection's repository. Hence, this URI will always point to the most up to date information known about the herbarium specimen.</p> <p>- jpegURL, tiffURL, jsonURL: URL's pointing straight to the respective image and data files themselves, to facilitate (selective) batch downloads.</p> <p>- pngSegAllURL and pngSegSelURL: Segmented overlays of the herbarium specimens indicating the location of different labels and reference material on the sheet ("All") and their content ("Sel"). More information can be found in the paper (in prep) associated with this data publication and the individual depositions themselves.</p> <p>- DOI: The DOI of the deposition of images and data of these specimens on Zenodo. DOI's point to the most up-to-date version of these depositions at the time of the publication of this CSV file. As a rule, this CSV file will be updated should any changes happen to any of the depositions.</p> <p>- jpegURL2, tiffURL2: A few herbarium sheets had labels on the back and consisted therefore of two scans. As a rule, the label scans are in this category.</p>
Zellige example dataset: synthetic image dataset
<p><strong>Phantom 3D image containing three distinct and superimposed synthetic surfaces. </strong></p> <p>It models a typical stack of confocal images of epithelial and non-epithelial structures.The surfaces generated are of two types: “solid” surfaces, presenting a homogeneous signal over the entire surface, or surfaces presenting a signal restricted to a polygonal mesh mimicking the mesh of apical cellular junctions of an epithelium observed at its surface. This dataset contains both the ground-truth height maps and the height maps generated with Zellige. The Zellige parameters used are:</p> <p><span class="math-tex">\(T_{A}=16, T_{otsu}=12, S_{min}=5, \sigma_{xy}=4, \sigma_{z}=2, T_{OSE1}=0.9, R_{1}=5, C_{1}=0.1, T_{OSE2}=0.1, R_{2}=10, C_{2}=0.8.\)</span></p> <p>Nota: the ground-truth height maps can be directly compared to Zellige height maps by subtraction.</p> <p>See the accompanying paper: Extracting multiple surfaces from 3D microscopy images in complex biological tissues with the Zellige software tool. Trébeau <em>et al.</em> 2022: <a href="https://doi.org/10.1101/2022.04.05.485876">https://doi.org/10.1101/2022.04.05.485876</a></p>
Dataset: Halving of Swiss glacier volume since 1931 observed from terrestrial image photogrammetry
<p>This is supplementary data for the article currently in review for The Cryosphere, titled "Halving of Swiss glacier volume since 1931 observed from terrestrial image photogrammetry".</p> <p><a href="https://doi.org/10.5194/tc-2022-14">See the preprint here</a></p> <p> </p>
Dataset of polygons with the contour of 900 juniper shrubs used to track shrub growth from 1977 to 2020 in Sierra Nevada (Spain) using very high resolution aerial and satellite RGB images.
<p><strong>This database provides as polygons the contours of 900 juniper shrubs (<em>Juniperus communis L.</em> and <em>Juniperus sabina L.</em>) along 5 decades (years 1977, 1984, 2001, 2010 and 2020). The contour of each of 900 shrubs manually mapped using the Google Satellite composite for the year 2020) was tracked back in time using orthophotos provided by REDIAM. Contours were obtained by manual annotation as polygon shapefiles in QGIS 3.10.3. Additionally, for the year 2020, the polygons were characterized with five attributes that gather ecological information: Morphotype (Hemispherical, Striped, Senescent, With rock), Presence of surrounding vegetation (Bare Soil, Surrounding Vegetation), Presence of nearby human land-uses (Surrounded by human facilities within 250 meters, Non-anthropized environment) Health status (as percentage of canopy cover with brown foliage: values between 0-5, where 0 corresponds to 100% photosynthetically active cover, decreasing the photosynthetically active cover until category 5 which corresponds to 100% damaged cover), and the subjective annotation certainty of the GIS technician (values between 0-5, where the value 0 corresponds to a very uncertain annotation up to the value 5 which corresponds to a fairly certain annotation). </strong></p>
DECIMER Image classifier dataset
<p>Images dataset divided into train (10905114 images), validation (2115528 images) and test (544946 images) folders containing a balanced number of images for two classes (chemical structures and non-chemical structures).</p> <p>The chemical structures were generated using RanDepict to random picked compounds from the ChEMBL30 database and the COCONUT database.</p> <p>The non-chemical structures were generated using Python or they were retrieved from several public datasets:</p> <p>COCO dataset, MIT Places-205 dataset, Visual Genome dataset, Google Open labeled Images, MMU-OCR-21 (kaggle), HandWritten_Character (kaggle), CoronaHack -Chest X-Ray-dataset (kaggle), PANDAS Augmented Images (kaggle), Bacterial_Colony (kaggle), Ceylon Epigraphy Periods (kaggle), Chinese Calligraphy Styles by Calligraphers (kaggle), Graphs Dataset (kaggle), Function_Graphs Polynomial (kaggle), sketches (kaggle), Person Face Sketches (kaggle), Art Pictograms (kaggle), Russian handwritten letters (kaggle), Handwritten Russian Letters (kaggle), Covid-19 Misinformation Tweets Labeled Dataset (kaggle) and grapheme-imgs-224x224 (kaggle).</p> <p>This data was used to build a CNN classification model using as a base model EfficienNetB0 and fine tuning it. The model is available on <a href="https://github.com/Iagea/CNN_chem_not_chem">Github</a>.</p>
Datasets for Background and Shading Correction of Optical Microscopy Images by BaSiC -- Downsampled Version
<p>This repository holds downsampled example data for publication: "<strong>A BaSiC tool for background and shading correction of optical microscopy images, Nature Communications (2017)</strong>" DOI: <a href="https://doi.org/10.1038/ncomms14836">https://doi.org/10.1038/ncomms14836</a>. For full-resolution testing data, please refer to Zenodo repository at DOI: <a href="https://zenodo.org/record/6334810#.YvD6zHZBxD8">10.5281/zenodo.6334810</a>.</p>
4 image lysozyme dataset recorded on the Jungfrau 16M detector at SwissFEL and formatted as a NeXus file
<p>This is a 4 image lysozyme datasets derived from https://doi.org/10.5281/zenodo.3352357. The specific 4 images are able to be processed by the software package DIALS using commands in the linked dataset above. The images were rounded to integer and compressed to save file space using this script:</p> <pre><code class="language-python">import shutil, h5py import numpy as np shutil.copyfile('../lyso009a_0087.JF07T32V01_master.h5', 'lyso009a_0087.JF07T32V01_master_4img.h5') h5 = h5py.File('lyso009a_0087.JF07T32V01_master_4img.h5', 'r+') data = h5["entry/data/data"][()] del h5["entry/data/data"] h = h5["entry/data"] subset = data[5:9].astype(np.int32) h.create_dataset("data", subset.shape, subset.dtype, subset, compression="gzip", compression_opts=9) h5.close()</code></pre> <p>The .expt file was created by dials.import and is useful for regression testing in DIALS.</p>
Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging – dataset
<p>Raw volumetric data used in the work "Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging" (<a href="http://doi.org/10.1364/BOE.466403">doi.org/10.1364/BOE.466403</a>). The data is packaged using the FIJI BigStitcher into HDF5 file. The file is split into 89 parts in ZIP format. Additionally we provide XML file needed for opening the data with BigStitcher and the TXT file with the nominal locations of the volumes based on the readings from the X-Y translation stage. The volumes inside the HDF5 file are already registered for stitching using the BigStitcher pairwise registration and global optimization procedure. Using the BigStitcher option "Resave to TIFF" one can access the raw data that we processed in the work. The processing code which operates on TIFF files is available here: <a href="https://github.com/biopto/QPI-stitching-2D-3D">https://github.com/biopto/QPI-stitching-2D-3D</a>.</p>
BioSR+: Dataset Extension of biological images for super-resolution microscopy
<p>BioSR+ dataset is an extension of our pre-published BioSR dataset of biological images for super-resolution microscopy, currently including image pairs of low-and-high resolution images of five biology structures (CCPs, ER, MTs, F-actin, Myosin-IIA) and 8 signal levels for each ROI. The BioSR+ dataset is related to our Nature Methods paper "Evaluation and development of deep neural networks for image super-resolution in optical microscopy" (DOI: 10.1038/s41592-020-01048-5) and Nature Biotechnology paper "Rationalized deep learning super-resolution <br> microscopy for sustained live imaging of rapid subcellular processes" (DOI:10.1038/s41587-022-01471-3). Both BioSR and BioSR+ are freely available and can be used for non-commercial purposes with proper citations of above two papers.</p>
Small-angle X-ray scattering datasets for imaging crossing fibers in mouse, pig, monkey, and human brain
<p>Small-angle X-ray scattering datasets for resolving crossing fibers (myelinated neuronal axon bundles), as described in</p> <p>"<strong><em>Imaging crossing fibers in mouse, pig, monkey, and human brain using small-angle X-ray scattering</em></strong>"</p> <p>deposited in bioRxiv:</p> <p>https://doi.org/10.1101/2022.09.30.510198</p>
TUT Acoustic scenes 2017, Evaluation & Development datasets, processed image
<p>Unseparated Pulse Energy Spectrogram</p> <p>Processed audio data.</p> <p>Sound source separation is a <strong>preliminary</strong> for <strong>acoustic scene classification</strong>. It can be argued that rare sound detection can be performed without separation, but in most cases it also depends on it.</p> <p>I have come up with the theory that the full <strong>time-domain</strong>, or if assumptions are made on the amplitude-waveform or the phase-profile, even the <strong>sequence</strong> of events can be <strong>discarded</strong> for acoustic scene classification.</p> <p>For short time frame bins, a <strong>statistical representation</strong> should be enough to correctly identify the scene. Even more so, if deep learning methods are applied.</p> <p>I have also come up with the theory that <strong>energy</strong> scalograms are applied <strong>pulse-length</strong> or waveform/profile-length wise. This can enhance the input representation for machine learning.</p> <p>Furthermore I have used derivatives of the time signal and applied similar signal processing methods to them. For visualisation I have added them to the original scalogram in different colors. The use of <strong>derivatives</strong> is very much <strong>distorted</strong>, if the sound is not separated.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.