Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.7.1
Dataset results
166 results for “image classification”
Crowd4SDG - Crowdsourced image classification and damage assessment
<p>This data set contains crowdsourced classification and damage assessment of images of an earthquake extracted from social media. <br> <br> A data set of 907 images posted on Twitter related to the 2019 Albanian Earthquake, that are filtered and pre-classified using an automated technique is cross-validated for accuracy by two different crowds. One, digital humanitarian volunteers using the crowdsourcing platform <a href="http://www.crowd4ems.org">CROWD4EMS</a> and another, paid micro-taskers of the Amazon Mechanical Turk. In order to compare and evaluate the efficiency and accuracy of the volunteers and the paid micro taskers, ground truth is established with the help of a team of experts, who validated the same set of data. <br> <br> <strong>Parameters considered for volunteer contributions:</strong> The dataset was imported to the Crowd4EMS platform for Crowd contribution. In the forum, each volunteer will see the image to be validated along with the tweet text and the link to the original tweet. The user has to validate whether the given image is <em>relevant or</em> <em>irrelevant</em> to the disaster. In case of doubt, the user can refer to the tutorial explaining the relevance or skip the task. Once the image's relevance is validated, the user will be asked to label the <em>severity</em> of the impact, as seen in the image.</p> <p>The Automated algorithm has pre-classified the images as <em>severe </em>and <em>minimal </em>damage. The Crowd4EMS platform lets the volunteer label them as '<em>severe damage</em>,' <em>moderate damage'</em>,' <em>minimal damage', </em>and' <em>no damage'.</em> Each task has to be answered <em>at least three times</em>, and the final consensus is taken as per the<em> inter-rater agreement. </em><br> <br> <strong>Parameters considered for micro-taskers contribution:</strong> The dataset was imported to the <em>Amazon Mechanical Turk</em> platform for Crowd contribution. In the platform, each worker will see only the image that is to be categorised as follows: The user has to validate whether the given image depicts <em>severe damage, moderate damage, minimal damage, no damage </em>or <em>irrelevant</em> to the disaster. Each task has to be answered <em>at least ten times</em>, and the final consensus is taken as per the<em> inter-rater agreement. </em><br> <br> <strong>Acknowledgements:</strong> We want to thank Muhammad Imran of Qatar Computing Research Institute for sharing their pre-filtered social media imagery dataset on the Albanian earthquake from the Artificial Intelligence for Disaster Response (AIDR) Platform. We would also like to extend our gratitude to the volunteers for their contribution on the Crowd4EMS Platform.<br> </p>
Infection Inspection: Classifications and images of ciprofloxacin-treated Escherichia coli clinical isolates
<p>This dataset includes a .csv file with the image metadata and a folder of RGB images of <i>E. coli</i> grown from clinical isolates with varying concentrations of the antibiotic ciprofloxacin and varying minimum inhibitory concentrations. The <i>E. coli</i> cell membranes are stained with Nile Red and the DNA is stained with DAPI. The details of the image data collection are included in: https://doi.org/10.1038/s42003-023-05524-4. The classification data come from a Zooniverse citizen science project, Infection Inspection. (https://www.zooniverse.org/projects/conor-feehily/infection-inspection) Volunteers learned how to interpret ciprofloxacin response phenotypes as antibiotic-sensitive or antibiotic-resistant, and their classifications are included in the Metadata.csv file.</p><p>This dataset could be used for further analysis into the volunteer classifications, or the image data could be used for further image feature analysis of the ciprofloxacin response phenotypes.</p>
Training data for: CoastSat image classification
<p><strong>CoastSat image classification training data </strong></p> <p>CoastSat is an open-source global shoreline mapping toolbox, available at https://github.com/kvos/CoastSat, which enables users to extract time-series of shoreline change from 30+ years of publicly available satellite imagery (Landsat 5, 7, 8 and Sentinel-2).</p> <p>The automated shoreline extraction relies on a classifier (Multilayer Perceptron from scikit-learn) which labels each pixels on the images with one of four classes: sand, water, white-water and other land features.</p> <p>The data used to train the classifier is stored here, the README.md file provides information on the data organisation and content of each file.</p>
Pre-training with simulated ultrasound images for breast mass segmentation and classification - dataset
<p>Dataset assosiated with the MICCAI Workshop on Data Engineering in Medical Imaging paper: "Pre-training with Simulated Ultrasound Images for Breast Mass Segmentation and Classification"</p>
Data release for "OrchID: a Generalized Framework for Taxonomic Classification of Images Using Evolved Artificial Neural Networks"
<p><strong>Abstract</strong></p> <p>Taxonomic expertise for the identification of species is rare and costly. On-going advances in computer vision and machine learning have led to the development of numerous semi- and fully automated species identification systems. However, these systems are rarely agnostic to specific morphology, rarely can perform taxonomic “approximation” (by which we mean partial identification at least to higher taxonomic level if not to species), and frequently rely on costly scientific imaging technologies.</p> <p>We present a generic, hierarchical identification system for automated taxonomic approximation of organisms from images. We assessed the effectiveness of this system using photographs of slipper orchids (Cypripedioideae), for which we implemented image pre-processing, segmentation, and colour and shape feature extraction algorithms to obtain digital phenotypes for 116 species. The identification system trained on these digital phenotypes uses a nested hierarchy of artificial neural networks for pattern recognition and automated classification that mirrors the Linnean taxonomy, such that user-submitted photos can be assigned a genus, section, and species classification by traversing this hierarchy.</p> <p>Performance of the identification system varied depending on photo quality, number of species included for training, and desired taxonomic level for identification. High quality photos were scarce for some taxa and were under-represented in the training set, resulting in imbalanced network training. The image features used for training were sufficient to reliably identify photos to the correct genus but less so to the correct section and species.</p> <p>The outcomes of this project include a library of feature extraction algorithms called <em>ImgPheno</em>, a collection of scripts for neural network training called <em>NBClassify</em>, a library for evolutionary optimization of artificial neural network construction called <em>AI::FANN::Evolving</em> and a planned web application called <em>OrchID</em> for identification of user-submitted images. All project outcomes are open source and freely available.</p> <p><strong>About this release</strong></p> <p>This release corresponds belongs with our response to the reviewers of PLoS One. At this stage of the review cycle the manuscript is assessed as 'minor revision'. Consequently, we don't anticipate making more releases until publication.</p>
1988-2009 time-series of land-use/land-cover maps for the Mar Menor / Campo de Cartagena watershed by means of supervised classification of Landsat images.
<p>Serie de mapas de usos y coberturas de la cuenca del Mar Menor (SE España): 2009, 2000, 1997 y 1998. Así como el documento completo de tesis en las que se generaron y analizaron.</p> <p>Time-series of land-use / land-cover maps of Mar Menor watershed (SE Spain): 2009, 2000, 1997 y 1998. As well as the complete thesis document in which they were generated and analyzed.</p>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
[MedMNIST+] 18x Standardized Datasets for 2D and 3D Biomedical Image Classification with Multiple Size Options: 28 (MNIST-Like), 64, 128, and 224
<h2><strong>Code</strong> [<a href="https://github.com/MedMNIST/MedMNIST" target="_blank" rel="noopener">GitHub</a>] | <strong>Publication</strong> [<a href="https://doi.org/10.1038/s41597-022-01721-8" target="_blank" rel="noopener">Nature Scientific Data'23</a> / <a href="https://doi.org/10.1109/ISBI48211.2021.9434062" target="_blank" rel="noopener">ISBI'21</a>] | <strong>Preprint</strong> [<a href="https://arxiv.org/abs/2110.14795" target="_blank" rel="noopener">arXiv</a>]</h2> <p> </p> <p><strong>Abstract</strong></p> <p>We introduce MedMNIST, a large-scale MNIST-like collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D. All images are pre-processed into 28x28 (2D) or 28x28x28 (3D) with the corresponding classification labels, so that no background knowledge is required for users. Covering primary data modalities in biomedical images, MedMNIST is designed to perform classification on lightweight 2D and 3D images with various data scales (from 100 to 100,000) and diverse tasks (binary/multi-class, ordinal regression and multi-label). The resulting dataset, consisting of approximately 708K 2D images and 10K 3D images in total, could support numerous research and educational purposes in biomedical image analysis, computer vision and machine learning. We benchmark several baseline methods on MedMNIST, including 2D / 3D neural networks and open-source / commercial AutoML tools. The data and code are publicly available at <a href="https://medmnist.com/">https://medmnist.com/</a>.</p> <p><em><strong>Disclaimer</strong></em>: The only official distribution link for the MedMNIST dataset is <a href="https://doi.org/10.5281/zenodo.10519652">Zenodo</a>. We kindly request users to refer to this original dataset link for accurate and up-to-date data.</p> <p><strong><em>Update</em>:</strong> We are thrilled to release <a href="https://github.com/MedMNIST/MedMNIST/blob/main/on_medmnist_plus.md">MedMNIST+</a> with larger sizes: 64x64, 128x128, and 224x224 for 2D, and 64x64x64 for 3D. As a complement to the previous 28-size MedMNIST, the large-size version could serve as a standardized benchmark for medical foundation models. Install the latest API to try it out!</p> <p> </p> <p><strong>Python Usage</strong></p> <p>We recommend our official <a href="https://github.com/MedMNIST/MedMNIST">code</a> to download, parse and use the MedMNIST dataset:</p> <blockquote> <pre>% pip install medmnist<br>% python</pre> <div> <div>To use the standard 28-size (MNIST-like) version utilizing the downloaded files:</div> <br> <div>>>> from medmnist import PathMNIST</div> <div>>>> train_dataset = PathMNIST(split="train")</div> <br> <div>To enable automatic downloading by setting `download=True`:</div> <br> <div>>>> from medmnist import NoduleMNIST3D</div> <div>>>> val_dataset = NoduleMNIST3D(split="val", download=True)</div> <br> <div>Alternatively, you can access MedMNIST+ with larger image sizes by specifying the `size` parameter:</div> <br> <div>>>> from medmnist import ChestMNIST</div> <div>>>> test_dataset = ChestMNIST(split="test", download=True, size=224)</div> </div> </blockquote> <p> </p> <p><strong>Citation</strong></p> <p>If you find this project useful, please cite both v1 and v2 paper as:</p> <blockquote> <p>Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, Bingbing Ni. Yang, Jiancheng, et al. "MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification." Scientific Data, 2023.</p> <p>Jiancheng Yang, Rui Shi, Bingbing Ni. "MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis". IEEE 18th International Symposium on Biomedical Imaging (ISBI), 2021.</p> </blockquote> <p>or using bibtex:</p> <blockquote> <pre>@article{medmnistv2, title={MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification}, author={Yang, Jiancheng and Shi, Rui and Wei, Donglai and Liu, Zequan and Zhao, Lin and Ke, Bilian and Pfister, Hanspeter and Ni, Bingbing}, journal={Scientific Data}, volume={10}, number={1}, pages={41}, year={2023}, publisher={Nature Publishing Group UK London} } @inproceedings{medmnistv1, title={MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis}, author={Yang, Jiancheng and Shi, Rui and Ni, Bingbing}, booktitle={IEEE 18th International Symposium on Biomedical Imaging (ISBI)}, pages={191--195}, year={2021} }</pre> </blockquote> <p>Please also cite the corresponding paper(s) of source data if you use any subset of MedMNIST as per the description on the <a href="https://medmnist.github.io/">project website</a>.</p> <p> </p> <p><strong>License</strong></p> <p>The MedMNIST dataset is licensed under <em>Creative Commons Attribution 4.0 International</em> (<a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>), except DermaMNIST under <em>Creative Commons Attribution-NonCommercial 4.0 International</em> (<a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC 4.0</a>).</p> <p>The code is under <a href="https://github.com/MedMNIST/MedMNIST/blob/main/LICENSE">Apache-2.0 License</a>.</p> <p> </p> <p><strong>Changelog</strong></p> <p><a href="https://doi.org/10.5281/zenodo.10519652">v3.0</a> (this repository): Released MedMNIST+ featuring larger sizes: 64x64, 128x128, and 224x224 for 2D, and 64x64x64 for 3D.</p> <p><a href="https://doi.org/10.5281/zenodo.10519195">v2.2</a>: Removed a small number of mistakenly included blank samples in OrganAMNIST, OrganCMNIST, OrganSMNIST, OrganMNIST3D, and VesselMNIST3D. </p> <p><a href="https://doi.org/10.5281/zenodo.6496656">v2.1</a>: Addressed an issue in the NoduleMNIST3D file (i.e., nodulemnist3d.npz). Further details can be found in this <a href="https://github.com/MedMNIST/MedMNIST/issues/22#issuecomment-1103438191">issue</a>.</p> <p><a href="https://doi.org/10.5281/zenodo.5208230">v2.0</a>: Launched the initial repository of MedMNIST v2, adding 6 datasets for 3D and 2 for 2D.</p> <p><a href="https://doi.org/10.5281/zenodo.4269852">v1.0</a>: Established the initial repository (in a separate repository) of MedMNIST v1, featuring 10 datasets for 2D.</p> <p> </p> <p><strong>Note</strong>: This dataset is <strong>NOT</strong> intended for clinical use.</p>
Jacobaea vulgaris and meadow image classification dataset (binary)
<h3>General Information</h3> <p>Instances in the Jacobaea vulgaris class: 895<br>Instances in the Meadow class: 9141<br>Image sizes from 77x77 to 817x817 pixels on three color channels (RGB)</p> <p> </p> <h3>Data Generation and Source</h3> <p>The images in this dataset were taken as part of the project “UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene” (engl. “UAV-based Grassland Monitoring at Population and Individual Plant Level”), financed by the Authority for Economy, Transport, and Innovation of Hamburg. <br>In September 2018, flights with an octocopter were conducted over two extensively used grassland areas in the urban area of Hamburg. The multicopter flew in a height of circa 11 meters and took pictures with a ground resolution of approximately 3,18 mm/pixel. Additional information about the process of image generation for this dataset are to be found in the relevant papers written by P. Zacharias: 1) <a href="https://archiv.geomv.de/geoforum/2019/doc/Tagungsband_GeoForum-MV-2019_eBook.pdf" target="_blank" rel="noopener">UAV-basiertes Grünland-Monitoring und Schadpflanzenkartierung mit offenen Geodaten</a> [p. 45–53] and 2) <a href="https://www.auf.uni-rostock.de/storages/uni-rostock/Alle_AUF/AUF/GG/PDF/gruenlandmonitoring/2019-12-12-FHH-Workshop_Vortrag_Zacharias.pdf" target="_blank" rel="noopener">UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene</a>.</p> <p>Additionally, to the images of Jacobaea vulgaris taken by the UAV, the dataset includes images of Jacobaea vulgaris plants from the internet (included in the total 895 images; e.g. images 'jkk0523.jpg', 'jkk0527.jpg'). Furthermore, some of the images of the Jacobaea vulgaris plants have been rotated, further cropped or a filter has been applied. The exact number of augmentations made is unknown. As there are augmented images included in the datasets -which makes the dataset useful for training and validation- a use of the dataset for testing purposes is not recommended due to the risk of data leakage.</p> <h3>Data License</h3> <p>The dataset is licensed under the license CC BY 4.0. The attributor of the data is the Chair of Geodesy and Geoinformatics at the University of Rostock. The data was created within the scope of the project 'UAV-based Grassland Monitoring at Population and Individual Plant Level', financed by the Authority for Economy, Transport, and Innovation of Hamburg.</p> <p> </p>
Jacobaea vulgaris and meadow Augmented image classification dataset (binary)
<h3>General Information</h3> <p>Total instances: 117008<br>Instances in the Jacobaea vulgaris class: 58504 <br>Instances in the Meadow class: 58504<br>Image sizes from 224x224 pixels on three color channels (RGB)</p> <p><br>Performance increase training a ResNet50 on the base dataset versus the same architecture on the augmented data set shared here: +3,79 percent points in ROC AUC on an independent test set with 240 instances.<br><br></p> <h3>Data Generation and Source</h3> <p>The initial images in this dataset were taken as part of the project “UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene” (engl. “UAV-based Grassland Monitoring at Population and Individual Plant Level”), financed by the Authority for Economy, Transport, and Innovation of Hamburg. <br>In September 2018, flights with an octocopter were conducted over two extensively used grassland areas in the urban area of Hamburg.</p> <p>In my master's thesis at <a href="https://www.tu.berlin/dams">DAMS Lab</a> at TU Berlin, I evaluated the effect of different augmentation strategies for Jacobaea vulgaris image classification on the several performance metrics (most importantly the ROC AUC score). The identified augmentation strategies are -besides to performance based selection- also selected based on domain knowledge, which I acquired during the research for my master thesis. </p> <p>Additional information about the initial image generation process is to be found <a href="https://archiv.geomv.de/geoforum/2019/doc/Tagungsband_GeoForum-MV-2019_eBook.pdf">here </a> [p. 45–53] and <a href="https://www.auf.uni-rostock.de/storages/uni-rostock/Alle_AUF/AUF/GG/PDF/gruenlandmonitoring/2019-12-12-FHH-Workshop_Vortrag_Zacharias.pdf">here</a>. </p> <h3> </h3> <h3>Augmentations applied</h3> <ul> <li>Gaussian Noise: For the Gaussian noise augmentation, the mean of the added noise is set to zero. The lower and upper bounds for the random variance of the noise are 20.4663 and 54.0395 respectively. The bounds were identified by hyperparameter tuning. The search space for the lower bound was set from 5 to 30 and for the upper bound from 31 to 100. Those two search spaces were defined by visual inspection of the effects of applying Gaussian noise with different variance values to images of both classes. The Gaussian noise is sampled for each color channel individually. </li> <li>Random Brightness and Contrast: The brightness will randomly be increased or decreased by a factor ranging from 0.7010 to 1.2990. The The contrast will also be randomly increased by a factor ranging from 0.5775 to 1.4225. Those two ranges were identified using hyperparameter tuning. The search space for the maximal percentual increase or decrease of brightness and contrast was individually set from 1% to maximally 50% increase or decrease.</li> <li>Cutout Dropout: In this augmentation method a certain percentage of the input image is getting covered by black patches. The patches have a certain size in pixels, the implementation of this technique in this thesis uses square patches. The black patches are then randomly introduced into the image, by randomly alloacting the<br>patches across the image and then setting the corresponding pixel values to zero. The iamge is getting covered with patches until the cover percentage is reached. We<br>set percentage of the image to be randomly covered by black patches to 56.76%. The size of the patches, which randomly cover the image, is set to 4 pixels. A<br>good illustration of this is found in figure 4.2. The augmentation technique is inspired by the research proposed by Devries et al.[8]. Both values were identified by hyperparameter tuning. The search space for the patch size in pixels is categorical and includes the values [1, 2, 4, 7, 8, 14, 16, 28]. Those values all are multiples of 224, which is the image width and height in pixels. The patch size needs to be a multiple of the width and height in order to be suitable for the algorithm implementation. The search space for the cover percentage of the image had been set from 1% to 60%. This search space limits narrows the search down to a space where still a big part of the image is uncovered. The algorithm rearranges the image into a two dimensional grid and randomly masks rows of this grid by setting the pixel values in this row to zero. Then, the image gets rearranged, now with the randomly generated patches included.</li> <li>Random Saturation: The saturation of each pixel is randomly getting shifted. The upper bound for randomly shifting<span> </span>the saturation value of each pixel is set to 231.689%. This value was identified using hyperparameter tuning. An upper limit of the maximal saturation shift had been set to 40% shift in either direction for hyperparameter tuning.</li> <li>Horizontal Flip: The image gets flipped along the horizontal axis. </li> <li>Vertical Flip: The image gets flipped along the vertical axis.</li> <li>Random Rotation 90 degrees: Randomly rotates the image by a k-fold of 90 degrees, whereby k = {0, 1, 2, 3}.</li> </ul> <p> </p> <p>All augmentation methods and with their tuned augmentation hyperparameters (if existent) are applied to an image from the test set in figure 4.2. With the seven identified<br>augmentation techniques a dataset of 800% the size of the original dataset is created. The Augment model is trained on exactly this dataset. Of course next to the augmented images, the dataset still includes the original, unaugmented images. TensorFlow, along with additional libraries including Optuna for hyperparameter optimization and Albumentations for image augmentation, were used in for the implementation of this project.</p> <p> </p> <h3>Rational behind the augmentations applied</h3> <ul> <li>Random Rotation, Vertical and Horizontal Flip: These three augmentation strategies were chosen to make the classifier less sensitive to the orientation of the plant. The goal is to train a model that can classify plants regardless of their orientation. In order to achieve this effectively across different orientations, vertical flips, horizontal flips, and random 90-degree rotations are chosen for evaluation.</li> <li>Random Saturation: The varying saturation of the images simulates different levels of chlorophyll in the leaves, which is responsible for the green color of the<br>leaves and the intensity of this color. The color of the plant parts (leaves, stems, and flowers) is also influenced by factors such as soil, sun, weed density and pressure, location, and water availability. Varying the saturation of the images simulates changes in these factors.</li> <li>Gaussian Noise: By adding noise, in this case Gaussian noise, different lighting conditions are simulated when capturing the images. We specifically chose Gaussian<br>noise because it is common in many real-world scenarios and is based on the Central Limit Theorem, which states that the sum of many independent random variables.<br>tends to be normally distributed. This makes Gaussian noise a logical choice for simulating real-world random noise.</li> <li>Random Brightness Contrast: The Random Brightness and Random Contrast Augmentation uses brightness to mimic varying lighting conditions and contrast to highlight differences between plants by contrasting them more strongly, thereby highlighting their edges. This approach for highlighting edges is of course much more subtle than the canny edge detection augmentation. This augmentation method combines a weak focus on edges with variations in lighting conditions in one approach. The random contrast is a much softer approach for highlighting edges of plants, compared to the Canny edge detection augmentation. The other features in the images do not get changed that much, compared to the changes from edge detection augmentation.</li> <li>Cutout Dropout: The cutout augmentation simulates random occlusion by other plants. These occlusions are common and expected. Jacobaea vulgaris plants may be partially or completely obscured by other plants during image capturing. This augmentation technique makes the models more robust to random occlusion.</li> </ul> <h3> </h3> <h3>Data License</h3> <p>The dataset is licensed under the license CC BY 4.0. The attributor of the data is the Chair of Geodesy and Geoinformatics at the University of Rostock. The data was created within the scope of the project 'UAV-based Grassland Monitoring at Population and Individual Plant Level', financed by the Authority for Economy, Transport, and Innovation of Hamburg.</p>
PlantVillage Disease Classification Challenge - Color Images
<p><br> This work is licensed under a <a href="http://creativecommons.org/licenses/by-sa/3.0/us/">Creative Commons Attribution-ShareAlike 3.0 United States License</a>.<br> <br> # Data origins<br> The dataset is originally hosted at <a href="https://www.crowdai.org/challenges/plantvillage-disease-classification-challenge">PlantVillage Disease Classification Challenge</a>.<br> We use the modified version in <a href="https://github.com/salathegroup/plantvillage_deeplearning_paper_dataset">this github repository</a> to do controlled experiments.<br> We only use the raw color images dataset and delete the unconventional characters in the classes directory name and `.csv` filenames.<br> <br> # Directory explanation<br> The `80-20` direcotry has multiple `.txt` files which contain the training (~80%), validation(~10%) and testing (~10%) datasets instances filenames and the corresponding label indexes. The validation dataset quantity is `5430` in all data separation. In our experiment code (not included in this archive), the validation and testing dataset are merged together.<br> <br> # Data usage<br> ## Replicate our experiments<br> We have used this dataset in writing our paper. The reference information can be seen at https://<a href="https://gitlab.com/huix/leaf-disease-plant-village">gitlab.com/huix/leaf-disease-plant-village</a>.<br> <br> ### Steps<br> 1. `cd` to the direcotry (e.g. `/home/usrname/plantvillage_deeplearning_paper_dataset`) that contains the `color` directory.<br> 2. run `python change_filename_prefix.py --prefix /home/usrname/plantvillage_deeplearning_paper_dataset` to modify the prefix path (which is `/home/h/plantvillage_deeplearning_paper_dataset` in our former generated datasets).<br> 3. Fin. You can use our <a href="https://gitlab.com/huix/leaf-disease-plant-village">opens ource codes repository</a> to do the later experiments.<br> <br> ## Generate your own training/validation/testing datasets<br> This data separation generating code isn't included in the dataset archive, it is in our open source code. Please see our <a href="https://gitlab.com/huix/leaf-disease-plant-village">open source code repository</a> for the detailed information.<br> If you have any questions, you can contact the author through email.<br> The email address is a QR code in the archive.</p>
EOL computer vision pipelines: Classification for Image Tagging: Flower Fruit
<p>Angiosperms: Stats from Colab:</p> <ul> <li>Number of positive identified reproductive structures: 490</li> <li>Number of possible identified reproductive structures: 4611</li> <li>Number of negative identified reproductive structures: 14833</li> </ul> <p> </p>
EOL computer vision pipelines: Classification for Image Tagging: Image Type: Anura
<p>Produced by the EOL Image Type Classifier. Classifies images as map, phylogeny, illustration, herbarium sheet, or none. Dataset generated for EOL Anura images. See model on the CV for <a href="https://www.kaggle.com/models/eolorg/image-quality-rating-bad-vs-good" target="_blank" rel="noopener">EOL Images Model Zoo on Kaggle</a>.</p>
EOL computer vision pipelines: Classification for Image Tagging: Image Rating: Chiroptera
<p>Produced by the EOL Image Rating Classifier. Classifies images as bad or good quality (used for image gallery sorting). Dataset generated for EOL Chiroptera images. See model on the CV for <a href="https://www.kaggle.com/models/eolorg/image-quality-rating-bad-or-good" target="_blank" rel="noopener">EOL Images Model Zoo on Kaggle</a>.</p>
Cosmic Muon Images open classification data
<p>This is a reduced, anonymised dataset containing the classifications made in the Cosmic Muon Images demonstrator project available on Zooniverse during the implementation period - from the 19th of October, 2021, to the 23rd of October, 2023 - as part of the REINFORCE project.</p>
Image-based Classification of Intense Radio Bursts from Spectrograms: An Application to Saturn Kilometric Radiation
<p>A catalogue of 4874 of the Low Frequency Extensions (LFEs) of Saturn Kilometric Radiation (SKR) detected by Cassini/RPWS from the beginning of 2004 until mission end in 2017. The LFEs presented in this catalogue were identified using a modified U-Net architecture that applied semantic segmentation to spectrogram images in order to extract the exact frequency-time coordinates of the LFE. The files consist of a .json file with the coordinates of each LFE in Time Frequency Catalogue (TFCat) format (Cecconi et. al. 2023). We also include a .csv file with the start and stop times of each LFE in the form of python datetime timestamps, with the average predicted probability per LFE as an accompanying column. </p>
Galaxy Zoo DESI: Detailed Morphology Classifications for 8.7M Galaxies in the DESI Legacy Imaging Surveys
<p>This repository contains the data released in the paper "Galaxy Zoo DESI: Detailed Morphology Classifications for 8.7M Galaxies in the DESI Legacy Imaging Surveys" <em>(DOI to follow on publication).</em></p> <p>We release detailed morphology measurements for bright (<em>r </em>< 19) galaxies in the DESI Legacy Imaging Surveys footprint. These measurements estimate the presence of bars, spirals arms, ongoing mergers, and more.</p> <p>---</p> <p><strong>GZ DESI Detailed Morphology Catalogs</strong></p> <p>These catalogs are created by training deep learning models on Galaxy Zoo volunteer responses, to predict what volunteers might say for new galaxies. The models are available at [www.github.com/mwalmsley/zoobot](www.github.com/mwalmsley/zoobot). Our measurements are predicted vote fractions i.e. the fraction of volunteers expected to select a given answer for a given question.</p> <p>We share two catalog versions containing the same morphology measurements but presented in different ways.</p> <p>gz_desi_deep_learning_catalog_friendly.parquet contains the morphology measurements</p> <p>gz_desi_deep_learning_catalog_advanced.parquet contains the same measurements, and additional information:</p> <p>- _friendly includes only relevant vote fractions, defined as vote fractions to answers of questions that a majority of volunteers would have been asked. This removes predicted vote fractions for e.g. the fraction of volunteers answering "2 spiral arms" to a galaxy with no spiral arms. _advanced includes all vote fractions and instead reports the (column "proportion_asked"). The user must select which vote fractions they consider relevant (we suggest proportion_asked > 0.5, which recovers the _friendly fractions).</p> <p>- _advanced includes columns with estimated credible intervals (error bars) around each vote fraction. These are calculated from the vote fraction posterior predicted by our models.</p> <p>Finally, we separately present volunteer votes collected for 96k galaxies during the GZD-8 campaign, i.e. after the release of GZ DECaLS but before this (GZ DESI) release. These are split into the _core and _extended catalogs, where _extended includes galaxies which received five or more votes for "artifact". The models above were trained on these votes as well as votes from GZ DECaLS.</p> <p>---</p> <p><strong>External Catalog</strong></p> <p>For convenience, we also include an additional catalog of non-morphology measurements created by other authors (external_catalog.parquet) cross-matched to our morphology catalogs. Please credit those authors if you use this catalog (references are in the GZ DESI paper).</p> <p>A particularly important external measurement is redshift. Morphology is increasingly hard to resolve at higher redshift and so <strong>distant galaxies appear less featured</strong>. external_catalog.parquet includes the column "redshift", which is the SDSS spectroscopic redshift where available and a photometric redshift estimate otherwise (again, see the GZ DESI paper for references and credit). You may want to select only galaxies at lower redshifts.</p> <p>---</p> <p><strong>Data Notes</strong></p> <p>Parquet is a fast csv-like format which can be read with pd.read_parquet(loc, columns=[some columns]). Parquet files are read column-by-column (rather than row-by-row) and so you can chose which columns to load. You can easily check which columns are available using columns=['foo'] and reading the error message. We suggest loading only the columns you need when working with the larger catalogs. This will require much less memory than loading every column.</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI to follow on publication) when using the data in this repository.</p> <p>---</p> <p><strong>History</strong></p> <p>v0.0.1 - closed pre-release for internal review</p> <p>v1.0.0 - draft public release. Removed low-z pre-filtered catalogs.</p> <p>v1.0.1 - first public release. Added .csv version of _friendly catalog. Tweaked catalog formatting for clarity and consistency.</p>
Raw images used for colony classification
<p>Raw images of plates with yeast colonies used in Figure 2 of "Carl et al. A fully automated deep learning pipeline for high-throughput colony segmentation and classification" (Biology Open 2020 : bio.052936 doi: 10.1242/bio.052936 Published 2 June 2020). The paper describes the development of a novel computational pipeline for colony segmentation and classification that achieves accuracy comparable to human performance.</p> <p>The actual data was generated in a project that was published earlier (Duempelmann, L. e<em>t al. </em>Inheritance of a Phenotypically Neutral Epimutation Evokes Gene Silencing in Later Generations. <em>Molecular Cell</em> <strong>74, 3</strong> (2019).) The experiments were testing trans-generational inheritance of <em>ade6<sup>+</sup> </em>silencing in <em>Schizosaccharomyces pombe</em>. <em>ade6<sup>+</sup></em> silencing was first induced by expression of small interfering RNAs (siRNAs) that are complementary to the <em>ade6<sup>+</sup></em> gene in a <em>paf1-Q264Stop</em> nonsense-mutant background, leading to red colonies. Paf1 is a subunit of the Paf1 complex (Paf1C), which represses siRNA-induced heterochromatin formation in <em>S. pombe</em>.</p>
Dataset: Whole blood count, used in: "AIDeveloper: deep learning image classification in life science and beyond"
<p>Real-time deformability cytometry (RT-DC) data of whole blood measurements.<br> Data was used to train and validate a neural net to perform a blood count based on brightfield images of RT-DC.</p> <p>01_Model: Contains the final model as well as an AIDeveloper meta-file that allows to reproduce the training procedure. The metafile preciesely defines which dataset was used for training and which for validation as well as all parameters that were set in AIDeveloper.</p> <p>The following folders contain data that was used for training (and validation):</p> <ul> <li>Cambr</li> <li>KIK</li> <li>20190306_DextranBlood_AI_DataSet</li> <li>Gs_Blood_Train</li> </ul> <p>Testing data is stored on figshare:<br> https://figshare.com/articles/Krater_et_al_2020_Data_zip/9902636</p>
Automatic plankton image classification - can capsules and filters help coping with data set shift?
<p>This data set is related to the article 'Automatic plankton image classification - can capsules and filters help coping with data set shift?' published in 'Limnology and Oceanography: Methods' by Plonus <em>et al.</em> (2021).</p> <p>The images belong to the trainings set used to train the models in the aforementioned paper (training_) and three different additional data sets which were used to evaluate the performance of the trained models in application mode (fs446_; fs466_; fs534_). The Python-Script 'separate_files.py' can be used to move all the images in different folders for each data set and class respectively.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.