Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,250
datasets available to search
ShareScore release 0.7.1
Dataset results
6,250 results for “classification”
Machine learning classifiers for species classification of fungi using error-prone long-reads on extended metabarcodes
<p>Machine learning models used in the decision tree of linked machine learning models (<a href="https://github.com/teenjes/fungal_ML">https://github.com/teenjes/fungal_ML</a>)</p>
Raw data for the article "Games on Climate Change: Identifying Development Potentials through Advanced Classification and Game Characteristics Mapping"
<p>Raw data used for the article "Gerber, Andreas, Markus Ulrich, Flurin X. Wäger, Marta Roca-Puigròs, João S.V. Gonçalves, and Patrick Wäger. 2021. "Games on Climate Change: Identifying Development Potentials through Advanced Classification and Game Characteristics Mapping" <em>Sustainability</em> 13, no. 4: 1997. <a href="https://doi.org/10.3390/su13041997">https://doi.org/10.3390/su13041997</a>"</p> <p>The documents include the raw data (both as .csv and .xlsx files with the same content), as well as the publication (.pdf file). The data collection process and the data itself are described in the publication. The data is published as "supplementary material" on the publisher's homepage.</p>
Social Innovations for Circularity in the Built Environment – a Scoping Review and Classification. Supplementary Data to the Bibliometric Review.
<p>To transition the built environment (BE) towards circularity, i.e., maximizing the time resources spend in the BE, thus minimizing negative environmental impacts of resource usage, social innovations (SI) – understood as new ways of doing, organizing, framing, and knowing – are just as important as technological advancements. This article provides an overview of the state of the knowledge regarding SI that contribute to circularity in the BE, and proposes a framework for coherently classifying such SI in terms of their main categories and effects.</p> <p>To identify and understand the current knowledge regarding social innovations (SI) that contribute to circularity in the built environment (BE), a bibliometric review of scientific literature is conducted. It shows that the term social innovation is not frequently used in this contexts although the buzzwords circularity and circular economy are themselves often framed as SI. To assess the characteristics and contribution of SI to circularity in the BE, a scoping review that includes grey literature into the context was conducted and a framework developed to classify predominant SI using concepts from transition studies as well as the systems thinking approach. The framework is designed to help assess the potential of SI and to identify research and/or action gaps. Key findings are (1) There is a broad diversity of SIs that contribute to circularity in the BE already in the focus of research, although they are not always identified as such; (2) Most SI focus on either the design or the demolition phase (i.e., market related phases), whereas user-centered SI are less frequently discussed; and (3) It is crucial to also consider potential sustainability goal conflicts in order to guide policies that address SI as a solution.</p> <p>These datasets are the basis to the bibliometric literature review.</p>
Tomato Classification using Mass Spectrometry-Machine Learning Technique: a Food Safety-enhancing Platform
<p>Food safety and quality assessment mechanisms are unmet needs that industries and countries have been continuously facing in recent years. Our study aimed at developing a platform using Machine Learning algorithms to analyze Mass Spectrometry data for classification of tomatoes on organic and non-organic. Tomato samples were analyzed using silica gel plates and direct-infusion electrospray-ionization mass spectrometry technique. Decision Tree algorithm was tailored for data analysis. This model achieved 92% accuracy, 94% sensitivity and 90% precision in determining to which group each fruit belonged. Potential biomarkers evidenced differences in treatment and production for each group.</p>
SeasoNet: A Seasonal Scene Classification, Segmentation and Retrieval Dataset for Satellite Imagery over Germany
<p>This dataset consists of 1,759,830 multi-spectral image patches from the Sentinel-2 mission, annotated with image- and pixel-level land cover and land usage labels from the German land cover model LBM-DE2018 with land cover classes based on the CORINE Land Cover database (CLC) 2018. It includes pixel synchronous examples from each of the four seasons, plus an additional snowy set, spanning the time from April 2018 to February 2019. The patches were taken from 519,547 unique locations, covering the whole surface area of Germany, with each patch covering an area of 1.2km x 1.2km. The set is split into two overlapping grids, consisting of roughly 880,000 samples each, which are shifted by half the patch size in both dimensions. The images in each of the both grids themselves do not overlap.</p> <p><strong>Contents</strong></p> <p>Each sample includes:</p> <ul> <li>3 10m resolution bands (RGB), 120px x 120px</li> <li>1 10m resolution band (infrared), 120px x 120px</li> <li>6 20m resolution bands, 60px x 60px</li> <li>2 60m resolution bands, 20xp x 20px</li> <li>1 pixel-level label map</li> <li>2 binary masks for cloud and snow coverage</li> <li>2 binary masks for easy and medium segmentation difficulties, marks areas <300px and <100px respectively</li> <li>1 JSON-file containing additional meta-information</li> </ul> <p>The meta.csv contains the following information about each sample:</p> <ul> <li>Which season it belongs to</li> <li>Which of the two grids it belongs to</li> <li>Coordinates of the patch center</li> <li>Whether it was acquired from Sentinel-2 Satellite A or B</li> <li>Date and time of image acquisition</li> <li>Snow and cloud coverage percentages</li> <li>Image-level multi-class labels</li> <li>Three additional image-level urbanization labels, based on the center pixel (details below)</li> <li>The path to the sample</li> </ul> <p><strong>Classes</strong></p> <table> <thead> <tr> <th scope="col">ID</th> <th scope="col">Class</th> </tr> </thead> <tbody> <tr> <td>1</td> <td>Continuous urban fabric</td> </tr> <tr> <td>2</td> <td>Discontinuous urban fabric</td> </tr> <tr> <td>3</td> <td>Industrial or commercial units</td> </tr> <tr> <td>4</td> <td>Road and rail networks and associated land</td> </tr> <tr> <td>5</td> <td>Port areas</td> </tr> <tr> <td>6</td> <td>Airports</td> </tr> <tr> <td>7</td> <td>Mineral extraction sites</td> </tr> <tr> <td>8</td> <td>Dump sites</td> </tr> <tr> <td>9</td> <td>Construction sites</td> </tr> <tr> <td>10</td> <td>Green urban areas</td> </tr> <tr> <td>11</td> <td>Sport and leisure facilities</td> </tr> <tr> <td>12</td> <td>Non-irrigated arable land</td> </tr> <tr> <td>13</td> <td>Vineyards</td> </tr> <tr> <td>14</td> <td>Fruit trees and berry plantations</td> </tr> <tr> <td>15</td> <td>Pastures</td> </tr> <tr> <td>16</td> <td>Broad-leaved forest</td> </tr> <tr> <td>17</td> <td>Coniferous forest</td> </tr> <tr> <td>18</td> <td>Mixed forest</td> </tr> <tr> <td>19</td> <td>Natural grasslands</td> </tr> <tr> <td>20</td> <td>Moors and heathland</td> </tr> <tr> <td>21</td> <td>Transitional woodland/shrub</td> </tr> <tr> <td>22</td> <td>Beaches, dunes, sands</td> </tr> <tr> <td>23</td> <td>Bare rock</td> </tr> <tr> <td>24</td> <td>Sparsely vegetated areas</td> </tr> <tr> <td>25</td> <td>Inland marshes</td> </tr> <tr> <td>26</td> <td>Peat bogs</td> </tr> <tr> <td>27</td> <td>Salt marshes</td> </tr> <tr> <td>28</td> <td>Intertidal flats</td> </tr> <tr> <td>29</td> <td>Water courses</td> </tr> <tr> <td>30</td> <td>Water bodies</td> </tr> <tr> <td>31</td> <td>Coastal lagoons</td> </tr> <tr> <td>32</td> <td>Estuaries</td> </tr> <tr> <td>33</td> <td>Sea and ocean</td> </tr> </tbody> </table> <p><strong>Urbanization classes</strong></p> <ul> <li><strong>SLRAUM</strong> <ul> <li>0: None</li> <li>1: Ländlicher Raum (~ rural area)</li> <li>2: Städtischer Raum (~ urban area)</li> </ul> </li> <li><strong>RTYP3</strong> <ul> <li>0: None</li> <li>1: Ländliche Regionen (~ rural areas)</li> <li>2: Regionen mit Verstädterungsansätzen (~ urbanizing areas)</li> <li>3: Städtische Regionen (~ urban areas)</li> </ul> </li> <li><strong>KTYP4</strong> <ul> <li>0: None</li> <li>1: Dünn besiedelte ländliche Kreise</li> <li>2: Kreisfreie Großstädte</li> <li>3: Ländliche Kreise mit Verdichtungsansätzen</li> <li>4: Städtische Kreise</li> </ul> </li> </ul> <p>Further information on the urbanization classes can be found here:</p> <p><strong>SLRAUM</strong></p> <p><a href="https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/kreise/staedtischer-laendlicher-raum/kreistypen.html">https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/kreise/staedtischer-laendlicher-raum/kreistypen.html</a></p> <p><strong>RTYP3</strong></p> <p><a href="https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/regionen/siedlungsstrukturelle-regionstypen/regionstypen.html">https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/regionen/siedlungsstrukturelle-regionstypen/regionstypen.html</a></p> <p><strong>KTYP4</strong></p> <p><a href="https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/kreise/siedlungsstrukturelle-kreistypen/kreistypen.html">https://www.bbsr.bund.de/BBSR/DE/forschung/raumbeobachtung/Raumabgrenzungen/deutschland/kreise/siedlungsstrukturelle-kreistypen/kreistypen.html</a></p> <p><strong>License of landcover model</strong></p> <p>Bundesamt für Kartographie und Geodäsie</p> <p>dl-de/by-2-0 from <a href="https://www.govdata.de/dl-de/by-2-0">https://www.govdata.de/dl-de/by-2-0</a></p> <p>© GeoBasis-DE / <strong>BKG</strong> 2022</p> <p><strong>Source of landcover model</strong></p> <p><a href="https://gdz.bkg.bund.de/index.php/default/catalog/product/view/id/1071/s/corine-land-cover-5-ha-stand-2018-clc5-2018/">https://gdz.bkg.bund.de/index.php/default/catalog/product/view/id/1071/s/corine-land-cover-5-ha-stand-2018-clc5-2018/</a></p>
Climate-representative locations according to the ASHRAE 169-2020 climate classification and within the WMO Region VI (Europe)
<p>This dataset comprises climate-representative locations according to the ASHRAE 169-2020 climate classification and within the WMO Region VI (Europe). This is, for each climate zone of the ASHRAE 169-2020 standard within the WMO Region VI (Europe), one location in close agreement with its centroid was selected based on the classification criteria. The recent typical meteorological year (TMYx.2007-2021 or TMYx.2004-2018) for each location is freely provided by Climate.One.Building.Org repository (https://climate.onebuilding.org/).</p>
Explainable AI for unveiling deep learning pollen classification model - Pollen dataset
<p>Dataset consists automatic particle detector Rapid-E measurements of pollen grains from 12 classes: Acer, Alnus, Alopecurus, Carex, Cupressus, Dactylis, Juglans, Morus, Platanus, Populus, Salix and Ulmus. Data are available i json format.</p> <p>Dataset also contains preprocessed data packed into csv files of 3 modalities: spectrum, lifetime, scattering and additional features from scattering and lifetime data are also available. These are ready to be used with machine learning models. Labels 0, 1, 2, ... 11 correspond to alphabetical order of examined pollen classes Acer, Alnus, Alopecurus ... Ulmus.</p> <p> </p>
Classification of Artificial Intelligence and eXplainable Artificial Intelligence publications in Air Traffic Management
<p>v1.0 version used and partially published in "A Survey on Artificial Intelligence (AI) and eXplainable AI in Air Traffic Management: Current Trends and Development with Future Research Trajectory". In this version, it references mainly Transportation Reasearch Part C, ICRAT, Journal of ATM, and ATM Seminar, IEEE transaction on ITS, but not only</p>
TCAB: Text Classification Attack Benchmark Dataset
<p>TCAB is a large collection of successful adversarial attacks on state-of-the-art text classification models trained on multiple sentiment and abuse domain datasets.</p> <p>The dataset is broken up into 2 files: <em>train.csv and</em> <em>val.csv</em>. The training set contains 1,448,751 instances (552,364 are "clean" unperturbed instances) and the validation set contains 482,914 instances (178,607 are "clean"). Each instance contains the following attributes:</p> <p><strong>scenario</strong>: Domain, either <em>abuse</em> or <em>sentiment</em>.</p> <p><strong>target_model_dataset</strong>: Dataset being attacked.</p> <p><strong>target_model_train_dataset</strong>: Dataset the target model trained on.</p> <p><strong>target_model</strong>: Type of victim model (e.g., <em>bert</em>, <em>roberta</em>, <em>xlnet</em>).</p> <p><strong>attack_toolchain</strong>: Open-source attack toolchain, either TextAttack or OpenAttack.</p> <p><strong>attack_name</strong>: Name of the attack method.</p> <p><strong>original_text</strong>: Original input text.</p> <p><strong>original_output</strong>: Prediction probabilities of the target model on the original text.</p> <p><strong>ground_truth</strong>: Encoded label for the original task of the domain dataset. 1 and 0 means toxic and toxic for abuse datasets, respectively. 1 and 0 means positive and negative sentiment for sentiment datasets. If there is a neutral sentiment, then 2, 1, 0 means positive, neutral, and negative sentiment.</p> <p><strong>status</strong>: Unperturbed example if "clean"; successful adversarial attack if "success".</p> <p><strong>perturbed_text</strong>: Text after it has been perturbed by an attack.</p> <p><strong>perturbed_output</strong>: Prediction probabilities of the target model on the perturbed text.</p> <p><strong>attack_time</strong>: Time taken to execute the attack.</p> <p><strong>num_queries</strong>: Number of queries performed while attacking.</p> <p><strong>frac_words_changed</strong>: Fraction of words changed due to an attack.</p> <p><strong>test_index</strong>: Index of each unique source example (original instance) (LEGACY - necessary for backwards compatibility).</p> <p><strong>original_text_identifier</strong>: Index of each unique source example (original instance).</p> <p><strong>unique_src_instance_identifier</strong>: Primary key to uniquely identify to every source instance; comprised of (<em>target_model_dataset</em>, <em>test_index</em>, <em>original_text_identifier</em>).</p> <p><strong>pk</strong>: Primary key to uniquely identify every attack instance; comprised of (<em>attack_name</em>, <em>attack_toolchain</em>, <em>original_text_identifier</em>, <em>scenario</em>, <em>target_model</em>, <em>target_model_dataset</em>, <em>test_index).</em></p>
Experiment on the performance of different machine learning algorithms for classification - Results
<h2>Results of a short performance study of machine learning algorithms</h2> <h3>Context and methodology</h3> <ul> <li>This data was produced while performing a university project to examine the performance of various machine learning algorithms on different prediction datasets</li> <li>The data serves the purpose of comparing the metrics of performing the different tasks</li> <li>The dataset contains a number of matrices for every classifier and every dataset</li> <li>The data was produced with python scripts provided further down and with the usage of the external datasets: <ul> <li>Membership Woes Dataset (OpenML): <a href="https://api.openml.org/d/44224">https://api.openml.org/d/44224</a></li> <li>Zoo dataset (UCI): <a href="https://doi.org/10.24432/C5R59V">https://doi.org/10.24432/C5R59V</a></li> <li>Breast Cancer Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> <li>Loan Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> </ul> </li> </ul> <h3>Technical details</h3> <ul> <li>The data consists of one JSON file</li> <li>The source code for producing this data is available at <a href="https://doi.org/10.5281/zenodo.11085222">https://doi.org/10.5281/zenodo.11085222</a></li> </ul> <h3>Structure of the data</h3> <p>[ {"classifier": ...,<br>"dataset": ...,<br>"hyper_parameters": ...,<br>"cross_validation_results": {<br> "fit_time": {} ,<br> "score_time": ...,<br> "metrics": {}<br>}, <br>"holdout_test_results": ...}, ]</p>
MUSCLE: MUlti Session Classification Longterm Electromyography
<p>The <strong><em>MUSCLE</em></strong> dataset, acronym for <em><strong>MU</strong>lti <strong>S</strong>ession <strong>C</strong>lassification <strong>L</strong>ongterm <strong>E</strong>lectromyography</em>, represents a significant resource for evaluating the robustness of upper limb prosthesis control. It features 10 recordings of High-Density surface electromyography (HD-sEMG) data gathered over an extensive timeframe. Utilizing a novel HD-sEMG setup, the dataset comprises 64 EMG signals during 10 repetitions of 10 distinct gestures, mirroring the functionality of a 5 Degrees of Freedom upper limb prosthetic device. This comprehensive dataset, accompanied by the release of control model code, not only enhances transparency but also fosters advancements in prosthetics research. For a deeper understanding of the HD-sEMG acquisition setup and control method implementation, readers are encouraged to explore the related publication titled "Non-Invasive High-Density Electromyography for Real-Time Control of 5 DoFs Hand Prosthesis", D. Di Domenico*, M. Iacono* et al., (2024).</p>
Classification and frequency of climate change drivers and responses of small-scale fishers found in literature review
<p>Climate change hazards were classified into resource availability and fishing operations or both following the framework proposed by Cheung et al (2012). Response units were firstly classified into overarching responses and then categorized as suggested by the adaptive-transformative framework of Barnes et al.<sup> </sup>(2020). Adaptation units that did not represent an active adaptation response were classified as remaining. More than one hazard could be attributed to each fishers' response.</p>
Jacobaea vulgaris and meadow image classification dataset (binary)
<h3>General Information</h3> <p>Instances in the Jacobaea vulgaris class: 895<br>Instances in the Meadow class: 9141<br>Image sizes from 77x77 to 817x817 pixels on three color channels (RGB)</p> <p> </p> <h3>Data Generation and Source</h3> <p>The images in this dataset were taken as part of the project “UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene” (engl. “UAV-based Grassland Monitoring at Population and Individual Plant Level”), financed by the Authority for Economy, Transport, and Innovation of Hamburg. <br>In September 2018, flights with an octocopter were conducted over two extensively used grassland areas in the urban area of Hamburg. The multicopter flew in a height of circa 11 meters and took pictures with a ground resolution of approximately 3,18 mm/pixel. Additional information about the process of image generation for this dataset are to be found in the relevant papers written by P. Zacharias: 1) <a href="https://archiv.geomv.de/geoforum/2019/doc/Tagungsband_GeoForum-MV-2019_eBook.pdf" target="_blank" rel="noopener">UAV-basiertes Grünland-Monitoring und Schadpflanzenkartierung mit offenen Geodaten</a> [p. 45–53] and 2) <a href="https://www.auf.uni-rostock.de/storages/uni-rostock/Alle_AUF/AUF/GG/PDF/gruenlandmonitoring/2019-12-12-FHH-Workshop_Vortrag_Zacharias.pdf" target="_blank" rel="noopener">UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene</a>.</p> <p>Additionally, to the images of Jacobaea vulgaris taken by the UAV, the dataset includes images of Jacobaea vulgaris plants from the internet (included in the total 895 images; e.g. images 'jkk0523.jpg', 'jkk0527.jpg'). Furthermore, some of the images of the Jacobaea vulgaris plants have been rotated, further cropped or a filter has been applied. The exact number of augmentations made is unknown. As there are augmented images included in the datasets -which makes the dataset useful for training and validation- a use of the dataset for testing purposes is not recommended due to the risk of data leakage.</p> <h3>Data License</h3> <p>The dataset is licensed under the license CC BY 4.0. The attributor of the data is the Chair of Geodesy and Geoinformatics at the University of Rostock. The data was created within the scope of the project 'UAV-based Grassland Monitoring at Population and Individual Plant Level', financed by the Authority for Economy, Transport, and Innovation of Hamburg.</p> <p> </p>
Jacobaea vulgaris and meadow Augmented image classification dataset (binary)
<h3>General Information</h3> <p>Total instances: 117008<br>Instances in the Jacobaea vulgaris class: 58504 <br>Instances in the Meadow class: 58504<br>Image sizes from 224x224 pixels on three color channels (RGB)</p> <p><br>Performance increase training a ResNet50 on the base dataset versus the same architecture on the augmented data set shared here: +3,79 percent points in ROC AUC on an independent test set with 240 instances.<br><br></p> <h3>Data Generation and Source</h3> <p>The initial images in this dataset were taken as part of the project “UAV-basiertes Grünlandmonitoring auf Bestands- und Einzelpflanzenebene” (engl. “UAV-based Grassland Monitoring at Population and Individual Plant Level”), financed by the Authority for Economy, Transport, and Innovation of Hamburg. <br>In September 2018, flights with an octocopter were conducted over two extensively used grassland areas in the urban area of Hamburg.</p> <p>In my master's thesis at <a href="https://www.tu.berlin/dams">DAMS Lab</a> at TU Berlin, I evaluated the effect of different augmentation strategies for Jacobaea vulgaris image classification on the several performance metrics (most importantly the ROC AUC score). The identified augmentation strategies are -besides to performance based selection- also selected based on domain knowledge, which I acquired during the research for my master thesis. </p> <p>Additional information about the initial image generation process is to be found <a href="https://archiv.geomv.de/geoforum/2019/doc/Tagungsband_GeoForum-MV-2019_eBook.pdf">here </a> [p. 45–53] and <a href="https://www.auf.uni-rostock.de/storages/uni-rostock/Alle_AUF/AUF/GG/PDF/gruenlandmonitoring/2019-12-12-FHH-Workshop_Vortrag_Zacharias.pdf">here</a>. </p> <h3> </h3> <h3>Augmentations applied</h3> <ul> <li>Gaussian Noise: For the Gaussian noise augmentation, the mean of the added noise is set to zero. The lower and upper bounds for the random variance of the noise are 20.4663 and 54.0395 respectively. The bounds were identified by hyperparameter tuning. The search space for the lower bound was set from 5 to 30 and for the upper bound from 31 to 100. Those two search spaces were defined by visual inspection of the effects of applying Gaussian noise with different variance values to images of both classes. The Gaussian noise is sampled for each color channel individually. </li> <li>Random Brightness and Contrast: The brightness will randomly be increased or decreased by a factor ranging from 0.7010 to 1.2990. The The contrast will also be randomly increased by a factor ranging from 0.5775 to 1.4225. Those two ranges were identified using hyperparameter tuning. The search space for the maximal percentual increase or decrease of brightness and contrast was individually set from 1% to maximally 50% increase or decrease.</li> <li>Cutout Dropout: In this augmentation method a certain percentage of the input image is getting covered by black patches. The patches have a certain size in pixels, the implementation of this technique in this thesis uses square patches. The black patches are then randomly introduced into the image, by randomly alloacting the<br>patches across the image and then setting the corresponding pixel values to zero. The iamge is getting covered with patches until the cover percentage is reached. We<br>set percentage of the image to be randomly covered by black patches to 56.76%. The size of the patches, which randomly cover the image, is set to 4 pixels. A<br>good illustration of this is found in figure 4.2. The augmentation technique is inspired by the research proposed by Devries et al.[8]. Both values were identified by hyperparameter tuning. The search space for the patch size in pixels is categorical and includes the values [1, 2, 4, 7, 8, 14, 16, 28]. Those values all are multiples of 224, which is the image width and height in pixels. The patch size needs to be a multiple of the width and height in order to be suitable for the algorithm implementation. The search space for the cover percentage of the image had been set from 1% to 60%. This search space limits narrows the search down to a space where still a big part of the image is uncovered. The algorithm rearranges the image into a two dimensional grid and randomly masks rows of this grid by setting the pixel values in this row to zero. Then, the image gets rearranged, now with the randomly generated patches included.</li> <li>Random Saturation: The saturation of each pixel is randomly getting shifted. The upper bound for randomly shifting<span> </span>the saturation value of each pixel is set to 231.689%. This value was identified using hyperparameter tuning. An upper limit of the maximal saturation shift had been set to 40% shift in either direction for hyperparameter tuning.</li> <li>Horizontal Flip: The image gets flipped along the horizontal axis. </li> <li>Vertical Flip: The image gets flipped along the vertical axis.</li> <li>Random Rotation 90 degrees: Randomly rotates the image by a k-fold of 90 degrees, whereby k = {0, 1, 2, 3}.</li> </ul> <p> </p> <p>All augmentation methods and with their tuned augmentation hyperparameters (if existent) are applied to an image from the test set in figure 4.2. With the seven identified<br>augmentation techniques a dataset of 800% the size of the original dataset is created. The Augment model is trained on exactly this dataset. Of course next to the augmented images, the dataset still includes the original, unaugmented images. TensorFlow, along with additional libraries including Optuna for hyperparameter optimization and Albumentations for image augmentation, were used in for the implementation of this project.</p> <p> </p> <h3>Rational behind the augmentations applied</h3> <ul> <li>Random Rotation, Vertical and Horizontal Flip: These three augmentation strategies were chosen to make the classifier less sensitive to the orientation of the plant. The goal is to train a model that can classify plants regardless of their orientation. In order to achieve this effectively across different orientations, vertical flips, horizontal flips, and random 90-degree rotations are chosen for evaluation.</li> <li>Random Saturation: The varying saturation of the images simulates different levels of chlorophyll in the leaves, which is responsible for the green color of the<br>leaves and the intensity of this color. The color of the plant parts (leaves, stems, and flowers) is also influenced by factors such as soil, sun, weed density and pressure, location, and water availability. Varying the saturation of the images simulates changes in these factors.</li> <li>Gaussian Noise: By adding noise, in this case Gaussian noise, different lighting conditions are simulated when capturing the images. We specifically chose Gaussian<br>noise because it is common in many real-world scenarios and is based on the Central Limit Theorem, which states that the sum of many independent random variables.<br>tends to be normally distributed. This makes Gaussian noise a logical choice for simulating real-world random noise.</li> <li>Random Brightness Contrast: The Random Brightness and Random Contrast Augmentation uses brightness to mimic varying lighting conditions and contrast to highlight differences between plants by contrasting them more strongly, thereby highlighting their edges. This approach for highlighting edges is of course much more subtle than the canny edge detection augmentation. This augmentation method combines a weak focus on edges with variations in lighting conditions in one approach. The random contrast is a much softer approach for highlighting edges of plants, compared to the Canny edge detection augmentation. The other features in the images do not get changed that much, compared to the changes from edge detection augmentation.</li> <li>Cutout Dropout: The cutout augmentation simulates random occlusion by other plants. These occlusions are common and expected. Jacobaea vulgaris plants may be partially or completely obscured by other plants during image capturing. This augmentation technique makes the models more robust to random occlusion.</li> </ul> <h3> </h3> <h3>Data License</h3> <p>The dataset is licensed under the license CC BY 4.0. The attributor of the data is the Chair of Geodesy and Geoinformatics at the University of Rostock. The data was created within the scope of the project 'UAV-based Grassland Monitoring at Population and Individual Plant Level', financed by the Authority for Economy, Transport, and Innovation of Hamburg.</p>
Supplementary dataset to the publication "Bi, S., and Hieronymi, M. (2024). Holistic optical water type classification for ocean, coastal, and inland waters. Limnology & Oceanography"
<p>The NetCDF data files contain the training dataset used to develop the Optical Water Type (OWT) framework proposed by Bi and Hieronymi (2024). The dataset is available in two spectral versions:</p> <p> 1. <code>owt_BH2024_training_data_hyper.nc</code>: This file includes training data with a spectral resolution of 2 nm, ranging from 400 to 900 nm.<br> 2. <code>owt_BH2024_training_data_olci.nc</code>: This file contains data formatted similarly to the hyperspectral version but aligned with the nominal Sentinel-3 OLCI wavebands.</p> <h2>Contents of the Dataset</h2> <p>For each version, the dataset includes spectral inherent and apparent optical properties such as:</p> <p> • Remote Sensing Reflectance (Rrs)<br> • Pure Water Absorption (aw)<br> • Absorption Coefficient of Detritus (ad)<br> • Total Absorption Coefficient without Pure Water (agp)<br> • Absorption Coefficient of Phytoplankton (aph)<br> • Backscattering Coefficient of Total Particulate Matter (bbp)<br> • Scattering Coefficient of Total Particulate Matter (bp)<br> • Scattering Coefficient of Pure Water (bw)</p> <p>Additionally, the dataset includes various environmental and biological parameters:</p> <p> • Chlorophyll a Concentration (Chl)<br> • Inorganic Suspended Matter Concentration (ISM)<br> • Colored Dissolved Organic Matter Absorption at 440 nm (ag440)<br> • Single-Scattering Albedo of Detritus at 550 nm (A_d)<br> • Power Law Exponent of Detritus Attenuation (G_d)<br> • Water Salinity (Sal)<br> • Water Temperature (Temp)<br> • Fraction for Diminished Coccolithophore Absorption (a_frac)<br> • Fraction of Coccolithophore Group (cocco_frac)</p> <h2>Optical Water Types</h2> <p>The training dataset includes 10 pre-defined optical water types, with 10,000 samples for each type. Detailed descriptions of these water types can be found in Table 1 of Bi and Hieronymi (2024) or as follows,</p> <table> <tbody> <tr> <td>OWT</td> <td>Desciption</td> </tr> <tr> <td>1</td> <td>Extremely clear and oligotrophic indigo-blue waters with high reflectance in the short visible wavelengths.</td> </tr> <tr> <td>2</td> <td>Blue waters with similar biomass level as OWT 1 but with slightly higher detritus and CDOM content.</td> </tr> <tr> <td>3a</td> <td>Turquoise waters with slightly higher phytoplankton, detritus, and CDOM compared to the first two types.</td> </tr> <tr> <td>3b</td> <td>A special case of OWT 3a with similar detritus and CDOM distribution but with strong scattering and little absorbing particles like in the case of Coccolithophore blooms. This type usually appears brighter and exhibits a remarkable ~490 nm reflectance peak.</td> </tr> <tr> <td>4a</td> <td>Greenish water found in coastal and inland environments, with higher biomass compared to the previous water types. Reflectance in short wavelengths is usually depressed by the absorption of particles and CDOM.</td> </tr> <tr> <td>4b</td> <td>A special case of OWT 4a, sharing similar detritus and CDOM distribution, exhibiting phytoplankton blooms with higher scattering coefficients, e.g., Coccolithophore bloom. The color of this type shows a very bright green.</td> </tr> <tr> <td>5a</td> <td>Green eutrophic water, with significantly higher phytoplankton biomass, exhibiting a bimodal reflectance shape with typical peaks at ~560 and ~709 nm.</td> </tr> <tr> <td>5b</td> <td>Green hyper-eutrophic water, with even higher biomass than that of OWT 5a (over several orders of magnitude), displaying a reflectance plateau in the Near Infrared Region, NIR (vegetation-like spectrum).</td> </tr> <tr> <td>6</td> <td>Bright brown water with high detritus concentrations, which has a high reflectance determined by scattering.</td> </tr> <tr> <td>7</td> <td>Dark brown to black water with very high CDOM concentration, which has low reflectance in the entire visible range and is dominated by absorption.</td> </tr> </tbody> </table> <h2>Additional Information</h2> <p>The detailed description of the data simulation can be found in the supporting information of Bi and Hieronymi (2024). The models used for simulating the data are available on GitHub:</p> <p> • Component IOP Model: <a href="https://github.com/bishun945/IOPmodel" target="_blank" rel="noopener">Bio-geo-optical modelling of natural waters by Bi, Hieronymi, and Röttgers (2023)</a><br> • OWT Package: <a href="https://github.com/bishun945/pyOWT" target="_blank" rel="noopener">pyOWT</a></p> <h2>References</h2> <p> 1. OWT Framework: Bi, S., and Hieronymi, M. (2024). Holistic optical water type classification for ocean, coastal, and inland waters. Limnology & Oceanography, lno.12606. doi: 10.1002/lno.12606<br> 2. Component IOP Model: Bi, S., Hieronymi, M., and Röttgers, R. (2023). Bio-geo-optical modelling of natural waters. Front. Mar. Sci. 10, 1196352. doi: 10.3389/fmars.2023.1196352<br> 3. Pure Water IOP Model: Röttgers, R., Doerffer, R., McKee, D., and Schönfeld, W. (2016). The Water Optical Properties Processor (WOPP): Pure Water Spectral Absorption, Scattering and Real Part of Refractive Index Model. Technical Report No WOPP-ATBD/WRD6. Available at: https://calvalportal.ceos.org/tools<br> 4. Rrs Model: Lee, Z., Du, K., Voss, K. J., Zibordi, G., Lubac, B., Arnone, R., et al. (2011). An inherent-optical-property-centered approach to correct the angular effects in water-leaving radiance. Appl. Opt. 50, 3155. doi: 10.1364/AO.50.003155</p> <h2>Authors and Contact</h2> <p> • Author: Shun Bi, Martin Hieronymi, Rüdiger Röttgers<br> • Creator: Shun Bi, Shun.Bi@hereon.de</p> <h2>Example Python Code to Read Data</h2> <p>Here is an example of how to read the NetCDF data using Python and the <code>xarray</code> library:</p> <pre><code>import xarray as xr # Load the dataset data_hyper = xr.open_dataset("path_to_your_file/owt_BH2024_training_data_hyper.nc") # Print the dataset to see its structure print(data_hyper) # Access a specific variable, e.g., remote sensing reflectance (Rrs) rrs = data_hyper['Rrs'] # Plot a sample of Rrs import matplotlib.pyplot as plt # Select a sample ID, for example the first sample sample_id = 0 plt.plot(data_hyper['wavelen'], rrs[sample_id, :]) plt.xlabel('Wavelength (nm)') plt.ylabel('Rrs (1/sr)') plt.title(f'Remote Sensing Reflectance for Sample ID {sample_id}') plt.show()</code></pre>
CLDF dataset derived from Blum et al.'s "A phylolinguistic classification of the Quechua language family" from 2023
<p>Cite the source of the dataset as:</p> <blockquote> <p>Blum, Frederic, Carlos Barrientos, Adriano Ingunza & Zoe Poirier. 2023. A phylolinguistic classification of the Quechua language family. INDIANA - Anthropological Studies on Latin America and the Caribbean 40(1). 29–-54. DOI: https://doi.org/10.18441/IND.V40I1.29-54.</p> </blockquote>
CLDF dataset derived from Chacon's "A revised proposal of Proto-Tukanoan consonants and Tukanoan family classification" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>Thiago Chacon. (2014). A revised proposal of Proto-Tukanoan consonants and Tukanoan family classification. Journal of American Linguistics 80.3, pp. 275–322. doi: https://doi.org/10.1086/676393</p> </blockquote>
CLDF dataset derived from Peiros' "Genetic classification of Austro-Asiatic languages" from 2004
<p>Cite the source of the dataset as:</p> <blockquote> <p>Peiros, I. I. (2004): Genetičeskaja klassifikacija avstroaziatskix jazykov [Genetic classification of Austro-Asiatic languages]. Russian State University for the Humanities, Russian State University for the Humanities, Moscow.</p> </blockquote>
CLDF dataset derived from Robinson and Holton's "Internal Classification of the Alor-Pantar Language Family" from 2012
<p>Cite the source of the dataset as:</p> <blockquote> <p>Robinson, Laura C. and Holton, Gary (2012): Internal Classification of the Alor-Pantar Language Family Using Computational Methods Applied to the Lexicon. Language Dynamics and Change 2.2. 123-149.</p> </blockquote>
A recent overview of the state-of-the-art elements of text classification - dataset
<p>The two available datasets were used to conduct the quantitative analysis of the text classification area. The set, such as:</p> <ol> <li>biblio.bib contains all articles that are grouped in categories</li> <li>biblio.csv contains processed records from biblio.bib, based on it were built the statistics presented in the article</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.