Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
598
datasets available to search
ShareScore release 0.9.0
Dataset results
598 results for “classifier”
Dataset of paper: Supervised and Dynamic Neuro-Fuzzy Systems to Classify Physiological Responses in Robot-Assisted Neurorehabilitation (PLOS One)
<p>The data set contains number of user, user's physiological signals (Pulse, SCL, SCR, Respiration rate, Skin temperature), Label, Difficulty level from relax to stress. Label is codified from 1 to 5 corresponding to the Difficulty level.</p>
QIIME2 2.0.5 feauture-classifier tutorial data
<p>Data for the QIIME2 2.0.5 feature-classifier tutorial.</p> <p>Some data reproduced from here: http://qiime.org/home_static/dataFiles.html</p>
Issue close time: datasets + prediction classifiers
<p>This project contains experiments on predicting the amount of time required to close issue reports in software repositories. Namely, it contains (a) issue lifetime datasets from 10 large software projects and (b) experiment scripts to generate decision tree classifiers that predict issue close time.</p> <p><strong>To run the cross-validation experiment:</strong><br> 1. Compile the Java classes by running "make" or "make compile-java" on the command line<br> 2. Configure the experimental setup by changing the variables at the top of run.sh<br> 3. Run "bash run.sh" on the command line<br> 4. Results can be found in out/</p> <p><strong>To run the round robin experiment:</strong><br> 1. Compile the Java classes by running "make" or "make compile-java" on the command line<br> 2. Run "bash roundRobin.sh" on the command line<br> 3. Results can be found in out/roundRobin</p> <p> </p> <p>The latest version of this project can be found on GitHub: https://github.com/reesjones/issueCloseTime</p>
Classifying the Correctness of Generated White-Box Tests: An Exploratory Study
<p>White-box test generator tools rely only on the code under test to select test inputs, and capture the implementation's output as assertions. If there is a fault in the implementation, it could get encoded in the generated tests. Tool evaluations usually measure fault-detection capability using the number of such fault-encoding tests. However, these faults are only detected, if the developer can recognize that the encoded behavior is faulty. We designed an exploratory study to investigate how developers perform in classifying generated white-box test as faulty or correct. We carried out the study in a laboratory setting with 54 graduate students. The tests were generated for two open-source projects with the help of the IntelliTest tool. The performance of the participants were analyzed using binary classification metrics and by coding their observed activities. The results showed that participants incorrectly classified a large number of both fault-encoding and correct tests (with median misclassification rate 33% and 25% respectively). Thus the real fault-detection capability of test generators could be much lower than typically reported, and we suggest to take this human factor into account when evaluating generated white-box tests.<br> <br> This material contains the dataset and the recorded videos of the study.</p>
Automated bio-AFM generation of large mechanome data set and their analysis by machine learning to classify prostatic cell lines_Training base 100 PC3-GFP
Open the record for dataset details and reuse information.
Labelled dataset to classify direct deforestation drivers in Cameroon
<p><strong>Overview</strong></p> <p>This dataset includes the images (visible bands for Landsat-8 or NICFI PlanetScope), auxiliary data (infrared, NCEP, forest gain, OpenStreetMap, SRTM, GFW), and data about forest loss (Global Forest Change) used to train, validate and test a model to classify direct deforestation drivers in Cameroon. </p> <p><strong>Description of the files</strong></p> <ul> <li>'my_examples_landsat_final_detailed.zip': Landsat-8 images, auxiliary data and forest loss data used to train, validate and test a model for a detailed classification of deforestation drivers in Cameroon (15 classes: ‘Oil palm plantation’, ‘Timber plantation’, ‘Fruit plantation (e.g. banana)’, ‘Rubber plantation’, ‘Other large-scale plantation (e.g. tea, sugarcane)’, ‘Grassland/Shrubland’, ‘Small-scale oil palm plantation’, ‘Small-scale maize plantation’, ‘Other small-scale agriculture’, ‘Mining’, ‘Selective logging’, ‘Infrastructure’, ‘Wildfire’, ‘Hunting’, ‘Other’)</li> <li>'my_examples_planet_final_detailed.zip': NICFI PlanetScope images, auxiliary data and forest loss data used to train, validate and test a model for a detailed classification of deforestation drivers in Cameroon (15 classes)</li> <li>'my_examples_landsat_final.zip': Landsat-8 images, auxiliary data and forest loss data used to train, validate and test a model for a classification of deforestation drivers by groups in Cameroon (4 classes: 'Plantation', 'Grassland/Shrubland', 'Smallholder agriculture', 'Other')</li> <li>'my_examples_planet_final.zip': NICFI PlanetScope images, auxiliary data and forest loss data used to train, validate and test a model for a classification of deforestation drivers by groups in Cameroon (4 classes)</li> <li>'my_examples_landsat_detailed_timeseries.zip': Landsat-8 images, auxiliary data and forest loss data used to test a model for a detailed classification of deforestation drivers in Cameroon (15 classes) using multiple images and a time series analysis </li> <li>'my_examples_planet_detailed_timeseries.zip': NICFI PlanetScope images, auxiliary data and forest loss data used to test a model for a detailed classification of deforestation drivers in Cameroon (15 classes) using multiple images and a time series analysis</li> <li> <p>‘labels.zip’: in csv files, the labels for each image in each folder described above (image identified by folder and coordinates or ‘path’) and matches the format of the csv files used as inputs to train, validate and test our classification model</p> <p>For ‘labels.zip’, we have subfolders for Landsat and PlanetScope. Then, for each type of imagery, we have subfolders for ‘detailed’, ‘groups’ and ‘time series’ which correspond to the different ‘my_examples’ folders listed above. </p> <p>For each folder, subfolders named with the coordinates of the centre of the images contain each:<br>• A folder ‘images’, with a sub-folder ‘visible’ containing the PNG RGB image; and a sub-folder ‘infrared’ containing the infrared bands in a NPY file.<br>• A folder ‘auxiliary’ with topographic and forest gain information in a NPY format, OpenStreetMap and peat data in a JSON format, and a sub-folder ‘ncep’ containing all data from NCEP in a NPY format.<br>• The forest loss pickle file delimiting the area of forest loss.</p> </li> </ul> <p><strong>Details about the images</strong></p> <ul> <li> <p>For Landsat-8 data (courtesy of the U.S. Geological Survey), this dataset contains 332x 332 pixels RGB calibrated top-of-atmosphere (TOA) reflectance images pan-sharpened to a 15 m resolution (less than 20% cloud cover)</p> </li> <li> <p>For NICFI PlanetScope data (catalog owner: Planet), this dataset contains 332x 332 pixels monthly RGB composite with a 4.77 m resolution</p> </li> </ul> <p><strong>Details about the auxiliary data</strong></p> <ul> <li>Forest gain from GFC: 30-m resolution, yearly data for 2000-2021, downloaded via Google Earth Engine</li> <li>Near infrared, shortwave infrared 1 and 2 bands from Landsat-8 TOA: 30-m resolution, data every 16 days for 2013-2023, downloaded via Google Earth Engine and selected using the same process as for Landsat-8 RGB images</li> <li> From NCEP Climate Forecast System Version 2 (CFSv2) 6-hourly Products: surface level albedo and volumetric soil moisture content (depths: 0.1 m, 0.4 m, 1.0 m, 2.0m) in 0.01%; radiative fluxes (clear-sky longwave flux downward and upward, clear-sky solar flux downward and upward, direct evaporation from bare soil, longwave and shortwave radiation flux downward and upward, latent, ground and sensible heat net flux), potential evaporation rate, and sublimation in W/m²; humidity (specific, maximum specific, minimum specific) in 10-4 kg/kg; ground level precipitation in 0.1 mm; air pressure at surface level in 10 Pa; wind level (u and v component) in 0.01 m/s, water runoff at surface level in 232.01 kg/ m²; temperature in K: 22264-m resolution, available four times a day for 2011-2023, downloaded directly from the NOAA website and selected the mean of the monthly mean over 5 years before the forest loss event, the monthly maximum over 5 years before the forest loss event, and the monthly minimum over 5 years before the forest loss event for each parameter</li> <li>Closest street and closest city from OpenStreetMap in km: directly downloaded with the Nominatim API</li> <li>Altitude in m, slope and aspect in 0.01° from Shuttle Radar Topography Mission (SRTM): 30-m resolution, measured for 2000, downloaded via Google Earth Engine</li> <li>Presence of peat from GFW: 232-m resolution, measured for 2017, directly downloaded on the GFW website</li> </ul> <p><strong>Details about Global Forest Change</strong></p> <p>For each image, there is a corresponding 'forest_loss_region' .pkl file delimiting a forest loss region polygon from Global Forest Change (GFC). GFC consists of annual maps of forest cover loss with a 30-m resolution. </p> <p><strong>License</strong></p> <p>The NICFI PlanetScope images fall under the same license as the <a href="https://assets.planet.com/docs/Planet_ParticipantLicenseAgreement_NICFI.pdf">NICFI data program license agreement</a> (data in 'my_examples_planet_final.zip', 'my_examples_planet_final_detailed.zip', 'my_examples_planet_detailed_timeseries.zip': subfolders '[coordinates]'>'images'>'visible'). </p> <p><a href="https://wiki.osmfoundation.org/wiki/Licence/Attribution_Guidelines">OpenStreetMap®</a> is open data, licensed under the <a href="https://opendatacommons.org/licenses/odbl/1-0/">Open Data Commons Open Database License (ODbL)</a> by the OpenStreetMap Foundation (OSMF) (data in all 'my_examples' folders: subfolders '[coordinates]'>'auxiliary'>'closest_city.json'/'closest_street.json'). The documentation is licensed under the <a href="https://creativecommons.org/licenses/by-sa/2.0/">Creative Commons Attribution-ShareAlike 2.0 license (CC BY-SA 2.0)</a>.</p> <p>The rest of the data is under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>. The data has been transformed following the code that can be found via this link: <a href="https://github.com/aedebus/Cam-ForestNet">https://github.com/aedebus/Cam-ForestNet</a> (in 'prepare_files').</p> <p> </p>
A Taxonomy classifying UI components in desktop interfaces
<p><strong>Taxonomy</strong></p> <p>This publication presents a structured taxonomy for desktop User Interface (UI) Components. The development methodology foundation is extensively detailed in [1]. This taxonomy categorizes UI components into hierarchical levels of abstraction, as described in [2]: <em>System, Application, UI Group, </em>and <em>UI Element</em>.</p> <p>The taxonomy is presented in a table format, organized into the following columns:</p> <ul> <li><strong>UI Component Type</strong>: Identifies the specific UI Component being described, with examples including Icons, Taskbars, etc.</li> <li><strong>Level</strong>: Specifies the component's hierarchical position by identifying the level of abstraction to which it belongs.</li> <li><strong>Detected by</strong>: Lists the model employed to detect a particular type of UI Component, based on a multi-modal approach for detecting UI Hierarchies from desktop screenshots, as discussed in [1], where M1, M2, M3, M4 and M5 are described.</li> <li><strong>Description</strong>: Provides an explanation of the criteria for classifying a specific UI Component type.</li> </ul> <p>This taxonomy facilitates a deeper understanding and navigation of the complexities inherent in desktop UIs.</p> <p><strong>References</strong></p> <p> 1. A. Martínez-Rojas, A. Rodríguez-Ruíz, J.G. Enríquez, and A. Jiménez-Ramírez (2024, September). What’s Behind the Screen? Unveiling UI Hierarchies in Process-Related UI logs. In <em>International Conference on Business Process Management</em>. Cham: Springer International Publishing.</p> <p> 2. Abb, L., Rehse, J.R.: A reference data model for process-related user interaction logs. In: International<br>Conference on Business Process Management. pp. 57–74. Springer (2022)</p>
Features matrices to classify mosquito-associated viruses
<p>Feature matrices associated with the manuscript Predicting novel mosquito-associated viruses from metatranscriptomic dark matter (in press)</p>
HiRISE and MOC-NA Images Used in "Mapping of Western Valles Marineris Light-toned Layered Deposits and Newly Classified Rim Deposits"
<p>These files contain the images used in the paper "Mapping of Western Valles Marineris Light-toned Layered Deposits and Newly Classified Rim Deposits"</p> <p>All images examined, and all images used in the creation of figures are listed within these documents. They are seperated by region, and type of image (HiRISE, MOC-NA)</p>
SHARAD Tracks Mapped in "Mapping of Western Valles Marineris Light-toned Layered Deposits and Newly Classified Rim Deposits"
<p>These three .CSV files contain the SHARAD tracks, as exported by JMARS.</p> <p>This data contains overlaps in each file, as they are sorted by the location in which they were examined, for example, if the SHARAD track crosses all three plateaus examined in the study, that SHARAD track will be listed in all three files. </p> <p>The Radar_Track_LOCATION_and_DETECTIONS.xlxs file contains information on if a basal detection was located in a particular track, and on what plateau the detection was located. </p>
Optimizing Deep Learning Models for Aflatoxin Detection: A Case of Artificial Intelligence-Driven Classified Groundnut Image Datasets for Postharvest Management
<p><strong>DATASET DESCRIPTION </strong><br>This dataset comprises a curated collection of classified groundnut images, specifically designed for deep learning applications in aflatoxin detection. The dataset is organized into four distinct categories: Healthy, Moldy, Insect-Infested, and Physiological Disorder, making it a vital resource for training AI and machine learning models aimed at advancing agricultural research. These classifications are crucial for the development of AI-driven solutions addressing aflatoxin contamination, enhancing crop quality assessments, and improving postharvest management practices.<br>The dataset has been developed to support research in agricultural Artificial Intelligence (AI), machine learning (ML), and food safety, with a focus on aiding resource-constrained regions in combating postharvest losses due to contamination. By leveraging this dataset, researchers can contribute to safeguarding public health, promoting food security, and supporting smallholder farmers.</p> <p><strong>POTENTIAL APPLICATIONS</strong><br>This dataset provides numerous opportunities for innovation in agriculture through AI and deep learning technologies. Its key applications include:<br><strong>Early Aflatoxin Detection</strong>: Facilitates the development of AI-powered models for prompt identification of aflatoxins in groundnuts, helping mitigate associated health risks.<br><strong>Postharvest Management Improvement</strong>: Enables the creation of innovative solutions to enhance storage, handling, and processing, reducing contamination and losses.<br><strong>Food Safety and Quality Assurance</strong>: Strengthens agricultural value chains by supporting the production of safe and high-quality food products.</p> <p><strong>BROADER IMPACT</strong><br>This resource is invaluable for fostering AI innovation in agriculture, particularly in resource-limited environments. It addresses critical challenges such as postharvest losses and food contamination while contributing to global efforts in sustainable agricultural development. By utilizing this dataset, researchers can improve food security, support smallholder farmers, and drive advancements in agricultural practices that benefit both local and global communities.</p>
Tracking and classifying Amazon fire events in near-real time
<p><strong>Summary</strong></p> <p>Time-series (2018-2024) of the Amazon dashboard, including minor updates to the methods.</p> <p>The Amazon dashboard data product tracks individual fire events across most of South America (10N - 25S, 85W - 30W) in near-real time. The model classifies fires into four key fire types (deforestation, forest, small clearing and agricultural, and savanna and grassland fires) and provides estimates of individual fire carbon emissions. Methods are described in Andela et al. (2022). Near-real time estimates are provided at https://amzfire.servirglobal.net/ and here we archive historic time-series.</p> <p><strong>Methods</strong></p> <p>The data archived here (v1.1) include several small updates.</p> <p>Two updates relate to the use of VIIRS active fire detections. First, VIIRS active fire detections have been updated from collection 1 to collection 2. Second, any full day of missing data from either the VIIRS instrument onboard NOAA-20 or Suomi NPP is now replaced by data of the other instrument. This "gap" filling helps reduce the impact of periods with instrument outage, like those of Suomi NPP VIIRS during the 2024 burning season. </p> <p>The other two updates relate to the emissions calculations. First, to convert dry matter burned to carbon emissions, we have introduced fire type specific emissions factors instead of the earlier assumption of 50% carbon content for all fire types. Second, as part of the Sense4Fire project (https://sense4fire.eu/), we provide daily gridded emissions estimates of Dry Matter (DM), C, CO2, CO, and NOx at 0.1 degree resolution. We used emissions factors provided by Andrea (2019) for savanna and grassland fires as well as small clearing and agricultural fires while for forest and deforestation fires we reviewed the literature to select the most relevant emissions factors (Table 1). </p> <p>Table 1: Emissions factors (gram species per kg dry matter burned) used to calculate C, CO2, CO, and NOx emissions. </p> <table> <tbody> <tr> <td>Fire type / trace gas emissions</td> <td>C</td> <td>CO2</td> <td>CO</td> <td>NOx</td> </tr> <tr> <td>Savanna and grassland</td> <td>480</td> <td>1656</td> <td>69.2</td> <td>2.5</td> </tr> <tr> <td>Small clearing and agricultural</td> <td>430</td> <td>1431</td> <td>76.2</td> <td>2.4</td> </tr> <tr> <td>Forest</td> <td>480</td> <td>1561</td> <td>104.0</td> <td>2.0</td> </tr> <tr> <td>Deforestation</td> <td>490</td> <td>1641</td> <td>95.5</td> <td>1.7</td> </tr> </tbody> </table> <p> </p> <p><strong>Dataset description<br></strong></p> <p>For full detail, please see Andela et al. (2022). The tables below (Tables 2 - 4) describe the content of the fire event (polygon) and active fire detections (point) shapefiles as well as the gridded emissions product. The active fire detections and associated estimates of dry matter burned can be combined with emissions factors (Table 1) to derive daily trace gas emissions time series for species and areas of interest.</p> <p>Table 2: Explanation of fire event shapefile attribute table.</p> <table> <tbody> <tr> <td>Attribute class</td> <td>Attribute</td> <td>Explanation / units</td> </tr> <tr> <td>Fire type classification</td> <td>Fire type</td> <td>(1) savanna and grassland, (2) small clearing and<br>agriculture, (3) forest, and (4) deforestation fires</td> </tr> <tr> <td> </td> <td>Confidence</td> <td>(1) low, (2) moderate, and (3) high</td> </tr> <tr> <td>Fire Atlas</td> <td>Size</td> <td>Fire size in km2</td> </tr> <tr> <td> </td> <td>Start day</td> <td>Day of new fire start as day of year (1-366)</td> </tr> <tr> <td> </td> <td>Duration</td> <td>Fire duration in days</td> </tr> <tr> <td> </td> <td>C Emissions</td> <td>Fire carbon emissions (ton C)</td> </tr> <tr> <td>Fire characterization</td> <td>Tree cover</td> <td>Average tree cover fraction within perimeter (%)</td> </tr> <tr> <td> </td> <td>Biomass</td> <td>Average biomass within fire perimeter (ton ha-1)</td> </tr> <tr> <td> </td> <td>Deforestation </td> <td>Fraction of 550 m grid cells with historic<br>deforestation (five years prior to fire) within fire perimeter (%)</td> </tr> <tr> <td> </td> <td>FRP</td> <td>Average fire radiative power (FRP) for all fire<br>detections within fire perimeter (MW)</td> </tr> <tr> <td> </td> <td>Persistence</td> <td>Average fire persistence across 550 m grid cells<br>within fire perimeter (days)</td> </tr> <tr> <td> </td> <td>Progression</td> <td>Average fire progression fraction across 550 m<br>grid cells within perimeter (%)</td> </tr> <tr> <td> </td> <td>Daytime</td> <td>Fraction of 1:30 pm detections (%) for all fire<br>detections within fire perimeter</td> </tr> <tr> <td> </td> <td>Detections</td> <td>Total active fire detections within fire perimeter</td> </tr> </tbody> </table> <p> </p> <p>Table 3: Explanation of active fire detection shapefile attribute table.</p> <table> <tbody> <tr> <td>Attribute class</td> <td>Attribute</td> <td>Explanation / units</td> </tr> <tr> <td>VIIRS active fire detections</td> <td>FRP</td> <td>Fire radiative power (MW)</td> </tr> <tr> <td> </td> <td>DOY</td> <td>Day of year (1-366)</td> </tr> <tr> <td>Fire type classification</td> <td>Fire type</td> <td>(1) savanna and grassland, (2) small clearing and agriculture, (3) forest, and (4) deforestation fires</td> </tr> <tr> <td> </td> <td>Confidence</td> <td>(1) low, (2) moderate, and (3) high</td> </tr> <tr> <td>Emissions</td> <td>C Emissions</td> <td>Fire carbon emissions (ton C) associated with each active fire detection</td> </tr> <tr> <td> </td> <td>DM Emissions</td> <td>Dry matter burned (ton) associated with each active fire detection</td> </tr> </tbody> </table> <p> </p> <p>Table 4: Content of daily gridded (0.1 degree resolution) emissions netcdf files. The daily emissions product provides emissions estimates of dry matter, C, CO2, CO, and NOx. For DM and CO partitioned emissions are also provided by fire type, for other species these can be derived by multiplying the dry matter burned (DM) estimates with trace gas specific emissions factors (Table 1). Values of each grid cell can be multiplied by the number of seconds per day and grid cell area to calculate total emissions (convert "kg species m-2 s-1" to "kg species day-1 per grid cell").</p> <table> <tbody> <tr> <td>/ancill</td> <td>grid_cell_area</td> </tr> <tr> <td>/partitioned_DM_emissions</td> <td>Deforestation emissions</td> </tr> <tr> <td> </td> <td>Forest emissions</td> </tr> <tr> <td> </td> <td>Savanna and grassland emissions</td> </tr> <tr> <td> </td> <td>Small clearing and agricultural emissions</td> </tr> <tr> <td>/partitioned_CO_emissions</td> <td>Deforestation emissions</td> </tr> <tr> <td> </td> <td>Forest emissions</td> </tr> <tr> <td> </td> <td>Savanna and grassland emissions</td> </tr> <tr> <td> </td> <td>Small clearing and agricultural emissions</td> </tr> <tr> <td>/total_emissions</td> <td>DM emissions</td> </tr> <tr> <td> </td> <td>C emissions</td> </tr> <tr> <td> </td> <td>CO2 emissions</td> </tr> <tr> <td> </td> <td>CO emissions</td> </tr> <tr> <td> </td> <td>NOx emissions</td> </tr> </tbody> </table> <p> </p> <p><strong>Results</strong></p> <p>Despite the various small improvements to the dataset, the data are largely consistent with the original dataset published for 2019-2020 (Table 5). </p> <p>Table 5: Comparison of model versions (original from Andela et al., 2022 and v1.1 published here) for April-December 2019 (equator-25S, 85W - 30W). Note that the current version (v1.1) is complete for 2019, but the original dataset had missing data due to incomplete active fire detections from NOAA-20 VIIRS at that time.</p> <table> <tbody> <tr> <td>Dataset</td> <td>Fire type</td> <td>Fire detections (x1,000)</td> <td>Mean fire radiative power (MW)</td> <td>Number of events (x1,000)</td> <td>Emissions (Tg C)</td> </tr> <tr> <td>Original</td> <td>Deforestation</td> <td>756.65</td> <td>15.15</td> <td>24.24</td> <td>99.18</td> </tr> <tr> <td>Original</td> <td>Forest</td> <td>637.58</td> <td>12.73</td> <td>5.28</td> <td>85.46</td> </tr> <tr> <td>Original</td> <td>Small clearing and agricultural</td> <td>348.49</td> <td>10.91</td> <td>154.68</td> <td>10.55</td> </tr> <tr> <td>Original</td> <td>Savanna and grassland</td> <td>1935.06</td> <td>12.11</td> <td>296.42</td> <td>71.75</td> </tr> <tr> <td>v1.1</td> <td>Deforestation</td> <td>742.64</td> <td>14.81</td> <td>24.02</td> <td>97.18</td> </tr> <tr> <td>v1.1</td> <td>Forest</td> <td>626.92</td> <td>12.06</td> <td>5.16</td> <td>77.84</td> </tr> <tr> <td>v1.1</td> <td>Small clearing and agricultural</td> <td>350.7</td> <td>10.89</td> <td>155.56</td> <td>9.27</td> </tr> <tr> <td>v1.1</td> <td>Savanna and grassland</td> <td>1877.16</td> <td>11.89</td> <td>299.17</td> <td>70.55</td> </tr> </tbody> </table> <p> </p> <p><strong>Acknowledgements</strong></p> <p>The Sense4Fire project is funded by ESA under ESA Contract Number: 4000134840/21/I-NB. </p> <p><strong>References</strong></p> <p>Andela, N., Morton, D.C., Schroeder, W., Chen, Y., Brando, P.M. and Randerson, J.T., 2022. Tracking and classifying Amazon fire events in near real time. Science advances, 8, eabd2713. https://doi.org/10.1126/sciadv.abd2713.</p> <p>Andreae, M.O., 2019. Emission of trace gases and aerosols from biomass burning–an updated assessment. Atmospheric Chemistry and Physics, 19, 8523-8546. https://doi.org/10.5194/acp-19-8523-2019.</p>
Heron Island Satellite Imagery Classified Benthic Data and Halo Analyses and Models
<p>This dataset includes satellite imagery data from Heron Island, Australia downloaded from Google Earth Pro in 2023, with image from Maxar Technologies dated 2016 clipped to the shallow lagoon layer from the Allen Coral Atlas shape file classified into benthic categories: corals, algae and sand using a combination of unsupervised machine learning spectral classification and manual training and assignment of classes. This dataset also includes scoring of selected coral patch reefs for isolated halos across time using historical aerial imagery.</p> <p>We also include 2 notebooks with code used to generate figures and run analyses for data, geometric, and consumer-resource models for coral halo patterns supporting the work entitled, "Consumer-resource interactions reflected in coral halo patterns" by the authors listed. A knitted html for R Markdown file is also included.</p>
A Diagnostic Classifier for Gene Expression-Based Identification of Early Lyme Disease
<p>This data repository contains the source code and pertinent training and test data sets used in the manuscript, <em>A Diagnostic Classifier for Gene Expression-Based Identification of Early Lyme Disease</em>. </p>
Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies"
<p>This supplementary material for the article"Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies. Dell’Anna, D.; Aydemir, F. B.; and Dalpiaz, F. Empirical Software Engineering. 2022" includes</p> <ul> <li>ECSER-ExploratoryStudy.csv: The annotated meta-data of the papers that have been published in ICSE between 2019 and 2021.</li> <li>ECSER_ROCplots+StatTest.ipynb: A python notebook that compares classifiers adn checks the statistical significance of the comparison results.</li> <li>ECSER_SummaryOfReplicationSteps.pdf: This table presents a summary of ECSER steps for the two original studies and our applications on ECSER.</li> <li>ECSER_RE: The directory that holds the datasets and code for the replication of Hay et al. [1] and additional runs on the new data sets.</li> <li>ECSER_FF: The directory that holds the code and data for the replication of Alshammari et al. [2]</li> <li>README.md presents the structure of the supplementary materials.</li> <li>requirements.txt lists the dependencies needed to run the code</li> </ul> <p>In the ECSER_RE directory, the code for multiple classifiers that are compared are kept in the "Classifiers" directory. The public data sets are shared in the Datasets directory. "ECSER_RE_Compare_Classifiers.ipynb" python notebook includes the code that runs each classifier. The results are presented in "ECSER_RE_results-Promise-vs-all.csv".</p> <p>In the ECSER_FF directory, the data sets are presented directly under the main directory. The notebook "ECSER-FF-Compare_Classifiers.ipynb" compares the classifiers of the original study and the results are kept under the "ECSER_FF_results" directory.</p> <p><strong>How to cite this repository </strong><br>If you use this repository, please cite the reference paper, and the repository, as below:</p> <p>Dell’Anna, Davide, Fatma Başak Aydemir, and Fabiano Dalpiaz. "Evaluating classifiers in SE research: the ECSER pipeline and two replication studies." Empirical Software Engineering 28.1 (2023): 3.</p> <p>Davide Dell'Anna, Fatma Başak Aydemir, & Fabiano Dalpiaz. (2021). Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies" [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6266675</p> <p> </p> <p> </p> <p> </p> <ol> <li>Tobias Hey, Jan Keim, Anne Koziolek, and Walter F. Tichy. 2020. SupplementaryMaterial of "NoRBERT: Transfer Learning for Requirements Classification". https://doi.org/10.5281/zenodo.3874137</li> <li>Abdulrahman Alshammari, Christopher Morris, Michael Hilton, and JonathanBell. 2021. FlakeFlagger: Predicting Flakiness Without Rerunning Tests. In43rdIEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid,Spain, 22-30 May 2021. IEEE, 1572–1584. https://doi.org/10.1109/ICSE43902.2021.00140</li> </ol>
Dataset used for training IoT C&C classifier
<p>This dataset was used for training the IoT C&C classifier. It is provided in the form of extended bidirectional flow data. The flow data were generated by <a href="https://github.com/CESNET/ipfixprobe">ipfixprobe</a> flow exporter and converted into CSV files. Apart from traditional flow information (IP addresses, ports, amount of transferred data), ipfixprobe was set with default timeouts (5 minutes active, 30 s inactive) to generate per-packet information for the first 30 packets. The flow records were then aggregated into 5-minute intervals - when the flow was split due to inactivity, the aggregator then stitched the flow back into a single one.</p> <p>The column headers in provided CSV files stand for:</p> <table> <thead> <tr> <th>Column Name</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td>ipaddr DST_IP</td> <td>Source IP address</td> </tr> <tr> <td>ipaddr SRC_IP</td> <td>Destination IP address</td> </tr> <tr> <td>uint64 BYTES</td> <td>The number of transmitted bytes from SRC->DST</td> </tr> <tr> <td>uint64 BYTES_REV</td> <td>The number of transmitted bytes from DST->SRC</td> </tr> <tr> <td>time TIME_FIRST</td> <td>Timestamp of the first packet in the flow in format YYYY-MM-DDTHH-MM-SS</td> </tr> <tr> <td>time TIME_LAST</td> <td>Timestamp of the last packet in the flow in format YYYY-MM-DDTHH-MM-SS</td> </tr> <tr> <td>macaddr DST_MAC</td> <td>Destination MAC address</td> </tr> <tr> <td>macaddr SRC_MAC</td> <td>Source MAC address</td> </tr> <tr> <td>uint32 COUNT</td> <td>Number of aggregated flow records</td> </tr> <tr> <td>uint32 PACKETS</td> <td>The number of packets transmitted from Source to Destination</td> </tr> <tr> <td>uint32 PACKETS_REV</td> <td>The number of packets transmitted from Destination to Source</td> </tr> <tr> <td>uint16 DST_PORT</td> <td>Destination port</td> </tr> <tr> <td>uint16 SRC_PORT</td> <td>Source port</td> </tr> <tr> <td>uint8 DIR_BIT_FIELD</td> <td>Flag for distinguishin WAN(1)/LAN(0)</td> </tr> <tr> <td>uint8 PROTOCOL</td> <td>The number of transport protocol</td> </tr> <tr> <td>uint8 TCP_FLAGS</td> <td>Logic OR across all TCP flags in the packets transmitted SRC->DST</td> </tr> <tr> <td>uint8 TCP_FLAGS_REV</td> <td>Logic OR across all TCP flags in the packets transmitted DST->SRC</td> </tr> <tr> <td>int8* PPI_PKT_DIRECTIONS</td> <td>Array with packets' direction (1)- SRC->DST, (-1)-DST->SRC</td> </tr> <tr> <td>uint8* PPI_PKT_FLAGS</td> <td>Array with packets' TCP flags</td> </tr> <tr> <td>uint16* PPI_PKT_LENGTHS</td> <td>Array with packets' payload lengths</td> </tr> <tr> <td>time* PPI_PKT_TIMES</td> <td>Array with packets' timestamps</td> </tr> </tbody> </table> <p>Dataset consists of two parts: a benign part captured on the real ISP network and a malicious part captured in a lab environment.</p> <p><strong>Bening part captured on the real ISP network</strong><br> This part was created by packet capturing on the metering points located at the perimeter of the CESNET2 network. The metering points monitor 100 Gbps backbone peering lines used by approximately half a million users. We performed packet filtering based on ports for the capture. The CESNET training capture was used as benign traffic in the C&C model training and testing pipeline to cover potential nuances and variability of benign data seen in the ISP-level network. Since we deal with data from the production network,<br> we cannot guarantee a benign nature of all captured communication. However, we verified every IP address according to the internal blocklist of the CESNET association and external ones. We used <a href="https://www.abuseipdb.com">AbuseIPDB</a> and <a href="https://urlhaus.abuse.ch/)">URLhaus</a> blocklists.</p> <p>Since we are dealing with the real captures, the IP addresses, and MAC addresses<br> were anonymized.</p> <p><br> <strong>Malicious part created in the controlled lab-created environment</strong><br> From leaked source codes, we picked one variant from each of the most prevalent client-server IoT botnet families: (1) Tsunami, (2) Gafgyt, (3) Mirai. Each implements a distinct communication protocol; Tsunami is an example of an IRC bot; Gafgyt<br> uses a simple text-based protocol; Mirai implements a custom binary protocol. Afterward, we prepared virtualized testing environment.</p> <p>We deployed the malware in a controlled manner, filtering out its scanning and exploiting activities. The dataset covers the most notable C&C behavior. As previously recognized, the C&C communication consists of C&C heartbeat and<br> bot commands. Thus, for each of the three prepared malware variants, we first imagine the malware running with no received commands. That includes the initiation of the TCP connection to the C&C server, which continues for one hour. And then, we imagine the malware receiving commands from its C&C server. The position of the command packets is chosen arbitrarily relative to the background heartbeat packets because, in the real-world scenario, the timing of the commands is tied to a random human action.</p> <p><br> <strong>Directory tree of provided dataset</strong><br> </p> <pre><code>. ├── README.md ├── benign │ ├── AN_p20-21-25-143-3389.agg.head.csv │ ├── AN_p22.agg.head.csv │ ├── AN_p443.agg.head.csv │ ├── AN_p80.agg.head.csv │ └── AN_p8080.agg.head.csv └── cnc ├── kaiten │ ├── cnc.csv │ ├── command-01.csv │ ├── command-02.csv │ ├── command-03.csv │ ├── command-04.csv │ ├── command-05.csv │ ├── command-06.csv │ ├── command-07.csv │ └── command-08.csv ├── mirai │ ├── cnc.csv │ ├── command-01.csv │ ├── command-02.csv │ ├── command-03.csv │ ├── command-04.csv │ ├── command-05.csv │ ├── command-06.csv │ ├── command-07.csv │ └── command-08.csv └── qbot ├── cnc.csv ├── command-01.csv ├── command-02.csv ├── command-03.csv └── command-04.csv </code></pre> <p><strong>Acknowledgment</strong><br> This research was funded by the Ministry of Interior of the Czech Republic,<br> grant No. VJ02010024: Flow-Based Encrypted Traffic Analysis and also by the<br> Grant Agency of the CTU in Prague, grant No. SGS20/210/OHK3/3T/18 funded by<br> the MEYS of the Czech Republic.</p> <p> </p>
Processed data for MethylBoostER: an XGBoost model to classify kidney cancer subtypes
<p>This is a repository containing processed data for MethylBoostER, an XGBoost model that classifies kidney cancer subtypes. The open-source code can be found here: https://github.com/ss-lab-cancerunit/MethylBoostER.</p>
Understanding challenges of GPU programming by classifying and analyzing Stack Overflow posts
<p>This dataset includes a dataset of posts related to GPU programming and supplemental materials including the complete analyzed results of our paper "Understanding challenges of GPU programming by classifying and analyzing Stack Overflow posts".</p>
Planar Graph Classifier - Class Information
<p>CSV files associating image files from the Planar Graph Classifier - Generated Graph Images (<a href="https://doi.org/10.5281/zenodo.6510092">10.5281/zenodo.6510092</a>) with planarity information.</p> <p>Class 0: graph drawn with networkx spring layout<br> Class 1: graph drawn with networkx planar layout </p>
Planar Graph Classifier - Set Split
<p>Set with class information (Planar Graph Classifier - Class Information - 10.5281/zenodo.6510113) is split into training and test set for model fitting and evaluation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.