Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “Classification algorithm”
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
Experiment on the performance of different machine learning algorithms for classification - Results
<h2>Results of a short performance study of machine learning algorithms</h2> <h3>Context and methodology</h3> <ul> <li>This data was produced while performing a university project to examine the performance of various machine learning algorithms on different prediction datasets</li> <li>The data serves the purpose of comparing the metrics of performing the different tasks</li> <li>The dataset contains a number of matrices for every classifier and every dataset</li> <li>The data was produced with python scripts provided further down and with the usage of the external datasets: <ul> <li>Membership Woes Dataset (OpenML): <a href="https://api.openml.org/d/44224">https://api.openml.org/d/44224</a></li> <li>Zoo dataset (UCI): <a href="https://doi.org/10.24432/C5R59V">https://doi.org/10.24432/C5R59V</a></li> <li>Breast Cancer Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> <li>Loan Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> </ul> </li> </ul> <h3>Technical details</h3> <ul> <li>The data consists of one JSON file</li> <li>The source code for producing this data is available at <a href="https://doi.org/10.5281/zenodo.11085222">https://doi.org/10.5281/zenodo.11085222</a></li> </ul> <h3>Structure of the data</h3> <p>[ {"classifier": ...,<br>"dataset": ...,<br>"hyper_parameters": ...,<br>"cross_validation_results": {<br> "fit_time": {} ,<br> "score_time": ...,<br> "metrics": {}<br>}, <br>"holdout_test_results": ...}, ]</p>
Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm
<p>Dataset from the results of data crawling via Twitter which discusses the Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>
A Point Cloud Dataset of Vehicles Passing Through a Toll Station for use in Training Classification Algorithms
<p>This work presents a point cloud dataset of vehicles passing through a toll station in Colombia to be used to train artificial vision and computational intelligence algorithms. This article details the process of creating the dataset, covering initial data acquisition, range information preprocessing, point cloud validation, and vehicle labeling. Additionally, a detailed description of the structure and content of the dataset is provided, along with some potential applications of its use. The dataset consists of 36,026 total object classes: 31,432 cars, campers, vans and 2-axle trucks with a single tire on the rear axle, 452 minibuses with a single tire on the rear axle, 1158 buses, 1179 2-axle small trucks, 797 2-axle large trucks, and 1008 trucks with 3 or more axles. The point clouds were captured using a LiDAR sensor and Doppler effect speed sensors. The dataset can be used to train and evaluate algorithms for range data processing, vehicle classification, vehicle counting, and traffic flow analysis. The dataset can also be used to develop new applications for intelligent transportation systems.</p> <table> <tbody> <tr> <td>Type</td> <td>Description</td> <td>Quantity</td> </tr> <tr> <td>1</td> <td>Cars, campers, vans and 2-axle trucks with<br>a single tire on the rear axle</td> <td>31,432</td> </tr> <tr> <td>2</td> <td>Minibuses with a single tire on the rear axle</td> <td>452</td> </tr> <tr> <td>3</td> <td>Buses</td> <td>1,158</td> </tr> <tr> <td>4</td> <td>Trucks with 3 or more axles</td> <td>1,008</td> </tr> <tr> <td>5</td> <td>2-axle small trucks</td> <td>1,179</td> </tr> <tr> <td>6</td> <td>2-axle large truck</td> <td>797</td> </tr> <tr> <td>Total</td> <td> </td> <td>36,026</td> </tr> </tbody> </table>
Figure 19. Dynamics of the filtered images with the 9 algorithms and the 2 methods of classification-Efficient Filtering of Noisy Fingerprint Images
<p>The classification (Malik, Gautam, Sahai, Jha & Singh, 2013) and ranking stage can be visualized in the Figure 19, the summary of the filtered images is shown in Table 3 and the pseudocode of the current step can be visualized in Figure 18. The overall results show that the two selection criterion: fuzzy and aggregation indicate that the most efficient algorithm is A6 and according to each criterion there can be made certain decisions to choose the best filters for each situation. Also, the results are influenced by the parameters set to calibrate the filtering, the fuzzy profiles, the weighted sum or the vicinity approach.</p>
Investigation of machine learning algorithms for taxonomic classification of marine metagenomes
<p>Training, testing, and blind datasets used for machine learning algorithms for taxonomic classification of marine metagenomes:</p> <ol> <li><strong>K12.kmers.txt</strong> - 12bp k-mer vocabulary constructed by Jellyfish v1.1.11 from 47,894 genomes in GTDB release 202</li> <li><strong>MarRef_1.6.tsv</strong> - Metadata file downloaded from MarRef v1.6</li> <li><strong>MarRef.genustrain.fasta</strong> - Training set from MarRef v1.6 (seed=808) used for genus classification</li> <li><strong>MarRef.genustest.fasta</strong> - Testing set from MarRef v1.6 (seed=747) used for genus classification </li> <li><strong>MarRef.speciestrain.fasta</strong> - Training set from MarRef v1.6 (seed=808) used for species classification</li> <li><strong>MarRef.speciestest.fasta</strong> - Testing set from MarRef v1.6 (seed=747) used for species classification</li> <li><strong>MarRef.traintest.key.tsv</strong> - Table containing MarRef accession, GenBank accession, GenBank taxonomy ID, taxonomic information, and labels used for species and genus testing and training</li> <li><strong>anonymous_reads_*.fq</strong> - Blind datasets (1-10) in interleaved fastq format</li> <li><strong>reads_mapping_*.tsv</strong> - Key for blind datasets 1-10. Each sequence header is mapped to its corresponding MarRef accession and NCBI taxonomic ID.</li> </ol>
Supporting datasets PubFig05 for: "Heterogeneous Ensemble Combination Search using Genetic Algorithm for Class Imbalanced Data Classification"
<p><strong>Faces Dataset: PubFig05</strong></p> <p>This is a subset of the ''PubFig83'' dataset [1] which provides 100 images each of 5 most difficult celebrities to recognise (referred as class in the classification problem). For each celebrity persons, we took 100 images and separated them into training and testing sets of 90 and 10 images, respectively:</p> <p><strong>Person: </strong>Jenifer Lopez; Katherine Heigl; Scarlett Johansson; Mariah Carey; Jessica Alba</p> <p> </p> <p><strong>Feature Extraction</strong></p> <p>To extract features from images, we have applied the HT-L3-model as described in [2] and obtained 25600 features.</p> <p><strong>Feature Selection</strong></p> <p>Details about feature selection followed in brief as follows:</p> <ol> <li> <p><strong>Entropy Filtering:</strong> First we apply an implementation of Fayyad and Irani's [3] entropy base heuristic to discretise the dataset and discarded features using the minimum description length (MDL) principle and only 4878 passed this entropy based filtering method.</p> </li> <li> <p><strong>Class-Distribution Balancing:</strong> Next, we have converted the dataset to binary-class problem by separating into 5 binary-class datasets using one-vs-all setup. Hence, these datasets became <em>imbalanced</em> at a ratio of 1:4. Then we converted them into <em>balanced binary-class</em> datasets using random sub-sampled method. Further processing of the dataset has been described in the paper.</p> </li> <li> <p><strong>(alpha,beta)-k Feature selection:</strong> To get a good feature set for training the classifier, we select the features using the approach based on the (alpha,beta)-k feature selection [4] problem. It selects a minimum subset of features that maximise both within class similarity and dissimilarity in different classes. We applied the entropy filtering and (alpha,beta)-k feature subset selection methods in three ways and obtained different numbers of features (in the Table below) after consolidating them into binary class dataset.</p> </li> </ol> <ul> <li> <p><strong>UAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>union</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>IAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>intersection</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>UEAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets. Then, we applied the entropy filtering and (alpha,beta)-k feature set selection method on each of the balanced binary-class datasets. Finally, we took the <em>union</em> of selected features for each <em>balanced binary-class</em> datasets and get a set of features.</p> </li> </ul> <p>All of these datasets are inside the compressed folder. It also contains the document describing the process detail.</p> <p> </p> <p><strong>References</strong></p> <p>[1] Pinto, N., Stone, Z., Zickler, T., & Cox, D. (2011). Scaling up biologically-inspired computer vision: A case study in unconstrained face recognition on facebook. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2011 IEEE Computer Society Conference on (pp. 35–42).</p> <p>[2] Cox, D., & Pinto, N. (2011). Beyond simple features: A large-scale feature search approach to unconstrained face recognition. In Automatic Face Gesture Recognition and Workshops (FG 2011), 2011 IEEE International Conference on (pp. 8–15).</p> <p>[3] Fayyad, U. M., & Irani, K. B. (1993). Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning. In International Joint Conference on Artificial Intelligence (pp. 1022–1029).</p> <p>[4] Berretta, R., Mendes, A., & Moscato, P. (2005). Integer programming models and algorithms for molecular classification of cancer from microarray data. In Proceedings of the Twenty-eighth Australasian conference on Computer Science - Volume 38 (pp. 361–370). 1082201: Australian Computer Society, Inc.</p> <p> </p>
The fully executable procedure of the U-Net model combined with the Multi-textRG algorithm to achieve fine ice-water classification ---- another 332 scenes of data-fused SIC labels.
<p>This data source is related to the manuscript titled "Combining the U-Net model and a Multi-textRG algorithm for fine SAR ice-water classification", which will be submitted to the journal---The Cryosphere. </p> <ul> <li>The"ready-to-train-fused_01.zip" to "ready-to-train-fused_10.zip" includes 200 scenes of data-fused SIC labels accessible with doi: 10.5281/zenodo.10973107, https://zenodo.org/records/10973107. </li> <li>The "ready-to-train-fused_11.zip" to "ready-to-train-fused_21.zip" includes another 332 scenes of data-fused SIC labels. </li> </ul>
Classification of Eye Images by Personal Details With Transfer Learning Algorithms
<p>During the data collection phase of the research, first of all, a brief information was given to the participants about the study, and how the data would be used and what to do. Photographs of the eye area were collected from participants consisting of a total of 96 different people aged between 3-64. It has been clearly stated that there will be no situations that will define them during the photo shoot. Then, at least ten images of the right eye area of each person were taken. In addition, at least ten photographs of the left eye area were taken. Along with these photographs, no data other than the age and gender of the persons was recorded. Below are images of two people of different genders.</p> <p>A total of 1980 images were obtained from the participants, as in the figure above. More than ten images were obtained from some people. For this reason, there is a difference in the number of photos of people. Care has been taken to use different angles and lights so that each photograph does not form the same frame. Thus, photographs that were not, all the same, were collected. In order for each photograph not to be confused with another photograph, a naming rule has been developed to express the person, age, gender and the number of the photograph taken. An underscore ("_") character is inserted between each expression. Each expression used in the naming convention is given below in order.</p> <ul> <li>Person ID: It is a unique code value for each person photographed. This value ranges from 1 to 100.</li> <li>Age: The age is written directly as a number to express how old the person is. This value varies between 3-64.</li> <li>Gender ID: The value of 1 is expressed if the person photographed is male, and the value of 0 if it is a woman.</li> <li>Photo ID: Due to the fact that more than one photo was taken for each person, each photo was numbered sequentially from 1-10.</li> </ul> <p>If this dataset is used, reference should be made to the article below.</p> <ul> <li>Aktürk, C., Aydemir, E., Hama Rashid, Y. M. 2022. Classification of eye images according to person details with the transfer learning algorithms. Acta Informatica Pragensia, DOI: 10.18267/j.aip.190</li> </ul>
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
<p><span>Kill rates are a central parameter to assess the impact of predation on prey species. An accurate estimation of kill rates requires correct identification of kill sites, often achieved by field-checking GPS location clusters (GLCs). However, there are potential sources of error included in kill site identification, such as failing to detect GLCs that are kill sites and misclassifying the generated GLCs (e.g. kill for non-kill) that were not field-checked. Here, we address these two sources of error using a large GPS dataset of collared Eurasian lynx, an apex predator of conservation concern in Europe, in three multi-prey systems, with different combinations of wild, semi-domestic, and domestic prey. We first used a subsampling approach to investigate how different GPS-fix schedules affect the detection of GLCs indicating kill sites. Then, we evaluated the potential of the random forest algorithm to classify GLCs as non-kills, small prey kills, and ungulate kills. We show that the number of fixes can be reduced to from 7 to 3 fixes/night without missing more than 5% of the ungulate kills, in a system composed of wild prey. Reducing the number of fixes per 24-h decreased the probability of detecting GLCs connected with kill sites, particularly those of semi-domestic or domestic prey, and small prey. Random forest successfully predicted between 73%-90% of ungulate kills but failed to classify most small prey in all systems, with sensitivity (true positive rate) lower than 65%. Additionally, removing domestic prey improved the algorithm's overall accuracy. We provide a set of recommendations for studies focusing on kill site detection, which can be considered for other large carnivore species besides the Eurasian lynx. We recommend caution when working in systems including domestic prey, as the odds of underestimating kill rates are higher.</span></p>
Generating a Labeled Dataset to Train Machine Learning Algorithms for Lithological Classification of Drill Cuttings
<p>This dataset contains 16,700 fully labeled SEM images of rock chips isolated from 14 thin sections of drill cutting samples. These samples come from a low-permeability reservoir in western Canada.</p>
Out-of-distribution detection algorithms for robust insect classification dataset and models
<p>This folder contains trained models and datasets for reproducing the results in the paper on out-of-distribution detection algorithms for robust insect classification. Specifically, it contains the following folders: </p> <p> </p> <ul> <li>OODInsect (out-of-distribution data)</li> <li>MSP, MAH, and EBM trained models, each wrapped around the three classifiers of ResNet50, RegNet32, and VGG11, and different combinations of ID and OOD test data for reproducing RQ1, RQ2, and RQ3.</li> <li>ID3 (in-distribution test data)</li> </ul>
Data and code for "Des-q: a quantum algorithm to construct and efficiently retrain decision trees for regression and binary classification"
<p>It contains the code and the data to reproduce the figures in the paper "Des-q: a quantum algorithm to construct and efficiently retrain decision trees for regression and binary classification" published in arXiv: https://arxiv.org/abs/2309.09976</p>
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
Open the record for dataset details and reuse information.
Data from: Early detection of encroaching woody Juniperus virginiana and its classification in multi-species forest using UAS imagery and semantic segmentation algorithms
Open the record for dataset details and reuse information.
Illustration of IMTs Behavior in Recurrent Expansion Algorithm for Complex Classification Problems
<p>This video provides a visual representation of the Inputs Mappings and estimated Targets (IMTs) behavior within the Recurrent Expansion Algorithm when addressing highly complex classification problems. Specific details about this experiment and the dataset used can be downloaded from [1]. The video serves as a valuable illustrative resource for understanding the algorithm's performance and its approach to complex classification challenges. It also serves as supplementary material aiding in the understanding of Figure 9 from [1]. For more detailed information, please refer to the accompanying references. Please cite our paper.</p> <p>[1] Berghout, T. & Benbouzid, M. (2024). Empirical Analysis of Aeroengine Inter-Shaft Bearing Faults: Multiverse Recurrent Expansion and Data Quality Enhancement Strategies. 1–21. https://doi.org/http://dx.doi.org/10.2139/ssrn.4833247</p>
K-Nearest-Neighbor algorithm to predict the survival time and classification of various stages of Oral Cancer: A machine learning approach
<p>This project predicts the survival time of a cancer patient in terms of the number of days and also classifies the dataset into various stages of cancer</p> <p>This is executed on the SPYDER platform using Python 3.7 on Anaconda Navigator.</p> <p>The dataset includes the oral cancer patient's record of 4 countries</p> <p> </p>
Implementation of K-Nearest Neighbor Algorithm and Gray Level Co-Occurance Matrix Method in Mushroom Type Classification
<p>This material has presented on 2nd International Conference on Advanced Research in Engineering and Technology in October 25, 2023.</p>
Comparison and assessment of different object-based classifications using machine learning algorithms and UAVs multispectral imagery in the framework of precision agriculture
<p>Supplementary material of the paper</p>
Algorithm performance comparison in a classification task
<p>Results of algorithm performance comparison experiment.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.