Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
106
datasets available to search
ShareScore release 0.7.1
Dataset results
106 results for “Classification model”
Dataset for training the Surrogate Model of microlaser neurons on the reduced MNIST classification task
<p>This dataset was used to train a surrogate multilayer perceptron surrogate model of microlaser neurons.</p> <p>It is in csv format. It was generated using the Yamada Model as found in </p> <p><span>Selmi F, Braive R, Beaudoin G, Sagnes I, Kuszelewicz R and Barbay S 2014 Relative Refractory Period in an Excitable Semiconductor Laser <em>Phys. Rev. Lett.</em> <strong>112</strong> 183902</span>.</p>
Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"
<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>
Explainable AI for unveiling deep learning pollen classification model - Pollen dataset
<p>Dataset consists automatic particle detector Rapid-E measurements of pollen grains from 12 classes: Acer, Alnus, Alopecurus, Carex, Cupressus, Dactylis, Juglans, Morus, Platanus, Populus, Salix and Ulmus. Data are available i json format.</p> <p>Dataset also contains preprocessed data packed into csv files of 3 modalities: spectrum, lifetime, scattering and additional features from scattering and lifetime data are also available. These are ready to be used with machine learning models. Labels 0, 1, 2, ... 11 correspond to alphabetical order of examined pollen classes Acer, Alnus, Alopecurus ... Ulmus.</p> <p> </p>
A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model
<p>A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model</p>
Broad-Coverage German Sentiment Classification Model and Dataset for Dialog Systems
<p><a href="http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.202.pdf"><strong>Training a Broad-Coverage German Sentiment Classification Model for Dialog Systems</strong></a></p> <p>This paper describes the training of a general-purpose German sentiment classification model. Sentiment classification is an important aspect of general text analytics. Furthermore, it plays a vital role in dialogue systems and voice interfaces that depend on the ability of the system to pick up and understand emotional signals from user utterances. The presented study outlines how we have collected a new German sentiment corpus and then combined this corpus with existing resources to train a broad-coverage German sentiment model. The resulting data set contains 5.4 million labelled samples. We have used the data to train both, a simple convolutional and a transformer-based classification model and compared the results achieved on various training configurations. The model and the data set will be published along with this paper.</p> <p>You can find the code for training testing the models, that was published along with the paper in this <a href="https://github.com/oliverguhr/german-sentiment">repository</a>.</p> <p>The <a href="https://github.com/oliverguhr/german-sentiment-lib"><em>germansentiment</em></a> Python package contains a easy to use interface for the model that was published with this paper.</p> <p> </p> <p> </p>
Convolutional Neural Net (CNN) models for epigenomic landscapes in epidermal differentiation - Basset architecture, classification and regression
<p>Deep learning models trained on epigenomic landscapes in keratinocyte differentiation. The models are Basset convolutional neural networks (Kelley, et al 2016). The dataset used to train these models can be found at https://doi.org/10.5281/zenodo.4062509. The file `nn.ggr.models.basset.clf.tar.gz` contains 10 cross-validated models that were pretrained using ENCODE-Roadmap trained model weights as initialization weights and also 10 cross-validated models that were initialized with random weights. Similarly, the file `nn.ggr.models.basset.regr.tar.gz` contains 10 cross-validated models that were pretrained using the classification model weights as initialization weights and also 10 cross-validated models that were initialized with random weights.</p>
Training CNNs with Low-Rank Filters for Efficient Image Classification: Trained Models
<p>Models from experiments referenced in the paper "Training CNNs with Low-Rank Filters for Efficient Image Classification", https://arxiv.org/abs/1511.06744</p> <p>Model names differ from those in the paper, but the csv files for each set of experiments relates the paper's name for the model and the real name of the model here:</p> <ul> <li>cifarma.csv: Network-in-Network CIFAR10 Models</li> <li>mitma.csv: MIT Places Models</li> <li>googlenetma.csv: GoogLeNet ILSVRC2012 Models</li> <li>vggma.csv: VGG-11 ILSVRC2012 Models</li> </ul> <p> </p> <p> </p>
Dataset used in the publication "Using of Transformers Models for Text Classification to Mobile Educational Applications"
<p>Dataset used in the publication "Using of Transformers Models for Text Classification to Mobile Educational Applications".</p> <p>More info about the dataset can be found in the published article.</p>
Train and Evaluation Code, Road Classification Models and Test set of the paper "Impact of Image Resolution and Image Overlap on the Prediction Performance of Convolutional Neural Networks Trained for Road Classification"
<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road classification models corresponding to the paper "Impact of Image Resolution and Image Overlap on the Prediction Performance of Convolutional Neural Networks Trained for Road Classification". The scripts make use of the Tensorflow with Keras framework and the additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (https://zenodo.org/records/6482346) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 546 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area of 28.5 km * 18.5 km and features binary road labels. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>
A tempοral Deep Convolutional Neural Network model on Sentinel-1 Image Time Series for pixel-wise Flood Classification (dataset)
<p>This is a dataset which has been designed to be used for flood time series classification. Each time series is annotated as flood or no-flood and represents a pixel-wise time series derived from stack of Sentinel-1 IW GRD images that have been pre-processed according to <a href="http://doi.org/10.5281/zenodo.6510223">https://doi.org/10.5281/zenodo.6510223</a>.</p>
4-way Tabla Stroke Classification with Models Adapted from ADT
<p>Recordings and 4-way stroke category annotations of tabla playing of solo compositions and accompaniment to vocals (tabla recorded in isolation) released with the following conference paper:</p> <p>M. A. Rohit, A. Bhattacharjee, and P. Rao, “Four-way Classification of Tabla Strokes with Models Adapted from Automatic Drum Transcription”, in Proc. of the 22nd Int. Society for Music Information Retrieval Conf., Online, 2021.</p> <p>The dataset is split into test and train sets. The test set consists of 10 pieces of only the tabla accompaniment recorded in perfect isolation to prerecorded solo Hindustani vocal tracks. It contains 20 minutes of audio and nearly 4,500 strokes. These recordings, made on 3 unique tabla sets by 2 different artists, are diverse in terms of tuning, tala (metre), and tempo. The training set consists of solo compositions and common theka patterns recorded from 10 different tabla-sets. The total audio duration is about 1.25 hours and there are 26,600 strokes.</p> <p>The test set was annotated by first running an automatic onset detector to obtain stroke onsets, followed by manually assigning the four-way labels by listening to the audio and visually inspecting the spectrogram. The training set was annotated by automatically aligning the composition score (supplied by artists) with the audios, and replacing the bols with corresponding target stroke categories. Given the imperfect score-stroke matching, labels were manually verified to assign the same category to similar sounding bols.</p> <p>Onsets are separated into folders for each stroke category. Each '.onsets' file in these folders corresponds to an audio ('.wav') file of the same name and is a text file containing a list of time instants where an onset of a particular stroke category occurs.</p>
Index based dataset for training ML classification models
<p>This dataset contains 202122 rows of data containing 61 unique indices from different world urban areas.</p>
OpenAlex Topic Classification v1 Model Artifacts and Training Data
<p>This is all data used to train the topic classification model and also the model artifacts to deploy the model. Please see the github repo for more information:</p> <p>https://github.com/ourresearch/openalex-topic-classification</p>
Train and Evaluation Code, Road Classification Models and Test set of the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography"
<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road segmentation models corresponding to the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography". The scripts make use of the Tensorflow with Keras framework and their additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (<a href="../records/6482346">https://zenodo.org/records/6482346</a>) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 492 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area from Palencia (Spain) and features 18 million pixels labelled with the positive "Road" class. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>
Data from: Performance of unmarked abundance models with data from machine-learning classification of passive acoustic recordings
<p>The ability to conduct cost-effective wildlife monitoring at scale is rapidly increasing due to availability of inexpensive autonomous recording units (ARUs) and automated species recognition, presenting a variety of advantages over human-based surveys. However, estimating abundance with such data collection techniques remains challenging because most abundance models require data that are difficult for low-cost monoaural ARUs to gather (e.g., counts of individuals, distance to individuals), especially when using the output of automated species recognition. Statistical models that do not require counting or measuring distances to target individuals in combination with low-cost ARUs provide a promising way of obtaining abundance estimates for large-scale wildlife monitoring projects but remain untested. We present a case study using avian field data collected in forests of Pennsylvania during the Spring of 2020 and 2021 using both traditional point counts and passive acoustic monitoring at the same locations. We tested the ability of the Royle-Nichols and time-to-detection models to estimate abundance of two species from detection histories generated by applying a machine-learning classifier to ARU-gathered data. We compared abundance estimates from these models to estimates from the same models fit using point-count data and to two additional models appropriate for point counts, the N-mixture model and distance models. We found that the Royle-Nichols and time-to-detection models can be used with ARU data to produce abundance estimates similar to those generated by a point-count based study but with greater precision. ARU-based models produced confidence or credible intervals that were on average 31.9% ( 11.9 SE) smaller than their point-count counterpart. Our findings were consistent across two species with differing relative abundance and habitat use patterns. The higher precision of models fit using ARU data is likely due to higher cumulative detection probability, which itself may be the result of greater survey effort using ARUs and machine-learning classifiers to sample significantly more time for focal species at any given point. Our results provide preliminary support the use of ARUs in abundance-based study applications, and thus may afford researchers a better understanding of habitat quality and population trends, while allowing them to make more informed conservation actions and recommendations.</p>
Figure 7. Neural Network model-Classification of Human Emotion from Deap EEG Signal Using Hybrid Improved Neural Networks with Cuckoo Search
<p>In Probabilistic Neural Network the operations are organized into a multilayer feed forward<br> neural network with four layers like input layer, hidden layer, pattern layer and output layer. PNN<br> use the Euclidean distance measure the difference between one neuron to other neurons. The actual<br> target values are stored in the hidden neuron and the optimized weighted values are fed into the<br> same category hidden neuron. Then finally the output layer compared the weighted votes of each<br> target values and the target votes are used to predict the emotions.</p>
BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 6. The result of building a 3D model based on RF and SVM classification with "Important features".
<p>From the chart of figure 6, we found that "Important Features" gave the best 3D model, which fits with the object in the image. The pattern is close to 90% compared with the true size. Apply classification algorithm RF increases the accuracy of the results and reduces computing time for the program. There are many methods for data classifying. One of them is the method of the support vector machine (SVM). The SVM method is represented by Vladimir N. Vapnik (1995) in Support Vector Machines (SVM) - a set of learning algorithms similar with the supervisor has two main tasks: the classification and the regression analysis. In this article we use the method of the SVM classification problem for the size of the human body with 5 classes to compare the performance between SVM methods and Random Forest algorithm. </p>
BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 2. Attributes of the classification models used in the experiments
<p>The authors used for their experiments a data set (UCI, 2016) containing 756 records about persons with thyroid dysfunctions. The classification model has 22 attributes; the class attribute is the target and it has three possible values: hypothyroidism, hyperthyroidism and normal. The current data set was extracted and preprocessed from the original file. A description of the attributes used in the experiments is given in Figure 2 (an extract from thyroid.arff test file). </p>
Large Language Model-Based Classification of Flash Flood Impacts Across the United States
<p>This repository contains the data sets used for the publication of the journal article titled <em>Large Language Model-Based Classification of Flash Flood Impacts Across the United States</em>.</p> <p>This is the first release of the data with a Zenodo DOI attached to the README.md file. </p> <p>Further information about the data can be found in the GitHub repository's README.md file.</p>
Sample data for "Classification Modeling for Hazardous Rip Current Prediction" Notebook
<p>This sample dataset is used in the notebook "Classification Modeling for Hazardous Rip Current Prediction" to demonstrate the application of using machine learning to identify hazardous rip current. The notebook is available in the NOAA Center for Artificial Intelligence GitHub Learning Journey repository (https://github.com/noaa-ncai/learning-journey). The full dataset is available via NOAA.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.