Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

106

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

106 results for “Classification model”

Learn how ShareScore rates datasets ↗
zenodo36/100

A model-based approach to characterize enzyme-mediated response to antibiotic treatments: towards a model-guided classification

<p>This dataset, taken together with the scripts at <a href="https://gitlab.inria.fr/Public/InBio/esbl-escape">https://gitlab.inria.fr/Public/InBio/esbl-escape</a>, allows one to reproduce the analyses and figures of the article &quot;A model-based approach to characterize enzyme-mediated response to antibiotic treatments: towards a model-guided classification&quot;.</p>

opencc-by-4.0Jul 2021View details →
dryad36/100

Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids

<p>Animal-attached devices can be used on cryptic species to measure their movement and behaviour, enabling unprecedented insights into fundamental aspects of animal ecology and behaviour. However, direct observations of subjects are often still necessary to translate biologging data accurately into meaningful behaviours. As many elusive species cannot easily be observed in the wild, captive or domestic surrogates are typically used to calibrate data from devices. However, the utility of this approach remains equivocal. </p> <p>Here, we assess the validity of using captive conspecifics, and phylogenetically-similar domesticated counterparts (surrogate species) for calibrating behaviour classification. Tri-axial accelerometers and tri-axial magnetometers were used with behavioural observations to build random forest models to predict the behaviours. We applied these methods using captive Alpine ibex (Capra ibex) and a domestic counterpart, pygmy goats (Capra aegagrus hircus), to predict the behaviour including terrain slope for locomotion behaviours of captive Alpine ibex. </p> <p>Behavioural classification of captive Alpine ibex and domestic pygmy goats was highly accurate (&gt; 98%). Model performance was reduced when using data split per individual, i.e., classifying behaviour of individuals not used to train models (mean ± sd = 56.1 ± 11%). Behavioural classifications using domestic counterparts, i.e., pygmy goat observations to predict ibex behaviour, however, were not sufficient to predict all behaviours of a phylogenetically similar species accurately (&gt; 55%).</p> <p>We demonstrate methods to refine the use of random forest models to classify behaviours of both captive and free-living animal species. We suggest there are two main reasons for reduced accuracy when using a domestic counterpart to predict the behaviour of a wild species in captivity; domestication leading to morphological differences and the terrain of the environment in which the animals were observed. We also identify limitations when behaviour is predicted in individuals that are not used to train models. Our results demonstrate that biologging device calibration needs to be conducted using: (i) with similar conspecifics, and (ii) in an area where they can perform behaviours on terrain that reflects that of species in the wild.</p>

opencc-zeroDec 2020View details →
zenodo36/100

Image dataset for training of an insect classification model

<p>&nbsp;</p><p><strong>This version is deprecated! Please use the updated </strong><a href="https://doi.org/10.5281/zenodo.8325383"><strong>Insect Detect - insect classification dataset v2</strong></a><strong> with more images and classes.</strong></p><p>&nbsp;</p><p>This dataset contains images of various insects and some other arthropods, sitting on or flying above an artificial flower platform. All images were automatically recorded with the <a href="https://maxsitt.github.io/insect-detect-docs/">Insect Detect DIY camera trap</a>, a hardware combination of the Luxonis OAK-1, Raspberry Pi Zero 2 W and PiJuice Zero pHAT for automated insect monitoring (<a href="https://doi.org/10.1101/2023.12.05.570242">bioRxiv preprint</a>).</p><p>This classification dataset contains the cropped bounding boxes, exported from the <a href="https://universe.roboflow.com/maximilian-sittinger/insect_detect_detection/dataset/6">Insect_Detect_detection</a> dataset together with 290 new images of <i>Episyrphus balteatus</i>.</p><h2>Classes</h2><p>The following classes were annotated in this dataset:</p><ul><li><strong>wasp</strong> (mostly <i>Vespula</i> sp.)</li><li><strong>hbee</strong> (<i>Apis mellifera</i>)</li><li><strong>fly</strong> (mostly Brachycera)</li><li><strong>hovfly</strong> (various Syrphidae, e.g. <i>Eupeodes corollae</i>,<i> Scaeva pyrastri</i>)</li><li><strong>episyr_balt</strong> (<i>Episyrphus balteatus</i>)</li><li><strong>other</strong> (all Arthropods with insufficient occurences, e.g. various Hymenoptera, true bugs, beetles)</li><li><strong>shadow</strong> (shadows of the recorded insects)</li></ul><p>View the <a href="https://universe.roboflow.com/maximilian-sittinger/insect_detect_classification/health">Health Check</a> for more info on class balance.</p><h2>Deployment</h2><p>You can use this dataset as starting point to train your own insect classification models. Check the <a href="https://maxsitt.github.io/insect-detect-docs/modeltraining/train_classification/">model training instructions</a> for more information.</p><p>To deploy the image classification model (ONNX format) on your PC for fast CPU inference, follow the provided <a href="https://maxsitt.github.io/insect-detect-docs/deployment/classification/">Step by Step instructions</a>. Open source Python scripts to deploy the trained model can be found at the <a href="https://github.com/maxsitt/insect-detect-ml">insect-detect-ml GitHub repo</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Data obtained during classification of intertidal habitats using UAV imagery in the Galapagos Archipelago (Orthophotos, digital elevation models (DEM) and orthophoto-draped 3D models)

<p>In the repository 5 folders exist. 1) Digital elevation models (DEMs), 2) Intertidal habitat map, 3) Othophoto&nbsp;draped 3D models, 4) Orthophotos, and 5) Processing reports. The data has been collected&nbsp;in Puerto Ayora at Santa Cruz in August 2017, the most urbanized island of the Galapagos Archipelago.&nbsp;The purpose of this study was to investigate the image classification opportunities for these intertidal habitats using Uncrewed Aerial Vehicle (UAV) imagery. This dataset is cited in&nbsp;an open-access publication: &nbsp;https://doi.org/10.3390/drones7070416.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Out-of-distribution detection algorithms for robust insect classification dataset and models

<p>This folder contains trained models and datasets for reproducing the results in the paper on out-of-distribution detection algorithms for robust insect classification. Specifically, it contains the following folders:&nbsp;</p> <p>&nbsp;</p> <ul> <li>OODInsect (out-of-distribution data)</li> <li>MSP, MAH, and EBM trained models, each wrapped around the three classifiers of ResNet50, RegNet32, and VGG11, and different combinations of ID and OOD test data for reproducing RQ1, RQ2, and RQ3.</li> <li>ID3 (in-distribution test data)</li> </ul>

opencc-by-4.0May 2023View details →
zenodo36/100

BioGraph Quality Score and Genotype Classifer Models

<p>BioGraph Quality Score and Genotype Classifer Models v7.1.0&nbsp;for use with BioGraph software available at&nbsp;https://github.com/spiralgenetics/biograph</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Comparative modeling of Vickery's faceted classification and the œuvre of S. R. Ranganathan...

<p>&nbsp;In an effort to make sense of both Ranganathan&rsquo;s work and Vickery&rsquo;s we modeled the<br> process involved in classification using the IDEF0 (Integrated Definition for Function<br> Modeling) formalism. This allows us to see five distinct parts of the classification process:<br> actions, inputs, outputs, mechanisms, and constraints.&nbsp;</p>

opencc-by-4.0Jul 2011View details →
ClinicalTrials.gov36/100

Detection and Classification of Diabetic Retinopathy From Posterior Pole Images With A Deep Learning Model

ClinicalTrials.gov study NCT04805541. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
dryad36/100

Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad36/100

MVCNN++: CAD model shape classification and retrieval using multi-view convolutional neural networks

Open the record for dataset details and reuse information.

publicAug 2020View details →
zenodo32/100

Pairwise Multi-Class Document Classification for Semantic Relations between Wikipedia Articles (Dataset, Models & Code)

<p>Many digital libraries recommend literature to their users considering the similarity between a query document and their repository. However, they often fail to distinguish what is the relationship that makes two documents alike. In this paper, we model the problem of finding the relationship between two documents as a pairwise document classification task. To find the semantic relation between documents, we apply a series of techniques, such as GloVe, Paragraph-Vectors, BERT, and XLNet under different configurations (e.g., sequence length, vector concatenation scheme), including a Siamese architecture for the Transformer-based systems. We perform our experiments on a newly proposed dataset of 32,168 Wikipedia article pairs and Wikidata properties that define the semantic document relations. Our results show vanilla BERT as the best performing system with an F1-score of 0.93,<br> which we manually examine to better understand its applicability to other domains. Our findings suggest that classifying semantic relations between documents is a solvable task and motivates the development of recommender systems based on the evaluated techniques. The discussions in this paper serve as first steps in the exploration of documents through SPARQL-like queries such that one could find documents that are similar in one aspect but dissimilar in another.</p> <p>Additional information can be found on <a href="https://github.com/malteos/semantic-document-relations/">GitHub</a>.</p> <p>The following data is supplemental to the experiments described in our research paper. The data consists of:</p> <ul> <li>Datasets (articles, class labels, cross-validation splits)</li> <li>Pretrained models (Transformers, GloVe, Doc2vec)</li> <li>Model output (prediction) for the best performing models</li> </ul> <p><strong>Dataset</strong></p> <p>The Wikipedia article corpus is available in <code>enwiki-20191101-pages-articles.weighted.10k.jsonl.bz2</code>. The original data have been downloaded as <a href="https://dumps.wikimedia.org/enwiki/">XML dump</a>, and the corresponding articles were extracted as plain-text with <a href="https://radimrehurek.com/gensim/scripts/segment_wiki.html">gensim.scripts.segment_wiki</a>. The archive contains only articles that are available in training or test data.</p> <p>The actual dataset is provided as used in the stratified k-fold with <code>k=4</code> in <code>train_testdata__4folds.tar.gz</code>.</p> <pre><code>├── 1 │   ├── test.csv │   └── train.csv ├── 2 │   ├── test.csv │   └── train.csv ├── 3 │   ├── test.csv │   └── train.csv └── 4 ├── test.csv └── train.csv 4 directories, 8 files </code></pre> <p>Pretrained models</p> <p>PyTorch: vanilla and Siamese BERT + XLNet</p> <p>Pretrained model for each fold is available in the corresponding model archives:</p> <pre><code># Vanilla model_wiki.bert_base__joint__seq512.tar.gz model_wiki.xlnet_base__joint__seq512.tar.gz # Siamese model_wiki.bert_base__siamese__seq512__4d.tar.gz model_wiki.xlnet_base__siamese__seq512__4d.tar.gz </code></pre>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Automatic Classification of Non-functional Requirements in App User Reviews Based on System Model and Artificial Intelligence

<p>This is the replication package for the paper: &quot;Automatic Classification of Non-functional Requirements in App User Reviews Based on System Model and Artificial Intelligence&quot;.&nbsp;It contains the dataset of our experiment for the&nbsp;replication&nbsp;by&nbsp;other&nbsp;researchers. In the meanwhile, we provide brief description of the files in the replication&nbsp;package in the following.</p> <p><strong>1. dataset folder</strong></p> <ul> <li>dataset_user_reviews.xlsx&nbsp; contains 1278 labelled non-requirement user reviews.</li> <li>readme.txt describes the meaning of the data in&nbsp;dataset_user_reviews.xlsx in detail.</li> </ul>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Protein Subcellular localization prediction data used in the article entitled "MSclassifier: Median-Supplement model-based Classification tool for automated knowledge discovery"

<p>This repository contains data used to obtain results from a 5-fold cross-validation testing of how MSclassifier and other packages accurately predict protein subcellular localization in the software article entitled &quot;MSclassifier: median-supplement model-based classification tool for automated knowledge discovery.&quot; The data used in the software article is derived from data generated in &quot;G. K. Acquaah-Mensah, S. M. Leach, and C. Guda, Predicting the subcellular localization of human proteins using machine learning and exploratory data analysis, Genomics Proteomics Bioinformatics, 4(2):120-133, 2006, <a href="https://doi.org/10.1016/S1672-0229(06)60023-5">https://doi.org/10.1016/S1672-0229(06)60023-5</a>&quot;</p>

opencc-by-nc-sa-3.0Jul 2020View details →
zenodo32/100

A classification scheme to determine wildfires from the satellite record in the cool grasslands of southern Canada: considerations for fire occurrence modelling and warning criteria

<p>This&nbsp;data set can be used to reproduce the results from the following paper: &quot;A classification scheme to determine wildfires from the satellite record in the cool grasslands of southern Canada: considerations for fire occurrence modelling and warning criteria&quot;.</p> <p>The Landsat Images are a series of pngs obtained from&nbsp;<a href="https://landbrowser.airc.aist.go.jp/hotarea/">https://landbrowser.airc.aist.go.jp/hotarea/</a>&nbsp;representing agricultural fires in our study area (Kato et al., 2018).</p> <p><a href="https://zenodo.org/api/files/b1449276-d0e7-4ede-aa97-6c976e27215f/Model_Input_MODIS_Hotspot_Clusters_Revised.csv">Model_Input_MODIS_Hotspot_Clusters_Revised.csv</a>&nbsp;contains a list of&nbsp;clusters of MODIS hotspots and associated attributes classified as either agricultural or grassland wildfires and is used to create a GAM to explore the conditions of these fires.</p> <p><a href="https://zenodo.org/api/files/b1449276-d0e7-4ede-aa97-6c976e27215f/Model_Predicted_MODIS_Hotspot_Clusters_Revised.csv">Model_Predicted_MODIS_Hotspot_Clusters_Revised.csv</a>&nbsp;contains the complete list of MODIS hotspot clusters and associated attributes for our study area from 2002-2018.&nbsp; These clusters have been classified as either agricultural or grassland wildfires using the GAM mentioned above.</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

PyTorch deep learning models for landscape classification (PyLC)

<pre><em>Pytorch pretrained models for use by the Python Landscape Classification Tool (PyLC) </em><em> Reference: An evaluation of deep learning semantic segmentation </em><em> for land cover classification of oblique ground-based photography, </em><em> MSc. Thesis 2020. </em><em> &lt;http://hdl.handle.net/1828/12156&gt; </em><em>Spencer Rose &lt;spencerrose@uvic.ca&gt;, June 2020 </em><em>University of Victoria</em></pre>

opencc-by-4.0Nov 2020View details →
zenodo32/100

PyTorch model for taxonomic classification

<p>This is a trained PyTorch model for classifying an amino acid sequence&#39;s (preferably of length 100) taxonomic domain as viral (class 0), bacterial (class 1) or mammalian (class 2).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo32/100

Supplementary material 1 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049

Table S1. Exotic species located in Quadrant 1 (see Figure 3) and impacts on biodiversity, agriculture and cattle raisng

opencc-zeroDec 2020View details →
zenodo32/100

Dataset for Ha and Aylward 'Automated classification of giant virus genomes using a random forest model built on trademark protein families'

<ul><li>Genome sets used for model training and testing</li><li>Custom Python script that generated fragmented genomes at random completeness levels</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Best learned models : Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features

<p>Best learned models (Gaussian Processes, Random Forest, Multilayer Perceptron and Lightweight Temporal Self-Attention models) for each region based on the classification data set DS-A with seed 0 (see description <a href="https://zenodo.org/deposit/7099785">here</a>).&nbsp;</p><p>Models: GP non spatial, GP spatial (sum), GP spatial (product),RF non spatial, RF spatial, MLP non spatial, MLP spatial, LTAE non spatial, LTAE spatial</p><p>For further details see section VI-C of the pre-print article "Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features ". This article is available <a href="https://hal.archives-ouvertes.fr/hal-03781332">here</a>.</p><p>The implementation of the models is available in the <a href="https://gitlab.cesbio.omp.eu/belletv/land_cover_southfrance_gp">open source repository</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Diversity is Definitely Needed: Improving Model-Agnostic Zero-shot Classification via Stable Diffusion

<p>In this work, we investigate the problem of Model-Agnostic Zero-Shot Classification (MA-ZSC), which refers to training non-specific classification architectures (downstream models) to classify real images without using any real images during training. Recent research has demonstrated that generating synthetic training images using diffusion models provides a potential solution to address MA-ZSC. However, the performance of this approach currently falls short of that achieved by large-scale vision-language models. One possible explanation is a potential significant domain gap between synthetic and real images. Our work offers a fresh perspective on the problem by providing initial insights that MA-ZSC performance can be improved by improving the diversity of images in the generated dataset. We propose a set of modifications to the text-to-image generation process using a pre-trained diffusion model to enhance diversity, which we refer to as our <strong>bag of tricks</strong>. Our approach shows notable improvements in various classification architectures, with results comparable to state-of-the-art models such as CLIP. To validate our approach, we conduct experiments on CIFAR10, CIFAR100, and EuroSAT, which is particularly difficult for zero-shot classification due to its satellite image domain. We evaluate our approach with five classification architectures, including ResNet and ViT. Our findings provide initial insights into the problem of MA-ZSC using diffusion models.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record