Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

106

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

106 results for “Classification model”

Learn how ShareScore rates datasets ↗
zenodo32/100

Rock physics models of gas hydrate bearing sediments – the classification, simulation workflow, and challenges

<p>This study reviews the rock physics models for simulating the elastic properties of gas hydrate bearing sediments. Considering that it is confusing to select the appropriate model for a specific study from the various models, we classify the models into five categories according to different principles. We also summarize a general workflow of the modeling process, elaborate the possible models in each step and bring up the potential sources of uncertainties. Besides, we explicate the general problems of the current models and raise several potential research directions. This study provides us a clear view of the rock physics models, the associated uncertainties, as well as the general modeling workflow of gas hydrate bearing sediments, and also provides some implications for future studies.</p>

opencc-by-4.0Oct 2021View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

<p>Occupancy modeling is used to evaluate avian distributions and habitat associations, yet it typically requires extensive survey effort because a minimum of three repeat samples are required for accurate parameter estimation. Autonomous recording units (ARUs) can reduce the need for surveyors on site, yet ARUs utility were limited by hardware costs and the time required to manually annotate recordings. Software that identifies bird vocalizations may reduce expert time needed, if classification is sufficiently accurate. We assessed the performance of BirdNET – an automated classifier capable of identifying vocalizations from &gt;900 North American and European bird species – by comparing automated to manual annotations of recordings of 13 breeding bird species collected in northwestern California. We compared the parameter estimates of occupancy models evaluating habitat associations supplied with manually annotated data (9 min recording segments) to output from models supplied with BirdNET detections. We used three sets of BirdNET output to evaluate the duration of automatic annotation needed to approach manually annotated model parameter estimates: 9-min, 87-min, and 87-min of high-confidence detections. We incorporated 100 3-sec manually validated BirdNET detections per species to estimate true and false positive rates within an occupancy model. BirdNET correctly identified 90% and 65% of the bird species a human detected when data were restricted to detections exceeding a low or high confidence score threshold, respectively. Occupancy estimates, including habitat associations, were similar regardless of method. Precision (proportion of true positives to all detections) was &gt;0.70 for 9 of 13 species, and a low of 0.29. However, processing of longer recordings was needed to rival manually annotated data. We conclude that BirdNET is suitable for annotating multispecies recordings for occupancy modeling when extended recording durations are used. Together, ARUs and BirdNET may benefit monitoring and, ultimately, conservation of bird populations by greatly increasing monitoring opportunities.   </p>

opencc-zeroFeb 2022View details →
zenodo32/100

Classification of Mobile Application Reviews using Deep Language Models

<p>supplementary material for ASE 2022&nbsp; &quot;Classification of Mobile Application Reviews using Deep Language Models&quot;</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Examining the Influence of Political Bias on Large Language Model Performance in Stance Classification

<p>Code and dataset for paper "Examining the Influence of Political Bias on Large Language Model Performance in Stance Classification". ICWSM 2025</p> <p>Preprint: https://arxiv.org/abs/2407.17688</p> <p>Citation:&nbsp;</p> <p>@misc{ng2024examininginfluencepoliticalbias,<br>&nbsp; &nbsp; &nbsp; title={Examining the Influence of Political Bias on Large Language Model Performance in Stance Classification},&nbsp;<br>&nbsp; &nbsp; &nbsp; author={Lynnette Hui Xian Ng and Iain Cruickshank and Roy Ka-Wei Lee},<br>&nbsp; &nbsp; &nbsp; year={2024},<br>&nbsp; &nbsp; &nbsp; eprint={2407.17688},<br>&nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br>&nbsp; &nbsp; &nbsp; primaryClass={cs.CL},<br>&nbsp; &nbsp; &nbsp; url={https://arxiv.org/abs/2407.17688},&nbsp;<br>}</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Model and sample data for MNIST classification

<p>Model and sample data for MNIST classification. The data is used in conjunction with <a href="https://github.com/freitaglab/LightToInformation">https://github.com/freitaglab/LightToInformation</a>.</p>

opencc-by-4.0Jul 2019View details →
zenodo32/100

NextVir: Enabling Classification of Tumor-Causing Viruses with Genomic Foundation Models

<p>These are a collection of 150bp reads synthesized using ART. Viral reference genomes were downloaded using iCAV, and the primary assemblies of GRCh38.p14 were used for the human reference.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Explainable Deep Learning for Automatic Rock Classification: High Accuracy Does Not Mean Great Model Performance <Dataset>

<p>This is the dataset of manuscript entitled &quot;Explainable Deep Learning for Automatic Rock Classification: High Accuracy Does Not Mean Great Model Performance&quot;. The manuscript is currently under review. Full access of this dataset will be released once the manuscript is accepted.</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Best Learned Models: End-to-end Learning for Land Cover Classification using Irregular and Unaligned SITS by Combining Attention-Based Interpolation with Sparse Variational Gaussian Processes

<p>Best learned model learned with the dataset available <a href="http://https://doi.org/10.5281/zenodo.8033058">here</a> for the mTAN-GP, mTAN-MLP, mTAN-LTAE, and raw-LTAE.</p> <p>For further details see the pre-print article &quot;End-to-end Learning for Land Cover Classification using Irregular and Unaligned SITS by Combining Attention-Based Interpolation with Sparse Variational Gaussian Processes &quot;. This article is available : <a href="https://hal.science/hal-04112115">here</a>.</p> <p>The implementation of the models is available in the <a href="https://gitlab.cesbio.omp.eu/belletv/land_cover_southfrance_mtan_gp_irregular_sits">open source repository</a>.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Dataset, models and code for "Automating global landslide detection with heterogeneous ensemble deep-learning classification"

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
ClinicalTrials.gov32/100

Conversational AI Models in Periodontitis Classification

ClinicalTrials.gov study NCT05926999. IPD Sharing: Not stated. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Establishment of Molecular Classification Models for Early Diagnosis of Digestive System Cancers

ClinicalTrials.gov study NCT05431621. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
dryad32/100

Systematic review of validation of supervised machine learning models in accelerometer-based animal behaviour classification literature

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad32/100

LeWoS: A universal leaf‐wood classification method to facilitate the 3D modelling of large tropical trees using terrestrial LiDAR

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

Open the record for dataset details and reuse information.

publicFeb 2022View details →
zenodo28/100

Supplementary material 2 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049

Table S2. Basic information obtained for 49 exotic plants in Chile

opencc-zeroDec 2020View details →
zenodo28/100

Supplementary material 3 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049

Map of the species

opencc-zeroDec 2020View details →
zenodo28/100

Replication Package for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"

<p>Replication package for our analysis of Semantic Versioning in Ansible Galaxy role repositories.</p> <p>This replication package consists of three parts:</p> <ul> <li> <p>Classification Model: Contains Jupyter notebooks used to train and evaluate a Random Forest classification model based on structural features. Training and evaluation data is included.</p> </li> <li> <p>Quantitative Notebooks: Contains Jupyter notebooks used to perform quantitative analyses of versions and changes.</p> </li> <li> <p>data: CSV files of the data used in the Quantitative Notebooks, and the source data for the classification model. Should be downloaded separately fromthe classification model. Should be downloaded separately from <a href="https://doi.org/10.5281/zenodo.4991955">https://doi.org/10.5281/zenodo.4991955</a>.</p> </li> </ul> <p>The data is under the Creative Commons Attribution Share-Alike 4.0 license. The source code is under the GNU General Public License.</p>

openother-openJun 2021View details →
dryad28/100

Data from: Quantifying and modelling decay in forecast proficiency indicates the limits of transferability in land-cover classification

1. The ability to provide reliable projections for the current and future distribution patterns of land-covers is fundamental if we wish to protect and manage our diminishing natural resources. Two inter-related revolutions made map productions feasible at unprecedented resolutions- the availability of high-resolution remotely-sensed data and the development of machine-learning algorithms. However, the ground-truth data needed for training models is in most cases spatially and temporally clustered. Therefore, map production requires extrapolation of models from one place to another and the uncertainty cost of such extrapolation is rarely explored. In other words, we focus mainly on projections, and less on quantifying how reliable they are. 2. Following the concept of 'forecast horizon', we suggest that the predictability of land-cover classification models should be methodologically explored with quantitative tools as a continuum against distances measured along multiple dimensions. Focusing on ten agricultural sites from England and using models specifically designed to predict multivariate decay-curves we ask: how does a model's predictive performance decay with distance? More specifically, we explored if we could predict the proficiency (kappa statistics) of a model trained in one site when making predictions in another site based on the spatial, temporal, spectral and environmental distances between sites. 3. We found that model proficiency decays with spatial, temporal, spectral and environmental distance between sites. More importantly, we found for the first time that it is possible to predict the performance a model transferred to or from a novel site will have, based on its distances from known sites. The spatial distance variables where the most important when predicting model transferability. 4. Exploring model transferability as a continuum may have multiple usages including predicting uncertainty values in space and time, prioritization of strategies for ground-truth data collection, and optimizing model characteristics for defined tasks.

opencc-zeroDec 2016View details →
zenodo28/100

HiPHD: Hierarchical Classification for Protein Remote Homology Detection using Graph Neural Networks and Language Models

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo28/100

CATH-KinFams: CATH Protein Kinase classification alignments and Hidden Markov Models

<p>CATH KinFams are protein kinase domain families classified according to functional similarity based on SDP. In this deposition we make available 2,210 KinFams sequence alignments alongside Hidden Markov Models built from them to be used with HMMER3.</p> <p>A concatenated library &#39;kinases_4.3-FF-seed.hmm&#39; is also available to scan against the whole KinFams dataset.</p> <p>The Zenodo deposition contains:</p> <p>kinfams-cath-4.3-seed-alignments.tar.gz&nbsp; - KinFams FASTA file alignments with headers &#39;UniProt_ID/start-stop&#39; i.e. A8XMX4/281-587</p> <p>kinfams-cath-4.3-seed-hmms.tar.gz - HMMs for each individual KinFam and concatenated in a HMM library.</p> <p>kinfams-cath-4.3-seed-mda-strings - Multi-Domain-Architecture string assignment for each sequence in the KinFams dataset.</p> <p>human_kinfams_af2_models_cif.tar.gz - Chopped mmCIF files containing Human Kinases AlphaFold2 Models.</p>

opencc-by-4.0Jan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record