Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “Software Defect Prediction”

Learn how ShareScore rates datasets ↗
zenodo40/100

SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features

<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer&rsquo;s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit&nbsp;<a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Paper Repository and References for "Early software defect prediction: A systematic map and review"

<p>Context: Software defect prediction is a trending research topic, and a wide variety of the published papers focus on coding phase or after. A limited number of papers, however, includes the prior (early) phases of the software&nbsp;development lifecycle (SDLC).<br> Objective: The goal of this study is to obtain a general view of the characteristics and usefulness of Early Software&nbsp;Defect Prediction (ESDP) models reported in scientific literature.&nbsp;<br> Method: A systematic mapping and systematic literature review study has been conducted. We searched for the&nbsp;studies reported between 2000 and 2016. We reviewed 52 studies and analyzed the trend and demographics,&nbsp;maturity of state-of-research, in-depth characteristics, success and benefits of ESDP models.&nbsp;<br> Results: We found that categorical models that rely on requirement and design phase metrics, and few continuous&nbsp;models including metrics from requirements phase are very successful. We also found that most studies&nbsp;reported qualitative benefits of using ESDP models.<br> Conclusion: We have highlighted the most preferred prediction methods, metrics, datasets and performance&nbsp;evaluation methods, as well as the addressed SDLC phases. We expect the results will be useful for software&nbsp;teams by guiding them to use early predictors effectively in practice, and for researchers in directing their future&nbsp;efforts.</p>

opencc-by-4.0Oct 2017View details →
zenodo36/100

An Audit of Machine Learning Experiments on Software Defect Prediction - Dataset

<p><strong>ML_Audit_20250328_anon.csv</strong>:<br>This CSV file contains anonymized data used in the audit of machine learning experiments on software defect prediction. The dataset includes variables and performance metrics extracted from studies published between 2019 and 2023. It supports the audit's evaluation of study reproducibility and issues related to experimental design and statistical analysis. This data can be used for replication and further analysis of the trends and reproducibility issues identified in the paper.</p> <p><strong>ML_Audit_March2025.Rmd</strong>:<br>This RMarkdown file contains the analysis script used for the statistical analysis and audit of the machine learning experiments reviewed in the study. It includes the procedures for data preprocessing, statistical evaluations, and reproducibility assessments. The script is integral for replicating the audit results presented in the paper and can be used by other researchers to perform similar audits or extend the analysis on different datasets.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Towards Developing and Analysing The Metric-Based Software Defect Severity Prediction Model

<p>This is a metric based approach to solve software defect severity prediction problem. In addition to that, this work proposes a new evaluation scheme that comprised of five metrics to analyze the performances.</p>

opencc-by-4.0Jun 2022View details →
zenodo28/100

Datasets for Software Defect Number Prediction

<p>27 Datasets with ARFF format&nbsp;for Software Defect Number Prediction</p>

opencc-bySep 2019View details →
zenodo20/100

Assessing software defection prediction performance: why using the Matthews correlation coefficient matters

<p>The documents include raw data and survey papers for our research.&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo20/100

Costs and Benefits of Machine Learning Software Defect Prediction

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record