Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Software Defect Prediction”
SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features
<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer’s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit <a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>
Paper Repository and References for "Early software defect prediction: A systematic map and review"
<p>Context: Software defect prediction is a trending research topic, and a wide variety of the published papers focus on coding phase or after. A limited number of papers, however, includes the prior (early) phases of the software development lifecycle (SDLC).<br> Objective: The goal of this study is to obtain a general view of the characteristics and usefulness of Early Software Defect Prediction (ESDP) models reported in scientific literature. <br> Method: A systematic mapping and systematic literature review study has been conducted. We searched for the studies reported between 2000 and 2016. We reviewed 52 studies and analyzed the trend and demographics, maturity of state-of-research, in-depth characteristics, success and benefits of ESDP models. <br> Results: We found that categorical models that rely on requirement and design phase metrics, and few continuous models including metrics from requirements phase are very successful. We also found that most studies reported qualitative benefits of using ESDP models.<br> Conclusion: We have highlighted the most preferred prediction methods, metrics, datasets and performance evaluation methods, as well as the addressed SDLC phases. We expect the results will be useful for software teams by guiding them to use early predictors effectively in practice, and for researchers in directing their future efforts.</p>
An Audit of Machine Learning Experiments on Software Defect Prediction - Dataset
<p><strong>ML_Audit_20250328_anon.csv</strong>:<br>This CSV file contains anonymized data used in the audit of machine learning experiments on software defect prediction. The dataset includes variables and performance metrics extracted from studies published between 2019 and 2023. It supports the audit's evaluation of study reproducibility and issues related to experimental design and statistical analysis. This data can be used for replication and further analysis of the trends and reproducibility issues identified in the paper.</p> <p><strong>ML_Audit_March2025.Rmd</strong>:<br>This RMarkdown file contains the analysis script used for the statistical analysis and audit of the machine learning experiments reviewed in the study. It includes the procedures for data preprocessing, statistical evaluations, and reproducibility assessments. The script is integral for replicating the audit results presented in the paper and can be used by other researchers to perform similar audits or extend the analysis on different datasets.</p>
Towards Developing and Analysing The Metric-Based Software Defect Severity Prediction Model
<p>This is a metric based approach to solve software defect severity prediction problem. In addition to that, this work proposes a new evaluation scheme that comprised of five metrics to analyze the performances.</p>
Datasets for Software Defect Number Prediction
<p>27 Datasets with ARFF format for Software Defect Number Prediction</p>
Assessing software defection prediction performance: why using the Matthews correlation coefficient matters
<p>The documents include raw data and survey papers for our research. </p>
Costs and Benefits of Machine Learning Software Defect Prediction
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.