Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,890
datasets available to search
ShareScore release 0.9.0
Dataset results
1,890 results for “Defects”
Defect Prediction: SPE
<p>Software process evaluation is essential to improve software development and the quality of software products in an organization. Conventional approaches based on manual qualitative evaluations (e.g., artifacts inspection) are deficient in the sense that (i) they are time-consuming, (ii) they suffer from the authority constraints, and (iii) they are often subjective. To overcome these limitations, this paper presents a novel semi-automated approach to software process evaluation using machine learning techniques. In particular, we formulate the problem as a sequence classification task, which is solved by applying machine learning algorithms. Based on the framework, we define a new quantitative indicator to objectively evaluate the quality and performance of a software process. To validate the efficacy of our approach, we apply it to evaluate the defect management process performed in four real industrial software projects. Our empirical results show that our approach is effective and promising in providing an objective and quantitative measurement for software process evaluation.</p> <p><strong>Reference: </strong>Chen, Ning, Steven CH Hoi, and Xiaokui Xiao. "Software process evaluation: A machine learning approach." <em>Proceedings of the 2011 26th IEEE/ACM International Conference on Automated Software Engineering</em>. IEEE Computer Society, 2011.</p>
Eclipse and Mozilla defect tracking dataset
<p>A dataset with over 200.000 reported bugs extracted from the Eclipse and Mozilla projects (respectively 47.000 and 168.000 reported bugs). Besides providing a single snapshot of a bug report, we also include all the incremental modifications as performed during the lifetime of the bug report.</p>
Data set for ``Why is Differential Evolution Better than Grid Search for Tuning Defect Predictors?''
<p>One of the black arts of data mining is learning the magic parameters that control the learners. In software analytics, at least for defect prediction, several methods, like grid search and differential evolution(DE), have been proposed to learn those parameters. They’ve been proved to be able to improve learner performance.</p> <p>We want to evaluate which method can find better parameters in terms of performance score and runtime. This paper compares grid search to differential evolution, which is an evolutionary algorithm that makes extensive use of stochastic jumps around the search space. We find that the seemingly complete approach of grid search does no better, and sometimes worse, than the stochastic search. Yet, when repeated 20 times to check for conclusion validity, DE was over 210 times faster (6.2 hours for DE vs 54 days for grid search when both tuning Random Forest over 17 test data sets with F-measure as optimization objective).</p> <p>These results are puzzling: why does a quick partial search be just as effective as a much slower, and much more, extensive search? To answer that question, we turned to the theoretical optimization literature. Bergstra and Bengio conjecture that grid search is not more effective than more randomized searchers if the underlying search space is inherently low dimensional. This is significant since recent results show that defect prediction exhibits very low intrinsic dimensionality– an observation that explains why a fast method like DE may work as well as a seemingly more thorough grid search. This suggests, as a future research direction, that it might be possible to peek at data sets before doing any optimization in order to match the optimization algorithm to the problem at hand.</p>
Dataset and source code for ICSME2017 paper "Supervised vs Unsupervised Models: A Holistic Look at Effort-Aware Just-in-Time Defect Prediction"
<p>Dataset and source code for ICSME2017 paper “Supervised vs Unsupervised Models: A Holistic Look at Effort-Aware Just-in-Time Defect Prediction”</p> <p>There are four different models in the paper (i.e., EALR, LT, CBS and OneWay). Each model was implemented in a single Java file in the model package. To reproduce the experiment results of each model in the paper, just run the main method in the corresponding Java file. </p> <p> </p>
Defect tolerance of lead-halide perovskite (100) surface relative to bulk: band bending, surface states, and characteristics of vacancies (dataset)
<p>This repository contains an input data set as well as the output data that were used in a study of surface and bulk defects in cubic CsPbI<sub>3</sub>. using a Vienna ab initio simulation package (VASP) and PyDEF 2 package for processing of defect calculations. More details about organization of the dataset can be found within REDME.txt files</p>
Supplementary material for the publication: Deep Learning of Crystalline Defects from TEM images: A Solution for the Problem of "Never Enough Training Data"
Open the record for dataset details and reuse information.
Data from: Severe enamel defects in wild Japanese macaques
<p>Plane-form enamel hypoplasia (PFEH) is a severe dental defect in which large areas of the crown are devoid of enamel. This condition is rare in humans and even rarer in wild primates. The etiology of PFEH has been linked to exposure to severe disease, malnutrition, environmental toxins, and associated with systemic conditions. In this study, we examined the prevalence of enamel hypoplasia in several populations of wild Japanese macaques (<em>Macaca fuscata</em>) with the aim of providing context for severe defects observed in macaques from Yakushima Island. We found that 10 of 21 individuals (48%) from Yakushima Island displayed uniform and significant PFEH; all 10 specimens were from two adjacent locations in the south of the island. In contrast, macaques from other islands and from mainland Japan have low prevalence of the more common types of enamel hypoplasia and none exhibit PFEH. In Yakushima macaques, every tooth type was affected to varying degrees except for first molars and primary teeth, and the mineral content of the remaining enamel in teeth with PFEH was normal (i.e., no hypo- or hyper mineralization). The aetiology of PFEH might be linked to extreme weather events or high rates of environmental fluoride causing enamel breakdown. However, given that the affected individuals underwent dental development during a period of substantial human-related habitat change, an anthropogenic related etiology seems most likely. Further research on living primate populations is needed to better understand the causes of PFEH in wild primates.</p>
Dataset for Article "On the experimental properties of the TS defect in 4H-SiC"
<p>This dataset contains raw data as well as evaluation scripts to the manuscript "On the experimental properties of the TS defect in 4H-SiC".</p>
Datasets and scripts for the publication "Insights into Defect Cluster Formation in Non-Stoichiometric Wustite (Fe1-xO) at Elevated Temperatures: Accurate force field from Deep Learning"
<div> <div> <div> <div> <p><strong>All the datasets and scripts for the publication"Insights into Defect Cluster Formation in Non-Stoichiometric Wustite (Fe<sub>1-x</sub>O) at Elevated Temperatures: Accurate force field from Deep Learning".</strong></p> <p>This database contains high-fidelity datasets for non-stoichiometric wüstite (Fe₁₋ₓO), including atomic coordinates, energies, and forces generated through ab initio molecular dynamics (AIMD) and refined using Deep Potential (DP) training. The dataset encompasses bulk phases, vacancy structures, and surface orientations, enabling accurate modeling of defect clusters and thermodynamic properties. It supports machine-learning force field development, offering insights into defect formation and large-scale simulations of Fe₁₋ₓO systems at elevated temperatures.</p> </div> </div> </div> </div> <div> <p>Description of the File Structure of Fe1-xO_DeepMD_Code_Datasets_Analysis.zip:</p> <p>1. `<code>init</code>` Folder <br>This folder contains the foundational datasets and inputs used for training and developing the machine-learning force field for Fe₁₋ₓO. </p> <blockquote> <p>1.1 `<code>01.train_data</code>` Subfolder <br>This folder organizes data related to the initial training of the Deep Potential (DP) model. <br>- `<code>dpmd_dataset</code>`: Processed dataset ready for DeepMD training, containing atomic configurations, forces, and energies.<br>- `<code>dpmd_rawfiles</code>`: Raw files from ab initio molecular dynamics (AIMD) simulations, serving as the source for generating training datasets.</p> <p>1.2 `<code>02.develop_data</code>` Subfolder<br>Contains `<code>.vasp</code>` files representing structural data used to develop and refine the force field. The structures include bulk, vacancy, and surface configurations of Fe₁₋ₓO. <br>- Files labeled `<code>bulk</code>` represent bulk Fe₁₋ₓO systems with varying lattice constants. <br>- Files labeled `<code>defect</code>` represent Fe and O vacancy structures (single and double vacancies). <br>- Files labeled `<code>surface</code>` represent Fe₁₋ₓO surface structures in various crystallographic orientations. <br><br></p> </blockquote> <p>2. `<code>run</code>` Folder<br>This folder contains files and logs generated during iterative training and testing of the DP force field, as well as subfolders for each iteration of the training process. </p> <blockquote> <p>2.1 Iteration Folders (`<code>iter.000000</code>` to `<code>iter.000024</code>`):<br>Each folder represents an iteration in the iterative refinement of the DP model, with three subfolders: <br>- `<code>00.train</code>`: Contains training data and outputs for the DP model during the current iteration. <br>- `<code>01.model_devi</code>`: Tracks deviations between DP predictions and ab initio results, guiding dataset selection for the next iteration. <br>- `<code>02.fp</code>`: Stores first-principles (FP) results from CP2K used to improve DP model accuracy. </p> <p>2.2 Other Key Files: <br>- `<code>cp2k.input</code>`: Input file for CP2K, used for performing ab initio calculations on configurations during the iterative process. <br>- `<code>dpdispatcher.log</code>`: Log file tracking the progress of data dispatching and task execution. <br>- `<code>dpgen.log</code>`: Log file recording operations of DPGEN during dataset generation and force field development. <br>- `<code>dpgen_nohup.sh</code>`: Script for running DPGEN in the background. <br>- `<code>machine_slurm_cp2k.json</code>`: Configuration file specifying computing resources for CP2K simulations in a cluster environment. <br>- `<code>param_cp2k.json</code>`: Parameter file for CP2K calculations, defining simulation settings. <br>- `<code>record.dpgen</code>`: Record of iterative processes, including input parameters and outputs for each stage.</p> </blockquote> <p>This organized structure ensures a systematic approach to dataset preparation, model training, and iterative refinement for developing accurate machine-learning potentials for Fe₁₋ₓO.</p> <p>graph-compress.0330.pb is the final compressed DeepMD potential parameters.</p> </div>
Dataset: The effect of a keyhole defect on strain localisation in an additive manufactured titanium alloy
<p><strong>This is the dataset used in the following publication: </strong></p> <div> <div> <div> <p>S. Cao, R. Thomas, A.D. Smith, P. Zhang, L. Meng, H. Liu, J. Guo, J. Donoghue, D. Lunt, The effect of a keyhole defect on strain localisation in an additive manufactured titanium alloy, Journal of Materials Research and Technology, https://doi.org/10.1016/j.jmrt.2024.11.237</p> </div> </div> </div> <p><strong>Contained in this dataset are:</strong></p> <p>A Jupyter notebook which uses the open-source DefDAP Python package (https://github.com/MechMicroMan/DefDAP) to open enclosed HRDIC and EBSD data for two regions in an SLM Ti64 sample, one around a keyhole defect and one ~1mm away in the bulk.</p> <p>Please use the 'master' version of DefDAP: <a href="https://github.com/MechMicroMan/DefDAP/tree/51074e158b0131c69358ddf7eee319e41cf582ca">https://github.com/MechMicroMan/DefDAP/</a></p> <p><strong>Publication abstract:</strong></p> <p>The influence of a keyhole defect on local deformation behaviour in additive manufactured Ti-6Al-4V was investigated by comparing it to a representative bulk region without a defect. High resolution digital image correlation (HRDIC) was used to measure the differences in strain localisation at the microstructural length-scale. A nanoscale speckle pattern was used to allow small changes in strain to be detected and resolved within a single individual lamella and at pre-existing crack locations around the defect. Strain localisation was observed around the defect and formed well below the macroscopic yield stress. In contrast, minimal deformation was found in the bulk at this stress level. Following further deformation into the plastic regime, the strain localisation around the keyhole became more heterogenous with a distinct strain field. A large amount of strain localisation and <c+a> slip was observed either side of the defect normal to the loading direction compared to relatively little in the regions close to the defect in line with the loading direction. This HRDIC observation was consistent with finite element analysis of the expected strain fields around the defect both below and above the yield point. Furthermore, micro-cracks were observed in αp/αp and αp/βt interfaces in both regions with the more pronounced strain fields around the defect leading to an increased number of long micro-cracks than in the bulk. The formation mechanisms of micro-cracks have been discussed, emphasising the role of localised strain caused by the defect.</p> <p> </p>
An Audit of Machine Learning Experiments on Software Defect Prediction - Dataset
<p><strong>ML_Audit_20250328_anon.csv</strong>:<br>This CSV file contains anonymized data used in the audit of machine learning experiments on software defect prediction. The dataset includes variables and performance metrics extracted from studies published between 2019 and 2023. It supports the audit's evaluation of study reproducibility and issues related to experimental design and statistical analysis. This data can be used for replication and further analysis of the trends and reproducibility issues identified in the paper.</p> <p><strong>ML_Audit_March2025.Rmd</strong>:<br>This RMarkdown file contains the analysis script used for the statistical analysis and audit of the machine learning experiments reviewed in the study. It includes the procedures for data preprocessing, statistical evaluations, and reproducibility assessments. The script is integral for replicating the audit results presented in the paper and can be used by other researchers to perform similar audits or extend the analysis on different datasets.</p>
7-Dehydrocholesterol-derived oxysterols cause neurogenic defects in Smith-Lemli-Opitz syndrome
<p>Defective 3beta-hydroxysterol-delta<sup>7 </sup>-reductase (DHCR7) in the developmental disorder, Smith-Lemli-Opitz syndrome (SLOS), results in deficiency in cholesterol and accumulation of its precursor, 7-dehydrocholesterol (7-DHC). Here, we show that loss of <i>DHCR7</i> causes accumulation of 7-DHC-derived oxysterol metabolites, premature neurogenesis, and perturbation of neuronal localization in developing murine or human cortical neural precursors, both <i>in vitro</i> and <i>in vivo</i>. We found that a major oxysterol, 3b,5a-dihydroxycholest-7-en-6-one (DHCEO), mediates these effects by initiating crosstalk between glucocorticoid receptor (GR) and neurotrophin receptor kinase TrkB. Either loss of <i>DHCR7</i> or direct exposure to DHCEO causes hyperactivation of GR and TrkB and their downstream MEK-ERK-C/EBP signaling pathway in cortical neural precursors. Moreover, direct inhibition of GR activation with an antagonist or inhibition of DHCEO accumulation with antioxidants rescues the premature neurogenesis phenotype caused by the loss of <i>DHCR7</i>. These results suggest that GR could be a new therapeutic target against the neurological defects observed in SLOS.</p>
Exploring point defects and trap states in undoped SrTiO3 single crystals
<p>This dataset includes the raw data for the publication "Exploring point defects and trap states in undoped SrTiO3 single crystals" by Siebenhofer et. al, published in the Journal of the European Ceramic Society. The originlab file includes all raw data for the original figures presented in the main paper and the supporting information, organized in labelled folders containing tabelled data and the corresponding figure. </p>
ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction
<p><strong>ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction</strong></p> <p>This archive contains the <strong>ApacheJIT</strong> dataset presented in the paper "ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction" as well as the replication package. The paper is submitted to <strong>MSR 2022 Data Showcase Track</strong>.</p> <p>The datasets are available under directory <em>dataset</em>. There are 4 datasets in this directory.</p> <ol> <li><strong>apachejit_total.csv</strong>: This file contains the entire dataset. Commits are specified by their identifier and a set of commit metrics that are explained in the paper are provided as features. Column <em>buggy</em> specifies whether or not the commit introduced any bug into the system.</li> <li><strong>apachejit_train.csv</strong>: This file is a subset of the entire dataset. It provides a balanced set that we recommend for models that are sensitive to class imbalance. This set is obtained from the first 14 years of data (2003 to 2016).</li> <li><strong>apachejit_test_large.csv</strong>: This file is a subset of the entire dataset. The commits in this file are the commits from the last 3 years of data. This set is not balanced to represent a real-life scenario in a JIT model evaluation where the model is trained on historical data to be applied on future data without any modification.</li> <li><strong>apachejit_test_small.csv</strong>: This file is a subset of the test file explained above. Since the test file has more than 30,000 commits, we also provide a smaller test set which is still unbalanced and from the last 3 years of data.</li> </ol> <p>In addition to the dataset, we also provide the scripts using which we built the dataset. These scripts are written in Python 3.8. Therefore, Python 3.8 or above is required. To set up the environment, we have provided a list of required packages in file <em>requirements.txt</em>. Additionally, one filtering step requires GumTree [1]. For Java, GumTree requires Java 11. For other languages, external tools are needed. Installation guide and more details can be found <a href="https://github.com/GumTreeDiff/gumtree/wiki/Getting-Started">here</a>.</p> <p>The scripts are comprised of Python scripts under directory <em>src</em> and Python notebooks under directory <em>notebooks</em>. The Python scripts are mainly responsible for conducting GitHub search via GitHub search API and collecting commits through PyDriller Package [2]. The notebooks link the fixed issue reports with their corresponding fixing commits and apply some filtering steps. The bug-inducing candidates then are filtered again using <em>gumtree.py</em> script that utilizes the GumTree package. Finally, the remaining bug-inducing candidates are combined with the clean commits in the <em>dataset_construction</em> notebook to form the entire dataset.</p> <p>More specifically, <em>git_token</em> handles GitHub API token that is necessary for requests to GitHub API. Script <em>collector</em> performs GitHub search. Tracing changed lines and git annotate is done in <em>gitminer</em> using PyDriller. Finally, <em>gumtree</em> applies 4 filtering steps (number of lines, number of files, language, and change significance).</p> <p>References:</p> <p><strong>1. GumTree</strong></p> <ul> <li> <p><a href="https://github.com/GumTreeDiff/gumtree">https://github.com/GumTreeDiff/gumtree</a></p> </li> <li> <p>Jean-Rémy Falleri, Floréal Morandat, Xavier Blanc, Matias Martinez, and Martin Monperrus. 2014. Fine-grained and accurate source code differencing. In ACM/IEEE International Conference on Automated Software Engineering, ASE ’14,Vasteras, Sweden - September 15 - 19, 2014. 313–324</p> </li> </ul> <p><strong>2. PyDriller</strong></p> <ul> <li> <p><a href="https://pydriller.readthedocs.io/en/latest/">https://pydriller.readthedocs.io/en/latest/</a></p> </li> <li> <p>Davide Spadini, Maurício Aniche, and Alberto Bacchelli. 2018. PyDriller: Python Framework for Mining Software Repositories. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering(Lake Buena Vista, FL, USA)(ESEC/FSE2018). Association for Computing Machinery, New York, NY, USA, 908–911</p> </li> </ul>
The Effect of Bone Graft Substitute in Healing Fractures with Bone Defects Through Examination of Alkaline Phosphatase and Radiology in the Murine Model (Rattus norvegicus) Wistar strain
<p>Raw data for manuscript with the title <strong>The Effect of Bone Graft Substitute in Healing Fractures with Bone Defects Through Examination of Alkaline Phosphatase and Radiology in the Murine Model (<em>Rattus norvegicus</em>) Wistar strain </strong></p>
Dataset: Theory of defect-mediated morphogenesis
<p>This archive contains the experimental and numerical data produced and analyzed for the article "Theory of defect-mediated morphogenesis". The archive is divded in two main folders:<br> - Experimental_Data<br> - Simulation_Data<br> <br> The folder "Experimental_Data" contains the row microscopy pictures of the domes taken in MDCK cells experiments a partially shown in Fig. 1(D) of the main text and in the Supplementary Informations. Each gif file shows two superimposed channels, namely E-cadherin (green) and Nucleus (blue). The stack outside the sample and move towards the monolayer. The plane-to-plane distance is 0.172 um and one pixel corresponds to 0.3451 um.</p> <p>The folder "Simulation_Data" contains the configurations and the raw data produced by means of numerical simulations.<br> The content is divided in subfolders and classified according to the Figure and panels where the data are shown. Simulation parameters are reported in the main text and in the Supplementary Information of the paper "Theory of defect-mediated morphogenesis".</p> <p>Simulation configurations are saved as vtm files. Each vtm file contains the values at each grid points of the polarization (Px,Py,Pz) and velocity field (ux,uy,uz) as well as the concentration field (phi). Vtm files can be visualized through the open-source visualization software Paraview. In folder Simulation_Data/Fig_2_panel_H the time series of the L2 distance of the measured profile with respect to the flat interface are found. The files is organized as follows: 1st column shows the time (iteration) at which the measurement is performed, 2nd column shows the L2 distance.</p>
The Effect of Stromal Vascular Fraction (SVF) & Scaffolds Application on Fracture Healing with Bone Defect as Assessed Through Osteocalcin and Bone Morphogenetic Protein-2 (BMP-2) Biomarker Examination: Experimental Study on Murine Model
<p>This data is the raw data for the manuscript with titled The Effect of Stromal Vascular Fraction (SVF) & Scaffolds Application on Fracture Healing with Bone Defect as Assessed Through Osteocalcin and Bone Morphogenetic Protein-2 (BMP-2) Biomarker Examination: Experimental Study on Murine Model.</p>
Towards Developing and Analysing The Metric-Based Software Defect Severity Prediction Model
<p>This is a metric based approach to solve software defect severity prediction problem. In addition to that, this work proposes a new evaluation scheme that comprised of five metrics to analyze the performances.</p>
Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steel battery tabs
<p>In this folder, excel files are stored with the results of signal processing that supported findings in the following paper:</p> <p>"Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steell battery tabs".</p> <p>Matlab scripts and orginal signals will be uploaded soon with more detailed description.</p> <p> </p>
Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects"
<p>Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects".</p> <p>The root folder contains 5 folders:</p> <ol> <li>FIBERS</li> <li>MEMBRANES_and_MICELLES</li> <li>NANOPARTICLES</li> <li>COMPARISON</li> <li>paper_images</li> </ol> <p>The folders 1. to 3. contain the data for every soft-matters architecture used to produce the results discussed in the main paper. Each of these folders contain additional sub-fordels: TRAJ, SOAP, PCA, CLUSTERING, containing the files discussed in the main paper.</p> <p>Folder 4. contains the data of the comparison between different classes of materials (SOAP, PCA, and CLUSTERING sub-folders).</p> <p>Folder 5. contains the images that are showed in the main paper and in the Supporting Information.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.