Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,870

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,870 results for “defect”

Learn how ShareScore rates datasets ↗
zenodo40/100

Exploring Design Smells for Smell-Based Defect Prediction

<p>The archived file datasets.zip includes the datasets used for supporting the conclusions in the article <em>Exploring Design Smells for Smell-Based Defect Prediction.</em></p> <p>In this paper, we answer two research questions:</p> <p><strong>RQ1.</strong> Do Design code smells contribute to the performance of defect prediction models trained with Traditional code smells?</p> <p><strong>RQ2. </strong>How do the different categories of Design smells impact the performance of the defect prediction models?</p> <p>Therefore, after extracting the archived file documents, you will find two sub-directories, respectively named &quot;RQ1&quot; and &quot;RQ2&quot;. They include the results obtained for each one of the research questions, thus supporting our conclusions.</p> <p>(You will also find a README.pdf file with these same instructions regarding the datasets.)</p> <p>Inside &quot;RQ1,&quot; you will find two directories, respectively named &quot;configuration_1&quot; and &quot;configuration_2&quot;. They represent the different configurations for the experiments. <strong>&quot;configuration_1&quot;</strong> contains the datasets with results for the ten classifiers configurations with the highest scores and <strong>&quot;configuration_2&quot; </strong>contains the datasets with the results classifier configuration with the overall best results - Support Vector Machine with C=0.1. Furthermore, within each directory, there are three sub-directories, respectively named &quot;designite,&quot; &quot;designite_traditional,&quot; and &quot;traditional.&quot; These have the datasets for each of the considered smell sets in our study. Inside &quot;RQ2,&quot; you will find four directories. Each corresponds to a category from the design smells for the dataset &quot;designite_traditional.&quot; These datasets were build from the same configuration as &quot;configuration_2&quot;.</p> <p>Then, within every directory, there are 97 sub-directories representing the 97 projects analyzed in this study.</p> <p>Every project folder follows the same structure, which we define as follows.</p> <ul> <li>The &quot;dataset&quot; directory contains the original training and testing dataset used.</li> <li>The &quot;oversamples&quot; directory contains the training dataset after oversampling for each of the feature selection approaches.</li> <li>The &quot;score_summary&quot; directory contains all classifier configurations considered, not only the 10 with the highest scores.</li> <li>The &quot;scores.csv&quot; file contains all the scores for the main classifier configurations studied in the particular experiment.</li> <li>The &quot;selected_features&quot; directory contains the selected features&#39; information and the selected features dataset for each feature_selection method.</li> <li>The &quot;selected_testing_X&quot; directory contains the testing datasets.</li> <li>The &quot;top_scores_summary&quot; directory contains the classifier configurations and hyper-parameter scores for the top 10 highest scores.</li> </ul>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Sensor Defect Detection Datasets

<p><strong>Deprecated</strong> - Use https://zenodo.org/record/48728 for a more comprehensive version.</p> <p>&nbsp;</p> <p>Two datasets of sensor values, with each dataset including one defect sensor that delivers incorrect values. The datasets where gathered during tests in a hazardous material storage demonstrator.</p> <p>The datasets are given as comma-separated values in text files. The first line in each file holds time stamps, while the following lines hold the sensor values. The first entry in every line gives the name of the sensor.</p> <p>The first dataset (data_scenario_1.txt) was recorded under normal operating conditions, with the sensor Temperature_Inside_8 delivering incorrect values. In the second scenario there is a leakage of fluid inside the hazardous material storage. At the same time the sensor Smoke_Inside_0 delivers incorrect values.</p>

opencc-by-4.0Nov 2015View details →
zenodo40/100

Sensor Defect Detection Datasets with Configuration

<p>Two datasets of sensor values, with each dataset including one defect sensor that delivers incorrect values. The datasets where gathered during tests in a hazardous material storage demonstrator.</p> <p>The datasets are given as comma-separated values in text files. The first line in each file holds time stamps, while the following lines hold the sensor values. The first entry in every line gives the name of the sensor.</p> <p>The first dataset (data_scenario_1.csv) was recorded under normal operating conditions, with the sensor Temperature_Inside_8 delivering incorrect values. In the second scenario&nbsp;(data_scenario_2.csv) there is a leakage of fluid inside the hazardous material storage. At the same time the sensor Smoke_Inside_0 delivers incorrect values.</p> <p>Additionally attached is configuration data (Configurations.pdf) for the sensor fusion approach that was used to classify the datasets.</p> <p>For more information please contact the uploader.</p>

opencc-by-4.0Mar 2016View details →
zenodo40/100

Defect Prediction: Xerces

<p>Background: This paper describes an analysis that was conducted on newly collected repository with 92 versions of 38 proprietary, open-source and academic projects. A preliminary study performed before showed the need for a further in-depth analysis in order to identify project clusters. <br> Aims: The goal of this research is to perform clustering on software projects in order to identify groups of software projects with similar characteristic from the defect prediction point of view. One defect prediction model should work well for all projects that belong to such group. The existence of those groups was investigated with statistical tests and by comparing the mean value of prediction efficiency. <br> Method: Hierarchical and k-means clustering, as well as Kohonen’s neural network was used to find groups of similar projects. The obtained clusters were investigated with the discriminant analysis. For each of the identified group a statistical analysis has been conducted in order to distinguish whether this group really exists. Two defect prediction models were created for each of the identified groups. The first one was based on the projects that belong to a given group, and the second one - on all the projects. Then, both models were applied to all versions of projects from the investigated group. If the predictions from the model based on projects that belong to the identified group are significantly better than the all-projects model (the mean values were compared and statistical tests were used), we conclude that the group really exists. <br> Results: Six different clusters were identified and the existence of two of them was statistically proven: 1) cluster proprietary B – T=19, p=0.035, r=0.40; 2) cluster proprietary/open - t(17)=3.18, p=0.05, r=0.59. The obtained effect sizes (r) represent large effects according to Cohen’s benchmark, which is a substantial finding. <br> Conclusions: The two identified clusters were described and compared with results obtained by other researchers. The results of this work makes next step towards defining formal methods of reuse defect prediction models by identifying groups of projects within which the same defect prediction model may be used. Furthermore, a method of clustering was suggested and applied.</p>

opencc-by-4.0Jul 2010View details →
zenodo40/100

Dataset from 'Topological interfaces crossed by defects and textures of continuous and discrete point group symmetries in spin-2 Bose-Einstein condensates'

<p>Dataset associated with the publication 'Topological interfaces crossed by defects and textures of continuous and discrete<br>point group symmetries in spin-2 Bose-Einstein condensates' in Physical Review Research.&nbsp; <br>Source data for Figures 3-8 in the manuscript.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

A Dataset for Corrosion Defects in Pipelines

<p>This dataset is helpful for the automated system's navigation part and can help with decision-making by correctly modeling how corrosion spreads in steel pipes in real life. This dataset can also be used with any system that has pipe damage.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Radiation-induced point defects in KBr: CRYSTAL calculations

<p>Calculation results on vibrational properties of radiation-induced point defects in potassium bromide, calculated using CRYSTAL17 software.</p>

opencc-by-4.0Mar 2024View details →
dryad40/100

Data for: Dysregulation of mTOR signaling mediates common neurite and migration defects in both idiopathic and 16p11.2 deletion autism neural precursor cells

<p>Autism spectrum disorder (ASD) is defined by common behavioral characteristics, raising the possibility of shared pathogenic mechanisms. Yet, vast clinical and etiological heterogeneity suggests personalized phenotypes. Surprisingly, our iPSC studies find that six individuals from two distinct ASD subtypes, idiopathic and 16p11.2 deletion, have common reductions in neural precursor cell (NPC) neurite outgrowth and migration even though whole genome sequencing demonstrates no genetic overlap between the datasets. To identify signaling differences that may contribute to these developmental defects, an unbiased phospho-(p)-proteome screen was performed. Surprisingly, despite the genetic heterogeneity, hundreds of shared p-peptides were identified between autism subtypes including the mTOR pathway. mTOR signaling alterations were confirmed in all NPCs across both ASD subtypes and mTOR modulation rescued ASD phenotypes and reproduced autism NPC-associated phenotypes in control NPCs. Thus, our studies demonstrate that genetically distinct ASD subtypes have common defects in neurite outgrowth and migration which are driven by the shared pathogenic mechanism of mTOR signaling dysregulation.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Simulation and laboratory eddy current testing data - modelling compound defects via perturbation theory

<p>This dataset serves to fit and validate a perturbation approach to the modelling eddy current signals of compound defects. It was obtained during the AIFRI project (Artificial Intelligence for Rail Inspection). The simulation data was generated with the Faraday software by INTEGRATED Engineering Software, using its BEM Solver. The simulation data is supplied as csv. The laboratory data was gathered by Rainer Pohl in the eddy current laboratory of BAM, section 8.4. It is supplied in the DICONDE data format. The first frame of the pixel array in the DICONDE files corresponds to the real part and the second frame corresponds to the imaginary part of the signal. The data set is analyzed in an upcoming article.</p> <p><span>&nbsp;</span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Dataset of the publication: Liquid‐Phase Fabrication of Janus 2D Materials: Defect‐Rich MoS2 Ultrathin Layers Asymmetrically Decorated with Au Nanoparticles. Small 2024, 2406599

<p>Dataset of the publication: Liquid‐Phase Fabrication of Janus 2D Materials: Defect‐Rich MoS2 Ultrathin Layers Asymmetrically Decorated with Au Nanoparticles</p> <p>N. V. Vassilyeva, A. Forment-Aliaga, E. Coronado, <em>Small</em> <strong>2024</strong>, e2406599.</p> <p>doi: 10.1002/smll.202406599</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

X-ray linear dichroic tomography of crystallographic and topological defects

<p>Open Data for "X-ray linear dichroic tomography of crystallographic and topological defects" published in <a href="https://www.nature.com/articles/s41586-024-08233-y">Nature <strong>636</strong>, 354 (2024) </a></p> <div>&nbsp;</div> <div>Full citation:</div> <div>A. Apseros, V. Scagnoli, M. Holler, M. Guizar-Sicairos, Z. Gao, C. Appel, L. J. Heyderman, C. Donnelly &amp; J. Ihli&nbsp;</div> <div>X-ray linear dichroic tomography of crystallographic and topological defects.</div> <div><em>Nature <strong>636</strong>, 354</em> (2024).</div> <div>https://www.nature.com/articles/s41586-024-08233-y</div>

opencc-by-4.0Dec 2024View details →
zenodo40/100

Research Data Supporting "Controlling Exchange Pathways in Dynamic Supramolecular Polymers by Controlling Defects"

<p>This repository contains the set of data shown in the paper &quot;<strong>Controlling Exchange Pathways in Dynamic Supramolecular Polymers by Controlling Defects</strong>&quot;, published on <strong>ACS Nano</strong> (DOI: 10.1021/acsnano.1c01398).</p> <p>The data stored in this section is organized as follow:</p> <p><strong>supervisedClustering/ :</strong> folder containing scripts for carrying out the supervised analysis as discussed in the main paper. Subfolders are divided in <strong><code>kmeans</code></strong><code><strong>/</strong></code> and <strong><code>spectral</code></strong><code>/</code> contain the different analysis as explained in the paper.</p> <p><strong>unbiasMD/ : </strong>folder containing the Gromacs input files for reproducing the unbias MD trajectories of the 3 types of fiber, as explained in the paper.</p> <p><strong>biasMD/</strong> <strong>:</strong> folder containing the PLUMED input files for reproducing enhanced sampling dynamics as discussed in the paper.</p> <p><strong>unsupervisedClustering/ : </strong>folder containing the directions and the original files to reproduce the unsupervised clustering results.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Dataset: HRSTEM Images of Defective and Non-Defective Quasi-Periodic Materials

<p>This is the image dataset and model used to produce the results reported in the following publication:&nbsp;<br> Dennler, N., Foncubierta-Rodriguez, A., Neupert, T., Sousa, M. (2021). <em>Learning-based defect recognition for quasi-periodic HRSTEM images</em>. Micron, 146(July 2020), 103069. https://doi.org/10.1016/j.micron.2021.103069<br> <br> For questions, please correspond with N. Dennler (n.dennler2<strong> </strong>at<strong> </strong>herts.ac.uk) or with M. Sousa (sou at zurich.ibm.com).</p> <p><strong>hrstem_defects_dataset.zip:</strong> These are the images and labels used to develop and test the algorithm proposed in the above-mentioned publication. They correspond to high resolution scanning transmission electron microscopy images obtained for various III-V films, namely InP, GaAs, InGaAs and InAlGaAs using a JEOL ARM200F microscope. The raw images have been converted in .tif format with the GMS 3 program from Digital Micrograph. The labels have been created by a microscopy expert. Black: main crystal symmetry (non-defective). Gray: secondary crystal symmetry (symmetry defect). White: blurred (amorphous region or beam defect)<br> <br> <strong>vgg16.zip: </strong>The trained neural network model as well as a detailed description of the training/testing dataset that was used to achieve the results reported in the above-mentioned publication.</p>

opencc-by-4.0May 2021View details →
zenodo40/100

Dataset Related to article "NEURODEVELOPMENTAL DISORDER AND LATE-ONSET DEGENERATIVE PARKINSONISM IN A PATIENT WITH A WDR45 DEFECT"

<p>NGS data (.vcf; .bam; .bam.bai; .csv; .fastq) of patient analysed and sanger sequencing (.abi) for confirm the mutation.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Imaging topological defects in a non-collinear antiferromagnet

<p>Data related to the publication, arXiv 2202.02243.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Structural Defects Improve the Memristive Characteristics of Epitaxial La0.8Sr0.2MnO3-based Devices

<p>Data files (.txt) of the data presented in the publication&nbsp;<a href="https://doi.org/10.1002/admi.202200498">10.1002/admi.202200498</a></p> <p>File names are referred to as the reference of the Figure and number in the main manuscript, followed by the sample (LSM/STO or LSM/LAO)&nbsp;and keywords of the figure.</p> <p>Each file corresponds only to the data of one sample.</p> <p>Regarding the series in a&nbsp;Figure, in the case there is very little data, the series are in&nbsp;a single file, arranged in columns. Otherwise, the&nbsp;series with large amounts of data have been split into single files named according to the series (legend).</p> <p>All the files include two headers (Name and Units) for each column. The files containing data from multiple series contain an extra header to precise the data series.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Dataset for "Engineering defect clustering in diamond-based materials for technological applications via quantum mechanical descriptors"

<p>The unique set of extreme physical properties makes diamond an ideal candidate for applications in the energy industry such as in high-power and high-frequency electronics as well as in electrochemistry and photovoltaics. Furthermore, dopant-vacancy complexes in diamond can be exploited for further development of quantum computers, single-photon emitters, high-precision magnetic field sensing and nanophotonic devices. While certain dopant-vacancy complexes are well-studied, studies of other dopant/vacancy clusters are focused mostly on defect detection while investigations on how to tune their electronic and optical properties for specific applications is mostly omitted. To this aim, we attempted to reveal coupled structural-electronic features and their effect on the band gap of such defects through first principle calculations. We investigated four different defect types: a) dopant-vacancy complexes (X-V), b) two dopants as nearest neighbours (X-X), c) two dopants separated by one carbon atom (X-C-X) and d) two dopants separated by a vacancy (X-V-X). For each of these configurations, we considered Al, B, N, P and Si as dopant atoms. This dataset contains input files needed to reproduce every ground state geometry used in our study.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Supplementary data for "Vacancy defect configurations in the metal-organic framework UiO-66: Energetics and electronic structure"

<p>Optimised structures for UiO-66 with various defect configurations in VASP POSCAR format. For naming see the associated paper (DOI:10.1039/c7ta11155j).</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

The loss of SMG1 causes defects in quality control pathways in Physcomitrella patens

<p>Nonsense-mediated mRNA decay (NMD) is important for both RNA quality control and gene regulation. NMD targets aberrant mRNA transcripts for decay and also directly influences the abundance of specific, non-aberrant transcripts in a wide range of eukaryotes. In animals, the PIK kinase, SMG1, plays an essential role in NMD, by phosphorylating the core NMD effector, UPF1. Despite being absent from the genome of the model plant, <em>Arabidopsis thaliana</em>, SMG1 is ubiquitous throughout the plant kingdom. Here we utilize RNA-seq to reveal the full range of processes involving <em>SMG1</em> in plants. In NMD-compromised, <em>smg1</em> mutant moss, 30% of multi-isoform genes produce NMD targeted transcript isoforms. Taking a machine learning approach, we show that an exon-exon junction downstream of a stop codon acts as the major feature to target transcripts to NMD and that retained intron isoforms are underrepresented among NMD targets. Furthermore, we show that <em>SMG1</em> is involved in other quality control pathways, affecting DNA repair and the unfolded protein response, in addition to its role in mRNA quality control. <em>smg1</em> plants have increased susceptibility to DNA damage, but increased tolerance to unfolded protein inducing agents. The involvement of SMG1 in RNA, DNA and protein quality control has major implications for the study of these processes in plants.</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features

<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer&rsquo;s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit&nbsp;<a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record