Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
562
datasets available to search
ShareScore release 0.9.0
Dataset results
562 results for “Faults”
A Systematic Literature Review of Machine Learning for Uncovering Software Faults and Failures
<p>This data set contains the results of an extensive, systematic literature review on the use of machine learning (ML) for uncovering software faults and failures. Covering the period of 2019 to 2022, this literature review identifies 874 relevant publications, classified into six distinct quality assurance tasks. Results show a compound annual growth rate (CAGR) of relevant publications of 38% over the last five years.</p> <p>This literature review particularly analyzed in how far these relevant papers leverage synergies between different quality assurance tasks. Results show that only 3% of all relevant papers leverage such synergies, indicating ample opportunities for future research. For example, a single type of quality assurance activity may not suffice to deliver the expected software quality. Ideally, one would use a suitable combination of different types of activities – such as combining dynamic testing with static code analysis. Also, leveraging the synergies between different quality assurance activities can increase the effectiveness of the individual activities. For example, having a good estimate of the fault density of a software component (e.g., using deep learning-driven fault prediction techniques) could help optimize and prioritize testing effort and budget.</p>
Supplementary Datasets and Movies for the Paper "Major California faults are smooth across multiple scales at seismogenic depth"
<p>Supplementary Datasets and Movies for the Paper<br> <strong>M</strong><strong>ajor California faults </strong><strong>are</strong><strong> smooth </strong><strong>across</strong><strong> multiple scales </strong><strong>at </strong><strong>seismogenic </strong><strong>depth</strong><br> by Anthony Lomax and Pierre Henry</p> <p>DOI: <a href="https://doi.org/10.26443/seismica.v2i1.324">https://doi.org/10.26443/seismica.v2i1.324</a></p> <p>Movies S1-2 Seismicity along the central San Andreas fault zone around Parkfield as Figure 1 in the main text. Shows rotating, lateral view around ~S40°E for (Movie S1) NCSS-DD and (Movie S2) NLL-SSST-coherence. See Figure 1 caption in the main text for key to figure elements.</p> <p>Movie S3-7 Animated, rotating later views of NLL-SSST-coherence relocations of seismicity other than Parkfield presented in the main paper and this supplement.<br> Movie_S1_Parkfield_2022_sect_DD_movie_20230401.mp4 Movie_S2_Parkfield_2022_sect_NLL-SSST-coherence_movie_20230401.mp4 Movie_S3_S_Calaveras_2022_NLL-SSST-coherence_movie_20230401.mp4 – Southern Calaveras Fault Zone<br> Movie_S4_Mendocino_2021_NLL-SSST-coherence_movie_20230401.mp4 – Cape Mendocino<br> Movie_S5_MountLewis_1986_NLL-SSST-coherence_movie_20230401.mp4 – Mount Lewis<br> Movie_S6_SW_SanFrancisco_NLL-SSST-coherence_movie_20230401.mp4 - Southwest of San Francisco<br> Movie_S7_Calipatria_2021_NLL-SSST-coherence_movie_20230401.mp4 - Calipatria sequence</p> <p>Datasets S1-5 CSV format catalogs of NLL-SSST-coherence relocation results presented in the main text.<br> CSV fields correspond to data in the <a href="http://alomax.net/nlloc/soft7.00/formats.html#_location_hypphs_">NLLoc Hypocenter-Phase file</a></p> <ul> <li>ds01_Parkfield_2004_NLL_SSST_coherence.csv - Parkfield</li> <li>ds02_S_Calaveras_2022_NLL_SSST_coherence.csv – Southern Calaveras Fault Zone</li> <li>ds03_Mendocino_2021_NLL_SSST_coherence,csv – Cape Mendocino</li> <li>ds04_MountLewis_1986_NLL_SSST_coherence.csv – Mount Lewis</li> <li>ds05_SW_SanFrancisco_2021_NLL_SSST_coherence.csv - Southwest of San Francisco</li> <li>ds06_Calipatria_2021_NLL_SSST_coherence.csv - Calipatria sequence</li> </ul> <p>File S1 (project_run_scripts.zip) Archive of run scripts and related set-up, configuration and other meta-data files for locations cases presented in this paper.</p>
Microstructure of experimental faults in Pāpaku Fault core samples (IODP site U1518)
<p>Backscattered electron images of experimentally deformed samples from the Pāpaku Fault, Hikurangi Margin, New Zealand (International Ocean Discovery Program, Site U1518). In the experiments, intact mini-cores extracted from drill core samples were deformed in a single-direct shear box at MARUM, University of Bremen. Details of the experiments and the experimental data are available from the Pangea data publisher at: </p> <p>The data set contains original mosaics of whole thin sections, cut parallel to the shear direction and perpendicular to the experimental fault. These are in TIFF format and named SAMPLE#.tif.</p> <p>Annotated images include interpretations of the deformed zone (in red shading) superimposed on images that have been enhanced for better contrast. These are in Adobe Illustrator format and named SAMPLE#_annotated.ai.</p> <p>For image analysis, traces of the inferred deformed zone were extracted and scaled in ImageJ. These traces are available in TIFF format and named SAMPLE#_dz_trace.tif. Our measurements based on these traces are tabulated in Papaku_experiments_microstructure_data.csv.<br> </p>
The effect of normal stress oscillations on fault slip behavior near the stability transition from stable to unstable motion
<p>Tectonic fault zones are subject to normal stress variations with a wide range of spatio-temporal scales. Stress perturbations cover a wide range of frequencies and amplitudes from high frequency seismic waves generated by earthquakes to low frequency transients associated with solid Earth tides. These perturbations can reactivate critically stressed faults and trigger earthquakes. Here, we describe lab experiments to illuminate the physics of such changes in friction and the mechanics of earthquake triggering and fault reactivation. Friction tests were done in a double direct shear configuration for conditions near the stability transition from stable to unstable motion. We studied simulated fault gouge composed of quartz powder and conducted experiments at reference normal stress from 10 to 13.5 MPa. After shearing to steady state sliding, we applied sinusoidal normal stress oscillations of amplitude 0.5 to 2 MPa, and period of 0.5 to 50 s. We performed numerical simulations using measured values of rate/state friction (RSF) parameters to assess our data. Our results show that low frequency stress oscillations cause a Coulomb-like response of shear strength that transitions from stable slip to slow lab earthquakes as frequency increases. At the critical frequency predicted by RSF we observe periodic stick-slip behavior. Perturbations of high amplitude and short period weaken the fault, while lower amplitudes strengthen the fault. We find that a modified RSF formulation is able to accurately match our laboratory data. Our findings highlight the complex effects of stress perturbations for fault strength and the mode of fault slip.</p> <p>The data are uploaded are structured as follow:</p> <p>1) For each experiment a .txt file of the datafile that is recorded from the machine (raw data) and a binary file containing the elaborated data (data_rp). The experiments information are listed in experiment_info.txt</p> <p>2) The folder <a href="https://zenodo.org/api/files/89fe30fb-a2cb-4fcb-b9df-db80583fc652/codes_results.zip">codes_results.zip</a> contain the codes of the data analysis and the related results </p> <p>The data are analyzed using rawPy that can be found at <a href="https://github.com/marcoscuderi/rawPy">https://github.com/marcoscuderi/rawPy</a></p> <p>For any additional information please do not hesitate to contact the corresponding author Federico Pignalberi at federico.pignalberi@uniroma1.it</p>
Hubbard Brook Experimental Forest Fault Zones: GIS Shapefile
This coverage was obtained in digital form from Chris Barton of the USGS. Bedrock geology in the Hubbard Brook Valley was mapped by C.C. Barton, R.H. Comerlo, and S.W. Bailey, August 1994 to August 1995. The Map is entitled "BEDROCK GEOLOGIC MAP OF HUBBARD BROOK EXPERIMENTAL FOREST AND MAPS OF FRACTURES AND GEOLOGY IN ROADCUTS ALONG INTERSTATE 93, GRAFTON COUNTY, NEW HAMPSHIRE" and was approved for publication on August 28, 1995. Data distributed as shapefile in Coordinate system EPSG:26919 - NAD83 / UTM zone 19N
Dataset of Real Faults in Deep Learning Systems
<p>This is a replication package for the "Taxonomy of Real Faults in Deep Learning Systems" paper.</p> <p>The dataset contains information on all the issues and real faults gathered in the course of the study.</p> <p>The dataset consists of three main folders: Manual_Labelling, Interviews and Survey.</p> <p>Manual_Labelling</p> <p>In this folder we have placed all the files associated with our manual labelling process. It is divided in 3 subfolders: SO_init, GitHub_init and Analysed_Artifacts.</p> <p>GitHub_init</p> <p>The GitHub mining process is explained in detail in Section 3.1.1 of the paper. The list of initially mined GitHub projects for each framework is presented in the file {frameworkname}_init.csv. Each line in the file provides project name, the link to the project and its general information such as number of commits, issues, stars and etc. As part of our procedure, we removed projects that do not represent real software systems. For such projects we provide an explanation on why it should be excluded in the column "Comment". The files {frameworkname}_after_init_removal.csv present the list of remaining projects after the exclusion of such systems. The files {frameworkname}_top_100.csv list the top 100 projects we selected for our final analysis.</p> <p>As explained in Section 3.1.1, to identify relevant commits and issues we used a vocabulary of related terms. The complete list of 11,968 stemmed words is presented in the file vocabulary_init.csv. The final list of 105 relevant words obtained by the exclusion of the words that appear less than 10 times and the further manual analysis is provided in the file vocabulary_final.csv.</p> <p>SO_init</p> <p>The extraction procedure from StackOverflow is explained in detail in Section 3.1.2 of the paper. In the file query.rtf we provide the code of the query we have used to extract discussions related to each of the analysed frameworks (Keras, Torch, Tensorflow) from StackOverflow. As a result, we got a csv file that lists discussions for each of the frameworks (files keras.csv, torch.csv and tensorflow.csv).</p> <p>Analysed_Artifacts</p> <p>Our manual analysis was conducted in 6 rounds (Table 1 in the paper). We provide information about each round in a separate csv file round_{roundnumber}.csv. In each file we provide the following information: (1) <strong>Entity Type</strong> - whether it is a StackOverflow or GitHub artifact and which framework it corresponds to; (2) <strong>Link</strong> to the artifact; (3) <strong>Number of Evaluators</strong> - how many evaluators were assigned to this artifact; (4) <strong>Conflict</strong> - whether there was a conflict between evaluators when assigning tag to this artifact, has a value 0 for no and 1 for yes; (5) <strong>Evaluator1..4</strong> and <strong>Tag1..4</strong> - the ID of each evaluator and the tag provided by each of the assigned evaluators; (6) <strong>Final Tag</strong> - final tag assigned to the artifact; (7) <strong>Taxonomy Tag</strong> - the tag in the final taxonomy;</p> <p>The detailed information on the statistics of each round can be found in file stats.xlsx.</p> <p>Interviews</p> <p>The background information about our interview participants is presented in the file interview_participant_info.xlsx. The interview guide we used to conduct the semi-structured interviews is in the file interview_guide.docx. In the subfolder "Transcriptions" we provide transcribed versions of all 20 interviews.</p> <p>We provide the details of open coding process for the interviews in the file interview_open_coding.xlsx. Each row corresponds to a part of the interview text to which at least one of the evaluators has assigned a tag. Therefore, each row contains the following information: (1) <strong>Interview Num</strong> - the interview number; (2) <strong>Evaluator 1</strong> - tag provided by the first evaluator; (3) <strong>Evaluator 2</strong> - tag provided by the second evaluator; (4) <strong>Moderator Tag</strong> - tag assigned by the moderator; (5) <strong>Tag</strong> - the final taxonomy tag; (6) <strong>Status</strong> - whether this tag has been added to the final taxonomy or not ("A" for yes and "R" for no).</p> <p>In the file Interview_Tags.xlsx we provide aggregated information about the tags obtained from interviews. The column "Number" shows the number of times the tag was extracted from the interviews. Similarly to the previous file, the column "Status" shows whether the tag became part of the final taxonomy or not.</p> <p>Survey</p> <p>We provide the survey form we have used for our validation study in the file survey_form.pdf. The file with the information about participants and their answers to the survey questions are in the file participant_info_and_responses.xlsx. The percentages reported in Table 2 in the paper are calculated in this file (last rows with a bold font).</p>
Numerical modeling of the seismic cycle for normal and reverse faulting earthquakes in Italy
<p>Results of the numerical models expressed in terms of nodal stresses, strains and displacements.</p> <p>Data Set S1. Nodal values of the modelled displacements for the L’Aquila 2009 earthquake.</p> <p>Data Set S2. Nodal values of the modelled strain tensor for the L’Aquila 2009 earthquake.</p> <p>Data Set S3. Nodal values of the modelled stress tensor for the L’Aquila 2009 earthquake.</p> <p>Data Set S4. Nodal values of the modelled displacements for the Norcia 2016 earthquake.</p> <p>Data Set S5. Nodal values of the modelled strain tensor for the Norcia 2016 earthquake.</p> <p>Data Set S6. Nodal values of the modelled stress tensor for the Norcia 2016 earthquake.</p> <p>Data Set S7. Nodal values of the modelled displacements for the Emilia 2012 earthquake.</p> <p>Data Set S8. Nodal values of the modelled strain tensor for the Emilia 2012 earthquake.</p> <p>Data Set S9. Nodal values of the modelled stress tensor for the Emilia 2012 earthquake.</p>
Data for "Volcano-tectonic interactions at Sabancaya volcano, Peru: Eruptions, magmatic inflation, moderate earthquakes, and fault creep"
<p>Data and models presented in the paper "Volcano-tectonic interactions at Sabancaya volcano, Peru: Eruptions, magmatic inflation, moderate earthquakes, and fault creep". See file "README.txt" for detailed descriptions of each item.</p>
Geophysical data set for San Ramon Fault master section
<p>This data set is a multivariable analysis carried out in the San Ramon Fault (SRF) along a master section perpendicular to the main fault scarp. These data include: (1) a ~ 1 km long Electrical resistivity tomography (ERT) with a dipolo-dipolo configuration, 48 channels every 20 meters (named as the master section). (2) Differential GPS data with the location of the ERT electrodes and gravity stations. (3) Data of the gravity stations, with a Garmin GPS data. (4) Seismic data of an active experiment of 24 channels every 5 m, with several hammer strikes, which are explained in a text-document inside each seismic file (TRV_seismic_data and SRF_seismic_data). (5) Time-series of the Nakamura stations located along the ERT master profile. (6) The voltage decay curve of a TEM-station in the western edge of the master section, which is ready for modeling.</p> <p> </p>
New fault slip distribution for the 2010 Mw 7.2 El Mayor Cucapah earthquake based on realistic 3D finite element inversions of coseismic displacements using space geodetic data
<p>The .csv files included in this repository contain the data used in the numerical model as input, while the .txt file is the output (slip on a regular grid of points on the fault planes from the joint inversion of the geodetic datasets.</p>
Series arcing faults voltage and current signatures in a DC low power network.
<p>The dataset contains series arc faults voltage and current signatures in a DC low power network.</p> <p>The first apparatus performs a series arc fault by opening contacts (ReadMe_test 1). After breaking, copper electrodes are maintained at a constant gap. The measurements are performed with direct current and resisve/Inductive loads.</p> <p>The second apparatus performs contact breaking under chosen voltage and current supply condions (ReadMe_test2). After breaking, electrodes are maintained at a constant gap during electrical measurement. The measurements are performed with direct current and resisve loads by separating electrodes.<br> </p>
F.A.I.R. open dataset of brushed DC motor faults for testing of AI algorithms
<p>Practical research in AI often lacks of available and reliable datasets so the practitioners can try different algorithms. The field of predictive maintenance is particularly challenging in this aspect as many researchers don't have access to full-size industrial equipment or there is not available datasets representing a rich information content in different evolutions of faults.</p> <p>This dataset presents the evolution of typical faults (commutator, winding and brush wear) in inexpensive DC motors under extensive monitoring (vibration, temperature, voltage, current and noise). These motors exhibit a particularly short useful life when operating out of nominal conditions (from 30 minutes to 6 hours) which make them very interesting to test different signal processing algorithms and introduce students and researchers into signal processing, fault detection and predictive maintenance.</p> <p>The data-set comprises two main elements:</p> <ul> <li>A spreadsheet with the processing of each raw data file</li> <li>4 folders with raw data files in HDF5 format</li> </ul> <p>The spread sheet contains the following columns</p> <ul> <li>filename of the raw data file</li> <li>timestamp of the raw data file</li> <li>speed of the motor in rpm</li> <li>speed of the motor in Hz</li> <li>Current of the moter (A)</li> <li>Voltage supply (V)</li> <li>surface motor temperature (ºC)</li> <li>ambient temperature (ºC)</li> <li>For each measured signal (Vibration, current, voltage) the vibration of the main harmonic (at the speed of the motor) and it's first 10 multiple.</li> <li>For each measured signal (Vibration, current, voltage) the vibration in 4 bands: 0-4kHz, 4kHz-8kHz, 8kHz-16kHz, 16kHz-26kHz</li> </ul> <p>The raw data files in HDF5 format contains the instantaneous measured vibration (g), current (A) and voltage (volts) of the DC motor at 51.200 Hz.</p>
The "Castelluccio-Amatrice" low-angle normal fault seismic dataset
<p>Seismology data (from INGV database) for idetification the geometry of “Castelluccio-Amatrice” low-angle normal fault. The seismic data are the same of central Italy sesmic sequence in 2014-2017 period. </p>
System Call Logs with Natural Random Faults: Experimental Design and Application
<p>In this work we present a large dataset for use in embedded systems research which has been gathered from a realistic development environment operating in the path of an accelerated neutron beam. The dataset contains traces with events from a real-time operating system as as user events from a safety critical application. All data is carefully timestamped and in human understandable form.</p>
Fault
<p><strong>Overview of Data</strong></p> <p>Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs</p> <p><strong>Paper Abstract</strong></p> <p>Rather than tediously writing unit tests manually, tools can be used to generate them automatically – sometimes even resulting in higher code coverage than manual testing. But how good are these tests at actually finding faults? To answer this question, we applied three state-of-the-art unit test generation tools for Java (Randoop, EvoSuite, and Agitar) to the 357 real faults in the Defects4J dataset and investigated how well the generated test suites perform at detecting these faults. Although the automatically generated test suites detected 55.7% of the faults overall, only 19.9% of all the individual test suites detected a fault. By studying the effectiveness and problems of the individual tools and the tests they generate, we derive insights to support the development of automated unit test generators that achieve a higher fault detection rate. These insights include 1) improving the obtained code coverage so that faulty statements are executed in the first instance, 2) improving the propagation of faulty program states to an observable output, coupled with the generation of more sensitive assertions, and 3) improving the simulation of the execution environment to detect faults that are dependent on external factors such as date and time.</p>
Generic, Scalable and Decentralized Fault Detection for Robot Swarms
<p>This raw data archive includes the data on fault detection in a simulated swarm of 20 e-puck robots. The data was used in the paper Generic, Scalable and Decentralized Fault Detection for Robot Swarms by D. Tarapore et al. (2017).</p> <p>See readme.txt for more details.</p>
EAST-WEST and VERTICAL deformation maps of Alto Tiberina Fault supersite: The Post-Proc service in Geohazard Exploitation Platform applied to Sentinel-1 dataset
<p>In the framework of the ESA funded project “MEMpHIS - Multi Scale and Multi Hazard Mapping Space based Solutions ”, the Istituto Nazionale di Geofisica e Vulcanologia (INGV) together with TRE-ALTAMIRA, generated the EAST-WEST and VERTICAL deformation maps of Alto Tiberina Fault supersite. The two maps of ground velocity were derived thanks to the adoption of a specific tool named "Post-Proc", implemented in MEMPHIS and with the support of Terradue, that is able to automatically re-project on the east-west and vertical directions the ascending and descending InSAR time series. In particular, the output refers to the deformation maps calculated by processing with SqueeSAR (TM) method, a large dataset acquired by the ESA Sentinel-1 mission.</p> <p>The Post-Proc tool also calculates the mean accelerations associated to each persistent scatterer in the scenes. Some additional features are also available from this tool:</p> <p>- Change the coherence threshold for selecting a subset of persistent scatterers</p> <p>- Activate a geometrical distortion filter to take into account the layover and foreshortening effects</p> <p>- Change the reference point (position and coherence)</p> <p>- Choose between two types of accelerations: 2<sup>nd</sup> order model or velocity derivative</p> <p>- Choose among three different projection: east-west, vertical, and downslope.</p> <p>The present dataset is composed of the EAST-WEST and VERTICAL acceleration maps. The data are in shapefile format: for each record (PS) the topography, velocity, acceleration, and InSAR coherence is reported.</p>
Dataset of Optimization Methods for Model-Implemented Fault Injection in Cyber-Physical Systems: A Systematic Literature Review
<p>Data set for the paper entitled “<strong>Optimization Methods for Model-Implemented Fault Injection in Cyber-Physical Systems: a Systematic Literature Review</strong>”</p> <p>In this repo, we have some pictures and Excel files.</p> <ul> <li>Pictures are screenshots from the Parsifal tool (https://parsif.al/) which we use for performing the SLR.</li> <li>Excel files are as follows:</li> </ul> <table style="border-collapse: collapse; width: 100%;"><colgroup><col style="width: 21.8789%;"><col style="width: 78.1211%;"></colgroup> <tbody> <tr> <td><strong>Excel’s file name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Keyword_analysis </td> <td>In this file, you can see the evolution of our keyword selection.</td> </tr> <tr> <td>Articles_InclusionExclusion_QA </td> <td>In this file, you can find all found papers until Feb. 27, 2025. In the last column of this excel file, we can see the status of each paper, if it has been included, or excluded by authors. For the included paper (their status is “Accepted”) you can see their quality score in the last column.</td> </tr> <tr> <td>Extracted_data </td> <td>In this file, we logged the result of data extraction from qualified paper. In the first sheet “Articles”, you can see a list of the read papers with corresponding data. Other sheets in this Excel file are driven from the “Article” sheet for data visualization. So, they are not important.</td> </tr> </tbody> </table> <p> <br>If you have any questions, you can read the corresponding paper and contact the authors.</p>
Fault tree reliability analysis via squarefree polynomials
<p>Artefact for the paper "Fault tree reliability analysis via squarefree polynomials" by Milan-Lopuhaä-Zwakenberg, MODELSWARD 2024.</p>
Dataset: Evolution of large Venusian coronae inferred from structural analyses and the presence of low-angle faults in chasmata
<p>The vector datasets mentioned in this article are available. You can find the Magellan SAR data and the Global Topography Data Records (GTDR) on NASA's Planetary Data System (PDS) website (specific links can be found on https://pdsgeosciences.wustl.edu/missions/magellan/index.htm). In addition, the stereo-derived topography dataset of Herrick et al. 2012 is also available on their personal website https://sites.google.com/alaska.edu/robertherrick/resources/stereo-derived-topography-for-venus.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.