Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “fuzzing”
Artifacts for ASE 2022 Paper -- FuzzerAid: Grouping Fuzzed Crashes Based On Fault Signatures
<p><strong>Artifacts for FuzzerAid: Grouping Fuzzed Crashes Based On Fault Signatures</strong></p> <p>Fuzzing has been an important approach for finding bugs and vulnerabilities in programs. Many fuzzers deployed in industry run daily and can generate an overwhelming number of crashes. Diagnosing such crashes can be very challenging and time consuming. Existing fuzzers typically employ heuristics such as code coverage or call stack hashes to weed out duplicate reporting of bugs. While these heuristics are cheap, they are often imprecise and end up still reporting many "unique" crashes corresponding to the same bug. In this paper, we present <em>FuzzerAid</em> that uses <em>fault signatures</em> to group crashes reported by the fuzzers. Fault signature is a small executable program and consists of a selection of necessary statements from the original program that can reproduce a bug. In our approach, we first generate a fault signature using a given crash. We then execute the fault signature with other crash inducing inputs. If the failure is reproduced, we classify the crashes into the group labeled with the fault signature; if not, we generate a new fault signature. After all the crash inducing inputs are classified, we further merge the fault signatures of the same root cause into a group. We implemented our approach in a tool called <em>FuzzerAid</em> and evaluated it on 3020 crashes generated from 15 real-world bugs and 4 large open source projects. Our evaluation shows that we are able to correctly group 99.1% of the crashes and reported only 17 (+2) "unique" bugs, outperforming the state-of-the-art fuzzers.</p> <p> </p> <p><strong>Change log for v1.0.1:</strong></p> <p>Fix wrong Bug ID for <em>sqlite</em> and add README clarification.</p> <p><strong>Change log for v1.0.2:</strong></p> <p>Added an example linking data in the repository to the table.</p>
BigFuzz: Efficient Fuzz Testing for Data Analytics using Framework Abstraction
<p>BigFuzz supplementary material.</p>
Python scripts for input and post-processing of fuzz sputtering TRI3DYN simulations
<p>The influence of a fuzzy surface on the physical sputtering of Mo in He plasmas has been studied with hyperspectral imaging (HSI) measurements and simulations that couple the TRI3DYN code with an impurity transport code. The 2D profiles of the Mo I line emission intensity from HSI images reveal that the sputtering yield, Y, is reduced to ~40% of the smooth-surface value due to the presence of a fuzz layer, while the angular distribution of the sputtered Mo atoms might not change significantly. The simulations reproduce the Y reduction successfully, but indicate that fuzz causes an increase in the small-angle distribution of sputtered atoms. However, the increase is too small to produce an observable change in the Mo I emission profiles. A simple analytical model that assumes a single collision mean free path for a fuzz layer and considers only the primary sputtering events qualitatively reproduces the Y reduction and the small-angle distribution enhancement, explaining the geometrical effect of fuzz on physical sputtering.</p>
ToneTwist AFx Dataset: Custom Dynamic Fuzz
<div> <div> <h3><strong>Settings</strong></h3> <table> <tbody> <tr> <td>Gain</td> <td>Sensitivity</td> <td>Attack</td> <td>Release</td> <td>Volume</td> </tr> <tr> <td>5</td> <td>10</td> <td>1ms</td> <td>2500ms</td> <td>10</td> </tr> </tbody> </table> <h3><strong>Dry with markers</strong></h3> </div> </div> <p>Dry inputs are a selection of clean guitar and bass recordings from different sources:</p> <ul> <li><a href="https://www.idmt.fraunhofer.de/en/publications/datasets/guitar.html" target="_blank" rel="noopener">IDMT-SMT-GUITAR</a> - dataset 2 (7:23 min)</li> <li><a href="https://www.idmt.fraunhofer.de/en/publications/datasets/guitar.html" target="_blank" rel="noopener">IDMT-SMT-GUITAR</a> - dataset 4 - Career SG (6:08 min)</li> <li><a href="https://www.idmt.fraunhofer.de/en/publications/datasets/guitar.html" target="_blank" rel="noopener">IDMT-SMT-GUITAR</a> - dataset 4 - Ibanez 2820 (5:14 min)</li> <li><a href="https://www.idmt.fraunhofer.de/en/publications/datasets/bass_lines.html" target="_blank" rel="noopener">IDMT-SMT-Bass-Single-Track</a> - (5:58 min)</li> <li><a href="https://github.com/sdatkinson/neural-amp-modeler?tab=readme-ov-file#download-audio-files" target="_blank" rel="noopener">NAM: Neural Amp Modeler</a> - (3:11 min)</li> <li>Private Guitar Data - (5:19 min)</li> <li>YouTube Bass Recordings - (10:09 min)</li> </ul> <p>Pre-processing:</p> <ul> <li>All: <ul> <li>synchronization markers (2 impulses) added at start and end of every file</li> </ul> </li> <li>IDMT-SMT-GUITAR - dataset 2:<br> <ul> <li>peak normalized to -6dBFS</li> </ul> </li> <li>NAM:<br> <ul> <li>no pre-processing</li> </ul> </li> <li>Others:<br> <ul> <li>peak normalized to -0.1dBFS</li> <li>signal multiplied by random number every 5 seconds (uniform distribution [0.1, 1.0] = [-20dB, 0dB])</li> </ul> </li> </ul> <div> <h3><strong>Authors</strong></h3> <p><a href="https://mcomunita.github.io/" target="_blank" rel="noopener">Marco Comunità</a> - <a href="http://c4dm.eecs.qmul.ac.uk/" target="_blank" rel="noopener">Centre for Digital Music</a>, Queen Mary University of London</p> <h3><strong>Github</strong></h3> <p><a href="https://github.com/mcomunita/tonetwist-afx-dataset" target="_blank" rel="noopener">https://github.com/mcomunita/tonetwist-afx-dataset</a></p> <h3><strong>Reference</strong></h3> <p>If you make use of this dataset, please cite the following publications:</p> <pre><code>@inproceedings{comunita2023modelling, title={Modelling black-box audio effects with time-varying feature modulation}, author={Comunit{\`a}, Marco and Steinmetz, Christian J and Phan, Huy and Reiss, Joshua D}, booktitle={ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, pages={1--5}, year={2023}, organization={IEEE} }</code></pre> <pre> </pre> <pre><code>@misc{comunità2025nablafxframeworkdifferentiableblackbox, title={NablAFx: A Framework for Differentiable Black-box and Gray-box Modeling of Audio Effects}, author={Marco Comunità and Christian J. Steinmetz and Joshua D. Reiss}, year={2025}, eprint={2502.11668}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2502.11668}, }</code></pre> </div>
Python scripts for input and post-processing of fuzz sputtering TRI3DYN simulations
Open the record for dataset details and reuse information.
BigFuzz: Efficient Fuzz Testing for Data Analytics using Framework Abstraction
<p>BigFuzz supplementary material.</p>
Supplemental Material for "ATNwalk: A Novel Approach for Grammar-Based Coverage-Guided Fuzzing"
<p>Contains source code, original data, additional graphs, and other material to reproduce the experiments, which were described in the paper.</p> <ul> <li><strong>corpus.tar.gz</strong> contains the seed corpus for every fuzzing campaign</li> <li><strong>crashes.tar.gz</strong> contains the crashes of each fuzzing campaign, filtered according to the technique described in the paper</li> <li><strong>data.zip</strong> contains CSV files that contain AFL++ and GCOV metrics that were used to generate the plots in the paper; it also contains the additional plots of other metrics which were referenced as supplemental material</li> <li><strong>fuzzing_20221116.tar.gz</strong> is the docker image that was used to perform the fuzzing campaigns and serves as a runtime environment</li> <li><strong>home.rocky.tar.gz</strong> contains the home folder that is mounted inside the container (see README.md for details)</li> <li><strong>plots.zip</strong> contains all plots from the paper and additional ones, like branches over time, or other box plots for AFL++ covered bits</li> </ul> <p>Consult the README.md to read on how to repeat the experiments.</p>
Artifact from "A Little Goes a Long Way: Tuning Configuration Selection for Continuous Kernel Fuzzing"
<p>Artifact from "A Little Goes a Long Way: Tuning Configuration Selection for Continuous Kernel Fuzzing"</p>
The artifact for the paper "Demystifying the Dependency Challenge in Kernel Fuzzing" in ICSE 2022 Technical Tracks.
<p>This artifact is for the paper "Demystifying the Dependency Challenge in Kernel Fuzzing" in ICSE 2022 Technical Tracks.</p> <p>More detail and update please refer to https://github.com/ZHYfeng/Dependency.</p>
Evaluation Data for the paper "Demystifying the Dependency Challenge in Kernel Fuzzing" in ICSE 2022 Technical Tracks.
<p>Evaluation Data for the paper "Demystifying the Dependency Challenge in Kernel Fuzzing" in ICSE 2022 Technical Tracks.</p> <p> </p> <p>2.1 results of static analysis.zip<br> - Metadata collected when fuzzing (protobuf format)<br> - the related protobuf files are in https://github.com/ZHYfeng/Dependency/05-proto<br> - Information to support manual analysis (generated by metadata and static analysis)<br> - with name "0xffffffffaddress.txt", including every information about this uncovered address.</p> <p>2.2 results of manual analysis.zip<br> - Information of non-UD and UD for each module<br> - Overall statistic data, for example, coverage, prevalent of UD<br> - Sample cases and accuracy checking of sampled UD and its WS<br> - Root causes of sampled UDs</p> <p>2.3 example case.zip<br> - example cases of different root causes</p>
Artifacts for Fuzzing SMT Solvers with Diversified Sub-formulas
<p><strong>Fuzzing SMT solvers with Diversified Sub-formulas</strong></p> <p>Table of Contents</p> <ul> <li>Background</li> <li>Install</li> <li>Usage</li> <li>Bugs</li> </ul> <p> </p> <p><strong>Background</strong></p> <p><strong>Octopus </strong>is the tool for detecting soundness bugs in SMT solvers. <br> We have submitted 10 valid bug reports for Z3 so far. 7 bugs are confirmed/fixed by developers among these reports.</p> <p> </p> <p><strong>Install</strong></p> <p>Octopus itself has few dependencies. It uses Python3 and Python-virtualenv.</p> <p>You can install Python-virtualenv using <code>pip install virtualenv</code></p> <p>Then install Octopus.</p> <p><code>virtualenv --python=/usr/bin/python3.6 virenv</code></p> <p><code>source virenv/bin/activate</code></p> <p><code>cd octopus</code></p> <p><code>python3 setup.py install</code> </p> <p> </p> <p><strong>Usage</strong></p> <p>You can download SMT instances in SMT COMP 2021 as benchmarks.</p> <p><a href="https://www.starexec.org/starexec/secure/explore/spaces.jsp?id=1">2021-05-26 - StarExec</a></p> <p>Then install and build the SMT solver you want to test.</p> <p>For example:</p> <p><code>git clone https://github.com/Z3Prover/z3.git</code></p> <p><code>python scripts/mk_make.py </code></p> <p><code>cd build; make</code></p> <p>Then Octopus can be used to validate it, for example:</p> <p><code>octopus --benchmark=/home/SMT2021 --solver=z3 --solverbin=../z3/build/z3 --theory=LIA</code></p> <p>To run Octopus in multiple cores:</p> <p><code>octopus --benchmark=/home/SMT2021 --solver=z3 --solverbin=../z3/build/z3 --theory=LIA --cores=20</code></p> <p> </p> <p><strong>Bugs</strong></p> <p>Octopus has detected many new refutational soundness bugs in Z3.</p> <p>Here is a list of issues we reported.</p> <p><a href="https://github.com/Z3Prover/z3/issues/5373">https://github.com/Z3Prover/z3/issues/5373</a> [confirmed]<br> <a href="https://github.com/Z3Prover/z3/issues/5443">https://github.com/Z3Prover/z3/issues/5443</a> [reported]<br> <a href="https://github.com/Z3Prover/z3/issues/5447">https://github.com/Z3Prover/z3/issues/5447</a> [fixed]<br> <a href="https://github.com/Z3Prover/z3/issues/5456">https://github.com/Z3Prover/z3/issues/5456</a> [fixed]<br> <a href="https://github.com/Z3Prover/z3/issues/5457">https://github.com/Z3Prover/z3/issues/5457</a> [fixed] <br> <a href="https://github.com/Z3Prover/z3/issues/5460">https://github.com/Z3Prover/z3/issues/5460</a> [fixed]<br> <a href="https://github.com/Z3Prover/z3/issues/5468">https://github.com/Z3Prover/z3/issues/5468</a> [fixed] <br> <a href="https://github.com/Z3Prover/z3/issues/5488">https://github.com/Z3Prover/z3/issues/5488</a> [fixed]<br> <a href="https://github.com/Z3Prover/z3/issues/5502">https://github.com/Z3Prover/z3/issues/5502</a> [duplicate]<br> <a href="https://github.com/Z3Prover/z3/issues/5508">https://github.com/Z3Prover/z3/issues/5508</a> [reported]<br> <a href="https://github.com/Z3Prover/z3/issues/5423">https://github.com/Z3Prover/z3/issues/5423</a> [invalid] </p>
JMLKelinci+: Detecting Semantic Bugs and Covering Branches with Valid Inputs using Coverage-Guided Fuzzing and Runtime Assertion Checking
<p>Testing to detect semantic bugs is essential, especially for critical systems. Coverage-guided fuzzing and runtime assertion checking (RAC) are two well-known approaches for detecting semantic bugs. Coverage-guided fuzzing aims to generate inputs tests with high code coverage. However, while coverage-guided fuzzers are equipped with sanitizers that can detect a fixed set of semantic bugs, they can otherwise only detect bugs that lead to a crash. Thus, the first problem we address is how to help fuzzers detect previously unknown semantic bugs that do not lead to a crash. Moreover, a coverage-guided fuzzer may not necessarily cover all branches with valid inputs, although invalid inputs are useless for detecting semantic bugs. So, the second problem is how to guide a fuzzer to cover all branches in a program using only valid inputs. On the other hand, RAC monitors the expected behavior of a program dynamically and can only detect a semantic bug when a valid input test shows that the program does not satisfy its specification. <br> Thus, the third problem is how to provide high-quality input tests for a RAC that can trigger potential bugs.<br> The combination of a coverage-guided fuzzer and RAC solves these problems and can cover branches with valid inputs and detect semantic bugs effectively. Our study uses RAC to guarantee that only valid inputs reach the program under test using the program's specified preconditions and it also uses RAC to detect semantic bugs using specified postconditions. A prototype tool was developed for this study, named JMLKelinci+. Our results show that combining a coverage-guided fuzzer with RAC will lead to executing the program under test only with valid inputs and that this technique can effectively detect semantic bugs. <br> Also, this idea improves the feedback given to a coverage-guided fuzzer, enabling it to cover all branches faster in programs with non-trivial preconditions.</p>
Dataset of "Poster - BugOss: Regression Bug Benchmark for Empirical Study of Regression Fuzzing Techniques"
<p>Dataset of "Poster - BugOss: Regression Bug Benchmark for Empirical Study of Regression Fuzzing Techniques"</p>
Experiment data of Automated Program Repair from Fuzzing Perspective
<p>This dataset contains the result of experiment of the paper called Automated Program Repair from Fuzzing Perspective.</p> <p>Each zip file contains the results of each APR tool.</p> <p>Our artifact evaluation is in <a href="https://doi.org/10.5281/zenodo.7972926">https://doi.org/10.5281/zenodo.7972926</a>.</p>
MorFuzz: Fuzzing Processor via Runtime Instruction Morphing enhanced Synchronizable Co-simulation
<p>This deposit maintains the inputs generated by DifuzzRTL and MorFuzz binaries.</p> <p>Following is the paper abstract:</p> <p>Modern processors are too complex to be bug free. Recently, a few hardware fuzzing techniques have shown promising results in verifying processor designs. However, due to the complexity of processors, they suffer from complex input grammar, deceptive mutation guidance, and model implementation differences. Therefore, how to effectively and efficiently verify processors is still an open problem.</p> <p>This paper proposes MorFuzz, a novel processor fuzzer that can efficiently discover software triggerable hardware bugs. The core idea behind MorFuzz is to use runtime information to generate instruction streams with valid formats and meaningful semantics. MorFuzz designs a new input structure to provide multi-level runtime mutation primitives and proposes the instruction morphing technique to mutate instruction dynamically. Besides, we also extend the co-simulation framework to various microarchitectures and develop the state synchronization technique to eliminate implementation differences. We evaluate MorFuzz on three popular open-source RISC-V processors: CVA6, Rocket, BOOM, and discover 17 new bugs (with 13 CVEs assigned). Our evaluation shows MorFuzz achieves 4.4× and 1.6× more state coverage than the state-of-the-art fuzzer, DifuzzRTL, and the famous constrained instruction generator, riscv-dv.</p>
Empirical Study Data for Test Program-Based Generative Fuzzing for Differential Testing of the Kotlin Compiler
<p>Empirical Study Dataset for MSc. Thesis titled "Test Program-Based Generative Fuzzing for Differential Testing of the Kotlin Compiler". The data contains automatically generated Kotlin files, the results of differentially testing the generated files, and aggregated data containing file information, and compilation features.</p>
Replication Data for "Grammar-based fuzzing of data integration parsers in computational materials science"
Open the record for dataset details and reuse information.
HeteroFuzz: Fuzz Testing to Detect Platform Dependent Divergence for Heterogeneous Applications
<p>HeteroFuzz supplementary material.</p>
WASMOI - Controlled Experiment Fuzzing Results
<div> <div>This repository contains the fuzzing results from the controlled experiment from WASMOI project. <em>directed_over_undirected.py</em> and <em>undirected_over_directed.py</em> scripts generates the results for the statistical test presented in Table 1 in the paper. </div> <div> </div> <div>You can simply run: <code>python3 directed_over_undirected.py</code> and <code>python3 undirected_over_directed.py</code></div> <div> </div> <div>In order to generate figures presented in Figure 4 in the paper and more information regarding the output directories please refer to the README.md file for more information.</div> </div>
GrayC: Greybox Fuzzing of Compilers and Analysers for C
<p>This contains the data and the tool to run the experiments and process the data.</p> <ol> <li>Anon. Bug Reports are in file GrayC-Artefact.zip</li> <li>Comparison with other fuzzers data in file AE-data.zip</li> <li>GrayC and EnhanCer code in file GrayC-Artefact.zip</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.