Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “ASE”
Artifacts for ASE 2022 Paper -- FuzzerAid: Grouping Fuzzed Crashes Based On Fault Signatures
<p><strong>Artifacts for FuzzerAid: Grouping Fuzzed Crashes Based On Fault Signatures</strong></p> <p>Fuzzing has been an important approach for finding bugs and vulnerabilities in programs. Many fuzzers deployed in industry run daily and can generate an overwhelming number of crashes. Diagnosing such crashes can be very challenging and time consuming. Existing fuzzers typically employ heuristics such as code coverage or call stack hashes to weed out duplicate reporting of bugs. While these heuristics are cheap, they are often imprecise and end up still reporting many "unique" crashes corresponding to the same bug. In this paper, we present <em>FuzzerAid</em> that uses <em>fault signatures</em> to group crashes reported by the fuzzers. Fault signature is a small executable program and consists of a selection of necessary statements from the original program that can reproduce a bug. In our approach, we first generate a fault signature using a given crash. We then execute the fault signature with other crash inducing inputs. If the failure is reproduced, we classify the crashes into the group labeled with the fault signature; if not, we generate a new fault signature. After all the crash inducing inputs are classified, we further merge the fault signatures of the same root cause into a group. We implemented our approach in a tool called <em>FuzzerAid</em> and evaluated it on 3020 crashes generated from 15 real-world bugs and 4 large open source projects. Our evaluation shows that we are able to correctly group 99.1% of the crashes and reported only 17 (+2) "unique" bugs, outperforming the state-of-the-art fuzzers.</p> <p> </p> <p><strong>Change log for v1.0.1:</strong></p> <p>Fix wrong Bug ID for <em>sqlite</em> and add README clarification.</p> <p><strong>Change log for v1.0.2:</strong></p> <p>Added an example linking data in the repository to the table.</p>
Linked collectors and determiners for: ASE - Herbário da Universidade Federal de Sergipe.
Natural history specimen data linked to collectors and determiners held within, "ASE - Herbário da Universidade Federal de Sergipe". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/76823beb-2f33-47a6-96a1-ccd2df47ac28">https://bionomia.net/dataset/76823beb-2f33-47a6-96a1-ccd2df47ac28</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/76823beb-2f33-47a6-96a1-ccd2df47ac28">https://gbif.org/dataset/76823beb-2f33-47a6-96a1-ccd2df47ac28</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: ASE herbarium - Universidade Federal de Sergipe - Herbário Virtual REFLORA.
Natural history specimen data linked to collectors and determiners held within, "ASE herbarium - Universidade Federal de Sergipe - Herbário Virtual REFLORA". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/ba06a193-82fc-425b-a7e5-f3df809fe378">https://bionomia.net/dataset/ba06a193-82fc-425b-a7e5-f3df809fe378</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/ba06a193-82fc-425b-a7e5-f3df809fe378">https://gbif.org/dataset/ba06a193-82fc-425b-a7e5-f3df809fe378</a>. Formatted as a Frictionless Data package.
The dataset of the ASE'20 paper titled "Automated Patch Correctness Assessment: How Far are We?"
<p>This is the experiment result of the ASE'20 paper titled "<strong>Automated Patch Correctness Assessment: How Far are We?</strong>".</p> <p>If you use our data for academic research, please cite our paper as:</p> <pre><code class="language-html">@inproceedings{wang2020automated, title={Automated Patch Correctness Assessment: How Far are We?}, author={Wang, Shangwen and Wen, Ming and Lin, Bo and Wu, Hongjun and Qin, Yihao and Zou, Deqing and Mao, Xiaoguang and Jin, Hai}, booktitle={Proceedings of the 35th International Conference on Automated Software Engineering (ASE)}, year={2020}, organization={ACM} }</code></pre> <p>The file <strong><em>Patches.zip</em></strong> includes all the patches we take into consideration in this study. Note that 269 patches come from "<a href="http://arxiv.org/pdf/1909.13694">Automated Patch Assessment for Program Repair at Scale</a> (Ye et al.), Technical report 1909.13694, arXiv, 2019".</p> <p>The file <em><strong>Patches_for_Static</strong></em> include all the class files we used for static method.</p> <p>The file <em><strong>Tests-oracle </strong></em>includes all the test cases generated by <strong>Evosuite</strong> and <strong>Randoop</strong> on the fixed version programs.</p> <p>The file <em><strong>Tests-buggy</strong></em> includes all the test cases generated by <strong>Evosuite</strong> and <strong>Randoop</strong> on the buggy version programs.</p> <p>The file <em><strong>DiffTGen-result</strong></em> includes ingredients and output information of <strong>DiffTGen</strong>.</p> <p>The file <em><strong>Daikon-output</strong></em> includes inferred invariants of each patch and its corresponding ground-truth.</p> <p>The file <em><strong>PATCH-SIM_result</strong></em> includes the output vector files from <strong>PATCH-SIM</strong> and<strong> E-PATCH-SIM</strong>.</p> <p>The file <em><strong>Training_result</strong></em> includes the output of six ML algorithms with or without oracle.</p> <pre><code class="language-xml">Chart: 1-26; Closure: 14, 18, 31, 33, 38, 40, 57, 62, 63, 70, 73, 86, 92, 93, 115, 123, 126; Lang: 6, 7, 10, 16, 20, 21, 22, 24, 26, 27, 33, 35, 38, 39, 41, 43, 44, 45, 50, 51, 55, 57, 58, 59, 60, 61, 63; Math: 2, 3, 4, 5, 6, 8, 20, 22, 25, 28, 30, 31, 32, 33, 34, 35, 39, 41, 49, 50, 53, 56, 57, 58, 59, 60, 61, 63, 65, 68, 70, 71, 73, 74, 75, 79, 80, 81, 82, 85, 86, 88, 89, 90, 93, 97, 98, 99, 104; Time: 4, 7, 11, 14, 15, 19. </code></pre> <p>Please note that for bugs in the above table, the Evosuite tests on the fixed version programs are reused from <a href="https://arxiv.org/abs/1909.13694">a previous study</a>. We thank <strong>He Ye</strong>, <strong>Matias Martinez</strong>, and <strong>Martin Monperrus</strong> so much for sharing their data.</p> <p> </p> <p><strong>Notice!</strong> For patches under the folder <em>Patches_ICSE</em>, those under <em>Ddifferent</em> and <em>Dsame</em> folders are all correct patches. <em>Different</em> and <em>Same</em> only indicate whether the patch is syntactically identical to the ground truth patch.</p> <p>Patches generated for Mockito project (2 in total): Kali-A-Mockito-10; Arja-Mockito-10</p> <p>Patches do not pass plausibility check (6 in total): Kali-Closure-133; kPAR-Chart-12; FixMiner-Chart-12; patch1-Lang-6-SketchFix-plausible; patch2-Lang-6-SketchFix-plausible; patch1-Math-2-SOFix</p> <p>Patches that are mistakenly labeled (12 in total): patch2-Lang-51-Jaid; patch1-Lang-43-CapGen; patch2-Lang-43-CapGen; patch2-Math-53-CapGen; patch2-Math-53-Jaid; jKali-Lang-7; ACS-Lang-35; Arja-Math-35; SimFix-Math-72; SimFix-Closure-19; Arja-Math-50; SimFix-Lang-60</p> <p>Detailed reasons for the mislabeled patches: 1. the ground-truth patch modifies multiple locations while the generated patch only modifies one of them (2/12, SimFix-Math-72, SimFix-Lang-60); 2. the edit points in the generated patch are different from those in ground-truth patch (8/12, patch2-Lang-51-Jaid, patch2-Math-53-Jaid, patch1-Lang-43-CapGen, patch2-Lang-43-CapGen, patch2-Math-53-CapGen, ACS-Lang-35, SimFix-Closure-19, Arja-Math-50); 3. the generated patch doesnot fulfill the intended function in ground-truth (2/12, jKali-Lang-7, Arja-Math-35).</p> <p>Take <em>Arja-Math-50</em> as an example, this patch deletes a conditional statement which deals with an unexpected input (<em>null</em>) in the method <em><strong>verifyBracketing</strong></em>.<em><strong> </strong></em>However, in the oracle program, this conditional statement still exists. Then, <strong>Randoop</strong> generated a test case by calling <strong><em>verifyBracketing</em></strong> with a <em>null</em> argument. This test passed on the ground-truthpatch but failed on the patch generated by Arja due to the removeof the exception handling statements. As a result, this patch is actually overfitting but mistakenly labeled as correct. We have confirmed this case with Kui Liu, the first author of the recent ICSE'20 paper (Title: <em>On the Efficiency of Test Suite based Program Repair</em>) which makes up our patch benchmark.</p> <p> </p> <p>Border line Patches (3 in total): ACS-Lang-7; kPAR-Lang-7; TBar-Lang-7. Reasons for overfitting: Evosuite generates some tests that fail on those patches, e.g., test049 in Seed 1; the Java documentation above the function states that it needs to deal with the situation where the input cannot be converted. Reasons for correct: it synthesizes the correct modification; currently, in the program, <em>createBigDecimal()</em> is not called directly in other part of the production code except <em>createNumber()</em> and the test code. In our paper, we consider these three patches as correct and that's why Evosuite has 3 false positives.</p>
Open anonymous repo hosting code and data for our submission in ASE 2020
<p>This repository presents sample publicly available anonymous source code and data for our submission in ASE 2020 conference.</p> <p>ProgressDroid source code is provided.</p> <p>Data for 10 top apps with the highest number of installs from our dataset are presented.</p> <p>For each app, we provide the following information:</p> <p>- The original APK file for the examined app.</p> <p>- The instrumented APK file using the extended Instrumenter module</p> <p>- Complete trace from running the extended AndroidSlicer tool on each app</p> <p>- Complete list of all UI update points in each specific app</p> <p>- List of slicing criteria for dynamic slicing </p> <p>- List of slices from the automated dynamic slicing analysis </p> <p>- List of all progress indicator occurrences for each trace </p> <p>- Complete runtime trace info including all events and states (including screenshots) </p> <p>Upon acceptance, we’ll complete the data sharing for all our dataset.</p>
ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities
<p>The dataset of paper---ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities.</p> <p>It includes 14,053 and 21,765 contributors, as well as 23,739 and 33,454 issues from the PyTorch and TensorFlow communities. </p>
ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities
<p>The dataset of paper---ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities.</p> <p>It includes 14,053 and 21,765 contributors, as well as 23,739 and 33,454 issues from the PyTorch and TensorFlow communities. </p>
ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities
<p>The dataset of paper---ASE_2022_A Tale of Two Cities: an Empirical Study on Deep Learning OSS Communities.</p> <p>It includes 23,739 and 33,454 issues from the PyTorch and TensorFlow communities. </p>
Test Instrument Example for ASE'18 paper 'Assessing the Type Annotation Burden" : Corrected Question Numbers
<p>Artifact for ASE'18 paper "Assessing the Type Annotation Burden". This PDF shows an example test instrument (a Qualtrics survey) with all 20 questions (20 code artifacts).</p> <p> </p> <p>A picture of the options in the drop-down box is shown on the last page (note: the order of elements in the drop-down was randomized for each test.</p>
Reproduction Package for ASE 2023 Submission `Improving Verification through Compiler Optimizations'
<p><strong>Artifact</strong></p> <p>In order to run this artifact please clone <a href="https://github.com/sosy-lab/sv-benchmarks">sv-benchmarks</a> into this folder. Afterwards you can just run the program using the running instructions down below. The results used in the paper can be found in the folder <code>transformation-for-verification-data</code>.</p> <p><strong>Setup</strong></p> <p>In order to setup this project, first initialize the submodules or clone this repository with the flag <code>--recursive</code>. Afterwards execute <code>python3 src/setup.py</code> in this directory, in order to add the required files to the submodules.</p> <p><strong>Running</strong></p> <p>In order to run this locally inside <code>./src</code></p> <pre><code>./main_bench.py --specification specification/path.prp program/to/verify.c</code></pre> <p>For example:</p> <pre><code>./main_bench.py --specification ../setup-files/test-run/unreach-call.prp ../setup-files/test-run/test_program.c </code></pre> <p>In order to execute with benchexec, execute the following insider <code>./src</code>, after adapting <code>bench.xml</code> to suit your purposes:</p> <pre><code>./benchmark_local.sh</code></pre>
CommitChronicle dataset from the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023
<pre>This is the CommitChronicle dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. For further details, see the attached README.md. <strong>Note.</strong> Also available on HuggingFace Hub: <a href="https://huggingface.co/datasets/JetBrains-Research/commit-chronicle">JetBrains-Research/commit-chronicle</a></pre>
ASE Database of CO2 Electro Capture on Redox-Active Metal-Organic Frameworks
<p>ASE database containing simulated structures for 1D, 2D and 3D conductive metal-organic frameworks.</p>
ASE_2022_How do code contexts evolve for software development tasks
<p>The dataset of paper---ASE_2022_How do code contexts evolve for software development tasks.</p> <p>It includes (1) working periods: the 1,375 interaction histories of development tasks and 4,219 working periods, and (2) results: results of our research and study.</p> <p>See README.md for more information.</p>
Artifacts for ASE 2022 Paper Submission # 1095
<p><strong>This data set is for ASE 2022 Paper Submission #1095</strong></p>
Dataset for ASE'24 Efficient Slicing of Feature Models via Projected d-DNNF Compilation
<p>This dataset includes various feature models as dimacs. Each dimacs includes a header indicating variables to be projected.</p> <p>The dataset was used for evaluating pd4 within the work Efficient Slicing of Feature Models via Projected d-DNNF Compilation at ASE'24.</p>
Replication Package for ASE 2023 Paper "Personalized First Issue Recommender for Newcomers in Open Source Projects"
<p>This replication package contains a replication package for ASE 2023 paper titled "Personalized First Issue Recommender for Newcomers in Open Source Projects." This package includes a dataset of 68,858 issues from 100 GitHub projects, records of 123 manually labeled issue samples, and Python scripts for analyzing the data and evaluating models. The package is also stored in the GitHub repository <a href="https://github.com/mcxwx123/PFIRec">https://github.com/mcxwx123/PFIRec</a>.</p> <p>Required Environment</p> <p>We recommend setting up the required environment on a commodity Linux machine with at least 1 CPU Core, 8GB Memory, and 100GB empty storage space. Our experiments were conducted on an Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>Files and Replicating Results</p> <p>We used the GFI-bot database and the GitHub GraphQL API to collect features of 68,858 candidate issues and restore historical states of resolvers of 11,615 FIs (first issues).</p> <p>The followings are the files and replicating results:</p> <p>Dataset:</p> <p>The raw data of newcomer-issue pairs' features are stored in <code>ReplicationPackage/data/dataset_{bertmodel}_{num}.pkl</code>, where {bertmodel} is one of the four BERT-based language models: SIMCSE, RoBERTa, CodeBERT, and BERTOverflow, corresponding to the dataset whose textual features are extracted by one of the four language models. And {num} is 0 to 19, corresponding to the 20 chronological folds. The training sets of the GFI-Bot approach are contained in <code>ReplicationPackage/data/training_set_recgfi_simcse_{num}.pkl</code>. <code>ReplicationPackage/data/newcomerdata.json</code> contains first issues' title and description and their resolvers' total commit number and number of commits in the latest month, and <code>ReplicationPackage/data/processeddata.pkl</code> contains the 37 developers' features for the empirical study. <code>ReplicationPackage/data/isstexts.json</code> contains issues titles and descriptions for Stanik et al.'s approach.</p> <p>Python scripts:</p> <p><code>ReplicationPackage/empirical.py</code> is the script for reproducing all the results in Section III of the paper. <code>ReplicationPackage/model.py</code> is the script for reproducing all the results in Section IV of the paper.</p> <p>Records:</p> <p><code>ReplicationPackage/PFIs.csv</code> records the manually labeled issues for the empirical study.</p> <p>Figures:</p> <p>By running <code>ReplicationPackage/empirical.py</code> and <code>ReplicationPackage/model.py</code>, you can get all the figures in the fold <code>ReplicationPackage/figures/</code>. Besides the figures in the paper, <code>ReplicationPackage/figures/</code> also contains <code>typedis_{num}.png</code>, and <code>domaindis_{num}.png</code>, {num} is 1 to 4, representing additional results of newcomer features for Figure 4 in the paper.</p>
ASE artifact
Open the record for dataset details and reuse information.
1-Piperidine Propionic Acid as an allosteric inhibitor of Prote-ase Activated Receptor-2
<p>Supplementary Figure S1 and Supplementary Figure S2 of the manuscript submitted to Antioxidants</p>
Genome-wide analysis of expression QTL (eQTL) and allele-specific expression (ASE) in pig muscle identifies candidate genes for meat quality traits
GEO Series GSE124315. Sus scrofa. 189 samples. Type: Expression profiling by high throughput sequencing.
Factors Driving Diversity in Gene Regulatory Networks at Genome Scale[ASE]
GEO Series GSE267878. Brachypodium distachyon. 24 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.