Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
677
datasets available to search
ShareScore release 0.9.0
Dataset results
677 results for “Replication package”
Replication package for: Stunting and wasting in a growing economy: biological living standards in Portugal during the twentieth century (updated)
<p>The replication materials contain a README file, STATA as well as .csv datasets, and do-files. This replication package for Cermeño et al. (2023) constructs the entire analysis from the data sources described in the published paper, using STATA. The replicator should expect the code to run for less than 10 minutes.</p> <p>Cermeño, A. L., Palma, N., and Pistola, R. (2023). Stunting and wasting in a growing economy: biological living standards in Portugal during the twentieth century. <em><strong>Economics & Human Biology</strong></em>, vol. 51: 101267</p>
Replication Package for "Benefits and pitfalls of token-level SZZ: An empirical study on OSS projects"
<p>Replication Package for "Benefits and pitfalls of token-level SZZ: An empirical study on OSS projects"</p><p>All materials are licensed under the MIT License (see LICENSE file). </p>
Replication Package of Deep Learning and Data Augmentation for Detecting Self-Admitted Technical Debt
<p>Self-Admitted Technical Debt (SATD) refers to circumstances where developers use code comments, issues, pull requests, or other textual artifacts to explain why the existing implementation is not optimal. Past research in detecting SATD has focused on either identifying SATD (classifying SATD instances as SATD or not) or categorizing SATD (labeling instances as SATD that pertain to requirements, design, code, test, etc.). However, the performance of such approaches remains suboptimal, particularly when dealing with specific types of SATD, such as test and requirement debt. This is mostly because the used datasets are extremely imbalanced.</p> <p>In this study, we utilize a data augmentation strategy to address the problem of imbalanced data. We also employ a two-step approach to identify and categorize SATD on various datasets derived from different artifacts. Based on earlier research, a deep learning architecture called BiLSTM is utilized for the binary identification of SATD. The BERT architecture is then utilized to categorize different types of SATD. We provide the dataset of balanced classes as a contribution for future SATD researchers, and we also show that the performance of SATD identification and categorization using deep learning and our two-step approach is significantly better than baseline approaches.</p> <p>Therefore, to showcase the effectiveness of our approach, we compared it against several existing approaches:</p> <ol> <li>Natural Language Processing (NLP) and Matches task Annotation Tags (MAT) [<a href="https://github.com/Naplues/MAT" target="_blank" rel="noopener">Github</a>]</li> <li>eXtreme Gradient Boosting+Synthetic Minority Oversampling Technique (XGBoost+SMOTE) [<a href="https://figshare.com/s/87a4b5002c7488822e60" target="_blank" rel="noopener">Figshare</a>]</li> <li>eXtreme Gradient Boosting+Easy Data Augmentation (XGBoost+EDA) [<a href="https://github.com/shenyuanduanzui/xgboost_satd" target="_blank" rel="noopener">Github</a>]</li> <li>MT-Text-CNN [<a href="https://github.com/yikun-li/satd-different-sources-data" target="_blank" rel="noopener">Github</a>]</li> </ol> <p> </p> <div> <div><strong>Structure of the Replication Package:</strong></div> </div> <p>In accordance with the original dataset, the dataset comprises four distinct CSV files delineated by the artifacts under consideration in this study. Each CSV file encompasses a text column and a class, which indicate classifications denoting specific types of SATD, namely code/design debt (C/D), documentation debt (DOC), test debt (TES), and requirement debt (REQ) or Not-SATD.</p> <div> <div><code>├── SATD Keywords</code></div> <div><code>│ ├── Keywords based on Source of Artifacts</code></div> <div><code>│ │ ├── Code comment.txt</code></div> <div><code>│ │ ├── Commit message.txt</code></div> <div><code>│ │ ├── Issue section.txt</code></div> <div><code>│ │ └── Pull section.txt</code></div> <div><code>│ ├── Keywords based on Types of SATD</code></div> <div><code>│ │ ├── code-design debt.txt</code></div> <div><code>│ │ ├── documentation debt.txt</code></div> <div><code>│ │ ├── requirement debt.txt</code></div> <div><code>│ │ └── test debt.txt</code></div> <div><code>├── src</code></div> <div><code>│ ├── bert.py</code></div> <div><code>│ ├── bilstm.py</code></div> <div><code>│ └── preprocessing.py</code></div> <div><code>├── data-augmentation-code_comments.csv</code></div> <div><code>├── data-augmentation-commit_messages.csv</code></div> <div><code>├── data-augmentation-issues.csv</code></div> <div><code>├── data-augmentation-pull_requests.csv</code></div> <div> <div><code>└── Supplementary Material.docx</code></div> <div> </div> </div> <div> </div> </div> <p><strong>Requirements:</strong></p> <div> <div><a href="https://nlp.stanford.edu/projects/glove/" target="_blank" rel="noopener">glove</a></div> <div>nltk</div> <div>transformers</div> <div>torch</div> <div>tensorflow</div> <div>keras</div> <div>langdetect</div> <div>inflect</div> <div>inflection</div> </div> <div> </div> <div> </div> <div> <div><strong>Project sources for each artifact are as follows:</strong></div> </div> <div> </div> <div> <table> <tbody> <tr> <td><strong>Source code comment</strong></td> <td><strong>Issue section</strong></td> <td><strong>Pull section</strong></td> <td><strong>Commit message</strong></td> </tr> <tr> <td>ant<br>argouml<br>columba<br>emf<br>hibernate<br>jedit<br>jfreechart<br>jmeter<br>jruby<br>squirrel</td> <td> <div>camel</div> <div>chromium</div> <div>gerrit</div> <div>hadoop</div> <div>hbase</div> <div>impala</div> <div>thrift</div> </td> <td> <div>accumulo</div> <div>activemq</div> <div>activemq-artemis</div> <div>airflow</div> <div>ambari</div> <div>apisix</div> <div>apisix-dashboard</div> <div>arrow</div> <div>attic-apex-core</div> <div>attic-apex-malhar</div> <div>attic-stratos</div> <div>avro</div> <div>beam</div> <div>bigtop</div> <div>bookkeeper</div> <div>brooklyn-server</div> <div>calcite</div> <div>camel</div> <div>camel-k</div> <div>camel-quarkus</div> <div>camel-website</div> <div>carbondata</div> <div>cassandra</div> <div>cloudstack</div> <div>commons-lang</div> <div>couchdb</div> <div>cxf</div> <div>daffodil</div> <div>drill</div> <div>druid</div> <div>dubbo</div> <div>echarts</div> <div>fineract</div> <div>flink</div> <div>fluo</div> <div>geode</div> <div>geode-native</div> <div>gobblin</div> <div>griffin</div> <div>groovy</div> <div>guacamole-client</div> <div>hadoop</div> <div>hawq</div> <div>hbase</div> <div>helix</div> <div>hive</div> <div>hudi</div> <div>iceberg</div> <div>ignite</div> <div>incubator-brooklyn</div> <div>incubator-dolphinscheduler</div> <div>incubator-doris</div> <div>incubator-heron</div> <div>incubator-hop</div> <div>incubator-mxnet</div> <div>incubator-pagespeed-ngx</div> <div>incubator-pinot</div> <div>incubator-weex</div> <div>infrastructure-puppet</div> <div>jena</div> <div>jmeter</div> <div>kafka</div> <div>karaf</div> <div>kylin</div> <div>lucene-solr</div> <div>madlib</div> <div>myfaces-tobago</div> <div>netbeans</div> <div>netbeans-website</div> <div>nifi</div> <div>nifi-minifi-cpp</div> <div>nutch</div> <div>openwhisk</div> <div>openwhisk-wskdeploy</div> <div>orc</div> <div>ozone</div> <div>parquet-mr</div> <div>phoenix</div> <div>pulsar</div> <div>qpid-dispatch</div> <div>reef</div> <div>rocketmq</div> <div>samza</div> <div>servicecomb-java-chassis</div> <div>shardingsphere</div> <div>shardingsphere-elasticjob</div> <div>skywalking</div> <div>spark</div> <div>storm</div> <div>streams</div> <div>superset</div> <div>systemds</div> <div>tajo</div> <div>thrift</div> <div>tinkerpop</div> <div>tomee</div> <div>trafficcontrol</div> <div>trafficserver</div> <div>trafodion</div> <div>tvm</div> <div>usergrid</div> <div>zeppelin</div> <div>zookeeper</div> </td> <td> <div>accumulo</div> <div>activemq</div> <div>activemq-artemis</div> <div>airflow</div> <div>ambari</div> <div>apisix</div> <div>apisix-dashboard</div> <div>arrow</div> <div>attic-apex-core</div> <div>attic-apex-malhar</div> <div>attic-stratos</div> <div>avro</div> <div>beam</div> <div>bigtop</div> <div>bookkeeper</div> <div>brooklyn-server</div> <div>calcite</div> <div>camel</div> <div>camel-k</div> <div>camel-quarkus</div> <div>camel-website</div> <div>carbondata</div> <div>cassandra</div> <div>cloudstack</div> <div>commons-lang</div> <div>couchdb</div> <div>cxf</div> <div>daffodil</div> <div>drill</div> <div>druid</div> <div>dubbo</div> <div>echarts</div> <div>fineract</div> <div>flink</div> <div>fluo</div> <div>geode</div> <div>geode-native</div> <div>gobblin</div> <div>griffin</div> <div>groovy</div> <div>guacamole-client</div> <div>hadoop</div> <div>hawq</div> <div>hbase</div> <div>helix</div> <div>hive</div> <div>hudi</div> <div>iceberg</div> <div>ignite</div> <div>incubator-brooklyn</div> <div>incubator-dolphinscheduler</div> <div>incubator-doris</div> <div>incubator-heron</div> <div>incubator-hop</div> <div>incubator-mxnet</div> <div>incubator-pagespeed-ngx</div> <div>incubator-pinot</div> <div>incubator-weex</div> <div>infrastructure-puppet</div> <div>jena</div> <div>jmeter</div> <div>kafka</div> <div>karaf</div> <div>kylin</div> <div>lucene-solr</div> <div>madlib</div> <div>myfaces-tobago</div> <div>netbeans</div> <div>netbeans-website</div> <div>nifi</div> <div>nifi-minifi-cpp</div> <div>nutch</div> <div>openwhisk</div> <div>openwhisk-wskdeploy</div> <div>orc</div> <div>ozone</div> <div>parquet-mr</div> <div>phoenix</div> <div>pulsar</div> <div>qpid-dispatch</div> <div>reef</div> <div>rocketmq</div> <div>samza</div> <div>servicecomb-java-chassis</div> <div>shardingsphere</div> <div>shardingsphere-elasticjob</div> <div>skywalking</div> <div>spark</div> <div>storm</div> <div>streams</div> <div>superset</div> <div>systemds</div> <div>tajo</div> <div>thrift</div> <div>tinkerpop</div> <div>tomee</div> <div>trafficcontrol</div> <div>trafficserver</div> <div>trafodion</div> <div>tvm</div> <div>usergrid</div> <div>zeppelin</div> <div>zookeeper</div> </td> </tr> </tbody> </table> </div> <div> <div> </div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <p>This dataset has undergone a data augmentation process using the <a href="https://arxiv.org/abs/2302.13007" target="_blank" rel="noopener">AugGPT</a> technique. Meanwhile, the original dataset can be downloaded via the following link: <a href="https://github.com/yikun-li/satd-different-sources-data">https://github.com/yikun-li/satd-different-sources-data</a></p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> <div> </div> </div> </div> </div> </div> </div>
Replication package - Potential Effectiveness and Efficiency Issues in Usability Evaluation within Digital Health: A Systematic Literature Review
<p>1. File: <strong>Maqbool_SLR_2023_JSS_Inclusion_610.xlsm.</strong></p> <p>There are two sheets in file. A. <strong>Final_Selected_papers</strong>, (sheet) aims to provide a comprehensive list of articles (n=610) selected for our SLR, whose process and data items specified and detailed in the article. </p> <p>B. <strong>Rejected_After_Full_Review</strong>, (sheet) aims to provide a comprehensive list of articles (n=153) rejected for our SLR based on inclusion or exclusion criteria after full article review process, whose process and data items specified and detailed in the article. </p> <p>2. File: <strong>Maqbool_SLR_2023_JSS_Data_Extraction_Form.pdf</strong></p> <p>This file aims to provide a comprehensive data extraction form, whose process and data items specified and detailed in the article. The form was used to elicit data relevant to answer the postulated research questions. This form served as the foundation for the additional information presented in the final paper.</p> <p>3. File: <strong>Bilal_SLR_JSS_Primary_Studies_References.pdf</strong></p> <p>This file contains the primary selected studies (n=610) for the systematic literature review. The systematic review aims to explore and analyse research literature related to usability evaluation methods and their effectiveness and efficiency in the context of digital health applications. This file will help to identity reference of the primary selected study that is cited in the paper using a prefix (S, e.g. S137). This file can be used for peer review, ensuring the reliability and correctness of findings.</p> <p>4. File: <strong>SLR_Analysis_updated_2023.nvp</strong></p> <p>The data extracted from each article was recorded in a worksheet (Excel) and then coded in NVivo 12/14 to categorise (classify) and compare extracted facets. Each data item's category and related paper id are coded in the given Excel file. Papers were not included in the NVivo project due to copyright concerns. Relevant papers can be tracked using the provided spreadsheet file (see Paper ID cell).<br>The file(s) are cleaned as much as reasonable and other raw data is removed. This file does not include the matrix tables or codes, which were produced and analysed run-time during the analysis phase. Although the given package allows for re-generation.</p> <p>-------- UPDATE: --------</p> <p>5. File: <strong>SLR_Analysis_updated_2023_for_MAC.nvpx</strong></p> <p>This is an extra copy of NVivo project, created for the MAC user.</p> <p> </p> <p>This replication package is produced and published here. Research conducted by Karlstad University researchers. We publish data sets to improve coverage and accessibility. For more info or concerns, contact us.</p> <p> </p> <p>Linked paper published at: Maqbool, Bilal, and Sebastian Herold. "Potential effectiveness and efficiency issues in usability evaluation within digital health: A systematic literature review." <em>Journal of Systems and Software</em> (2023): 111881.</p> <p>DOI: <a href="https://doi.org/10.1016/j.jss.2023.111881">https://doi.org/10.1016/j.jss.2023.111881</a><br> </p> <p>This work was funded, in parts, by Region Värmland through the DHINO project, Sweden (Grant: RUN/220266) and Vinnova through the DigitalWell Arena (DWA) project, Sweden (Grant: 2018-03025).</p>
Replication Package: Product-Line Engineering for Smart Manufacturing: A Systematic Mapping Study on Security Concepts
<p><strong>Welcome to the public repository for the additional content of the paper "Product-Line Engineering for Smart Manufacturing: A Systematic Mapping Study on Security Concepts", accepted at the ICSOFT 2024.</strong></p> <p>This repository provides additional information to the conducted mapping study, including the following file:</p> <ul> <li>analysis_sheet_ICSOFT2024.csv: sheet containing information regarding the analysis results of 43 included papers based on the extraction criteria.</li> </ul>
Replication Package for: International Comovement in the Global Production Network
<p>This package contains the data and code necessary to reproduce the figures and tables in Huo, Levchenko and Pandalai-Nayar (forthcoming). "International Comovement in the Global Production Network," Review of Economic Studies. Detailed instructions are given about accessing the raw data, and running the code to generate output.</p>
Replication package
<p>Replication package</p>
Anonymous Replication Package
<p>Anonymous Replication Package</p>
Replication package for article 'Parallel Program Analysis on Path Ranges'
<p>This Replication package contains all the results for the article "Parallel Program Analysis on Path Ranges"</p> <p>Abstract. Symbolic execution is a software verification technique symbolically running programs and thereby checking for bugs. <br> Ranged symbolic execution <br> performs symbolic execution on program parts, so called {\em path ranges}, in parallel.<br> Due to the parallelism, verification is accelerated and hence scales to larger programs.</p> <p>In this paper, we discuss a generalization of ranged symbolic execution to arbitrary program analyses.<br> More specifically, we present a verification approach that splits programs into path ranges and<br> then runs arbitrary analyses on the ranges in parallel. Our approach in particular allows to run {\em different}<br> analyses on different program parts.<br> We have implemented this generalization on top of the tool \textsc{CPAchecker} and evaluated it on programs from the SV-COMP benchmark. Our evaluation shows that verification can benefit from the parallelisation of the verification task,<br> but also needs a form of work stealing (between analysis) as to become efficient.</p>
Model Generation from Requirements with LLMs: an Exploratory Study - Replication Package
<p>This is a replication package for the paper "<span>Model Generation from Requirements </span><span>with LLMs: an Exploratory Study</span>", by Sallam Abualhaija, Chetan Arora, and Alessio Ferrari.</p> <p><strong>Abstract: </strong>Complementing natural language (NL) requirements with graphical models can improve stakeholders’ communication and provide directions for system design. However, creating models from requirements involves manual effort. The advent of generative large language models (LLMs), ChatGPT being a notable example, offers promising avenues for automated assistance in model generation. This paper investigates the reliability of ChatGPT in generating sequence diagrams from NL requirements. Specifically, we conduct a qualitative study examining the sequence diagrams generated by ChatGPT for 28 requirements documents of various types and from different domains. Our study aims to uncover potential issues that emerge in the models generated by ChatGPT, thereby hindering its applicability in practice. Observations have systematically been captured through evaluation logs, and categorized through thematic analysis. Our results indicate that, although the models generally conform to the standard and exhibit a reasonable level of understandability, their correctness with respect to the specified requirements often presents challenges. This issue is particularly pronounced in the presence of requirements smells, such as ambiguity and inconsistency. The insights derived from this study can influence the practical utilization of LLMs in the RE process, and open the door to novel RE-specific prompting strategies targeting effective model generation.</p> <p>The replication package consists of the following folders:</p> <p><strong>logs:</strong> includes the evaluation logs produced by each evaluator</p> <p><strong>original-documents: </strong>includes the original requirements documents used for the evaluation</p> <p><strong>RQ1 - quantitative analysis:</strong> includes the analysis made on the scores given to each model and model variant. It includes five files:</p> <p>- results.csv: numerical results of the evaluation for each criterion<br>- analysis-results.Rmd: R file used to perform the quantitative analysis (requires R Studio to be executed)<br>- analysis-results.html: html file produced by analysis-results.Rmd<br>- cross-check.csv: file with the cross-checking of the two assessors applied to a subset of the models<br>- symmary_results.xlsx: final output of the quantitative results in terms of Wilcoxon signed rank tests</p> <p><strong>RQ2 - thematic analysis: </strong>includes the codebook produced by the thematic analysis of the issues in generating models with ChatGPT</p>
Formalising a Gateway-based Blockchain Interoperability Solution with Event-B (Replication package)
<p>This repository holds the artefacts that raised the results of our first paper <em>Formalising a Gateway-based Blockchain Interoperability Solution with Event-B</em> to be presented at the <a href="https://icbc2024.ieee-icbc.org/workshop/crosschain" target="_blank" rel="noopener">ICBC Cross-chain workshop</a>.</p> <p>In this paper, we explore the formalisation of a gateway-based interoperability solution with Event-B. The results showed that the method was suitable and that a straightforward specification could be developed considering Ethereum and Hyperledger Fabric as the involved blockchains. The Event-B specification was assessed with three strategies that enabled its verification and validation. In particular, formal verification (e.g. safety properties), functional validation, and functional utility. These promising results constitute a step forward in the development of formal specifications for blockchain interoperability solutions.</p> <p>The artefacts generated in this research were:</p> <ul> <li>An Event-B specification of the gateway-based interoperability solution proposed by <a href="https://ieeexplore.ieee.org/document/10346168" target="_blank" rel="nofollow noreferrer noopener">Pandolfi et al.</a>. This specification is composed of three machines: an abstract machine and two refinements. One refinement describes the behaviour of the gateway and smart contracts involved when the source blockchain is Ethereum and the target blockchain is Hyperledger Fabric. The second refinement describes the behaviour of the gateway and smart contracts when the source is Hyperledger Fabric and the target Ethereum.</li> <li>Animations of the three specifications (i.e. abstract and refinements) that enabled us to validate the functional behaviour of the specification and provided the means to understand the gateways' behaviour without Event-B knowledge.</li> <li>An Event-B specficiation of a use case scenario that shows the utility of the specification.</li> </ul>
Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"
<h1>Replication Package for the Paper: “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis”</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.”</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>
Replication Package of the paper: "A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems"
<p># README<br>## _Replication Package for A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems_</p> <p>This package contains several files and a folder with both code and data used within the context of the study. </p> <p>## output_StaticAnalysis.csv<br>This file contains the results of the manual annotation aimed at verifying whether a static code analysis tool can be used to pinpoint the presence of SATD and, more specifically, DL-SATD. For each SATD instance, you can find a list of warnings impacting the area of the code where the SATD is, as well as the outcome of our manual annotation. </p> <p>## dependentsTensorflow.csv<br>This file contains the list of dependents of the TensorFlow DL framework together with the number of stars and the number of forks. </p> <p>## dependentsTorch.csv<br>This file contains the list of dependents of the PyTorch DL framework, together with the number of stars and the number of forks.</p> <p>## 100Projects.csv<br>This file contains the list of 100 open-source Python projects importing at least one among Tensorflow or PyTorch that we have used as the initial set for our study. </p> <p>## SATDTensorFlow.csv<br>This file contains the list of SATD detected for the 50 Python open-source projects relying on TensorFlow. Each line contains the name of the project, the path to the file in which the SATD has been detected, the line in the file where the SATD comment starts, and the SATD comment. </p> <p>## SATDTorch.csv<br>This file contains the list of SATD detected for the 50 Python open-source projects relying on PyTorch. Each line contains the name of the project, the path to the file in which the SATD has been detected, the line in the file where the SATD comment starts, and the SATD comment.</p> <p>## SampledSATD.csv<br>This file contains the list of the 443 SATD comments used to determine the DL-specific SATD Taxonomy. Each line contains the link to the SATD, followed by the SATD comment body, and two identifiers used to determine the context of the SATD comment (used to properly select the presence of static code analysis tools warnings). </p> <p>## FinalValidation.csv<br>This file contains the outcome of the manual validation of the 443 SATD in our sample. </p> <p>## MappingWithHumbatovaEtAl.xls<br>This file contains the outcome of the mapping between our DL-specific SATD taxonomy and the taxonomy of DL-bugs by Humbatova et al. </p> <p>## extractSATDComments.py<br>This file contains the source code used to analyze the 100 projects in our study. Specifically, for each project and each Python file within it, the script checks whether the file imports one of the two DL frameworks used in the context of the study. If this is the case, it extracts all comments within it and re-implements the KL-SATD to check whether the comment is a SATD candidate. </p> <p>## runAscat.py<br>This file contains the source code used to run Prospector as an aggregator of static code analysis tools for Python. </p> <p>## CompleteStaticAnalysisWarnings<br>This directory contains the outputs of Prospector (one for each studied project, in JSON format)</p>
Replication Package: Pandemic Startup Software Engineering: An Experience Report on the Development of a COVID-19 Certificate Verification System
<p><strong>Welcome to the public repository for the additional content of the paper "Pandemic Startup Software Engineering: An Experience Report on the Development of a COVID-19 Certificate Verification System" (Journal of Systems and Software)<br></strong></p> <p>This repository provides additional information to the experience report, including the following files:</p> <ul> <li>survey_questions_de.txt: sheet containing the online questionnaire in German (original language)</li> <li>survey_questions_en.txt: sheet containing the online questionnaire translated into English</li> <li>survey_answers_original.csv: sheet containing the extracted questionnaire data of the participants in German (original language)</li> <li>survey_analysis.csv: sheet containing the analysis of the extracted questionnaire data in English</li> </ul>
Replication package for "Chain Restaurant Calorie Posting Laws, Obesity, and Consumer Welfare"
<p>This package contains the data, programs, and instructions to replicate the manuscript "Chain Restaurant Calorie Posting Laws, Obesity, and Consumer Welfare" by Charles Courtemanche, David Frisvold, David Jimenez-Gomez, Marietou Ouayogode, and Michael Price, which is forthcoming at JEEA.</p>
Replication Package: Block-based or Graph-Based? Why Not Both? Designing a Hybrid Programming Environment for End-users
<p><strong>Block-based or Graph-based? Why Not Both? Designing a Hybrid Programming Environment for End-users: Replication Package</strong></p> <p>This repository contains supplementary materials for the paper "Block-based or Graph-based? Why Not Both? Designing a Hybrid Programming Environment for End-users". We provide this data for transparency reasons and to support replications of our experiments.</p> <p><strong>Summary of files contained in this package</strong></p> <p>This package contains two parts:</p> <ul> <li> <p>The <code>data-analysis/</code> folder contains the raw dataset we collected for our experiment in CSV format, as well as scripts we used for our analyses.</p> <ul> <li>Column <code>ID</code> contains a unique 4-digit identifier for each participant that they were assigned throughout our study.</li> <li>Column <code>Group</code> contains the group (Blocks/Graph) that participants were randomly assigned to.</li> <li>Columns <code>Task1Time</code> and <code>Task2Time</code> contain the time participants spent to complete the two programming tasks of our study in minutes.</li> <li>Columns <code>Task1Success</code> and <code>Task2Success</code> contain a boolean value indicating whether the participants successfully completed the given task. Note that participants had unlimited attempts until they timed out after a strict time limit of 30 minutes, so if a participant was unsuccessful the corresponding time value is 30.</li> <li>Columns <code>Task1Tests</code> and <code>Task2Tests</code> contain the number of times a participant executed their code throughout a task, including their final submission if they were successful.</li> <li>Columns <code>LearnTask</code>, <code>ReadTask</code> and <code>WriteTask</code> contain the scores that participants gave to the task editor component of their assigned programming environment. There are 3 scores for the categories "learnability", "readability" and "writability". Scores are on a 5-point scale from 1 (worst) to 5 (best).</li> <li>Columns <code>LearnTrig</code>, <code>ReadTrig</code> and <code>WriteTrig</code> contain the scores that participants gave to the trigger editor component of their assigned programming environment. There are 3 scores for the categories "learnability", "readability" and "writability". Scores are on a 5-point scale from 1 (worst) to 5 (best).</li> <li>Columns <code>LearnComp</code>, <code>ReadComp</code> and <code>WriteComp</code> contain the scores that participants gave to their assigned assigned programming environment in direct comparison to the other alternative. There are 3 scores for the categories "learnability", "readability" and "writability". Unlike in the paper, where scores are on a scale from -2 to 2, the raw scores here are on a 5-point scale from 1 (strong preference for other environment) to 5 (strong preference for own environment).</li> <li>The script <code>successplot.py</code> was used to generate the success rate plot used in a figure in the paper</li> <li>The script <code>survival.py</code> was used to perform the survival analysis presented in the paper and generate the related figure.</li> <li>The script <code>batplot.py</code> was used to generate the 3x3 grid of ratings used in a figure in the paper.</li> </ul> </li> <li> <p>The <code>materials/</code> folder contains the tutorials and task descriptions we presented to study participants. It also contains the exact wording of pre-screening and post-experiemental survey questions.</p> <ul> <li>The image <code>pre-screening.png</code> shows the three pre-screening questions we used to determine whether our participants could be included in our study.</li> <li>The images <code>tutorial1_instructions.png</code> and <code>tutorial1_sim.png</code> contain the instructions and initial simulator state we provided to participants for the first programming tutorial. This tutorial did not provide starter code and was identical for both participant groups.</li> <li>The images <code>tutorial2_instructions.png</code> and <code>tutorial2_sim.png</code> contain the instructions and initial simulator state we provided to participants for the second programming tutorial. This tutorial was identical for both participant groups and provided participants with starter code, which is shown in the images: <ul> <li><code>tutorial2_code_main.png</code> for the main program in the left canvas</li> <li><code>tutorial2_code_move.png</code> for the definition of "Move box to the right".</li> </ul> </li> <li>The images <code>tutorial3_instructions_blocks.png</code>/<code>tutorial3_instructions_graph.png</code> and <code>tutorial3_sim.png</code> contain the instructions and initial simulator state we provided to participants for the third programming tutorial. This tutorial also provided participants with starter code, which is shown in the images: <ul> <li><code>tutorial3_code_main.png</code> for the main program in the left canvas</li> <li><code>tutorial3_code_pick.png</code> for the definition of "Pick up box"</li> <li><code>tutorial3_code_place.png</code> for the definition of "Place box"</li> </ul> </li> <li>The images <code>task1_instructions.png</code> and <code>task1_sim.png</code> contain the instructions and initial simulator state we provided to participants for the first programming task. The task did not provide starter code and the instructions were identical for both participant groups.</li> <li>The images <code>task2_instructions.png</code> and <code>task2_sim.png</code> contain the instructions and initial simulator state we provided to participants for the second programming task. The instructions were identical for both groups. This task also provided participants with starter code, which is shown in the images: <ul> <li><code>task2_code_main.png</code> for the main program in the left canvas</li> <li><code>task2_code_pick_prog.png</code> for the definition of "Pick up block"</li> <li><code>task2_code_load_trig_blocks.png</code>/<code>task2_code_load_trig_graph.png</code> for the definition of the trigger "Ready to load machine"</li> <li><code>task2_code_load_prog.png</code> for the definition of "Load and activate machine"</li> <li><code>task2_code_finished_trig_blocks.png</code>/<code>task2_code_finished_trig_graph.png</code> for the definition of the trigger "Machine finished"</li> <li><code>task2_code_finished_prog1.png</code> for the definition of "Get block from machine"</li> <li><code>task2_code_finished_prog2.png</code> for the definition of "Place block in bin"</li> </ul> </li> <li> <div>The document <code>post_survey_full.pdf</code> contains a raw export of the comprehension questions and post-experimental survey as they were presented to participants </div> </li> <li>The image <code>usability.png</code> shows the usability questions we used to determine a participant's rating of their assigned programming environment. The questions were identical for both participant groups.</li> <li>The images <code>comprehension_blocks_1.png</code> and <code>comprehension_blocks_2.png</code> show the program comprehension questions we used to determine whether participants in the Blocks group could understand more complex triggers.</li> <li>The images <code>comprehension_graph_1.png</code> and <code>comprehension_graph_2.png</code> show the program comprehension questions we used to determine whether participants in the Graph group could understand more complex triggers.</li> <li>The images <code>comparison_blocks.png</code> and <code>comparison_graph.png</code> show the images of triggers in the alternative environment that we showed to our participants before choosing their preferred environment. The questions were identical for both participant groups.</li> <li>The image <code>comparison.png</code> shows the questions we used to determine a participant's preference between the two programming environment alternatives.</li> </ul> </li> </ul>
Replication Package for "Mitigating Automated Obfuscation Attacks on Software Plagiarism Detection Systems"
<p>This is the replication package for the doctoral dissertation titled "<em>Mitigating Automated Obfuscation Attacks on Software Plagiarism Detection Systems</em>".</p> <p>The contributions of the dissertation were also integrated into the source code plagiarism detection system <a href="https://github.com/jplag/JPlag/">JPlag</a> to ensure they are widely accessible.</p> <p><strong>Contents Overview:</strong></p> <p>- <strong>Datasets</strong>: the artifacts of the evaluation datasets.<br>- <strong>Raw Results</strong>: the measured results of our evaluation.<br>- <strong>Evaluation Scripts</strong>: the evaluation code for plotting and statistical tests.<br>- <strong>Implementation</strong>: the source code of the JPlag-based implementation and the prebuilt application as a JAR file.<br>- <strong>Other</strong>: additional plots.</p>
Replication Package: Asking Security Practitioners: Did You Find the Vulnerable (Mis)Configuration?
<p><strong>Welcome to the public repository for the additional content of the paper "Asking Security Practitioners: Did You Find the Vulnerable (Mis)Configuration?", accepted at the International Working Conference on Variability Modelling of Software-Intensive Systems (VAMOS) 2025.</strong></p> <p>This repository provides additional information to the conducted survey study on configuration-related vulnerabilities, including the following files:</p> <ul> <li>QUESTIONNAIRE_VAMOS2025.csv: sheet containing all questionnaire data and answer options</li> <li>DATA_VAMOS2025.csv: sheet containing data of the 41 participants, including additional codings of free-text answers</li> <li>README.txt: readMe file</li> </ul> <p><strong>Requirements for using the data<br></strong></p> <ul> <li>No requirements</li> </ul> <p><strong>License for using the data<br></strong></p> <p>Creative Commons Attribution 4.0 International</p> <p>The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.</p> <p>Further information: https://creativecommons.org/licenses/by/4.0/legalcode</p>
Replication package for "An Empirical Evaluation of Static, Dynamic, and Hybrid Slicing of WebAssembly Binaries"
<p>This is the replication package that accompanies the paper "An Empirical Evaluation of Static, Dynamic, and Hybrid Slicing of WebAssembly Binaries".</p> <p>It is structured as follows.</p> <p> - RQs.py is the script that generates the data that is included in the paper. For each research question, it generates statistics and plots.<br> - The .csv files and nok.txt are used by RQs.py. They contain the raw size and timing data for each slice. The headers of the .csv files indicate what each column represents.<br> - `sqlite-slices.csv` contains the raw data for RQ8, which is contained in the paper<br> - The .tar.gz files contain the source and the slices generated by each slicer, namely:<br> - `static_slices.tar.gz` contains the slices generated by CsE, the static slicer. For example, the file `static_slices/adpcm/adpcm_ah1_254_expr/static_adpcm.wat.slice` is the slice of the adpcm program taken with respect to the variable `ah1` at line `254` of the original `.c` program.<br> - `slice.tar.gz` contains slices from the dynamic slicers:<br> - slice-sce - the SCE slices<br> - slice-ces - the CES slices<br> - slice-cse - the CSE slices<br> - slice-sces - the SCES slices<br> - slice-cses - the CsES slices<br> - `src.tar.gz` includes the unsliced `C` and `wasm` source<br> - `bc.support.tar.gz` includes the complete `bc` source code <br> - `sqlite.tar.gz` contains the slices of RQ8, each in a directory named after the slicer used. For example, the file `sqlite/CES/avg.wat` contains the CES slice for the `avg` slicing criterion.<br> <br>In order to regenerate the data and plots that are in the paper, you should simply run:</p> <p>```<br>python RQs.py<br>```</p>
CombiANT Reader replication package
<p>This dataset and software accompany the article "CombiANT Reader - Deep learning-based automatic image processing and measurement of distances to robustly quantify antibiotic interactions" for reproducing the results. </p> <p>In the study, a deep learning-based image processing method is developed to analyze agar plates of the CombiANT combination antibiotics test. The package contains data and software to re-run the experiments, generate output metrics, and build the graphs in the article.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.