Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

677

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

677 results for “Replication package”

Learn how ShareScore rates datasets ↗
zenodo40/100

Replication Package for: Detecting Edgeworth Cycles (Version 2)

<p>This replication package contains the <strong>data </strong>and the <strong>code </strong>to generate the main results reported in <strong>"Detecting Edgeworth Cycles"</strong>&nbsp;by <strong>Timothy Holt, Mitsuru Igami, and Simon Scheidegger</strong>, to be published in the February 2024 issue of <em>The Journal of Law and Economics</em>.&nbsp;</p> <p>Additionally, this package also allows the users to<strong> apply these pre-trained models to new datasets</strong> of their choice. As an example of such a new dataset, we include the <strong>entire German dataset available at the time of our research (2014:Q4&ndash;2020:Q4)</strong>, including both manually labeled and unlabeled subsamples.</p> <p>Finally, this package includes <strong>tools to facilitate the acquisition and pre-processing of the most recent data from Germany</strong>, which is updated every day on the <em>Tankerkoenig</em> website (at the time of our preparation of this package).</p> <p><strong>[3/29/2024 update]</strong> The URL for the German antitrust authority's fuel-data website has changed to <a href="https://www.bundeskartellamt.de/EN/Tasks/markettransparencyunit_fuels/markettransparencyunit_fuels.html" target="_blank" rel="noopener">https://www.bundeskartellamt.de/EN/Tasks/markettransparencyunit_fuels/markettransparencyunit_fuels.html</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Replication package for the paper "A configurational approach to job quality analysis: forms of inequalities at work in Europe"

<p>The following replication package is appended to the article&nbsp;<span><span><span><span>&Eacute;tienne Penissat</span><span>, </span></span><span><span>C&eacute;cile Rodrigues</span><span> &amp; </span></span><span><span>Alexis Spire</span></span></span></span> <span>(2024)</span> "<span>A configurational approach to job quality analysis: forms of inequalities at work in Europe",</span> <span>European Societies,</span> <span>DOI: <a href="https://doi.org/10.1080/14616696.2024.2312950">10.1080/14616696.2024.2312950</a></span></p> <p>The scripts to be run in the following order are:</p> <p>- 1_Penissat_EuropeanSocieties_2023_DataPreparation.R : Recoding, formatting and scope of data used</p> <p>- 2_Penissat_EuropeanSocieties_2023_DataAnalysis.Rmd : Analysis and statistical results</p> <p>The data used in the article comes from the EWCS (2015) - European Working Condition Survey - provided by the European foundation for the improvement of living and working conditions. The data is not available on free access but can be obtained on request. Information on the survey wave used can be found here : https://www.eurofound.europa.eu/surveys/european-working-conditions-surveys/sixth-european-working-conditions-survey-2015</p> <p>- In the first "1_Penissat_EuropeanSocieties_DataPreparation.R" script, the file containing data named "ewcs_1991-2015.dta" is used. The file called "eseg2_trad.csv" contains english labels for the nomenclature of professional positions ESeG. As "ewcs_1991-2015.dta" is not freely available, it is not included in the package and "eseg2_trad.csv" is located in the "data" folder.</p> <p>- The first script creates the data file "Penissat_EuropeanSocieties_2023_EWCS15_cleanData.rds" in the "results" folder.</p> <p>- The second script called "2_Penissat_EuropeanSocieties_DataAnalysis.Rmd" uses the "Penissat_EuropeanSocieties_2023_EWCS15_cleanData.rds" data file and produces the "2_Penissat_EuropeanSocieties_DataAnalysis.html" file containing all code and results presented in the article.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Replication package: Dataset and stata-do-file for analysis in "Intragroup communication in social dilemmas: An artefactual public good field experiment in small-scale communities"

<p>This dataset was used for the analysis in "Intragroup communication in social dilemmas: An artefactual public good field experiment in small-scale communities". The data was collected in Namibia in 2017 as part of the SASSCAL research project by Nils Christian Hoenow and Adrian Pourviseh as members of the Chair for Development and Cooperative Economics at the University of Marburg. Funded by the Southern African Science Service Center for Climate Change and Adaptive Land-UseManagement (SASSCAL) through the German Federal Ministry for Education and Research (Grant No. 01LG1201B).</p> <p>&nbsp;</p> <p>Article Title: Intragroup communication in social dilemmas: An artefactual public good field experiment in small-scale communities&nbsp;</p> <p>Authors: Nils Christian Hoenow* and Adrian Pourviseh**</p> <p>&nbsp;</p> <p>*RWI &ndash; Leibniz Institute for Economic Research, Essen, Germany and &amp; School of Business and Economics, University of<br>Marburg, Marburg, Germany</p> <p>**School of Business and Economics, University of<br>Marburg, Marburg, Germany</p> <p>Abstract:&nbsp;<br>Communication is well-known to increase cooperation rates in social dilemma situations, but the exact mechanisms behind this remain largely unclear. This study examines the impact of communication on public good provisioning in an artefactual field experiment conducted with 216 villagers from small, rural communities in northern Namibia. In line with previous experimental findings, we observe a strong increase in cooperation when face-to-face communication is allowed before decision-making. We additionally introduce a condition in which participants cannot discuss the dilemma but talk to their group members about an unrelated topic prior to learning about the<br>public good game. It turns out that this condition already leads to higher cooperation rates, albeit not as high as in&nbsp;the condition in which discussions about the social dilemma are possible. The setting in small communities also allows investigating the effects of pre-existing social relationships between group members and their interaction with communication.We find that both types of communication are primarily effective among socially more distant group members, which suggests that communication and social ties work as substitutes in increasing cooperation. Further analyses rule out better comprehension of the game and increased mutual expectations of one&rsquo;s group members&rsquo; contributions as drivers for the communication effect. Finally, we discuss the role of personal and injunctive norms to keep commitments made during discussions.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Replication package for the paper: "A Study on the Pythonic Functional Constructs' Understandability"

<h1>Replication Package for "<em>A Study on the Pythonic Functional Constructs' Understandability</em>" to appear at ICSE 2024</h1> <ul> <li><strong>Authors</strong>: Cyrine Zid, Fiorella Zampetti, Giuliano Antoniol, Massimiliano Di penta</li> <li><strong>Article Preprint:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> <li><strong>License</strong>: <a>GPL V3.0</a></li> </ul> <p>This package contains folders and files with code and data used in the study described in the paper. In the following, we first provide all fields required for the submission, and then report a detailed description of all repository folders.</p> <h2>Artifact Description</h2> <h3>Purpose</h3> <p>The artifact is about a controlled experiment aimed at investigating the extent to which Pythonic functional constructs have an impact on source code understandability. The artifact archive contains:</p> <ol> <li>The material to allow replicating the study (see Section <a>Experimental-Material</a>)</li> <li>Raw quantitative results, working datasets, and scripts to replicate the statistical analyses reported in the paper. Specifically, the executable part of the replication package reproduces figures and tables of the quantitative analysis (RQ1 and RQ2) of the paper starting from the working datasets.</li> <li>Spreadsheets used for the qualitative analysis (RQ3).</li> </ol> <p>We apply for the following badges:</p> <ul> <li><strong>Available and reusable</strong>: because we provide all the material that can be used to replicate the experiment, but also to perform the statistical analyses and the qualitative analyses (spreadsheets, in this case)</li> </ul> <h3>Provenance</h3> <ul> <li><strong>Paper preprint link:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> </ul> <h3>Data</h3> <p>Results have been obtained by conducting the controlled experiment involving <a href="https://www.prolific.com/">Prolific</a>workers as participants. Data collection and processing followed a protocol approved by the University ethical board. Note that all data enclosed in the artifact is completely anonymized and does not contain sensible information.</p> <p>Further details about the provided dataset can be found in the Section <a>Results' directory and files</a></p> <h3>Setup and Usage (for executable artifacts):</h3> <p>See the Section <a>Scripts to reproduce the results, and instructions for running them</a></p> <h2><a>Experiment-Material/</a></h2> <p>Contains the material used for the experiment, and, specifically, the following subdirectories:</p> <h3><a>Google-Forms/</a></h3> <p>Contains (as PDF documents) the questionnaires submitted to the ten experimental groups.</p> <h3><a>Task-Sources/</a></h3> <p>Contains, for each experimental group (G-1...G-10), the sources used to produce the Google Forms, and, specifically: - The cover letter (Letter.docx). - A directory for each experimental task (Lambda 1, Lambda 2, Comp 1, Comp 2, MRF 1, MRF 2, Lambda Comparison, Comp Comparison, MRF Comparison). Each directory contains: (i) the exercise text (in both Word and .txt format), the source code snippet, and its .png image to be used in the form. <strong>Note:</strong> the "Comparison" tasks do not have any exercise as the purpose is always the same, i.e., to compare the (perceived) understandability of the snippets and return the results of the comparison.</p> <h3><a>Code-Examples-Table1/</a></h3> <p>Contains the source code snippets used as objects of the study (the same you can find under "Task-Sources/"), named as reported in Table 1.</p> <h2>Results' directory and files</h2> <p>&nbsp;</p> <h3><a>raw-responses/</a></h3> <p>Contains, as spreadsheets, the raw responses provided by the study participants through Google forms.</p> <h3><a>raw-results-RQ1/</a></h3> <p>Contains the raw results for RQ1. Specifically, the directory contains a subdirectory for each group (G1-G10). Each subdirectory contains: - For each user (named using their Prolific IDs, a directory containing, for each question (Q1-Q6) the produced python code (Qn.py) its output (QnR.txt) and its StdErr output (QnErr.txt). - "expected-outputs/": A directory containing the expected outputs for each task (Qn.txt).</p> <h4><a>working-results/RQ1-RQ2-files-for-statistical-analysis/</a></h4> <p>Contains three .csv files used as input for conducting the statistical analysis and drawing the graphs for addressing the first two research questions of the study. Specifically:</p> <ul> <li> <p><a>ConstructUsage.csv</a> contains the declared frequency usage of the three functional constructs object of the study. This file is used to draw Figure 4. The file contains an entry for each participant, reporting the (text-coded) frequency of construct usage for Comprehension, Lambda, and MRF.</p> </li> <li> <p><a>RQ1.csv</a> contains the collected data used for the mixed-effect logistic regression relating the use of functional constructs with the correctness of the change task, as well as the logistic regression relating the use of map/reduce/filter functions with the correctness of the change task. The csv file contains an entry for each answer provided by each subject, and features the following columns:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>Section</em>: section of the questionnaire (lambda, comp, or mrf)</li> <li><em>Construct</em>: specific construct being presented (same as "Section" for lambda and comp, for mrf it says whether it is a map, reduce, or filter)</li> <li><em>Question</em>: question id, from Q1 to Q6, indicate the ordering of the question</li> <li><em>MainFactor</em>: main factor treatment for the given question - "f" for functional, "p" for procedural counterpart</li> <li><em>Outcome</em>: TRUE if the task was correctly performed, FALSE otherwise</li> <li><em>Complexity</em>: cyclomatic complexity of the construct (empty for mrf)</li> <li><em>UsageFrequency</em>: usage frequency of the given construct</li> </ul> </li> <li> <p><a>RQ1Paired-RQ2.csv</a> contains the collected data used for the ordinal logistic regression of the relationship between the perceived ease of understanding of the functional constructs and (i) participants' usage frequency, and (ii) constructs' complexity (except for map/reduce/filter). The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>LambdaF</em>: result for the change task related to a lambda construct</li> <li><em>LambdaP</em>: result for the change task related to the procedural counterpart of a lambda construct</li> <li><em>CompF</em>: result for the change task related to a comprehension construct</li> <li><em>CompP</em>: result for the change task related to the procedural counterpart of a comprehension construct</li> <li><em>MrfF</em>: result for the change task related to an MRF construct</li> <li><em>MrfP</em>: result for the change task related to the procedural counterpart of a MRF construct</li> <li><em>LambdaComp</em>: perceived understandability level for the comparison task (RQ2) between a lambda and its procedural counterpart</li> <li><em>CompComp</em>: perceived understandability level for the comparison task (RQ2) between a comprehension and its procedural counterpart</li> <li><em>MrfComp</em>: perceived understandability level for the comparison task (RQ2) between a MRF and its procedural counterpart</li> <li><em>LambdaCompCplx</em>: cyclomatic complexity of the lambda construct involved in the comparison task (RQ2)</li> <li><em>CompCompCplx</em>: cyclomatic complexity of the comprehension construct involved in the comparison task (RQ2)</li> <li><em>MrfCompType</em>: type of MRF construct (map, reduce, or filter) used in the comparison task (RQ2)</li> <li><em>LambdaUsageFrequency</em>: self-declared usage frequency on lambda constructs</li> <li><em>CompUsageFrequency</em>: self-declared usage frequency on comprehension constructs</li> <li><em>MrfUsageFrequency</em>: self-declared usage frequency on MRF constructs</li> <li><em>LambdaComparisonAssessment</em>: outcome of the manual assessment of the answer to the "check question" required for the lambda comparison ("yes" means valid, "no" means wrong, "moderate<em>chatgpt" and "extreme</em>chatgpt" are the results of GPTZero)</li> <li><em>CompComparisonAssessment</em>: as above, but for comprehension</li> <li><em>MrfComparisonAssessment</em>: as above, but for MRF</li> </ul> </li> </ul> <h3><a>working-results/inter-rater-RQ3-files/</a></h3> <p>This directory contains four .csv files used as input for computing the inter-rater agreement for the manual labeling used for addressing RQ3. Specifically, you will find one file for each functional construct, i.e., comprehension.csv, lambda.csv, and mrf.csv, and a different file used for highlighting the reasons why participants prefer to use the procedural paradigm, i.e., procedural.csv.</p> <h3><a>working-results/RQ2ManualValidation.csv</a></h3> <p>This file contains the results of the manual validation being done to sanitize the answers provided by our participants used for addressing RQ2. Specifically, we coded the behaviour description using four different levels: (i) correct ("yes"), (ii) somewhat correct ("partial"), (iii) wrong ("no"), and (iv) automatically generated. The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID</li> <li><em>Comparison for lambda construct description</em>: answer provided by the user for the lambda comparison task</li> <li><em>Final Classification</em>: our assessment of the lambda comparison answer</li> <li><em>Comparison for comprehension description</em>: answer provided by the user for the comprehension comparison task</li> <li><em>Final Classification</em>: our assessment of the comprehension comparison answer</li> <li><em>Comparison for MRF description</em>: answer provided by the user for the MRF comparison task</li> <li><em>Final Classification</em>: our assessment of the MRF comparison answer</li> </ul> <h3><a>working-results/RQ3ManualValidation.xlsx</a></h3> <p>This file contains the results of the open coding applied to address our third research question. Specifically, you will find four sheets, one for each functional construct and one for the procedural paradigm. Each sheet reports the provided answers together with the categories assigned to them. Each sheet contains the following columns:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID (as in the tables from the quantitative analysis)</li> <li>: question asked to the user</li> <li><em>Final Classification</em>: The outcome of our categorization according to the taxonomy shown in Table 10.</li> </ul> <h2>Scripts to reproduce the results and instructions for running them</h2> <p>&nbsp;</p> <h3><a>FuncConstructs-Statistics.r</a></h3> <p>This file contains an R script that you can reuse to re-run all the analyses conducted and discussed in the paper.</p> <h3><a>FuncConstructs-Statistics.ipynb</a></h3> <p>This file contains the code to re-execute all the analysis conducted in the paper as a Jupyter Notebook (using the R Kernel).</p> <h3><a>run-analysis.sh</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Unix operating systems.</p> <h3><a>run-analysis.bat</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Windows operating systems (power shell recommended).</p> <h3><a>run-jupyter-container.sh</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Unix operating system.</p> <h3><a>run-jupyter-container.bat</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Windows operating systems (power shell recommended).</p> <h3>How to Run the scripts</h3> <p>There are four options to run the scripts. In all cases, one has first to open a shell terminal window (e.g., bash or sh in Unixes) in the replication package directory. For Windows, we suggest to use a Power Shell.</p> <ol> <li> <p><strong>Running the R script using Dockerized R installation</strong>: this is the simplest option, and it simply requires a running Docker engine. In <strong>Unix (MacOS, Linux)</strong>, to produce the results, one has to run the shell script "run-analysis.sh" (e.g., by typing <code>sh run-analysis.sh</code> or simply <code>./run-analysis.sh</code> after making it executable). In <strong>Windows</strong>, one has to run the script "run-analysis.bat" instead (by typing <code>.\run-analysis.bat</code>). This script (either .sh or .bat) will:</p> <ul> <li>Pull a docker image named <code>mdipenta/rexp</code> which contains an R installation with all required packages.</li> <li> <p>Run R from the container created from the image and produce the paper's results under a directory named <code>results/</code>.</p> </li> <li> <p><strong>Note:</strong> An alternative would be to run everything from inside the container, after running it in interactive mode. To this aim, please execute the following commands:</p> <ol> <li>In <strong>Unix</strong>: <code>docker run -v${PWD}:/data --rm -ti --name shell mdipenta/rexp:latest bash</code>in <strong>Windows</strong>: <code>docker run -v %cd%:/data --rm -ti --name shell mdipenta/rexp:latest bash</code></li> <li><code>cd data</code></li> <li><code>R --no-save &lt; FuncConstructs-Statistics.r</code> After exiting the container, the "results" directory will be again populated with the study results.</li> </ol> </li> </ul> </li> <li> <p><strong>Running the R script from own R installation</strong>: this option works if one has an R installation already (or wants to use an R installation) without relying on the Docker image. The steps to be followed are:</p> <ul> <li>Uncomment the <code>install.packages(..)</code> instruction in the first lines of the script. This will allow for the installation of the required packages.</li> <li>Just run, from the current directory, the script <code>FuncConstructs-Statistics.r</code>, using the command <code>Rscript FuncConstructs-Statistics.r</code> (making sure the directory containing Rscript is in your PATH, this should work fine in Unixes, it might require to modify the PATH environment variable in Windows). Should you experience problems with the first part of the script (installations), try to execute the <code>install.packages(..)</code> statement from your R GUI, and then run the script again.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook using a Dockerized Jupyter lab with R kernel</strong>: this option allows for opening the Jupyter Notebook with all results without having to install Jupyter with the R kernel, nor all the required R packages. The steps required are:</p> <ul> <li>Run the <code>./run-jupyter-container.sh</code> (<strong>Unix</strong>) or <code>.\run-jupyter-container.bat</code> (<strong>Windows</strong>). It will download the <code>mdipenta/myjupyter</code> image and run Jupyter lab from it.</li> <li>Open a browser on <a href="http://localhost:8888/">localhost:8888</a> (or if it does not work, <a href="http://127.0.0.1:8888/">127.0.0.1:8888)</a> and, when being asked for a password, type <code>docker</code>.</li> <li>From the Jupyter lab page, open the "FuncConstruct-Statistics.ipynb" notebook, and (if you wish) re-run it, or simply browse its results. Note: differently from options 1 and 2, results are not saved, but just displayed in the notebook.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook from your installation</strong>: this is similar to Option 3, but it can work if you have already Jupyter lab installed, with the R kernel enabled (for details see: <a href="https://github.com/IRkernel/IRkernel">https://github.com/IRkernel/IRkernel</a>). The steps to follow are:</p> <ul> <li>Run jupyter lab (e.g., jupyter lab from the command line) and open it on a webpage.</li> <li>Open the <code>FuncConstruct-Statistics.ipynb</code> notebook.</li> <li>If you want to re-execute it, uncomment the <code>install.packages()</code> line.</li> <li>Re-run it (if you wish).</li> </ul> </li> </ol> <h3>The output</h3> <p>If using Option 1 or 2, the <code>results</code> directory will contain the following files:</p> <ul> <li><strong>Figures 4 and 5</strong> as in the paper.</li> <li><strong>Tables 2-9</strong> as in the paper in various formats (csv, tex, and for Tables 3-5 also .txt). Some notes: The diagnostics (top part, up to "Fixed effects") for Tables 3-5 are shown in the .txt files only. However, these files do not report the "OR" columns that correspond to exp(Estimate). This is because the .txt file contains the statistics dump which does not include the ORs. The .csv and .tex tables report the Fixed effects as shown in the paper, including the ORs.</li> <li><strong>rq1-rq2-correlation</strong> (.tex and .csv) contains the correlation analysis between RQ1 and RQ2 results as discussed in the "Threats to construct validity" (Section 6).</li> <li><strong>rq3-inter-rater</strong> (.tex and .csv) contains the results of the inter-rater agreements analysis discussed in Section 3.6.</li> </ul>

opengpl-3.0-or-laterDec 2023View details →
zenodo40/100

Replication Package of Understanding Developers Well-Being and Productivity: a 2-year Longitudinal Analysis during the COVID-19 Pandemic

<p>The COVID-19 pandemic has brought significant and enduring shifts in various aspects of life, including increased flexibility in work arrangements. In a longitudinal study, spanning 24 months with six measurement points from April 2020 to April 2022, we explore changes in well-being, productivity, social contacts, and needs of software engineers during this time. Our findings indicate systematic changes in various variables. For example, well-being and quality of social contacts increased while emotional loneliness decreased as lockdown measures were relaxed. Conversely, people&#39;s boredom and productivity, remained stable. Furthermore, a preliminary investigation into the future of work at the end of the pandemic revealed a consensus among developers for a preference of hybrid work arrangements. We also discovered that prior job changes and low job satisfaction were consistently linked to intentions to change jobs if current work conditions do not meet developers&#39; needs. This highlights the need for software organizations to adapt to various work arrangements to remain competitive employers. Building upon our findings and the existing literature, we introduce the Integrated Job Demands-Resources and Self-Determination (IJARS) Model as a comprehensive framework to explain the well-being and productivity of software engineers during the COVID-19 pandemic.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Replication Package for "Software Quality Assurance Analytics: Enabling Software Engineers to Reflect on QA Practices" Paper (SCAM 2024)

<p>Welcome to our artifact!<br>In here we provide additional information for you to retrace our steps in the interview analysis.<br>It has the following contents:</p> <ul> <li><code>codebook.xlsx</code>: Our full codebook with our open codes, structured after the axial codes that emerged. <code>codebook-statistics.xlsx</code> lists for each code in which participant's interview it can be found.</li> <li><code>generate-figures</code>: The plain data and scripts used to generate the figures in the paper.</li> <li><code>survey.pdf</code>: An printout of our whole online questionnaire that guided the participants through the pretest-posttest study and the interview.</li> <li><code>survey-answers.xlsx</code>: The complete data for our participants answers in the online survey during the interviews.</li> <li><code>repoinsights-dashboard-software</code>: The code of our prototype repoinsights. As it is under active development, this is not yet documented for replicating the study setup or extending it. Still, we are providing the source code for transparency and will publish a version with comprehensive setup instructions later.</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo40/100

On-the-Fly Syntax Highlighting: Generalisation and Speed-ups - Replication Package

<p><strong>On-the-Fly Syntax Highlighting: Generalisation and Speed-ups</strong></p> <p>On-the-fly syntax highlighting involves the rapid association of visual secondary notation with each character of a language derivation. This task has grown in importance due to the widespread use of online software development tools, which frequently display source code and heavily rely on efficient syntax highlighting mechanisms. In this context, resolvers must address three key demands: speed, accuracy, and development costs. Speed constraints are crucial for ensuring usability, providing responsive feedback for end users and minimizing system overhead. At the same time, precise syntax highlighting is essential for improving code comprehension. Achieving such accuracy, however, requires the ability to perform grammatical analysis, even in cases of varying correctness. Additionally, the development costs associated with supporting multiple programming languages pose a significant challenge. The technical challenges in balancing these three aspects explain why developers today experience significantly worse code syntax highlighting online compared to what they have locally. The current state-of-the-art relies on leveraging programming languages' original lexers and parsers to generate syntax highlighting oracles, which are used to train base Recurrent Neural Network models. However, questions of generalisation remain. This paper addresses this gap by extending previous work validation dataset to six mainstream programming languages thus providing a more thorough evaluation. In response to limitations related to evaluation performance and training costs, this work introduces a novel Convolutional Neural Network (CNN) based model, specifically designed to mitigate these issues. Furthermore, this work addresses an area previously unexplored performance gains when deploying such models on GPUs. The evaluation demonstrates that the new CNN-based implementation is significantly faster than existing state-of-the-art methods, while still delivering the same near-perfect accuracy.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Replication package for Transport and urban growth in the First Industrial Revolution

<p><span>Replication package for Alvarez-Palau, E. J., Bogart, D., Satchell, A.E., Shaw-Taylor, L.</span><strong><span> </span></strong><span>&lsquo;Transport and urban growth in the First Industrial Revolution&rsquo; Economic Journal<span>&nbsp; </span></span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Replication Package of "Exploiting Vision-Language Models in GUI Reuse"

<p>This replication package is for the paper entitled "Exploiting Vision-Language Models in GUI Reuse". The authors remain anonymous for double-blind review purposes. The package contains six files. If the paper is accepted, then the authors will move the replication package to a public repository hosted by an institution.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Replication package for "Why do people persist in sea-level rise threatened coastal regions? Empirical evidence on risk aversion and place attachment"

<p><strong>Steps to replicate the tables and figures in &ldquo;Why do people persist in sea-level rise threatened coastal regions? Empirical evidence on risk aversion and place attachment&rdquo;</strong></p> <p><em>by Ivo Steimanis, Matthias Mayer and Bj&ouml;rn Vollan</em></p> <p><strong>General information:</strong></p> <ul> <li>Instructions for replication of the results using Stata. All do-files were created in Stata 16.</li> <li>There are 4 folders (DO-FILES, DTA-FILES, OUTPUT, XLS-FILES), in the replication package. Copy these folders to your computer in a common directory</li> </ul> <p>&nbsp;</p> <p><strong>Do-files:</strong></p> <ul> <li>In the DO-FILES folder run the <strong>&ldquo;00_master.do&rdquo;</strong> to replicate the results reported in the main manuscript and the supplementary materials. The results will be saved in the OUTPUT folder. All additional Stata packages will be automatically installed.</li> <li><strong>&ldquo;01_merge_generate.do&rdquo; </strong>merges the different datasets and creates additional variables using in the analysis</li> <li><strong>&ldquo;02_analysis.do&rdquo; </strong>provides the code to replicate all figures and tables reported in the main manuscript and supplementary materials</li> </ul> <p>&nbsp;</p> <p><strong>Data sets:</strong></p> <ul> <li>&ldquo;bd_combine.dta&rdquo;: cleaned survey data from Bangladesh</li> <li>&ldquo;vn_combine.dta&rdquo;: cleaned survey data from Vietnam</li> <li>&ldquo;data_analysis.dta&rdquo;: main data set with the survey data from Bangladesh and Vietnam merged</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Replication Package of: "From Anecdote to Evidence: The Relationship Between Personality and Need for Cognition"

<p>Several anecdotes suggests that software engineers enjoy engaging in solving puzzles and other cognitive efforts.&nbsp;This tendency to engage in and enjoy effortful thinking is referred to as a person&#39;s &#39;need for cognition.&#39;&nbsp;An open question is, however, whether developers differ from the general population in their scores of need for cognition.&nbsp;To address this question, we conducted a large-scale sample study of 483 software engineers. &nbsp;Personality plays a significant role in people&#39;s behavior and is stable over time, and is therefore considered a defining characteristic of individuals.&nbsp;We analyzed the data using multiple Bayesian linear regression analyses.&nbsp;The results indicate that ca. 33% of variation in developers&#39; need for cognition can be explained by personality traits. &nbsp;Given the importance of human factors for software developer performance in general, and problem solving skills in particular, answering this question has substantial implications, such as recruitment &amp; retention, working behaviour, and teaming.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Replication Package for "Evaluating SZZ Implementations Through a Developer-informed Oracle"

<p>This is the replication package for the paper&nbsp;&quot;Evaluating SZZ Implementations Through a Developer-informed Oracle&quot; published in the 43rd International Conference on Software Engineering (ICSE 2021).&nbsp;<a href="https://arxiv.org/abs/2102.03300">https://arxiv.org/abs/2102.03300</a></p>

openmit-licenseJan 2022View details →
zenodo40/100

Replication Package for ROSDiscover: Statically Detecting Run-Time Architecture Misconfigurations in Robotics Systems

<p><strong>Replication Package for ROSDiscover: Statically Detecting Run-Time Architecture Misconfigurations in Robotics Systems</strong></p> <p>This is the replication package for the paper, ROSDiscover: Statically Detecting Run-Time Architecture Misconfigurations in Robotics Systems, which has been accepted at the International Conference on Software Architecture (ICSA), 2021. A preprint of the paper is included in this replication package (paper.pdf).</p> <p>This artifact is archived on Zenodo with the following DOI: <a href="https://doi.org/10.5281/zenodo.5834633">https://doi.org/10.5281/zenodo.5834633</a></p> <p>The study associated with this artifact was carried out by the following investigators:</p> <ul> <li><a href="http://christimperley.co.uk">Christopher S. Timperley</a> (Carnegie Mellon University)</li> <li><a href="https://tobiasduerschmid.github.io">Tobias D&uuml;rschmid</a> (Carnegie Mellon University)</li> <li><a href="https://www.cs.cmu.edu/~schmerl">Bradley Schmerl</a> (Carnegie Mellon University)</li> <li><a href="http://www.cs.cmu.edu/~garlan">David Garlan</a> (Carnegie Mellon University)</li> <li><a href="https://clairelegoues.com">Claire Le Goues</a> (Carnegie Mellon University)</li> </ul> <p>If you have any questions regarding the research or the replication package, you should contact Christopher, Tobias, or Bradley.</p> <p><strong>Abstract</strong></p> <p>Robot systems are growing in importance and complexity. Ecosystems for robot software, such as the Robot Operating System (ROS), provide libraries of reusable software components that can be configured and composed into larger systems. To support compositionality, ROS uses late binding and architecture configuration via &ldquo;launch files&rdquo; that describe how to initialize the components in a system. However, late binding often leads to systems failing silently due to misconfiguration, for example by misrouting or dropping messages entirely.</p> <p>In this paper we present ROSDiscover, which statically recovers the run-time architecture of ROS systems to find such architecture misconfiguration bugs. First, ROSDiscover constructs component level architectural models (ports, parameters) from source code. Second, architecture configuration files are analyzed to compose the system from these component models and derive the connections in the system. Finally, the reconstructed architecture is checked against architectural rules described in first-order logic to identify potential misconfigurations.</p> <p>We present an evaluation of ROSDiscover on real world, off-the-shelf robotic systems, measuring the accuracy, effectiveness, and practicality of our approach. To that end, we collected the first data set of architecture configuration bugs in ROS from popular open-source systems and measure how effective our approach is for detecting configuration bugs in that set.</p>

openmit-licenseJan 2022View details →
zenodo40/100

Replication package for: Privatization and Productivity in China

<p>This replication package contains the data and the code to generate the results reported in &quot;Privatization and Productivity in China&quot; by Yuyu Chen, Mitsuru Igami, Masayuki Sawada, and Mo Xiao, published in The RAND Journal of Economics (Winter 2021), volume 52, issue 4, pages 884&ndash;916.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Replication Package for: "Overcoming the gap between structured dependencies and change coupling"

<p>This datasets contains example projects that resulted from the empirical analysis of the history of object-oriented cyber-physical<br> systems:&nbsp;Glucosio/glucosio-android<sup>[1]</sup>&nbsp;, isl-org/OpenBot<sup>[2]</sup>, eclipse/concierge<sup>[3]</sup>, and WPIRoboticsProjects/GRIP<sup>[4]</sup>. It was created using our prototype tool <em>callgraphCA</em>&nbsp;to discover metrics of software changes, based on<br> call graph analysis and evolution and serve to display change coupling of artifacts and call graph evolution.</p> <p>Contained in the dataset are the databases generated by <em>callgraphCA </em>as well as example Notebooks that show the project results and <em>callgraphCA </em>libraries that support the analysis. For detailed information about <em>callgraphCA </em>visit&nbsp;<a href="https://github.com/GLopezMUZH/callgraphCA">GitHub: GLopezMUZH/callgraphCA</a></p> <p>[1]&nbsp;&nbsp;https://github.com/Glucosio/glucosio-android</p> <p>[2] https://github.com/isl-org/OpenBot/</p> <p>[3]&nbsp;https://github.com/eclipse/concierge</p> <p>[4]&nbsp;https://github.com/WPIRoboticsProjects/GRIP</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Worldwide Gender Differences in Public Code Contributions - Replication Package

<p><strong>Worldwide Gender Differences in Public Code Contributions - Replication Package</strong></p> <p>This document describes how to replicate the findings of the paper: Davide Rossi and Stefano Zacchiroli, 2022,&nbsp;<em>Worldwide Gender Differences in Public Code Contributions</em>. In Software Engineering in Society (ICSE-SEIS&#39;22), May 21-29, 2022, Pittsburgh, PA, USA. ACM, New York, NY, USA, 12 pages.&nbsp;<a href="https://doi.org/10.1145/3510458.3513011">https://doi.org/10.1145/3510458.3513011</a></p> <p>This document comes with the software needed to mine and analyze the data presented in the paper.</p> <p><strong>Prerequisites</strong></p> <p>These instructions assume the use of the&nbsp;<a href="https://www.gnu.org/software/bash/">bash</a>&nbsp;shell, the&nbsp;<a href="https://www.python.org/">Python</a>&nbsp;programming language, the&nbsp;<a href="https://www.postgresql.org/">PosgreSQL</a>&nbsp;DBMS (version 11 or later), the&nbsp;<a href="https://facebook.github.io/zstd/">zstd</a>&nbsp;compression utility and various usual *nix shell utilities (cat, pv, ...), all of which are available for multiple architectures and OSs.<br> It is advisable to create a&nbsp;<a href="https://docs.python.org/3/tutorial/venv.html">Python virtual environment</a>&nbsp;and install the following PyPI packages:&nbsp;<code>click==8.0.3 cycler==0.10.0 gender-guesser==0.4.0 kiwisolver==1.3.2 matplotlib==3.4.3 numpy==1.21.3 pandas==1.3.4 patsy==0.5.2 Pillow==8.4.0 pyparsing==2.4.7 python-dateutil==2.8.2 pytz==2021.3 scipy==1.7.1 six==1.16.0 statsmodels==0.13.0</code></p> <p><strong>Initial data</strong></p> <ul> <li><code>swh-replica</code>, a PostgreSQL database containing a copy of Software Heritage data. The schema for the database is available at&nbsp;<a href="https://forge.softwareheritage.org/source/swh-storage/browse/master/swh/storage/sql/">https://forge.softwareheritage.org/source/swh-storage/browse/master/swh/storage/sql/</a>.<br> We retrieved these data from&nbsp;<a href="https://www.softwareheritage.org">Software Heritage</a>, in collaboration with the archive operators, taking an archive snapshot as of 2021-07-07. We cannot make these data available in full as part of the replication package due to both its volume and the presence in it of personal information such as user email addresses. However, equivalent data (stripped of email addresses) can be obtained from the Software Heritage archive dataset, as documented in the article: Antoine Pietri, Diomidis Spinellis, Stefano Zacchiroli,&nbsp;<em>The Software Heritage Graph Dataset: Public software development under one roof</em>. In proceedings of MSR 2019: The 16th International Conference on Mining Software Repositories, May 2019, Montreal, Canada. Pages 138-142, IEEE 2019.&nbsp;<a href="http://dx.doi.org/10.1109/MSR.2019.00030">http://dx.doi.org/10.1109/MSR.2019.00030</a>.<br> Once retrieved, the data can be loaded in PostgreSQL to populate&nbsp;<code>swh-replica</code>.</li> <li><code>names.tab</code>&nbsp;- forenames and surnames per country with their frequency</li> <li><code>zones.acc.tab</code>&nbsp;- countries/territories, timezones, population and world zones</li> <li><code>c_c.tab</code>&nbsp;- ccTDL entities - world zones matches</li> </ul> <p><strong>Data preparation</strong></p> <ul> <li>Export data from the&nbsp;<code>swh-replica</code>&nbsp;database to create&nbsp;<code>commits.csv.zst</code>&nbsp;and&nbsp;<code>authors.csv.zst</code>&nbsp;<code>sh&gt; ./export.sh</code></li> <li>Run the authors cleanup script to create&nbsp;<code>authors--clean.csv.zst</code>&nbsp;<code>sh&gt; ./cleanup.sh authors.csv.zst</code></li> <li>Filter out implausible names and create&nbsp;<code>authors--plausible.csv.zst</code>&nbsp;<code>sh&gt; pv authors--clean.csv.zst | unzstd | ./filter_names.py 2&gt; authors--plausible.csv.log | zstdmt &gt; authors--plausible.csv.zst</code></li> </ul> <p><strong>Gender detection</strong></p> <ul> <li>Run the gender guessing script to create&nbsp;<code>author-fullnames-gender.csv.zst</code>&nbsp;<code>sh&gt; pv authors--plausible.csv.zst | unzstd | ./guess_gender.py --fullname --field 2 | zstdmt &gt; author-fullnames-gender.csv.zst</code></li> </ul> <p><strong>Database creation and data ingestion</strong></p> <ul> <li> <p>Create the PostgreSQL DB&nbsp;<code>sh&gt; createdb gender-commit&nbsp;</code>Notice that from now on when prepending the&nbsp;<code>psql&gt;</code>&nbsp;prompt we assume the execution of psql on the&nbsp;<code>gender-commit</code>&nbsp;database.</p> </li> <li> <p>Import data into PostgreSQL DB&nbsp;<code>sh&gt; ./import_data.sh</code></p> </li> </ul> <p><strong>Zone detection</strong></p> <ul> <li>Extract commits data from the DB and create&nbsp;<code>commits.tab</code>, that is used as input for the gender detection script<br> <code>sh&gt; psql -f extract_commits.sql gender-commit</code></li> <li>Run the world zone detection script to create&nbsp;<code>commit_zones.tab.zst</code>&nbsp;<code>sh&gt; pv commits.tab | ./assign_world_zone.py -a -n names.tab -p zones.acc.tab -x -w 8 | zstdmt &gt; commit_zones.tab.zst&nbsp;</code>Use&nbsp;<code>./assign_world_zone.py --help</code>&nbsp;if you are interested in changing the script parameters.</li> <li>Read zones assignment data from the file into the DB<br> <code>psql&gt; \copy commit_culture from program &#39;zstdcat commit_zones.tab.zst | cut -f1,6 | grep -Ev &#39;&#39;\s$&#39;&#39;&#39;</code></li> </ul> <p><strong>Extraction and graphs</strong></p> <ul> <li>Run the script to execute the queries to extract the data to plot from the DB. This creates&nbsp;<code>commits_tz.tab</code>,&nbsp;<code>authors_tz.tab</code>,&nbsp;<code>commits_zones.tab</code>,&nbsp;<code>authors_zones.tab</code>, and&nbsp;<code>authors_zones_1620.tab</code>.<br> Edit&nbsp;<code>extract_data.sql</code>&nbsp;if you whish to modify extraction parameters (start/end year, sampling, ...).&nbsp;<code>sh&gt; ./extract_data.sh</code></li> <li>Run the script to create the graphs from all the previously extracted tabfiles. This will generate&nbsp;<code>commits_tzs.pdf</code>,&nbsp;<code>authors_tzs.pdf</code>,&nbsp;<code>commits_zones.pdf</code>,&nbsp;<code>authors_zones.pdf</code>, and&nbsp;<code>authors_zones_1620.pdf</code>.&nbsp;<code>sh&gt; ./create_charts.sh</code></li> </ul> <p><strong>Additional graphs</strong></p> <p>This package also includes some already-made graphs</p> <ul> <li><code>authors_zones_1.pdf</code>: stacked graphs showing the ratio of female authors per world zone through the years, considering all authors with at least one commit per period</li> <li><code>authors_zones_2.pdf</code>: ditto with at least two commits per period</li> <li><code>authors_zones_10.pdf</code>: ditto with at least ten commits per period</li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Replication package of "The Chemical Origins of Plasma Contraction and Thermalization in CO2 Microwave Discharges" by A.W. van de Steeg et al.

<p>Title: The Chemical Origins of Plasma Contraction and Thermalization in CO2 Microwave Discharges<br> Authors: A.W. van de Steeg, L. Vialetto, A.F. Sovelas da Silva, P. Viegas, P. Diomede, M.C.M. van de Sanden, G.J. van Rooij<br> Journal: The Journal of Physical Chemistry Letters<br> Date: 2022-02-11<br> Description:&nbsp;<br> In this letter first Thomson scattering measurements in pure CO2 plasma were reported. The measurements show a thermalazation of electron temperature and gas temperature occurs with increasing pressure, coinciding with plasma contraction.<br> A 1D model is coupled with these measurements to elucidate the mechanisms behind both phenomena. Thus, it is revealed that associative ionization of radicals plays a crucial role. It connects thermal properties (gas temperature, heat and mass transport) with electon properties.&nbsp;<br> Access rights: Open<br> License: Creative Commons Attribution 4.0</p> <p><br> This replication package consists of the raw data used for the creation of each figure, ie measurement data, together with the python scripts used for data treatment.</p> <p>Most raw data is in the form of .spe files. The python library (by D.v.d. Bekerom) calibrate_fiber_pos is always used to read these camera images.</p> <p>The modelling data can be found separately in the publication of L. Vialetto (ref 20 in the Letter).<br> &nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Replication Package for Maintainability Challenges in ML: A Systematic Literature Reveiw

<p>In order to ensure transparency and reproducibility, we have&nbsp;made all study artefacts publicly here. Replication package for <strong>Maintainability Challenges in ML : A Systematic Literature Review</strong>. This package contains the data used&nbsp;for this Systematic Literature Review process and additional data synthesised from this study.</p> <ul> <li>Coding Summary Report Generated from Nvivo project. (Coding Summary By Code Report.pdf)</li> <li>Detailed List of Selected papers for Literature Review papers(Literature Review Papers.xlsx)</li> <li>Papers relating to implications for developers and Researchers(Implication for Developer and Researchers.pdf)</li> <li>Search query from Different Database.pdf</li> <li>No of papers used in to answer RQ using the&nbsp;&nbsp;Literature review paper by categories (No of Literature review papers.pdf)</li> <li>Maintainability challenges in ML workflow from SLR.pdf</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Replication package for "To What Extent do Deep Learning-based Code Recommenders Generate Predictions by Cloning Code from the Training Set?

<p>Replication package for &quot;To What Extent do Deep Learning-based Code Recommenders Generate Predictions by Cloning Code from the Training Set?&quot;</p>

openmit-licenseApr 2022View details →
zenodo40/100

Replication package of `Can We Automatically Generate Class Comments in Pharo?`

<pre><code class="language-markdown"># RP Automatic-comment-generation This folder contains all the material needed to replicate the experiments. ## Content - [RP Automatic-comment-generation](#automatic-comment-generation) - Appendix.pdf - [Content](#content) - [Dataset/](#dataset/) - [Online evaluation/](#online-evaluation/) - [Results/](#results/) - [SI-Approach/](#SI-Approach/) ## Dataset/ This contains the data for the online evaluation - #### Online evaluation/ Contains set of questions used in online evaluation, results of these questions and lists of classes used for each online evaluation - [Class_understanding_questions.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Class_understanding_questions.xlsx) Questions used to determin if a participant understood the functionality of a class. - [Q_and_A_class_comment_characteristics.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Q_and_A_class_comment_characteristics) Questions and possible answers used for the evaluation of the characteristics adequacy, conciseness and comprehensibility of a generated class comment. - [Q_and_A_what_participants_write_and_look_for.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Q_and_A_what_participants_write_and_look_for.xlsx) Questions and possible answers used to determine what participants look for and write in class comments. - [Classes_used_in_evaluation1.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation1.xlsx) Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 1. - [Classes_used_in_evaluation2.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation2.xlsx) Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 2. - [Classes_used_in_evaluation3.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation3.xlsx) Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 3. - [Classes_used_in_evaluation4.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation4.xlsx) Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 4. - [How_often_participants_write_comments.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\How_often_participants_write_comments.xlsx) Results for distribution of how often participants claim to write comments. - [How_participants_follow_class_comment_template.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\How_participants_follow_class_comment_template.xlsx) Results for distribution of how participants follow the class comment template. - [What_participants_look_for_in_comments.xlsx](RP-Automatic-comment-generation\Dataset\Online_evaluation\What_participants_look_for_in_comments.xlsx) Results of evaluation for what participants look for in class comments. - [What_participants_write_in_comments.xlsx](\RP-Automatic-comment-generation\Dataset\Online_evaluation\What_participants_write_in_comments.xlsx) Results of evaluation for what participants write in class comments. ## Results/ - #### SI-Approach/ Contains all the data for the results related to SI-Approach and the scripts used to generate said data. - [Class_stereotype_distribution.xlsx](RP-Automatic-comment-generation\Results\SI-Approach\Class_stereotype_distribution.xlsx) Extracted class stereotypes for 350 random classes in the Pharo base image. Evaluated in 10 classes per step. - [Class_stereotype_script.txt](RP-Automatic-comment-generation\Results\SI-Approach\Class_stereotype_script.txt) Script used in the Playground of Pharo to return counts of how many classes get assigned the specific class stereotypes. Returns numbers for the class stereotypes in alphabetical order. Boundary =&gt; Small - [Method_stereotype_dstribution.xlsx](RP-Automatic-comment-generation\Results\SI-Approach\Method_stereotype_distribution.xlsx) Extracted method stereotypes for 500 random classes in the Pharo base image. Evaluated in 20 classes per step. - [Method_stereotype_script.txt](RP-Automatic-comment-generation\Results\SI-Approach\Method_stereotype_script.txt) Script used in the Playground of Pharo to return counts of how many methods get assigned the specific method stereotypes. Returns numbers for the method stereotypes in the order Accessors, Getters, Mutators, Setters, Collaborators, Controllers, Factories and Degenerate. </code></pre> <p>Replication package of &#39;Can We Automatically Generate Class Comments in Pharo?&#39;</p> <pre><code class="language-markdown"># RP Automatic-comment-generation This folder contains all the material needed to replicate the experiments. ## Content - [RP Automatic-comment-generation](#automatic-comment-generation)     - [Content](#content)     - [Dataset/](#dataset/)         - [Online evaluation/](#online-evaluation/)     - [Results/](#results/)         - [SI-Approach/](#SI-Approach/) ## Dataset/ This contains the data for the online evaluation - #### Online evaluation/     Contains set of questions used in online evaluation, results of these questions and lists of classes used for each online evaluation     - [Class_understanding_questions.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Class_understanding_questions.xlsx)         Questions used to determin if a participant understood the functionality of a class.     - [Q_and_A_class_comment_characteristics.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Q_and_A_class_comment_characteristics)         Questions and possible answers used for the evaluation of the characteristics adequacy, conciseness and comprehensibility of a generated class comment.     - [Q_and_A_what_participants_write_and_look_for.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Q_and_A_what_participants_write_and_look_for.xlsx)         Questions and possible answers used to determine what participants look for and write in class comments.     - [Classes_used_in_evaluation1.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation1.xlsx)         Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 1.      - [Classes_used_in_evaluation2.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation2.xlsx)         Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 2.      - [Classes_used_in_evaluation3.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation3.xlsx)         Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 3.      - [Classes_used_in_evaluation4.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\Classes_used_in_evaluation4.xlsx)         Extracted classes with information to LOC, instance variables, number of methods, number of classes using it and class stereotypes for evaluation 4.      - [How_often_participants_write_comments.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\How_often_participants_write_comments.xlsx)         Results for distribution of how often participants claim to write comments.      - [How_participants_follow_class_comment_template.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\How_participants_follow_class_comment_template.xlsx)         Results for distribution of how participants follow the class comment template.     - [What_participants_look_for_in_comments.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\What_participants_look_for_in_comments.xlsx)         Results of evaluation for what participants look for in class comments.     - [What_participants_write_in_comments.xlsx](thesis\RP-Automatic-comment-generation\Dataset\Online_evaluation\What_participants_write_in_comments.xlsx)         Results of evaluation for what participants write in class comments. ## Results/ - #### SI-Approach/     Contains all the data for the results related to SI-Approach and the scripts used to generate said data.     - [Class_stereotype_distribution.xlsx](thesis\RP-Automatic-comment-generation\Results\RQ2\Class_stereotype_distribution.xlsx)     Extracted class stereotypes for 350 random classes in the Pharo base image. Evaluated in 10 classes per step.     - [Class_stereotype_script.txt](thesis\RP-Automatic-comment-generation\Results\RQ2\Class_stereotype_script.txt)     Script used in the Playground of Pharo to return counts of how many classes get assigned the specific class stereotypes. Returns numbers for the class stereotypes in alphabetical order. Boundary =&gt; Small     - [Method_stereotype_dstribution.xlsx](thesis\RP-Automatic-comment-generation\Results\RQ2\Method_stereotype_distribution.xlsx)     Extracted method stereotypes for 500 random classes in the Pharo base image. Evaluated in 20 classes per step.     - [Method_stereotype_script.txt](thesis\RP-Automatic-comment-generation\Results\RQ2\Method_stereotype_script.txt)     Script used in the Playground of Pharo to return counts of how many methods get assigned the specific method stereotypes. Returns numbers for the method stereotypes in the order Accessors, Getters, Mutators, Setters, Collaborators, Controllers, Factories and Degenerate.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record