Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
192
datasets available to search
ShareScore release 0.7.1
Dataset results
192 results for “Software engineering”
The Role of Informal Communication in Building Shared Understanding of Non-Functional Requirements in Remote Continuous Software Engineering
<p><strong>Study Information</strong></p> <p>We conducted an ethnography-informed case study of a remote software organization that adopts CSE practices to explore how the organization builds a shared understanding of NFRs. Our study uses semi-structured interviews with a period of observations to answer the following research questions:</p> <p> </p> <ol> <li> <p>How does a remote software organization that adopts CSE practices reach a shared understanding of NFRs?</p> </li> <li> <p>What are the limitations to the shared understanding of NFRs in a remote software organization that adopts CSE practices?</p> </li> <li> <p>What organizational practices for remote collaboration supported a shared understanding of NFRs?</p> </li> </ol> <p> </p> <p>In our study, we refer to our partner organization as Alpha. We used ethnography-informed methods to study Alpha's practices and processes and how they approach a shared understanding of NFRs in their product development. </p> <p> </p> <p><strong>Data Analysis</strong></p> <p>We performed a qualitative study through semi-structured interviews and observations. We use the open, axial and selective coding approach from grounded theory [1] to create our codebook, which informed the results and discussion of our study. Two independent coders held agreement sessions to discuss the codes, consolidate the codes and calculate the inter-rater reliability using the Cohen Kappa's coefficient for measuring observer agreement for categorical data [2]. </p> <p> </p> <p><strong>Artifact Descriptions</strong></p> <p>Our replication package contains three artifacts:</p> <p>1. Codebook.csv: The codebook contains rows for the list of codes used, including the code name and the description of the codes. The codes are the final set of themes derived during the thematic analysis of the interview responses. For example, 'Gaps in communication' means when interview participants describe miscommunications due to team members making assumptions about a project/process or having unclear expectations for a project.</p> <p>2. kappa-scores.csv: This contains the associated kappa values for each round of inter-rater agreement sessions. For each agreement session, the Cohen Kappa's coefficient was calculated from the number of agreements and disagreements of codes within one or two interview transcripts. The Kappa values represent the level of agreement ranging from 0 to 1, where > 0.6 represents substantial agreement. </p> <p>3. Interview-questions.csv: This contains the interview questions used in the semi-structured interviews. Some of the interview questions varied depending on the interviewee’s role, experience and the flow of the interviews.</p> <p><strong> </strong></p> <p><strong>Usefulness</strong></p> <p>We recognize that the value and usefulness of our replication package are yet-to-be-determined. In the interest of transparency of open science, we published our artifacts. We hope that these artifacts are useful to either replicate our findings or to further analyze them to produce other enlightening results.</p> <p><strong> </strong></p> <p><strong>References</strong></p> <p>1. Rashina Hoda, James Noble, and Stuart Marshall. "Grounded theory for geeks". In: Proceedings of the 18th conference on pattern languages of programs. 2011, pp. 1–17.</p> <p>2. J Richard Landis and Gary G Koch. "The measurement of observer agreement for categorical data". In: biometrics (1977), pp. 159–174.</p> <p><strong> </strong></p> <p> </p>
The Daily Life of Software Engineers during the COVID-19 Pandemic -- Replication Package
<p>Following the onset of the COVID-19 pandemic and subsequent lockdowns, software engineers' daily life was disrupted and abruptly forced into remote working from home. This change deeply impacted typical working routines, affecting both well-being and productivity. Moreover, this pandemic will have long-lasting effects in the software industry, with several tech companies allowing their employees to work from home indefinitely if they wish to do so. Therefore, it is crucial to analyze and understand how a typical working day looks like when working from home and how individual activities affect software developers' well-being and productivity. We performed a two-wave longitudinal study involving almost 200 globally carefully selected software professionals, inferring daily activities with perceived well-being, productivity, and other relevant psychological and social variables. Results suggest that the time software engineers spent doing specific activities from home was similar when working in the office. (e.g., coding > emails > code review > networking). However, we also found some meaningful mean differences. The amount of time developers spent on each activity was unrelated to their well-being, perceived productivity, and other variables. We conclude that working remotely is not per se a challenge for organizations or developers.</p>
Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering
<p>Update: We updated the data set in March 2022 by adding newly published papers and by providing more insights on how we analyzed them. Details can be found in the file " SEnti-SMS.xlsx".</p> <p>----------</p> <p>Update: The updated version (-v2) contains the results of one more snowballing iteration and extracted information on the accuracy of the used methods.</p> <p>----------</p> <p>In 2020, we conducted a systematic literature review to explore the development and application of sentiment analysis tools in software engineering.</p> <p>Information on the execution of the SLR, its scope, the search string, etc. are presented in the paper linked below.</p> <p> </p> <p> </p>
Replication data [The who, what and how of the current research at the Brazilian Symposium on Software Engineering]
<p>Replication Data for the SBES paper <em>"The who, what and how of the current research at the Brazilian Symposium on Software Engineering"</em></p> <p>Dataset containing analysis (who, what and how) of 90 SBES papers: 27 from SBES’19, 43 from SBES’20, and 20 from SBES’21.</p>
Machine Learning for Software Engineering: A Tertiary Study
<p>Dataset of the research paper: <strong>Machine Learning for Software Engineering: A Tertiary Study</strong></p> <p>Machine learning (ML) techniques increase the effectiveness of software engineering (SE) lifecycle activities. We systematically collected, quality-assessed, summarized, and categorized 83 reviews in ML for SE published between 2009–2022, covering 6,117 primary studies. The SE areas most tackled with ML are software quality and testing, while human-centered areas appear more challenging for ML. We propose a number of ML for SE research challenges and actions including: conducting further empirical validation and industrial studies on ML; reconsidering deficient SE methods; documenting and automating data collection and pipeline processes; reexamining how industrial practitioners distribute their proprietary data; and implementing incremental ML approaches.</p> <p>The following data and source files are included.</p> <ul> <li><strong>review-protocol.md</strong>: The protocol employed in this tertiary study</li> </ul> <p><strong>data/</strong></p> <p><strong> dl-search/</strong></p> <p><strong> input/</strong></p> <ul> <li><strong>acm_comput_surveys_overviews.bib</strong>: Surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_surveys_overviews_titles.txt</strong>: Titles of surveys</li> <li><strong>acm_comput_ml_surveys.bib</strong>: Machine learning (ML)-related surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_ml_surveys_titles.txt</strong>: Titles of ML-related surveys</li> <li><strong>dl_search_queries.txt</strong>: Search queries applied to IEEE Xplore, ACM Digital Library, and Elsevier Scopus</li> <li><strong>ml_keywords.txt</strong>: ML-related keywords extracted from ML-related survey titles and used in the search queries</li> <li><strong>se_keywords.txt</strong>: Software Engineering (SE)-related keywords derived from the 15 SWEBOK Knowledge Areas (KAs—except for Computing Foundations, Mathematical Foundations, and Engineering Foundations) and used in the search queries</li> <li><strong>secondary_studies_keywords.txt</strong>: Survey-related keywords composed of the 15 keywords introduced in the tertiary study on SLRs in SE by Kitchenham <em>et al.</em> (2010), and the survey titles, and used in the search queries</li> </ul> <p><strong> output/</strong></p> <ul> <li><strong>acm/</strong> <ul> <li><strong>acm{1–9}.bib</strong>: Search results from ACM Digital Library</li> </ul> </li> <li><strong>ieee.csv</strong>: Search results from IEEE Xplore</li> <li><strong>scopus_analyze_year.csv</strong>: Yearly distribution of ML and SE documents extracted from Scopus's <em>Analyze search results</em> page</li> <li><strong>scopus.csv</strong>: Search results from Scopus</li> </ul> <p><strong> study-selection/</strong></p> <ul> <li><strong>backward_snowballing.csv</strong>: Additional secondary studies found through the backward snowballing process</li> <li><strong>backward_snowballing_references.csv</strong>: References of quality-accepted secondary studies</li> <li><strong>cohen_kappa_agreement.csv</strong>: Inter-rater reliability of reviewers in study selection</li> <li><strong>dl_search_results.csv</strong>: Aggregated search results of all three digital libraries</li> <li><strong>forward_snowballing_reviewer_{1,2}.csv</strong>: Divided forward snowballing citations of quality-accepted studies assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> <li><strong>study_selection_reviewer_{1,2}.csv</strong>: Divided search results assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> </ul> <p><strong> quality-assessment/</strong></p> <ul> <li><strong>dare_assessment.csv</strong>: Quality assessment (QA) of selected secondary studies based on the Database of Abstracts of Reviews of Effects (DARE) criteria by York University, Centre for Reviews and Dissemination</li> <li><strong>quality_accepted_studies.csv</strong>: Details of quality-accepted studies</li> <li><strong>studies_for_review.bib</strong>: Bibliography details and QA scores of selected secondary studies</li> </ul> <p><strong> data-extraction/</strong></p> <ul> <li><strong>further_research.csv</strong>: Recommendations for further research of quality-accepted studies</li> <li><strong>further_research_general.csv</strong>: The complete list of associated studies for each general recommendation</li> <li><strong>knowledge_areas.csv</strong>: Classification of quality-accepted studies using the SWEBOK KAs and subareas</li> <li><strong>ml_techniques.csv</strong>: Classification of the quality-accepted studies based on a four-axis ML classification scheme, along with extracted ML techniques employed in the studies</li> <li><strong>primary_studies.csv</strong>: Details of reviewed primary studies by the quality-accepted secondary</li> <li><strong>research_methods.csv</strong>: Citations of the research methods employed by the quality-accepted studies</li> <li><strong>research_types_methods.csv</strong>: Research types and methods employed by the quality-accepted studies</li> </ul> <p><strong>src/</strong></p> <ul> <li><strong>data-analysis.ipynb</strong>: Analysis of data extraction results (data preprocessing, top authors and institutions, study types, yearly distribution of publishers, QA scores, and SWEBOK KAs) and creation of all figures included in the study</li> <li><strong>scopus-year-analysis.ipynb</strong>: Yearly distribution of ML and SE publications retrieved from Elsevier Scopus</li> <li><strong>study-selection-preprocessing.ipynb</strong>: Processing of digital library search results to conduct the inter-rater reliability estimation and study selection process</li> </ul>
Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis
<p>Dataset of the research paper: <strong>Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis</strong></p> <p>Existing work on the practical impact of software engineering (SE) research examines industrial relevance rather than adoption of study results, hence the question of how results have been practically applied remains open. To answer this and investigate the outcomes of impactful research, we performed a quantitative and qualitative analysis of 4 354 SE patents citing 1 690 SE papers published in four leading SE venues between 1975–2017. Moreover, we conducted a survey on 475 authors of 593 top-cited and awarded publications, achieving 26% response rate. Overall, researchers have equipped practitioners with various tools, processes, and methods, and improved many existing products. SE practice values knowledge-seeking research and is impacted by diverse cross-disciplinary SE areas. Practitioner-oriented publication venues appear more impactful than researcher-oriented ones, while industry-related tracks in conferences could enhance their impact. Some research works did not reach a wide footprint due to limited funding resources or unfavorable cost-benefit trade-off of the proposed solutions. The need for higher SE research funding could be corroborated through a dedicated empirical study. In general, the assessment of impact is subject to its definition. Therefore, academia and industry could jointly agree on a formal description to set a common ground for subsequent research on the topic.</p> <p>The following data files are included.</p> <ul> <li><em>./fields</em>: <ul> <li><strong>engi-fields.csv</strong>: Publication and PhD dissertation counts of main engineering branches</li> <li><strong>engi-fields-queries.txt</strong>: Queries applied to Elsevier's Scopus and Open Access Theses and Dissertations databases to retrieve the publication and dissertation counts</li> </ul> </li> <li><em>./patents</em>: <ul> <li><strong>sample-se-references-verified.csv</strong>: Manual verification of a random sample of references by software engineering (SE) patents to SE papers</li> <li><strong>se-cpc.tsv</strong>: Manually-identified SE-related Cooperative Patent Classification (CPC) categories</li> <li><strong>se-references-in-patents.csv</strong>: SE references made by SE patents to SE papers</li> <li><em>./patents/litigation</em>: <ul> <li><strong>case-values.csv</strong>: Manually-retrieved litigation damages of citing SE patents</li> <li><strong>lit-per-paper.csv</strong>: Litigation cases of citing SE patents</li> </ul> </li> <li><em>./patents/maintenance</em>: <ul> <li><strong>maint-code-fee-mapping.csv</strong>: Mapping of patent maintenance fee codes to their fee values</li> <li><strong>maint-fees.csv</strong>: Fee values of maintenance fee codes</li> <li><strong>maint-per-paper.csv</strong>: Maintenance fee events of citing SE patents</li> </ul> </li> <li><em>./patents/reports</em>: <ul> <li><strong>lit-sum-per-paper.csv</strong>: Counts and total damages of litigation cases of patent-cited SE papers</li> <li><strong>maint-sum-per-paper.csv</strong>: Counts and total values of maintenance fee events of patent-cited SE papers</li> <li><strong>patent-ref-counts.csv</strong>: SE patent citation counts of patent-cited SE papers</li> </ul> </li> </ul> </li> <li><em>./survey</em>: <ul> <li><strong>emse-top.csv</strong>: Most-cited papers of the Empirical Software Engineering (EMSE) journal</li> <li><strong>icse-bp.csv</strong>: Distinguished papers of the International Conference of Software Engineering (ICSE)</li> <li><strong>icse-mip.csv</strong>: Most influential ICSE papers</li> <li><strong>icse-top.csv</strong>: Most-cited ICSE papers</li> <li><strong>survey-questionnaire-emse.pdf</strong>: The EMSE survey questionnaire</li> <li><strong>survey-questionnaire.pdf</strong>: The ICSE, TSE, and TOSEM survey questionnaire</li> <li><strong>survey-responses.csv</strong>: The anonymized survey responses</li> <li><strong>tosem-top.csv</strong>: Most-cited papers of the ACM Transactions on Software Engineering and Methodology (TOSEM)</li> <li><strong>tse-top.csv</strong>: Most-cited papers of the IEEE Transactions on Software Engineering (TSE)</li> <li><em>./survey/manual-coding</em>: <ul> <li><strong>feedback.txt</strong>: Manual coding of survey feedback</li> <li><strong>practical-impact.csv</strong>: Manual coding of responses about practical impact of work</li> <li><strong>practical-impact-lack.csv</strong>: Manual coding of responses about lack of practical impact</li> <li><strong>research-methods.csv</strong>: Manual coding of additional research methods of surveyed papers</li> <li><strong>state-of-practice.csv</strong>: Manual coding of responses about changes in state of practice</li> </ul> </li> </ul> </li> <li><em>./venues</em>: <ul> <li><strong>se-venues.csv</strong>: Top SE venues according to Google Scholar Metrics</li> <li><strong>se-venues-impact.csv</strong>: SE patent citations and patent-based impact factors of SE venues</li> <li><strong>se-venues-scopus-queries.txt</strong>: Queries applied to Scopus to retrieve the publication counts of the SE venues</li> </ul> </li> </ul>
Dataset: Economical Accommodations for Neurodivergent Students in Software Engineering Education: Experiences from an Intervention in Four Undergraduate Courses
<p>This dataset contains anonymised raw data and examples of accommodations made for neurodiverse students in four undergraduate courses in Computer Science and Software Engineering programmes. The dataset is published as a part of a book chapter in which we report the accommodations.</p> <p>Overall guidelines we followed, including their sources, are contained in <strong>guidelines.md.</strong></p> <p>The raw data for the two surveys is contained in the two Excel files <strong>survey1.xlsx</strong> and <strong>survey2.xlsx</strong>. Free-text answers have been aggregated by neurodiverse and neurotypical students and anonymised, and are available in the files<strong> survey1_freetext_neurodiverse.txt, survey1_freetext_neurotypical.txt, survey2_freetext_neurodiverse.txt, </strong>and<strong> survey2_freetext_neurotypical.txt.</strong></p> <p>The remaining files are examples of the adapted lecture slides and assignment texts. Here, files starting with WEBcourse are from a mandatory undergraduate course on web development, while files starting with SEcourse are from a mandatory undergraduate course giving an overview of Software Engineering.</p>
Student's logs and perceptions of an automated assessment tool in a software engineering MOOC specialization
<p>Our dataset contains students' perceptions and usage of an automated assessment tool (MOOCauto) for obtaining formative feedback in software engineering assignments that are part of a MOOC specialization at Universidad Politécnica de Madrid (Spain), delivered by the MiriadaX platform. The dataset has previously been used in a study to evaluate students' perceptions of the tool and to analyze their usage patterns using Growth Mixture Models <a href="https://www.computer.org/csdl/magazine/so/5555/01/10196480/1P9AhkBLYXK">(López-Pernas et al., 2023)</a>. The code of each of the assignments is available on Github: <a href="https://github.com/ging-moocs">https://github.com/ging-moocs</a>.</p> <p>Our dataset contains two files:</p> <h2>MOOCauto usage logs</h2> <p>The first file is called<strong> moocauto_logs.csv </strong>and it contains 9,108 anonymized logs of students' use of the automated assessment tool in the MOOC specialization assignments. The columns of the dataset are as follows:</p> <ul> <li><strong>MOOCid</strong>: Unique numeric identifier for the MOOC (1-4)</li> <li><strong>MOOC: </strong>Name of the MOOC: Frontend Development, Backend Development, Git & Github, Fullstack Development</li> <li><strong>AssignmentName</strong>: Name of the assignment.</li> <li><strong>AssignmentId</strong>: Unique identifier for each assignment (1-17)</li> <li><strong>user: </strong>Unique identifier of the student (it varies per assignment)</li> <li><strong>timestamp: </strong>Time in which the assessment was performed</li> <li><strong>score</strong>: Score obtained (0-10)</li> </ul> <h2>Students' perceptions of MOOCauto</h2> <p>The second file is called <strong>moocauto_questionnaire.csv</strong> and it contains 213 students' responses to the questionnaire conducted at the end of each MOOC in order to evaluate their opinion of the tool and perception on usefulness, ease of use, and other aspects related to the Technology Acceptance Model (TAM). The questions were as follows:</p> <ul> <li><strong>What is your general opinion of MOOCauto?</strong> (1 Horrible - 5 Excellent)</li> <li><strong>Indicate your level of agreement with the following statements </strong>(1 Strongly disagree - 5 Strongly agree) <ul> <li>MOOCauto has been easy to install</li> <li>MOOCauto has been easy to use</li> <li>The feedback provided by MOOCauto was easy to understand</li> <li>The feedback provided by MOOCauto was useful</li> <li>The feedback provided by MOOCauto helped me improve my assignments</li> <li>The documentation Of MOOCauto was useful</li> <li>MOOCauto has increased my motivation to work on the assignments</li> <li>I prefer the feedback from MOOCauto than from peer assessment</li> <li>I would like to have a bot like MOOCauto in other MOOCs</li> </ul> </li> <li><strong>How useful do you perceive the following features of MOOCauto?</strong> (1 Useless - 5 Very useful) <ul> <li>It works locally on my computer</li> <li>It allows to run the test suite as many times as I want</li> <li>It provides instantaneous feedback every time the test suite is executed</li> <li>It has documentation that explains its use and available options</li> </ul> </li> </ul>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>This is a course project and I collect the data in a rush.</p> <p>If you want to use this dataset and find any error, please contact me ;-)</p> <p>My email: echo.xiangchen@gmail.com</p>
Dataset and replication package for Temporal Discounting in Software Engineering: A Replication Study
<p>Dataset and replication package for the paper Temporal Discounting in Software Engineering: A Replication Study (Fagerholm, F., Becker, C., Chatzigeorgiou, A., Betz, S., Duboc, L., Penzenstadler, B., Mohanani, R., Venters, C. (2019). Temporal Discounting in Software Engineering: A Replication Study. 13th ACM/IEEE International Symposium of Empirical Software Engineering and Measurement (ESEM 2019)). The dataset consists of answers to a questionnaire on temporal discounting in a technical debt context. Two questionnaire templates illustrate how to gather the data for professional and student participants. An analysis script is provided which shows the details of the calculations and analyses performed for the paper. More information is given in the description file.</p>
Open dataset for publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems"
<p>This publication contains open dataset for the journal publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems".</p> <p>The dataset contains the data extracted from 280 selected primary studies.</p> <p>The dataset includes the following data:</p> <ul> <li>study metadata (title, venue, publication year, authors, authors’ affiliation, abstract);</li> <li>challenges to regulatory compliance (direct excerpts from studies);</li> <li>categories of challenges to compliance;</li> <li>principles and practices (direct excerpts from text);</li> <li>categories of principles and practices;</li> <li>types of automation of principles and practices;</li> <li>involved stakeholders (direct excerpts from studies);</li> <li>categories of involved stakeholders;</li> <li>phase of the principle and practice life cycle for which involvement of stakeholders was considered;</li> <li>SDLC process areas covered by the study;</li> <li>regulations considered in the study;</li> <li>fields of regulations that were considered;</li> <li>domains of application that were considered;</li> <li>assessment of rigor and relevance of the study.</li> </ul>
Software Engineering Education Knowledge versus Industrial Needs
<p>Dataset of the research paper: <strong>Software Engineering Education Knowledge versus Industrial Needs</strong></p> <p><em>Contribution</em>: Determine and analyze the gap between software practitioners’ education outlined in the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and industrial needs pointed by Wikipedia articles referenced in Stack Overflow (SO) posts.<br> <em>Background</em>: Previous work has uncovered deficiencies in the coverage of computer fundamentals, people skills, software processes, and human-computer interaction, suggesting rebalancing.<br> <em>Research Questions</em>: 1) To what extent are developers’ needs, in terms of Wikipedia articles referenced in SO posts, covered by the SEEK knowledge units? 2) How does the popularity of Wikipedia articles relate to their SEEK coverage? 3) What areas of computing knowledge can be better covered by the SEEK knowledge units? 4) Why are Wikipedia articles covered by the SEEK knowledge units cited on SO?<br> <em>Methodology</em>: Wikipedia articles were systematically collected from SO posts. The most cited were manually mapped to the SEEK knowledge units, assessed according to their degree of coverage. Articles insufficiently covered by the SEEK were classified by hand using the 2012 ACM Computing Classification System. A sample of posts referencing sufficiently covered articles was manually analyzed. A survey was conducted on software practitioners to validate the study findings.<br> <em>Findings</em>: SEEK appears to cover sufficiently computer science fundamentals, software design and mathematical concepts, but less so areas like the World Wide Web, software engineering components, and computer graphics. Developers seek advice, best practices and explanations about software topics, and code review assistance. Future SEEK models and the computing education could dive deeper in information systems, design, testing, security, and soft skills.</p> <p>The following data files are included.</p> <ul> <li><strong>wikipedia_articles.csv</strong>: Wikipedia articles mapped to the knowledge units of the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and the first and second level categories of the 2012 ACM Computing Classification System (CCS).</li> <li> <p><strong>posts_analysis.csv</strong>: Stack Overflow post data and metadata.</p> </li> <li> <p><strong>posts_aggregated_codes.csv</strong>: The aggregated codes that resulted from the manual analysis of the Stack Overflow posts by grouping individual keywords assigned to the posts.</p> </li> <li> <p><strong>survey_questionnaire.csv</strong>: The final survey questionnaire.</p> </li> <li> <p><strong>survey_responses.csv</strong>: Anonymized responses of the final survey questionnaire. (E-mail addresses have been excluded for privacy reasons.)</p> </li> </ul>
Diversity Awareness in Software Engineering Participant Research
<p>This dataset contains the result of a classification of three ICSE venues namely, ICSE 2019, 2020, and 2021 technical tracks, as stated in the methodology of the paper “Diversity awareness in software engineering participant studies” by Dutta et al. (2023).</p>
Rapid Review Dataset for Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team
<p>A collection of evidence briefings produced through a rapid literature review protocol the Department of Software Engineering and Research at Sandia National Laboratories. These briefings are described in our research paper, "Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team", which was accepted for publication at the 1st Annual Conference of the United States Research Software Engineer Association (US-RSE'23).</p> <p>Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525. SAND2023-06549O.</p>
A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective - Supplementary Material
<p>This repository contains the supplementary material of the paper "A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective" <br> (DOI https://doi.org/10.1145/3613372.3613407) accepted at the Research Track of the <br> XXXVII Brazilian Symposium on Software Engineering (SBES 2023).</p> <p>The artifacts are a result of a thematic synthesis of grey literature <br> performed to investigate the meaning, importance, practices, and effects of empathy <br> from the perspective of software practitioners. <br> The analysis was based on web articles from DEV, an online community used by software developers. <br> The data were collected and stored in the repository to preserve the evidence and ensure the study’s replicability.<br> <br> The repository contains the following material:</p> <p>1- <all codes.ods> and <all codes.xlsx><br> All codes generated in the data extraction process, considering research questions RQ1-RQ5:<br> The two files have the same content in different formats - ODS and XLSX.</p> <p>2 - <dataset.csv> <br> The list of web articles collected from the DEV in CSV format with all inclusion and <br> exclusion information, plus demographic data.</p> <p>3 - <empathy-framework.jpg><br> Figure 3 of the paper: A conceptual map of the meaning (boxes in orange) and <br> the value (boxes in blue) of empathy according to the software practitioners</p> <p>4 - <empathy-model.jpg> <br> Figure 4 of the paper: A conceptual framework for communication and collaboration (A), <br> management and leadership (B), coding (C), and code review (D).</p> <p>5 - <extraction.ods> and <extraction.xlsx>. The data extracted from the web articles, <br> including codes and quotes for each research question. <br> The two files have the same content in different formats - ODS and XLSX.</p>
Hackathons as a Pedagogical Strategy to Engage Students to Learn and to Adopt Software Engineering Practices
<p>Teaching Software Engineering is not a trivial duty since several pedagogical strategies can be used and sometimes the impact of these on students is uncertain. Hackathons are similar to marathons, however used to produce solutions to solve a specific problem in a short period of time and based on intense collaboration. Educational hackathons aim to promote learning in such an environment. The Undergraduate computing programs of PUCRS decided to use a hackathon as a pedagogical strategy aiming to motivate the students to practice the adoption of software development practices and to work in groups as a means to practice the development of social skills. Therefore, we conducted a case study to investigate: 1) The motivations to students to attend or not attend an educational hackathon, 2) The students perceptions about this hackathon, 3) The Software Engineering practices adopted by students. In this study, we identified factors that may affect students motivation to participate (e.g., improve the teamwork skills), some students expectations about the hackathon (e.g., work in teams), and the practices adopted by the students (e.g., pair programming). Some of our findings include that students enjoy participating in an informal educational environment (e.g., hackathons) to improve their technical skills and to build network with some colleagues. This study can provide insights to teachers that wants to organize some activity than traditional teaching and the students perspective about this kind of strategy.</p>
A Hybrid Feature Location Technique for Re-engineering Single Systems into Software Product Lines
<p>The dataset used for evaluating the hybrid feature location technique presented in the paper: "A Hybrid Feature Location Technique for Re-engineering Single Systems into Software Product Lines". This enables reproducibility, evaluation, and comparison of our study.</p> <p>_________________________________________________________________________________________________________</p> <p>Folder "Dataset" contains for each subject system used:</p> <p>(i) the artificial variants and their configurations;</p> <p>(ii) the ECCO repository containing the traces;</p> <p>(iii) the ground truth and composed variants;</p> <p>(iv) the metrics results.</p> <p>_________________________________________________________________________________________________________</p> <p>Folder "Scenarios" contains for each subject system used:</p> <p>(i) the videos recorded from exercising features on GUI.</p>
A Survey on the Adoption of Patterns for Engineering Software for the Cloud - Response Dataset
<p>This work takes as a starting point a collection of patterns for engineering software for the cloud and tries to find how they are regarded and adopted by professionals. We investigate (1) their relevance for professional software developers, (2) the extent to which product and company characteristics influence their adoption, and (3) how adopting some patterns might correlate with the likelihood of adopting others. For this purpose, we surveyed 102 practitioners using an online questionnaire. </p> <p>Amongst other findings, we conclude that most companies are using these patterns, with the overwhelming majority (97%) using at least one. We observe that the mean pattern adoption tends to increase as companies mature, namely when varying the product operation complexity, active monthly users, and company size. Finally, we establish clear correlations in the adoption of specific pairs of patterns, with conditional probabilities as high as 94%, which hints on how some practices are dependent or influence the adoption of others.</p>
Research Software Engineers Supporting Science: Survey Responses
<p>Raw survey data for "Not everyone can use git: Research Software Engineers’ recommendations for scientist-centred software support (and what researchers think of them)", a talk given by Caroline Jay at RSE16, Manchester, UK.</p>
Understanding the Building Blocks of Accountability in Software Engineering
<p>In the social and organizational sciences, accountability has been linked to the efficient operation of organizations. However, it has received limited attention in software engineering (SE) research, in spite of its central role in the most popular software development methods (i.e., Scrum). In this article, we seek to explore the mechanisms of accountability in SE environments and investigate the factors that foster software engineers' individual accountability within their teams through an interview study with 12 people. Our findings recognize two primary forms of accountability shaping software engineers individual senses of accountability: \emph{institutionalized} and \emph{grassroots}. While the former is directed by formal processes and mechanisms, like performance reviews, grassroots accountability arises organically within teams, driven by factors such as peers' expectations and intrinsic motivation. This organic form cultivates a shared sense of collective responsibility, emanating from shared team standards and individual engineers' inner commitment to their personal, professional values, and self-set standards. While institutionalized accountability relies on traditional "carrot and stick" approaches, such as financial incentives or denial of promotions, grassroots accountability operates on reciprocity with peers and intrinsic motivations, like maintaining one's reputation in the team.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.