Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
192
datasets available to search
ShareScore release 0.7.1
Dataset results
192 results for “Software engineering”
Unveiling Hurdles in Software Engineering Education: The Role of Learning Management Systems
<p>Learning management systems (LMSs) are established tools in higher education, especially in the field of software engineering (SE). The onset of the COVID-19 pandemic further amplified the utilization of these systems, which necessitated their integration into educational curricula for both lecturers and students. However, adopting LMSs within SE education has presented distinctive challenges impeding their seamless incorporation into the courses. This paper aims to scrutinize the challenges and requirements encountered by professors, lecturers, and students in the domain of SE education when using LMSs. We conducted an empirical study that included (i) a survey with 47 professors/lecturers and 133 students, (ii) an analysis of the ensuing data, and (iii) 18 additional interviews conducted with professors and lecturers to delve into nuanced variations in viewpoints. The findings derived from our study reveal that the challenges and requirements pertaining to LMSs are rather specific depending on the scope and size of the respective courses. Nevertheless, many participants have a consensus on numerous challenges and requirements for improving certain features of LMSs in order to improve their usage in SE education. The findings are valuable for advancing research and development in the field of LMSs and provide guidance for lecturers in SE education.</p>
ChatGPT's Aptitude in Utilizing UML Diagrams for Software Engineering Exercise Generation
<p>The integration of Artificial Intelligence (AI) technologies into educational settings has paved the way for innovative teaching and learning approaches. In Software Engineering (SE) education, using Unified Modeling Language (UML) diagrams is a fundamental teaching element for understanding complex software systems. This research addresses the ability of ChatGPT to utilize UML class and sequence diagrams for creating SE modeling exercises. We use ChatGPT to generate exercises based on the information from uploaded UML diagrams by analyzing textual UML representations such as Mermaid and graphical diagrams. The research explores ChatGPT's ability to synthesize UML-specific information from class and sequence diagrams, enabling the generation of various exercises tailored to strengthen conceptual understanding and practical application. Furthermore, we investigate generating graphical UML class and sequence diagrams based on natural language as input. By bridging the gap between AI-driven natural language understanding and the comprehension of UML diagrams, this study highlights the potential of ChatGPT to improve SE education. Our concise findings address educators, practitioners, and other researchers engaged in the field of SE education with a special focus on UML.</p>
Replication Package for "Software Quality Assurance Analytics: Enabling Software Engineers to Reflect on QA Practices" Paper (SCAM 2024)
<p>Welcome to our artifact!<br>In here we provide additional information for you to retrace our steps in the interview analysis.<br>It has the following contents:</p> <ul> <li><code>codebook.xlsx</code>: Our full codebook with our open codes, structured after the axial codes that emerged. <code>codebook-statistics.xlsx</code> lists for each code in which participant's interview it can be found.</li> <li><code>generate-figures</code>: The plain data and scripts used to generate the figures in the paper.</li> <li><code>survey.pdf</code>: An printout of our whole online questionnaire that guided the participants through the pretest-posttest study and the interview.</li> <li><code>survey-answers.xlsx</code>: The complete data for our participants answers in the online survey during the interviews.</li> <li><code>repoinsights-dashboard-software</code>: The code of our prototype repoinsights. As it is under active development, this is not yet documented for replicating the study setup or extending it. Still, we are providing the source code for transparency and will publish a version with comprehensive setup instructions later.</li> </ul>
Reference data and analysis software for "Four-color single-molecule imaging with engineered tags resolves the molecular architecture of signaling complexes in the plasma membrane"
<p>Reference data set for the single molecule co-tracking analysis presented in "Four-color single-molecule imaging with engineered tags resolves the molecular architecture of signaling complexes in the plasma membrane". Corresponding author for further inquiries:</p> <p>Prof. Dr. Jacob Piehler</p> <p>University of Osnabrück, Department of Biology/Chemistry, Division of Biophysics, Barbarastr. 11, 49076 Osnabrück, Germany</p> <p>https://www.biophysik.uni-osnabrueck.de/</p>
Gamification in Software Engineering: The Mediating Role of Developer Engagement and Job Satisfaction
<p>Replication package with covariance matrices (instead of original dataset) and R script.</p>
Grey Literature in Software Engineering: A Critical Review
<p>Replication package of the study "Grey Literature in Software Engineering: A Critical Review" published in Information and Software Technology (IST).</p>
Supplementary Material for Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study
<p>This repository delivers the supplementary material for the paper: <em>Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study.</em></p> <p>In the following, we present the abstract of the study:</p> <p><strong>Context:</strong> Agile Software Development (ASD) is a disruptive process compared to traditional software development. Therefore, traditional Requirements Engineering (RE) forms may not be the best way to do RE for ASD (RE-ASD). <strong>Objective:</strong> Working with ASD using traditional RE ways could limit ASD's potential. Thus, it is necessary to investigate what academia and industry have done in RE to take full advantage of all of the capabilities of ASD beyond traditional RE. <strong>Method: </strong>We conducted a Tertiary Study looking for solutions for RE-ASD using the Systematic Literature Review (SLR) protocol described by Kitchenham and Charters. We then categorized the solutions into families using Targeted Coding and Constant Comparison, tools from Socio-Technical Grounded Theory (STGT). Afterward, we classified the solutions as disruptive using our model based on the Hype Level Curve concept, assessing their hype (popularity) in the software engineering community using Google Trends and Google Colab tools. <strong>Results:</strong> After executing the SLR protocol, we accepted 37 studies and encountered 136 solutions used by academia and industry for RE-ASD. We categorized these solutions into 21 solution families, six of which we classified as disruptive. Design Thinking (DT) and Artificial Intelligence (AI) were the two families of solutions that stood out the most. We also identified the type of solution (e.g., process, method, technique, tool, model, framework) and domain (academia or industry). Furthermore, we cataloged the challenges presented by the solutions. <strong>Conclusion:</strong> We concluded that only a few solutions that have been used for RE-ASD have the power to successfully challenge the mainstream Agile Software Development process by using innovation (26 out of 106). There is a gap between academia and industry regarding these disruptive solutions, and some challenges still need to be addressed in using these solutions.</p> <p>The repository contains the following:</p> <ul> <li>Dataset from the Tertiary Study: <ul> <li>Data of the retrieved studies. It presents the classifications of the documents as 'Accepted,' 'Rejected' (with the indication of the step of the protocol the authors rejected the study), or 'Duplicated.'</li> <li>Data of all solutions retrieved from the accepted studies</li> </ul> </li> <li>Socio-Technical Grounded Theory (STGT) tools <ul> <li>Result of the use of Targeted Coding and Constant Comparison</li> </ul> </li> <li>The Google Colab Notebook <ul> <li>Code in python</li> <li>Results</li> </ul> </li> </ul> <p> </p>
Dataset and Replication Package for the View-Based Retriever Approach To Reverse Engineering Software Architecture Models
<div> <div><span>Dataset and replication package for the view-based Retriever approach to reverse engineering software architecture models. Each Dataset project is structured as follows:</span></div> <ul> <li><span>The .ruleengine.yml file contains the configuration for running the Retriever approach.</span> <ul> <li><span>The repository value is the ID of a GitHub repository.</span></li> <li><span>The current_version value is the latest version of the retriever approach used to build the architectural models.</span></li> <li><span>The rules values are the rules used to build the architectural models.</span></li> </ul> </li> <li><span>The model_re folder contains the architectural model of the system automatically generated by the Retriever approach.</span> <ul> <li><span>The pcm folder contains the Palladio Component Model (PCM) of the system.</span></li> <li><span>The uml folder contains the PlantUML model.</span></li> </ul> </li> <li><span>The model_gs folder contains our manual gold standards for the system.</span></li> </ul> <div><span>The easiest way to use our approach is to use the CLI application with the given parameters: ./eclipse -i /path/to/input/directory -o /path/to/output/directory -r supported_rules</span></div> </div>
Data - Teaching and Learning of Introduction to Software Engineering Experimentation to Distance-Learning Students: a Quasi-Experiment
<p>Data from a quasi-experiment on "Teaching and Learning of Introduction to Software Engineering Experimentation to Distance-Learning Students: a Quasi-Experiment"</p> <p> </p> <p>Funded by Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq) 311503/2022-5</p>
Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal
<p>This Research Artifact contains supplemental material from the study: "Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal". The supplementary material contains:</p> <ol> <li>The list of the 106 primary studies classified by journals.</li> <li>The list of the 12 research artifacts classified by conferences.</li> <li>The .xlsx file of the dataset used to analyze the RQs.</li> <li>The .xlsx file of the dataset used to analyze the research artifacts problems (Table 3). </li> <li>The list of the figures published in the scientific article.</li> </ol>
Word Embeddings for the Software Engineering Domain
<p>A .bin file for a word2vec model pre-trained on 15GB of Stack Overflow posts. </p> <p>For more details refer to the following paper:</p> <p>Efstathiou, V., Chatzilenas, C., Spinellis, D., 2018. "Word Embeddings for the Software Engineering Domain". In <em>Proceedings of the 15th International Conference on Mining Software Repositories.</em> ACM</p>
Frame Embeddings for Software and Requirements Engineering Domain
<p>This project is aimed to identify semantic relatedness of <a href="https://framenet2.icsi.berkeley.edu/">FrameNet </a>semantic frames in the domain of software and requirements engineering. The folder contains the frame embeddings that are obtained using the <strong>context-based method</strong> described in our ESEM paper*.</p> <p>Waad Alhoshan, Liping Zhao, and Riza Batista-Navarro. 2018. Using Semantic Frames to Identify Related Textual Requirements: An Initial Validation. In ACM / IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (ESEM ’18), October 11–12, 2018, Oulu, Finland. ACM, New York, NY, USA, 2 pages. https://doi.org/10.1145/3239235.3267441 </p> <p> </p> <p> </p>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Analysis of the DLR Knowledge Exchange Workshop Series on Software Engineering
<p>This repository is used to analyze the workshops of the DLR internal workshop series on software<br> engineering. These workshops are two-day events of the DLR software engineering community and<br> focus on different main topics every year.</p>
Levels of a Research Software Engineer
<p><strong>Levels of a Research Software Engineer: </strong>The diverse role of the RSE can be captured by the degree or level to which they work with researchers, and in what scope. Level 1 of RSE "domain" are closest to researchers, working directly on their behalf. Level 2 "generalist" RSE work on core technologies needed across the scientific community, and level 3 "researcher" take this a step further, researching the space or models underlying the software itself.</p>
Enhancing Motivation in Software Engineering Education through Gamified Agile Project-based Learning
<p>Project-based learning (PBL), e.g., student software development projects, is an essential part of today's Software Engineering (SE) education. They allow students to work on real-world projects and gain practical experience as a team. However, several challenges arise in such projects, including learning new technologies and dealing with communication and coordination issues within the team. These factors can lead to a lack of motivation to contribute to the project and a decrease in productivity, potentially resulting in an insufficient project outcome. This paper aims to promote student motivation in PBL and increase team productivity by applying gamification. We conducted a user and requirements analysis to identify the needs of students and supervisors of such projects. Based on the insights, we designed and implemented DinoDev, a gamified project management tool that combines project management features with gamification elements. The DinoDev concept was evaluated in a student project, indicating increased motivation and team productivity. The findings are valuable for advancing research on using gamification in PBL and for lecturers to improve their students' motivation and team productivity in SE education.</p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>We collect the data in a rush.</p> <p>If you want to use this dataset and find any errors, please contact us ;-)</p> <p> </p> <p>Our emails:</p> <ul> <li>echo.xiangchen@gmail.com</li> <li>zhifengyao731@gmail.com</li> </ul>
Research on Cognition in Software Engineering
<p>This dataset includes the primary studies selected for literature review on cognition in software engineering. </p>
Replication package for Workshop on Software Engineering 22' - What does the pytest plugins data say?
<p>This database stores the information used to run the experiment in the article: <strong>What does the pytest plugins data say?</strong></p>
Modelling Guidance in Software Engineering: A Systematic Literature Review
<p>The dataset contains three supporting documents covering the paper selection details (including selection criteria and extracted papers after each round) from three rounds of selection for the literature review titled- Modelling Guidance in Software Engineering: A Systematic Literature Review. There are three excel sheets named: <strong>search_s1</strong>,<strong> search_s2</strong> and <strong>snowballing_bpm</strong>. The file "<strong>search_s1</strong>" contains details of the search with the search string: {("modeling" OR "modelling" OR "model-driven" OR "model-based") AND ("guidelines" OR "training" OR "styles" OR "creation") AND ("software engineering")}. The second file, "<strong>search_s2</strong>", contains details of the search with the string: {("modeling" OR "modelling" OR "model-driven" OR "model-based") AND ("approach" OR "process" OR "method" OR "template") AND ("software engineering")}. The third file, "<strong>snowballing_bpm</strong>", contains the backward and forwards snowballing search results. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.