Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
199
datasets available to search
ShareScore release 0.9.0
Dataset results
199 results for “software studies”
The potential of low-cost UAVs and open-source photogrammetry software for high-resolution monitoring of alpine glaciers: A case study from the Kanderfirn (Swiss Alps)
<p>This dataset contains high-resolution orthophotos (5 x 5 cm) and digital surface models (25 x 25 cm) of the Kandernfirn Glacier located in the Swiss Alps. Aerial images were aquired with a self-developed fixed-wing Unmanned Aerial Vehicle during ten surveys on five different days in 2017 and 2018. The open-source photogrammetry software OpenDroneMap (version 0.4.1) was used for image processing.</p> <p>The orthophotos and digital surface models were validated through dGNSS point measurements of ground control points. Please refer to the corresponding paper for information on the horizontal and vertical accuracy of the files.</p>
Software Architecture Assessment for Sustainability: A Case Study
<h3>Replication Package: Software Architecture Assessment for Sustainability: A Case Study</h3> <p><br>This repository contains the supplementary material to support the paper published at the International Conference on Software Architecture (ECSA) 2024 titled, "Software Architecture Assessment for Sustainability: A Case Study". This repository can be used to replicate the study and carry out a Software Architecture Evaluation of other software systems.<br><br>The online version can be browsed on the linked <a href="https://github.com/S2-group/rep-pkg-ecsa-2024-SA-assessment-for-sustainability-a-case-study/tree/v1.0.0" target="_blank" rel="noopener">Github Repository</a></p> <p> </p> <p> </p>
Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering
<p>Update: We updated the data set in March 2022 by adding newly published papers and by providing more insights on how we analyzed them. Details can be found in the file " SEnti-SMS.xlsx".</p> <p>----------</p> <p>Update: The updated version (-v2) contains the results of one more snowballing iteration and extracted information on the accuracy of the used methods.</p> <p>----------</p> <p>In 2020, we conducted a systematic literature review to explore the development and application of sentiment analysis tools in software engineering.</p> <p>Information on the execution of the SLR, its scope, the search string, etc. are presented in the paper linked below.</p> <p> </p> <p> </p>
Machine Learning for Software Engineering: A Tertiary Study
<p>Dataset of the research paper: <strong>Machine Learning for Software Engineering: A Tertiary Study</strong></p> <p>Machine learning (ML) techniques increase the effectiveness of software engineering (SE) lifecycle activities. We systematically collected, quality-assessed, summarized, and categorized 83 reviews in ML for SE published between 2009–2022, covering 6,117 primary studies. The SE areas most tackled with ML are software quality and testing, while human-centered areas appear more challenging for ML. We propose a number of ML for SE research challenges and actions including: conducting further empirical validation and industrial studies on ML; reconsidering deficient SE methods; documenting and automating data collection and pipeline processes; reexamining how industrial practitioners distribute their proprietary data; and implementing incremental ML approaches.</p> <p>The following data and source files are included.</p> <ul> <li><strong>review-protocol.md</strong>: The protocol employed in this tertiary study</li> </ul> <p><strong>data/</strong></p> <p><strong> dl-search/</strong></p> <p><strong> input/</strong></p> <ul> <li><strong>acm_comput_surveys_overviews.bib</strong>: Surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_surveys_overviews_titles.txt</strong>: Titles of surveys</li> <li><strong>acm_comput_ml_surveys.bib</strong>: Machine learning (ML)-related surveys of ACM Computing Surveys journal</li> <li><strong>acm_comput_ml_surveys_titles.txt</strong>: Titles of ML-related surveys</li> <li><strong>dl_search_queries.txt</strong>: Search queries applied to IEEE Xplore, ACM Digital Library, and Elsevier Scopus</li> <li><strong>ml_keywords.txt</strong>: ML-related keywords extracted from ML-related survey titles and used in the search queries</li> <li><strong>se_keywords.txt</strong>: Software Engineering (SE)-related keywords derived from the 15 SWEBOK Knowledge Areas (KAs—except for Computing Foundations, Mathematical Foundations, and Engineering Foundations) and used in the search queries</li> <li><strong>secondary_studies_keywords.txt</strong>: Survey-related keywords composed of the 15 keywords introduced in the tertiary study on SLRs in SE by Kitchenham <em>et al.</em> (2010), and the survey titles, and used in the search queries</li> </ul> <p><strong> output/</strong></p> <ul> <li><strong>acm/</strong> <ul> <li><strong>acm{1–9}.bib</strong>: Search results from ACM Digital Library</li> </ul> </li> <li><strong>ieee.csv</strong>: Search results from IEEE Xplore</li> <li><strong>scopus_analyze_year.csv</strong>: Yearly distribution of ML and SE documents extracted from Scopus's <em>Analyze search results</em> page</li> <li><strong>scopus.csv</strong>: Search results from Scopus</li> </ul> <p><strong> study-selection/</strong></p> <ul> <li><strong>backward_snowballing.csv</strong>: Additional secondary studies found through the backward snowballing process</li> <li><strong>backward_snowballing_references.csv</strong>: References of quality-accepted secondary studies</li> <li><strong>cohen_kappa_agreement.csv</strong>: Inter-rater reliability of reviewers in study selection</li> <li><strong>dl_search_results.csv</strong>: Aggregated search results of all three digital libraries</li> <li><strong>forward_snowballing_reviewer_{1,2}.csv</strong>: Divided forward snowballing citations of quality-accepted studies assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> <li><strong>study_selection_reviewer_{1,2}.csv</strong>: Divided search results assessed by reviewer 1 and 2, correspondingly, based on IC/EC</li> </ul> <p><strong> quality-assessment/</strong></p> <ul> <li><strong>dare_assessment.csv</strong>: Quality assessment (QA) of selected secondary studies based on the Database of Abstracts of Reviews of Effects (DARE) criteria by York University, Centre for Reviews and Dissemination</li> <li><strong>quality_accepted_studies.csv</strong>: Details of quality-accepted studies</li> <li><strong>studies_for_review.bib</strong>: Bibliography details and QA scores of selected secondary studies</li> </ul> <p><strong> data-extraction/</strong></p> <ul> <li><strong>further_research.csv</strong>: Recommendations for further research of quality-accepted studies</li> <li><strong>further_research_general.csv</strong>: The complete list of associated studies for each general recommendation</li> <li><strong>knowledge_areas.csv</strong>: Classification of quality-accepted studies using the SWEBOK KAs and subareas</li> <li><strong>ml_techniques.csv</strong>: Classification of the quality-accepted studies based on a four-axis ML classification scheme, along with extracted ML techniques employed in the studies</li> <li><strong>primary_studies.csv</strong>: Details of reviewed primary studies by the quality-accepted secondary</li> <li><strong>research_methods.csv</strong>: Citations of the research methods employed by the quality-accepted studies</li> <li><strong>research_types_methods.csv</strong>: Research types and methods employed by the quality-accepted studies</li> </ul> <p><strong>src/</strong></p> <ul> <li><strong>data-analysis.ipynb</strong>: Analysis of data extraction results (data preprocessing, top authors and institutions, study types, yearly distribution of publishers, QA scores, and SWEBOK KAs) and creation of all figures included in the study</li> <li><strong>scopus-year-analysis.ipynb</strong>: Yearly distribution of ML and SE publications retrieved from Elsevier Scopus</li> <li><strong>study-selection-preprocessing.ipynb</strong>: Processing of digital library search results to conduct the inter-rater reliability estimation and study selection process</li> </ul>
Dataset and replication package for Temporal Discounting in Software Engineering: A Replication Study
<p>Dataset and replication package for the paper Temporal Discounting in Software Engineering: A Replication Study (Fagerholm, F., Becker, C., Chatzigeorgiou, A., Betz, S., Duboc, L., Penzenstadler, B., Mohanani, R., Venters, C. (2019). Temporal Discounting in Software Engineering: A Replication Study. 13th ACM/IEEE International Symposium of Empirical Software Engineering and Measurement (ESEM 2019)). The dataset consists of answers to a questionnaire on temporal discounting in a technical debt context. Two questionnaire templates illustrate how to gather the data for professional and student participants. An analysis script is provided which shows the details of the calculations and analyses performed for the paper. More information is given in the description file.</p>
Open dataset for publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems"
<p>This publication contains open dataset for the journal publication "Systematic Mapping Study on Requirements Engineering for Regulatory Compliance of Software Systems".</p> <p>The dataset contains the data extracted from 280 selected primary studies.</p> <p>The dataset includes the following data:</p> <ul> <li>study metadata (title, venue, publication year, authors, authors’ affiliation, abstract);</li> <li>challenges to regulatory compliance (direct excerpts from studies);</li> <li>categories of challenges to compliance;</li> <li>principles and practices (direct excerpts from text);</li> <li>categories of principles and practices;</li> <li>types of automation of principles and practices;</li> <li>involved stakeholders (direct excerpts from studies);</li> <li>categories of involved stakeholders;</li> <li>phase of the principle and practice life cycle for which involvement of stakeholders was considered;</li> <li>SDLC process areas covered by the study;</li> <li>regulations considered in the study;</li> <li>fields of regulations that were considered;</li> <li>domains of application that were considered;</li> <li>assessment of rigor and relevance of the study.</li> </ul>
A Study on Organizational IT Security in Mobile Software Ecosystems Literature
<p>Information security is a key topic for most organizations. With the digital revolution, smartphones have become popular not only for personal use but also within organizations where many employees use them for business purposes. As smartphones are increasingly present in organizations, it is necessary to understand what recommendations the literature provides for the safe use of such devices, helping organizations to protect themselves from threats. ISO 27000 is a well-known standard for information security in a business context. It provides a set of controls that must be observed to ensure more secure organizational information. Therefore, the goal of this study is to identify which controls presented in ISO 27000, more specifically ISO 27001, are present in the Mobile Software Ecosystem (MSECO) literature. To do so, we conducted a systematic mapping review supplemented by a snowballing process to identify studies in the field of MSECO that have addressed any subject that is present in ISO 27001. We found that 34 out of the 114 ISO 27001 controls are covered by the MSECO literature. Also, some of the ISO sections (e.g., Asset Management) have not yet been explored in the MSECO literature. Our results can inspire future and further studies on the topic of MSECO information security.</p>
Software Product Line Traceability and Product Configuration in Class and Sequence Diagrams: an Empirical Study
<p>Software Product Line Traceability and Product Configuration in Class and Sequence Diagrams: an Empirical Study</p>
Software Product Line Configuration and Traceability: an Empirical Study on SMarty Class and Component Diagrams
<p>Software Product Line Configuration and Traceability: an Empirical Study on SMarty Class and Component Diagrams</p>
The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study
<p>Online appendix of the paper entitled: "The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study". It contains all scripts and data required to replicate the four research questions of the study.</p> <p>Abstract: Software vulnerabilities are weaknesses in source code that can be potentially exploited to cause loss or harm. While researchers have been devising a number of methods to deal with vulnerabilities, there is still a noticeable lack of knowledge on their software engineering life cycle, for example how vulnerabilities are introduced and removed by developers. This information can be exploited to design more effective methods for vulnerability prevention and detection, as well as to understand the granularity that these methods should aim at. To investigate the life cycle of software vulnerabilities, we focus on how, when, and under which circumstances vulnerabilities are introduced in software projects, as well as whether, after how long, and how they are removed. We consider 4,097 vulnerabilities with public patches from the National Vulnerability Database—pertaining to 1,163 open-source software projects on GITHUB—and define a six-step process that involves both automated parts (e.g., using the SZZ algorithm to find the vulnerability-inducing commits) and manual analyses (e.g., how vulnerabilities were fixed). The investigated vulnerabilities can be classified in 148 categories, take on average 4.19 commits before being introduced, and remain unfixed for a median of 1,506.50 commits and 691.50 days. Most of them are introduced by developers with high workload, often when doing maintenance activities, and removed with mostly with the addition of new source code aiming at implementing further checks on inputs. We conclude by distilling practical implications on when and how vulnerability detectors should work to better assist developers in early detecting these issues.</p>
Dataset: We Do Not Understand What It Says -- Studying Student Perceptions of Software Modelling
<p>The dataset contains two supporting documents for the paper title, "We Do Not Understand What It Says -- Studying Student Perceptions of Software Modelling". The first one is an excel sheet containing interview transcripts of 13 of the participants of this study (who agreed to publish their statements) and the second is an appendix file containing the interview guide (questionnaire used for interviews with students and instructors) used in the case study. </p> <p>The interview transcripts are supported by "in-vivo coding" used by both authors separately during analysis. </p>
Supplementary Material for Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study
<p>This repository delivers the supplementary material for the paper: <em>Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study.</em></p> <p>In the following, we present the abstract of the study:</p> <p><strong>Context:</strong> Agile Software Development (ASD) is a disruptive process compared to traditional software development. Therefore, traditional Requirements Engineering (RE) forms may not be the best way to do RE for ASD (RE-ASD). <strong>Objective:</strong> Working with ASD using traditional RE ways could limit ASD's potential. Thus, it is necessary to investigate what academia and industry have done in RE to take full advantage of all of the capabilities of ASD beyond traditional RE. <strong>Method: </strong>We conducted a Tertiary Study looking for solutions for RE-ASD using the Systematic Literature Review (SLR) protocol described by Kitchenham and Charters. We then categorized the solutions into families using Targeted Coding and Constant Comparison, tools from Socio-Technical Grounded Theory (STGT). Afterward, we classified the solutions as disruptive using our model based on the Hype Level Curve concept, assessing their hype (popularity) in the software engineering community using Google Trends and Google Colab tools. <strong>Results:</strong> After executing the SLR protocol, we accepted 37 studies and encountered 136 solutions used by academia and industry for RE-ASD. We categorized these solutions into 21 solution families, six of which we classified as disruptive. Design Thinking (DT) and Artificial Intelligence (AI) were the two families of solutions that stood out the most. We also identified the type of solution (e.g., process, method, technique, tool, model, framework) and domain (academia or industry). Furthermore, we cataloged the challenges presented by the solutions. <strong>Conclusion:</strong> We concluded that only a few solutions that have been used for RE-ASD have the power to successfully challenge the mainstream Agile Software Development process by using innovation (26 out of 106). There is a gap between academia and industry regarding these disruptive solutions, and some challenges still need to be addressed in using these solutions.</p> <p>The repository contains the following:</p> <ul> <li>Dataset from the Tertiary Study: <ul> <li>Data of the retrieved studies. It presents the classifications of the documents as 'Accepted,' 'Rejected' (with the indication of the step of the protocol the authors rejected the study), or 'Duplicated.'</li> <li>Data of all solutions retrieved from the accepted studies</li> </ul> </li> <li>Socio-Technical Grounded Theory (STGT) tools <ul> <li>Result of the use of Targeted Coding and Constant Comparison</li> </ul> </li> <li>The Google Colab Notebook <ul> <li>Code in python</li> <li>Results</li> </ul> </li> </ul> <p> </p>
Replication Package of the study "Automated Identification and Qualitative Characterization of Safety Concerns Reported in UAV Software Platforms"
<p><strong>Description of the Dataset of the work "Automated Identification and Qualitative Characterization of Safety<br> Concerns Reported in UAV Software Platforms"</strong></p> <p><strong><em>"1_Safety-Dataset" folder: </em></strong>This folder contains the bugs data and row data of all analyzed projects.<br> Specifically, this folder contains the following relevant entries<br> <br> - "bugs" folder: It contains the bugs of all analyzed projects (PX4-merged.json.gz, dDronin-merged.json.gz, ardupilot-merged.json.gz)<br> of all sentences extracted from the project issues<br> - "Dataset-safety-bugs.csv": For all projects, it contains the raw data of the set of sentences classified as safety and non-safety related.<br> </p> <p><em><strong>"2_Scripts-and-generated-data (RQ1)" folder:</strong> </em>This folder contains the scripts and code used to preprocess and analyze the issue data in <br> the context of RQ1<br> Specifically, this folder contains the following relevant entries<br> <br> - "main-program.py" file: Main program executing all subscripts generating the data required for RQ1 (detailed in the following line)<br> - "utilities.R" file: (Utility) R script containing relevant functions for pre-processing/indexing text and issue data<br> - "1_Script-to-create-test-dataset.r" file: R script containing simple code for analyzing issue data<br> - "2_MainScript.r" file: Main R program orchestrating the scripts "utilities.R" and "1_Script-to-create-test-dataset.r" execution<br> - "files-setDirectory" folder: Folder where data are generated and stored from the "main-program.py"<br> - "fasttext" folder: Folder where data used as input from fastText (by "main-program.py") are reported<br> - "cross-project-analysis" folder: Folder with data used for the cross-project analysis</p> <p> - "main-program-grid-search.py" file: Main program executing all experiments for the grid search analysis</p> <p><em><strong>"3_Results" folder: </strong></em>This folder contains the results, scripts and figures used to discuss results of the study.<br> Specifically, this folder contains the following relevant entries<br> <br> - "RQ1" folder: This folder contains the results, scripts and figures used to discuss results of RQ1.<br> - "RQ2" folder: This folder contains the results, scripts and Tables used to discuss results of RQ2.</p>
AI Tool Use and Adoption in Software Development by Individuals and Organizations: A Grounded Theory Study
<div> <p>This data represents four artifacts from our research in studying what impacts AI adoption and use in SE. It includes our interview questions, the codebook with example quotes, survey questions, and table with the code to category generation.</p> <p>This page includes supplementary materials associated with our paper entitled "<span>AI Tool Use and Adoption in Software Development by </span><span>Individuals and Organizations: A Grounded Theory Study</span>".</p> </div>
Figure 2. Some screenshots from the software system-Design and Development of a Software System for Swarm Intelligence Based Research Studies
<p>All of the mentioned operations can be performed easily by using the provided controls over<br> the related interfaces – windows of each algorithm. It is also important that each algorithm interface<br> is supported by visual controls to view obtained results with typical iteration-based graphics or<br> problem oriented visual elements. For instance, resulting graph structures are automatically shown<br> by the algorithm interfaces after solving some specific, popular problems like Travelling Salesman<br> Problem (TSP), Vehicle Routing Problem (VCP)…etc. Visually improved using features and<br> functions of the software system are critical aspects to provide more effective and useful platform to<br> perform SI based research studies better.<br> Related to the designed and developed software system, some screenshots from the software<br> system [interfaces of two algorithms (IWDs and ABC)] are represented in Fig. 2.</p>
Figure 9. Experimental Page Rank dependency on Markov Chain length with balanced distribution-Study of a Random Navigation on the Web Using Software Simulation
<p>This paper explored different implementation choices for analyzing the most important<br> parameters about a web. Many researchers explored the use of new search engines for studying the<br> evolution of the web (Ntoulas, Cho and Olston, 2004). Another important research is realized about<br> the Link Structure Graph (LSG). The LSG captures a complete hyperlink structure from the web<br> and models link associations reflected in the page layout (Rodrigues, Milic-Frayling and Fortuna,<br> 2007). For further works ideas like extrapolation methods for accelerating page rank calculation can<br> be developed (Kamvar et al., 2003).</p>
Figure 7. Experimental Page Rank dependency on Markov Chain length-Study of a Random Navigation on the Web Using Software Simulation
<p>The next diagram proves that the values for Experimental Page Rank depend on the length<br> of the Markov Chain, while Algorithmic Page Rank remains constant.</p>
Figure 6. Outlinks degree distribution for all web sites-Study of a Random Navigation on the Web Using Software Simulation
<p>Some of the most important aspects of the analysis is obtaining the parameters which can<br> give the main information about a web. In the simulation implementation information as: page,<br> number of inlinks, number of outlinks, value for Algorithmic Page Rank and Experimental Page<br> Rank will be processed for obtaining the results of the analysis. For the first part it was necessary to<br> use experimental values as: inlinks, outlinks and in and out frequencies.</p>
Figure 1. Markov Chain Model&Figure 2. Transition matrix-Study of a Random Navigation on the Web Using Software Simulation
<p>For a good simulation it is very important to find methods for<br> navigating through the web (Levene and Wheeldon, 2004). John Kemeny and Laurie Snell have<br> proposed the use of Markov models for web simulations (Kemeny and Snell, 1960). Cadez et al. (2000)<br> used Markov models for classifying the sessions into different categories for browsers. Some other<br> proposed techniques choose to combine different order Markov models for obtaining low state<br> complexity and improving accuracy, as Deshpande and Karypis (2004). Dongshan and Junyi (2002)<br> used for predicting the access providing good scalability and high coverage a hybrid-order tree-like<br> Markov model. As an alternative to the Markov model Pitkow proposed a longest subsequence model<br> (Pitkow and Pirolli, 1999), also for predicting the next page accessed by the user Sarukkai chose<br> Markov models (Sarukkai, 2000).<br> Transitions are simulated using the Markov Chain nodes, Google matrix and an arbitrary initial<br> probability distribution. Examples can be seen in Figure 1 and Figure 2.</p>
Fig. 8. Algorithmic Page Rank-Study of a Random Navigation on the Web Using Software Simulation
<p>Another diagram shows that for each site there is only one constant value independent of the<br> length of Markov Chain. Algorithmic Page Rank only depends on the number of inlinks and<br> outlinks.Table 4 contains the values for Experimental Page Rank with balanced distribution. In the<br> following figure it is shown the diagram for the values obtained for Experimental Page Rank.<br> The Experimental Page Rank values oscillate between the same limits independently of the<br> change of Markov Chain length (N).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.