Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

164

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

164 results for “software development”

Learn how ShareScore rates datasets ↗
zenodo48/100

Software Developer Expertise GitHub and Stack Overflow data sets

<p>Cross-Platform Software Developer Expertise Learning by Norbert Eke</p> <p>This data set is part of my Master&#39;s thesis project on developer expertise learning by mining Stack Overflow (SOTorrent) and Github (GHTorrent) data. Check out my portfolio website at norberte.github.io</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Software developers are users of the Heureka microservice platform

<p>This data set contains the qualitative analysis of the &quot;software developers are users&quot; study, conducted to investigate the fitting of the Heureka microservice platform (http://doc.soteto.net) to support software developers in participating the&nbsp;change of socio-technical evolutionary-teal organizations.</p> <p>The study is part of the SOTETO project (http://soteto.net)&nbsp;and its results will be published and contextualized by the dissertation thesis of Johann Sell.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Bots in Software Development: A Systematic Literature Review [Data Set]

<p>This repository contains al the artifacts of the research: Bots in Software Development: &nbsp;A Systematic Literature Review&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering

<p>Update: We updated the data set in March 2022 by adding newly published papers and by providing more insights on how we analyzed them. Details can be found in the file &quot; SEnti-SMS.xlsx&quot;.</p> <p>----------</p> <p>Update: The updated version (-v2) contains the results of one more snowballing iteration and extracted information on the accuracy of the used methods.</p> <p>----------</p> <p>In 2020, we conducted a systematic literature review to explore the development and application of sentiment analysis tools in software engineering.</p> <p>Information on the execution of the SLR, its scope, the search string, etc. are presented in the paper linked below.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Dataset of Open-Source Software Developers Labeled by their Experience Level and Associated with their Software Metrics

<p>This dataset contains 703&nbsp;anonymized developers extracted from 17 open-source projects from GitHub. Projects were chosen because they use:</p> <ul> <li>the Java programming language</li> <li>the <a href="https://spring.io/projects/spring-framework">Spring framework</a></li> <li><a href="https://maven.apache.org/">Maven</a> / <a href="https://gradle.org/">Gradle</a> build tools</li> </ul> <p>For all these developers, 23 software metrics were calculated for each project to which they contribute. These metrics are either calculated by analyzing the source code or relative to project management metadata. Each of these developers then have been manually annotated. To do this, developers have been searched&nbsp; for in professionnal social media such as:</p> <ul> <li><a href="https://www.linkedin.com/">Linkedin</a></li> <li><a href="https://twitter.com/">Twitter</a></li> <li><a href="https://github.com/">Github</a></li> </ul> <p><strong>This dataset is published in the following journal article: </strong></p> <p><strong>Dataset of Open-Source Software Developers Labeled by their Experience Level in the Project and their Associated Software Metrics, Q. Perez, C. Urtado and </strong><strong>S. Vauttier, Data In Brief, </strong></p> <p><a href="https://www.sciencedirect.com/science/article/pii/S2352340922010459">https://www.sciencedirect.com/science/article/pii/S2352340922010459</a></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Perceptions on the utility of community question and answer websites like Stack Overflow to software developers (Replication package)

<p>Interview Questions on the perception of the utility of CQAs like Stack Overflow to software developers. In this study, we focused on the questions highlighted in yellow.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

A StackExchange Dataset of Developer Questions Related to Checked-in Secrets in Software Artifacts

<p>Throughout 2021, GitGuardian&#39;s monitoring of public GitHub repositories revealed a two-fold increase in the number of secrets (database credentials, API keys, and other credentials) exposed compared to 2020, accumulating more than six million secrets. To our knowledge, the challenges developers face to avoid checked-in secrets are not yet characterized. In our artifact, we provide a dataset containing 779 questions mined from three StackExchange sites asked by developers related to checked-in secrets from three StackExchange sites. In addition, we provide 434 accepted answers provided by the other users of StackExchange to mitigate the challenge of checked-in secrets.</p> <p>&nbsp;</p> <table> <caption>An overview of StackExchange artifact</caption> <thead> <tr> <th scope="col">Field Name</th> <th scope="col">Description</th> </tr> </thead> <tbody> <tr> <td>Id</td> <td>An unique identifier of the question.</td> </tr> <tr> <td>Title</td> <td>The title of the question.</td> </tr> <tr> <td>Body</td> <td>The description of the question.</td> </tr> <tr> <td>Tags</td> <td>The tags related to the question such as &quot;security&quot;, &quot;git&quot; and &quot;key-management&quot;.</td> </tr> <tr> <td>CreationDate</td> <td>The date when the question is posted.</td> </tr> <tr> <td>Score</td> <td>The count of upvotes in the question.</td> </tr> <tr> <td>ViewCount</td> <td>The number of users who viewed the question.</td> </tr> <tr> <td>AnswerCount</td> <td>The total number of answers posted in the question.</td> </tr> <tr> <td>CommentCount</td> <td>The total number of comments posted in the question.</td> </tr> <tr> <td>FavouriteCount</td> <td>The total number of users who marked the question as favourite.</td> </tr> <tr> <td>ClosedDate</td> <td>The date when the community marked the question as closed.&nbsp;</td> </tr> <tr> <td>URL</td> <td>The url of the question.</td> </tr> <tr> <td>AcceptedAnswerId</td> <td>The unique identifier of the accepted answer for the question.</td> </tr> <tr> <td>Answer</td> <td>The accepted answer of the question.</td> </tr> </tbody> </table>

openmit-licenseFeb 2023View details →
zenodo44/100

Phloem anatomy constraints root system architecture development: theoretical clues from in silico experiments [software and dataset]

<p>Simulation software and results for &quot;<strong>Phloem anatomy constraints root system architecture development: theoretical clues from in silico experiments</strong>&quot;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Gamification in Software Engineering: The Mediating Role of Developer Engagement and Job Satisfaction

<p>Replication package with covariance matrices (instead of original dataset) and R script.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Mining Fork-Including Software Development Traces

<p>This dataset relates&nbsp;to the paper: Mining Fork-Including Development Traces (abstract below)<br> Authors: Iris Reinhartz-Berger and Amir Tomer<br> Starting point: readme.txt</p> <p>Open-source software development is a common practice that encourages collaborative development and reuse across projects. Forking is a way to make a copy of an existing project and explore it for different purposes. Two types of forks are commonly mentioned in the literature: <em>contributing forks</em> which continue the development lines of the forked projects and aim at merging the contribution back to the forked projects; and <em>independently developed forks</em> which open new lines of development deviating from the forked projects. In this study, we aim to explore characteristics of fork-involving software development traces. Analyzing 880 Java projects and their related action and observation events, with process mining and statistical techniques, we found that the occurrence of certain event types may predict the fork type, while the creation of certain fork types increase the involvement of users in the forked projects.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Supplementary Material for Disruptive Solutions on Requirement Engineering for Agile Software Development: A tertiary study

<p>This repository delivers the supplementary material for the paper: <em>Disruptive&nbsp;Solutions on Requirement Engineering for Agile Software Development: A tertiary study.</em></p> <p>In the following, we present the abstract of the study:</p> <p><strong>Context:</strong> Agile Software Development (ASD) is a disruptive process compared to traditional software development. Therefore, traditional Requirements Engineering (RE) forms may not be the best way to do RE for ASD (RE-ASD). <strong>Objective:</strong> Working with ASD using traditional RE ways could limit ASD&#39;s potential. Thus, it is necessary to investigate what academia and industry have done in RE to take full advantage of all of the capabilities of ASD beyond traditional RE.&nbsp;<strong>Method:&nbsp;</strong>We conducted a Tertiary Study looking for solutions for RE-ASD using the Systematic Literature Review (SLR) protocol described by Kitchenham and Charters. We then categorized the solutions into families using Targeted Coding and Constant Comparison, tools from Socio-Technical Grounded Theory (STGT). Afterward, we classified the solutions as disruptive using our model based on the Hype Level Curve concept, assessing their hype (popularity) in the software engineering community using Google Trends and Google Colab tools. <strong>Results:</strong> After executing the SLR protocol, we accepted 37 studies and encountered 136 solutions used by academia and industry for RE-ASD. We categorized these solutions into 21 solution families, six of which we classified as disruptive. Design Thinking (DT) and Artificial Intelligence (AI) were the two families of solutions that stood out the most. We also identified the type of solution (e.g., process, method, technique, tool, model, framework) and domain (academia or industry). Furthermore, we cataloged the challenges presented by the solutions.&nbsp;<strong>Conclusion:</strong> We concluded that only a few solutions that have been used for RE-ASD have the power to successfully challenge the mainstream Agile Software Development process by using innovation (26 out of 106). There is a gap between academia and industry regarding these disruptive solutions, and some challenges still need to be addressed in using these solutions.</p> <p>The repository contains the following:</p> <ul> <li>Dataset from the Tertiary Study: <ul> <li>Data of the retrieved studies. It presents the classifications of the&nbsp;documents as &#39;Accepted,&#39; &#39;Rejected&#39; (with the indication of the step of the protocol the authors rejected the study), or &#39;Duplicated.&#39;</li> <li>Data&nbsp;of all solutions retrieved from the accepted studies</li> </ul> </li> <li>Socio-Technical Grounded Theory (STGT) tools <ul> <li>Result of the use of&nbsp;Targeted Coding and Constant Comparison</li> </ul> </li> <li>The Google Colab Notebook <ul> <li>Code in python</li> <li>Results</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

JOSSE: A Software Development Effort Dataset Annotated with Expert Estimates

<p>The JIRA Open-Source Software Effort (JOSSE) dataset consists of software development and maintenance tasks collected from the JIRA issue tracking system for Apache, JBoss, And Spring open-source projects. All the issues were annotated with actual effort and 19% of them were annotated with expert estimates. JOSSE is a task-based dataset with a textual attribute represented as a task description for each data point. This paper explains how the data were collected and details six data quality refinement procedures of the data points.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Are Neural Bug Detectors Comparable to Software Developers on Variable Misuse Bugs?

<p>Artifact for &quot;Are Neural Bug Detectors Comparable to Software Developers on Variable Misuse Bugs?&quot;</p> <p><strong>Abstract:</strong>&nbsp;</p> <p>Debugging, that is, identifying and fixing bugs in software, is a central part of software development. Developers are therefore often confronted with the task of deciding whether a given code snippet contains a bug, and if yes, where. Recently, data-driven methods have been employed to&nbsp;learn&nbsp;this task of bug detection, resulting (amongst others) in so called&nbsp;neural bug detectors. Neural bug detectors are trained on millions of buggy and correct code snippets.</p> <p>Given the &ldquo;neural learning&rdquo; procedure, it seems likely that neu- ral bug detectors &ndash; on the specific task of finding bugs &ndash; have a performance similar to human software developers. For this work, we set out to substantiate or refute such a hypothesis. We report on the results of an empirical study with over&nbsp;100&nbsp;software developers, targeting the comparison of humans and neural bug detectors. As detection task, we chose a specific form of bugs (variable misuse bugs) for which neural bug detectors have recently made significant progress. Our study shows that despite the fact that neural bug detectors see millions of such misuse bugs during training, software developers &ndash; when conducting bug detection as a majority decision &ndash; are slightly better than neural bug detectors on this class of bugs. Altogether, we find a large overlap in the performance, both for classifying code as buggy and for localizing the buggy line in the code. In comparison to developers, one of the two evaluated neural bug detectors, however, raises a higher number of false alarms in our study.</p> <p><strong>Content:</strong>&nbsp;The artifact includes the following components:</p> <ul> <li> <p><strong>Web UI:</strong>&nbsp;The developer survey was performed online in the browser of the participants. For this, we created a custom web interface tailored for our study task. We included both the implementation of the frontend (website) and backend implementation (buisness logic and database) in this artifact. Therefore, it is not only possible to replicate our survey with same interface and a new group of participants but it is also possible to extend the interface for future studies.&nbsp;</p> </li> <li> <p><strong>Neural bug detectors:&nbsp;</strong>We evaluate the performance of the developers against two neural bug detectors. In this artifact, we include the bug detectors (implementation + trained models) and the evaluation script used for producing our results. Besides the replication of our bug detector evaluation, the detectors can also be used in future projects for detecting variable misuse bugs in Java methods.</p> </li> <li> <p><strong>Analysis scripts</strong>:&nbsp;After collecting the raw results from the developers and neural bug detectors, we performed several analysis to gain insights how developers and bug detectors compare on the variable misuse task. We include all analysis steps in form of Jupyter notebooks in the artifact. With this, it is possible to reproduce all the figures of our paper.&nbsp;</p> </li> </ul> <p>In addition, we also provide further artifacts that were successfully evaluated at ASE 2022:</p> <p><strong>ASE 2022 Artifact:&nbsp;</strong><a href="https://doi.org/10.5281/zenodo.6958242">10.5281/zenodo.6958242</a></p> <p><strong>Virtual machine:&nbsp;</strong><a href="https://doi.org/10.5281/zenodo.6957849">10.5281/zenodo.6957849</a></p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

AI Tool Use and Adoption in Software Development by Individuals and Organizations: A Grounded Theory Study

<div> <p>This data represents four artifacts from our research in studying what impacts AI adoption and use in SE.&nbsp;It includes our interview questions, the codebook with example quotes, survey questions, and table with the code to category generation.</p> <p>This page includes supplementary materials associated with our paper entitled "<span>AI Tool Use and Adoption in Software Development by </span><span>Individuals and Organizations: A Grounded Theory Study</span>".</p> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

PREPCLIM software - additional material for publication in "Geoscientific Model Development" 2024

<p>Data set illustrating the software developed in the PREPCLIM project and used in the proposed paper:</p> <p>A Modeling System for Identification of Maize Ideotypes, optimal sowing dates and nitrogen<br>fertilization under climate change &ndash; PREPCLIM-v1</p> <p>https://doi.org/10.5194/gmd-2024-105<br>Preprint. Discussion started: 11 July 2024<br>c Author(s) 2024. CC BY 4.0 License.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Figure 2. Some screenshots from the software system-Design and Development of a Software System for Swarm Intelligence Based Research Studies

<p>All of the mentioned operations can be performed easily by using the provided controls over<br> the related interfaces &ndash; windows of each algorithm. It is also important that each algorithm interface<br> is supported by visual controls to view obtained results with typical iteration-based graphics or<br> problem oriented visual elements. For instance, resulting graph structures are automatically shown<br> by the algorithm interfaces after solving some specific, popular problems like Travelling Salesman<br> Problem (TSP), Vehicle Routing Problem (VCP)&hellip;etc. Visually improved using features and<br> functions of the software system are critical aspects to provide more effective and useful platform to<br> perform SI based research studies better.<br> Related to the designed and developed software system, some screenshots from the software<br> system [interfaces of two algorithms (IWDs and ABC)] are represented in Fig. 2.</p>

opencc-by-4.0Jun 2012View details →
zenodo40/100

Results of a research software programming and development survey at the University of Reading

<p>In 2017 an online survey of University of Reading staff active in or supporting research and registered PhD students was undertaken to assess the nature and extent of research programming and software development activities in the University, and to understand how the University might provide guidance, training and support. The survey was a administered by the Research Data Manager on behalf of the University&#39;s Research Data Management Steering Group. The survey ran from 1st November to 15th December 2017 and collected a total of 170 responses.</p> <p>The survey sought responses from anyone in the University who was involved in any of the following activities:</p> <ul> <li>writing code and using software for numerical and statistical analysis;</li> <li>creating and contributing to computational models or simulations;</li> <li>conducting Text and Data Mining (TDM) and content analysis;</li> <li>creating and contributing to software distributed as a product or implemented as a service;</li> <li>creating data visualisations;</li> <li>using markup languages to structure and render content.</li> </ul> <p>The survey was distributed using the Bristol Online Survey. A dataset of anonymised survey responses and a PDF of the survey questions are here included.</p>

opencc-by-4.0Feb 2018View details →
zenodo40/100

Survey and Interview Data from Mixed-Method Survey of Serverless Computing and Function-as-a-Service Software Development in Industrial Practice

<p>This dataset contains the almost-raw data resulting from two out of the three methods chosen by the researchers for their namesake study &laquo;A Mixed-Method Empirical Study of Function-as-a-Service Software Development in Industrial Practice&raquo;.&nbsp; Among the files are web survey questions, anonymised survey results, and interview guidelines. We encourage other researchers to perform open coding and other analysis techniques on the data to verify our claims and to generate new insights.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Replication Package for a Systematic Literature Mapping of Agility in Safety-Critical Software Development within the Aerospace Industry

<p>This file collection package facilitates the replication of a Systematic Literature Mapping (SLM) focused on Agility in Safety-Critical Software Development within the Aerospace Industry. Authored by J. Eduardo Ferreira Ribeiro, Jo&atilde;o Gabriel Silva, and Ademar Aguiar, this dataset is dedicated to improving transparency and reproducibility in this field of study and future research.</p> <p>Specifically, the package includes:</p> <ul> <li><a href="https://github.com/zemacedo99/Replication-Package-Builder/releases/tag/v1.0.2">Replication Package Builder Version 1.0.2</a></li> <li>A list of terms (both inclusion and exclusion) used to construct the research string.</li> <li>A list of venues unrelated to the research topic, to be excluded from the results.</li> <li>The inclusion and exclusion criteria applied during the study.</li> <li>Lists of publication results from indexing services like Scopus, IEEE Xplore, Science Direct, HAL Open Science, Springer Nature, and the ACM Digital Library are all provided in CSV file format.</li> <li>A list of all publications in CSV format, compiled after the automated exclusion phase using the established inclusion and exclusion criteria.</li> <li>Finally, a complete list of all publications, including those from Snowball sampling, in XLSX format was compiled after the manual exclusion phase using the established inclusion and exclusion criteria.</li> </ul> <p>Compiled and published on Saturday, September 14, 2024, this dataset is crucial for researchers seeking to replicate or extend the SLM's findings.</p> <p>Lastly, we thank J. Antonio Dantas Macedo for contributing to developing and providing this <a href="https://github.com/zemacedo99/Replication-Package-Builder">replication package builder</a>.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Replication Package for ICSE'21 paper - Representation of Developer Expertise in Open Source Software

<p>Replication package for ICSE&#39;21 paper: Representation of Developer Expertise in Open Source Software.</p> <p>See README for details.</p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record