Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

72

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

72 results for “stack overflow”

Learn how ShareScore rates datasets ↗
zenodo32/100

Supplemental materials for the study of Stack Overflow

<p>Supplemental materials for&nbsp;the study of Stack Overflow.</p> <p>-Source bib files of the studied papers</p> <p>-Abbreviations and the full names of the selected&nbsp;venues</p> <p>-Correspondences between the authors&#39; abbreviations and full names</p> <p>-Correspondences between the affiliations&#39; abbreviations and full names</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Dataset for Stack overflow Manual Results about challenges in developing Spark applications

<p>This dataset contains&nbsp;Stack Overflow manual study results for the paper &quot;An Empirical Study on the Challenges that Developers Encounter When Developing Apache Spark Applications&quot;.</p> <ul> <li>the <em>data</em>&nbsp;folder contains the&nbsp;<em>Stackoverflow Manual Results.csv&nbsp;</em>file that is the manual analysis result for the Stack Overflow posts. The CSV file contains information on the classification of the data, the reasons and the number of views, etc.&nbsp;</li> </ul> <p>&nbsp;</p> <ul> <li>the <em>scripts</em> folder&nbsp;contains the python and SQL files that are used for data collection and data analysis. <ul> <li><em>query_data.sql</em>&nbsp;is used to collect data from the Stack Exchange website.</li> <li><em>sample.py</em>&nbsp;is used to sample data for the manual analysis in the paper.</li> <li><em>common_issue.py</em>&nbsp;is used to study the percentage of common issues in rq1.&nbsp;</li> <li><em>popularity.py&nbsp;</em>is used to calculate the average of normalized view counts in rq2.</li> <li><em>popularity_difficulty.py</em>&nbsp;is used to calculate the average of raw view counts and the median hours to receive an answer in rq2.</li> <li><em>root_cuase.py</em>&nbsp;is used to study the percentage of root causes in rq3.&nbsp;</li> </ul> </li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo32/100

An Empirical Study of Package Management Issues via Stack Overflow

<p>The&nbsp;package manager (PM) is crucial to most technology stacks, acting as a broker to ensure that a verified dependency package is correctly installed, configured, or removed from an application. Diversity in technology stacks has led to dozens of PMs with various features. While our recent study indicates that package management features of PM are related to end-user experiences, it is unclear what those issues are and what information is required to resolve them.&nbsp; In this paper, we have investigated PM issues faced by end-users through an empirical study of content on Stack Overflow (SO). We carried out a qualitative analysis of 1,131 questions and their accepted answer posts for three popular PMs (i.e., Maven, npm, and NuGet ) to identify issue types, underlying causes, and their resolutions. Our results confirmed that end-users struggle with PM tool usage (approximately 64-72%). We observed that most issues are raised by end-users due to a lack of instructions and errors messages from PM tools. In terms of issue resolution, we observed that external link sharing is the most common practice to resolve PM issues. Additionally, we found that links pointing to useful resources (i.e., official documentation websites, tutorials, etc.) are the most frequently shared, indicating the potential for tool support and the ability to provide relevant information.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Replication materials for: Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow

<p>These are the replication materials for the paper:<br>Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow<br>By: Maria del Rio-Chanona, Nadzeya Laurentsyeva, and Johannes Wachs. &nbsp;<br>Preprint: https://arxiv.org/abs/2307.07367<br>Under Revision for PNAS Nexus</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Dataset of the Paper "Exploring the Problems, their Causes and Solutions of AI Pair Programming: A Study on GitHub and Stack Overflow"

<p>The dataset collected from GitHub Discussions, GitHub Issues, and Stack Overflow is used to conduct an empirical study on the problems, causes, and solutions of using GitHub Copilot in practice. A brief description of each document in the dataset is provided below:</p> <p><br><strong>1. Dataset(GitHub_Discussions).xlsx</strong></p> <p>contains the Discussion IDs and URLs in the Copilot category of GitHub Discussions, and the data extracted from the related discussions along with analysis results.</p> <p><strong>2. Dataset(GitHub_Issues).xlsx</strong></p> <p>contains the Issue IDs and URLs of the labelled issues which are related to Copilot from GitHub Issues, and the data extracted from the related issues along with analysis results.</p> <p><strong>3. Dataset(SO_Posts).xlsx</strong></p> <p>contains the SO Post IDs and URLs of the labelled posts which are related to Copilot from Stack Overflow, and the data extracted from the related posts along with analysis results.</p> <p><strong>4. Extracted_Data.xlsx</strong></p> <p>contains the final results of the data extracted from GitHub Discussions, GitHub Issues, and SO posts.</p> <p><strong>5. pilot labelling folder</strong></p> <p>contains three .xlsx files (i.e., Pilot_Labelling(GitHub_Discussions).xlsx, Pilot_Labelling(GitHub_Issues), and Pilot_Labelling(SO)), with each file corresponding to one of the three data sources and containing the pilot data labelling results with the Cohen's kappa value.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Exploring Accessibility Trends and Challenges in Mobile App Development: A Study of Stack Overflow Questions

<p>This is the dataset for the paper: Exploring Accessibility Trends and Challenges in Mobile App Development: A Study of Stack Overflow Questions</p> <p>This paper was accepted for publication at the 58th Hawaii International Conference on System Sciences (HICSS) - Software Technology Track</p> <p>Preprint: <a href="https://arxiv.org/abs/2409.07945">https://arxiv.org/abs/2409.07945</a></p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Mobile App Security Trends and Topics: An Examination of Questions From Stack Overflow

<p>This is the dataset for the paper: Mobile App Security Trends and Topics: An Examination of Questions From Stack Overflow</p> <p>This paper was accepted for publication at the 58th Hawaii International Conference on System Sciences (HICSS) - Software Technology Track</p> <p>Preprint: <a href="https://arxiv.org/abs/2409.07926">https://arxiv.org/abs/2409.07926</a></p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Stack Overflow's Hidden Nuances: How Does Zip Code Define User Contribution? – Replication Package

<p>Collective intelligence constitutes a foundational element within online community question-and-answering (CQA) platforms, such as Stack Overflow, being the source of most programming-related issues. Despite this relevance, concerns remain regarding issues surrounding user participation. Precedent research tends to focus on simple numerical measurements to analyse participation, which may sideline the inherent, subtler aspects.</p> <p>The proposed study aims to bridge this gap by operationalising 11 distinct metrics to represent user participation, behaviour, and community value across different regions of the USA. The study also conducts inductive content analysis to understand the impact of regional contextual factors on users' knowledge sharing patterns.</p> <p>This replication package is provided for those interested in further examining our research methodology.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

A Cross-Continental Analysis of How Regional Cues Shape Developers' Stack Overflow Contributions – Replication Package

<p>Stack Overflow provides a wide range of knowledge for the software development community. Despite the importance of these platforms, several studies have shown that digital information tends to cluster geographically, which limits knowledge access that is otherwise necessary for innovation.</p> <p>The proposed study highlights the dynamics of users from different geographical backgrounds within Stack Overflow, which entails intra-country interactions, predominant topics of discourse, as well as their communication patterns. Finally, the study highlights that regional behavioural variations stem beyond cultural factors, encompassing technological advancement, entrepreneurial ventures, and workforce composition.&nbsp;</p> <p>This replication package is provided for those interested in further examining our research methodology.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Stack Overflow's Hidden Nuances: How Does Zip Code Define User Contribution? – Replication Package

<p>Collective intelligence constitutes a foundational element within online community question-and-answering (CQA) platforms, such as Stack Overflow, being the source of most programming-related issues. Despite this relevance, concerns remain regarding issues surrounding user participation. Precedent research tends to focus on simple numerical measurements to analyse participation, which may sideline the inherent, subtler aspects.</p> <p>The proposed study aims to bridge this gap by operationalising 11 distinct metrics to represent user participation, behaviour, and community value across different regions of the USA. The study also conducts inductive content analysis to understand the impact of regional contextual factors on users' knowledge sharing patterns.</p> <p>This replication package is provided for those interested in further examining our research methodology.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Stack Overflow Chat Room Messages (up to Jul. 2019)

<p>Stack Overflow chat room dataset</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

A First Look at Information Highlighting in Stack Overflow Answers

<p>This dataset is used for token classification of words tagged with formatting tags.</p> <p>Training set: Used for training the spacy model.<br> Test set: Used for testing the models.<br> Test_sentences: Examples of some sentences for testing the results.</p>

opencc-by-4.0Mar 2023View details →
zenodo28/100

An Annotated Dataset of Stack Overflow Post Edits

<p>To improve software engineering, software repositories have been mined for code snippets and bug fixes. Typically, this mining takes place at the level of files or commits. To be able to dig deeper and to extract insights at a higher resolution, we hereby present an annotated dataset that contains over 7 million edits of code and text on Stack Overflow. Our preliminary study indicates that these edits might be a treasure trove for mining information about fine-grained patches, e.g., for the optimisation of non-functional properties.</p> <p>EDIT: In the more recent version I fixed&nbsp;<a href="https://zenodo.org/api/files/85e326db-29e3-432f-bf30-cb88deb89deb/GetEditContent.sql">GetEditContent.sql</a>, which had an ambiguous column name in one of the select statements.</p>

opencc-by-sa-4.0Apr 2020View details →
zenodo28/100

Dataset for the paper: Generating Question Titles for Stack Overflow from Mined Code Snippets

<p>This is the dataset for our paper:&nbsp;Generating Question Titles for Stack Overflow from Mined Code Snippets</p> <p>All the data are extracted from the Stack Overflow data dump, please feel free to use! :)</p>

opencc-by-4.0May 2020View details →
zenodo28/100

Are comments on Stack Overflow well organized for easy retrieval by developers?

<p>Many Stack Overflow answers have associated informative comments that can strengthen them and assist developers. A prior study found that comments can provide additional information to point out issues in their associated answer, such as the obsolescence of an answer. By showing more informative comments (e.g., the ones with higher scores) and hiding less informative ones, developers can more effectively retrieve information from the comments that are associated with an answer. Currently, Stack Overflow prioritizes the display of comments and as a result, 4.4 million comments (possibly including informative comments) are hidden by default from developers. In this study, we investigate whether this mechanism effectively organizes informative comments. We find that: 1) The current comment organization mechanism does not work well due to the large amount of tie-scored comments (e.g., 87% of the comments have 0-score). 2) In 97.3% of answers with hidden comments, at least one comment that is possibly informative is hidden while another comment with the same score is shown (i.e., unfairly hidden comments). The longest unfairly hidden comment is more likely to be informative than the shortest one. Our findings highlight that Stack Overflow should consider adjusting the comment organization mechanism to help developers effectively retrieve informative comments. Furthermore, we build a classifier that can effectively distinguish informative comments from uninformative comments. We also evaluate two alternative comment organization mechanisms (i.e., the <em>Length</em> mechanism and the <em>Random</em> mechanism) based on text similarity and the prediction of our classifier.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

tags-stack-overflow

<h3><strong>Overview</strong></h3><p>This dataset is derived from tags on Stack Overflow posts. Each hyperedge corresponds to all of the tags used in a post, and each node in a hyperedge corresponds to a tag. The timestamps of the posts are in millisecond resolution, are adjusted so that the time of the earliest tag starts at 0, and are in ISO8601 format.</p><h4><strong>Statistics</strong></h4><p>Some basic statistics of this dataset are:</p><ul><li>number of nodes: 49,998</li><li>number of timestamped hyperedges: 14,458,875</li><li>number of unique hyperedges: 5,675,497</li><li>Component sizes:</li></ul><p>Component size, number</p><ul><li>49931, 1</li><li>2, 7</li><li>1, 53</li></ul><h4><strong>Source of original data</strong></h4><ul><li><a href="https://www.cs.cornell.edu/~arb/data/tags-stack-overflow/">tags-stack-overflow dataset</a></li><li><a href="https://archive.org/details/stackexchange">StackExchange</a></li></ul><h4><strong>References</strong></h4><p>If you use this data, please cite the following paper:</p><ul><li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. Proceedings of the National Academy of Sciences (PNAS), 2018.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo28/100

Unveiling Challenges in Python Library: Insights from Stack Overflow

<p>A dataset of the SANER 2024 ERA track submission "Unveiling Challenges in Python Library: Insights from Stack Overflow" containing all the comments in Stack Overflow that were classified.</p>

opencc-by-4.0Nov 2023View details →
zenodo28/100

Retrieving API knowledge from Tutorials and Stack Overflow based on Natural Language Queries

<p>The replication package of PLAN</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

What Edits Are Done on The Highly Answered Questions in Stack Overflow? An Empirical Study

<p>This contains the data set and the code we take advantage of to process data and make analyzation.</p>

opencc-by-4.0Mar 2019View details →
zenodo28/100

Stack Overflow June 2017 Dump

<p>June 2017 Dump of stack overflow data</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record