Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.7.1
Dataset results
15 results for “Stack Exchange”
Stack Exchange Open Source site questions categorization
<p>This dataset contains the posts of Open Source Stack Exchange site, collected at the end of 2020, along with the categorization of the posts. For each post a category, and potentially a second one is indicated, along with the cluster (generic group) each category belongs to. The coding task of assigning each question to a category was performed by two independent coders for each question (the categorization of each coder is also provided in the dataset). The dataset contains also (in a separate file) a dictionary of the most correlated unigrams and bigrams per category.</p>
Dataset - What are the Machine Learning best practices reported by practitioners on Stack Exchange?
<p>The data correspond to the posts (questions and answers) retrieved by querying for posts related to the tag 'machine learning' and the phrase 'best practice(s).' The data were used as the basis for a study currently under review on discussing machine learning best practices as discussed by practitioners in question-and-answer communities such as Stack Exchange. The information from each type of post (i.e., questions and answers) is presented in multiple formats (i.e., .txt, .csv, and .xlsx).</p> <p> </p> <p><strong>Answers - Variables</strong></p> <ul> <li><strong>AID</strong>:<strong> </strong> Unique identification of the answer in the Q&A website.</li> <li><strong>ParentId</strong>: Unique identification of the question associated with the answer in the Q&A website </li> <li><strong>AcceptedAnswerId</strong> : In the case in which an answer is the most voted question associated with the <em>ParentId</em>, and it is different from the accepted answer, a different identifier from the <em>AID</em> is available. In the case in which the accepted question had a <em>score</em> lower than 1, a -1 is assigned. </li> <li><strong>ABody:</strong> HTML text of the answer.</li> <li><strong>Score:</strong> Upvotes - downvotes of the answer.</li> <li><strong>url_Answer:</strong> URL of the answer. The question URL can be from different websites. </li> <li><strong>type:</strong> best or accepted. Accepted in the case that the information belongs to the accepted answer of the <em>ParentId </em>question and best in the case in which it is the most voted question of the <em>ParentId </em>question.</li> <li><strong>Date: </strong>Creation date of the answer.</li> </ul> <p><strong>Questions - Variables</strong></p> <ul> <li><strong>QID</strong>: Unique identification of the question in the Q&A website. </li> <li><strong>AcceptedAnswerId</strong>: Unique identification of the accepted answer for a specific question in the Q&A website. In the case in which a question had a most-voted answer different from the accepted one, and the accepted one had a negative score, a -1 was assigned to the <em>AcceptedAnswerId</em><strong>. </strong></li> <li><strong>BestAnswerId</strong>: Unique identification of the most voted answer for a specific question in the Q&A website. In the case in which the most voted and accepted questions were the same, then a -1 was assigned to the <em>BestAnswerId</em>. </li> <li><strong>Qtitle</strong>: Title of the question.</li> <li><strong>QBody</strong>: HTML text of the question.</li> <li><strong>Score</strong>: Upvotes - downvotes of the questions.</li> <li><strong>QTags</strong>: Tags that are associated with each question.</li> <li><strong>url_question</strong>: URL of the question. The question URL can be from different websites. </li> <li><strong>Date</strong>: Creation date of the question</li> </ul> <p>This dataset is a subset of the Stack Exchange dump of 03.2021 (<a href="https://archive.org/details/stackexchange_20210301">https://archive.org/details/stackexchange_20210301</a>) in which a series of filters were applied to obtain the data used in the study.</p>
SANER 2022 - Industrial Track - Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange
<p>Dataset related to the paper Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange. </p> <p> </p> <p>Saner 2022 Industrial Track</p>
Mathematics Stack Exchange API Q&A Data
<p>This dataset was compiled as part of the ESPRC project "Example-driven machine-human collaboration in mathematics", for the purpose of doing text-based analysis of mathematical discourse and for the construction of a conversational mathematics bot.</p> <p>It consists of approximately 1 million mathematics questions and their respective answers, as well as markers of interaction quality (such as user-provided scoring of question and answer quality) and social dynamics (reputation scores, badges, etc).</p> <p>The data was obtained from the <a href="https://stackexchange.com/">StackExchange</a> website, by querying the <a href="https://api.stackexchange.com/">Stack Exchange API</a> according to its documentation.</p> <p> </p> <p> </p>
Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange
<p>Dataset used in the work 'Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange'.</p> <p>Authors not presented due to double blind review.</p>
Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange
<p>Dataset used in the work 'Investigating the Point of View of Project Management Practitioners on Technical Debt - A Preliminary Study on Stack Exchange'.</p>
Tendencia de topics y comunidades activas en Stack Exchange
<p>La plataforma Stack Exchange tiene diferentes foros de diversos temas. En estos dataset los que se busca es recorger las tendencias de los topics y, así, poder evaluar la actividad de las comunidades dentro de este sitio web. Esto nos permite una visión general sobre qué tecnologías y categorías generan más interacción en la plataforma. </p>
Investigating the Point of View of Project Management Practitioners on Technical Debt - A Study on Stack Exchange
<p>Investigating the Point of View of Project Management Practitioners on Technical Debt - A Study on Stack Exchange.</p> <p>Journal of Software Engineering Research and Development (JSERD) 2023.</p> <p> </p>
Replication Package for "The Double-edged Sword of Banning Generative AI on Online Question-and-Answer Communities: Evidence from Stack Exchange"
<p>This is a replication package for "The Double-edged Sword of Banning Generative AI on Online Question-and-Answer Communities: Evidence from Stack Exchange".</p>
Exploring the Black Box: Analyzing Explainable AI Challenges and Best Practices Through Stack Exchange Discussions
Open the record for dataset details and reuse information.
Investigating the use of Snowballing on Gray Literature Reviews: A Study on Stack Exchange
<p>Dataset and scripts for the papaer "Investigating the use of Snowballing on Gray Literature Reviews: A Study on Stack Exchange".</p>
Dataset of the Paper "Architecture Decisions in Quantum Software Systems: An Empirical Study on Stack Exchange and GitHub"
<p>This dataset was collected from GitHub and Stack Exchange (including Stack Overflow, Quantum Computing Stack Exchange, and Computer Science Stack Exchange) to conduct an empirical study on architecture decisions in quantum software systems. We provide below a brief description of each file:</p><p><strong>1. Dataset (GitHub).xlsx</strong></p><p>contains selected quantum software projects from GitHub with project names, issue IDs, and issue URLs and the data extracted from the GitHub issues that are related to architecture decisions in quantum software development.</p><p><strong>2. Dataset (SO).xlsx</strong></p><p>contains the IDs and URLs of Stack Overflow (SO) labeled posts and the extracted data from the Stack Overflow posts that are related to architecture decisions in quantum software development.</p><p><strong>3. Dataset (QC).xlsx</strong></p><p>contains the IDs and URLs of Quantum Computing (QC) Stack Exchange labeled posts and the extracted data from the Quantum Computing Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>4. Dataset (CS).xlsx</strong></p><p>contains the IDs and URLs of Computer Science (CS) Stack Exchange labeled posts and the extracted data from the Computer Science Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>5. Extracted Data (GitHub+SO+QC+CS).xlsx</strong></p><p>provides the final results of data extracted from the related GitHub issues, SO posts, QC posts, and CS posts.</p>
Replication Package for Automating Technical Debt Management: Insights from Practitioner Discussions in Stack Exchange
Open the record for dataset details and reuse information.
Technical Debt on Agile Projects: Managers' point of view at Stack Exchange
<p>Technical Debt on Agile Projects: Managers’ point of view at Stack Exchange</p> <p>XXI Simpósio Brasileiro de Qualidade de Software (SBQS '22), Nov 07--10, 2022, 2022, Curitiba, Paraná, Brasil.</p>
Architecture Decisions in Quantum Software Systems: An Empirical Study on Stack Exchange and GitHub
<p>This dataset was collected from GitHub and Stack Exchange (including Stack Overflow, Quantum Computing Stack Exchange, and Computer Science Stack Exchange) to conduct an empirical study on architecture decisions in quantum software systems. We provide below a brief description of each file:</p><p>1. Dataset (GitHub).xlsx</p><p>contains selected quantum software projects from GitHub with project names, issue IDs, and issue URLs and the data extracted from the GitHub issues that are related to architecture decisions in quantum software development.</p><p>2. Dataset (SO).xlsx</p><p>contains the IDs and URLs of Stack Overflow (SO) labeled posts and the extracted data from the Stack Overflow posts that are related to architecture decisions in quantum software development.</p><p>3. Dataset (QC).xlsx</p><p>contains the IDs and URLs of Quantum Computing (QC) Stack Exchange labeled posts and the extracted data from the Quantum Computing Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p>4. Dataset (CS).xlsx</p><p>contains the IDs and URLs of Computer Science (CS) Stack Exchange labeled posts and the extracted data from the Computer Science Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p>5. Extracted Data (GitHub+SO+QC+CS).xlsx</p><p>provides the final results of data extracted from the related GitHub issues, SO posts, QC posts, and CS posts.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.