Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.7.1
Dataset results
4 results for “Issue Trackers”
Technical Debt Classification in Issue Trackers using Natural Language Processing based on Transformers
<p>In order to ensure transparency and reproducibility, we have made everything available publicly here, including the Code, Models, Datasets and more. All the files and their functionality used in this paper are explained clearly in the <strong>README.md</strong> file.</p> <p>Background: Technical Debt (TD) needs to be controlled and tracked during software development. Support to automatically track TD in issue trackers is limited. </p> <p>Aim: We explore the usage of a large dataset of developer-labeled TD issues in combination with cutting-edge Natural Language Processing (NLP) approaches to automatically classify TD in issue trackers.</p> <p>Method: We mine and analyze more than 160GB of textual data from GitHub projects, collecting over 55,600 TD issues and consolidating them into a large dataset (GTD dataset). We use such datasets to train and test Transformer ML models. Then we test the model's generalization ability by testing them on six unseen projects. Finally, we re-train the models including part of the TD issues from the target project to test their adaptability. </p> <p>Results and Conclusion: (i) We create and release the GTD dataset, a comprehensive dataset including TD issues from 6,401 public repositories with various contexts; (ii) By training Transformers using the GTD dataset, we achieve performance metrics that are promising; (iii) Our results are a significant step forward towards supporting the automatic classification of TD in issue trackers, especially when the models are adapted to the context of unseen projects after fine-tuning.</p>
PHI-base 5 web display issue tracker Nov 2021 to Sept 2024 with closed issues
<p>Closed GitHub issues are listed for the PHI-base 5 web display for the period 08/11/2021 to 17/09/2024.</p>
Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers
<p>The online repository for the paper: Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers, published in ECSA 2024.</p> <p>The online repo has the following folders:</p> <p>1) Classifier: it contains the source code of the classifier and quantitative analysis. Furthermore, there are sufficient details on the classifier performance, and how to replicate the results.</p> <p>2) Dataset: it contains the datasets we used to perform our analysis. This involves dataset before and after BERT classification, as well as exported json files used for training and analysis. The dataset in a zip file. This is a MySQL database in a zip file. It can be opened separately or opened using the search tool.</p> <p>3) Qualitative analysis: it contains coding book of design decisions in mailing lists as well as coding book of methods to discuss ADDs between emails and issues, as well as precision charts for the applied similarity algorithms.</p> <p>4) Searching tool: it contains the jar file and source code of the searching tool. The tool used to annotate and search for emails. It can be used to open the dataset from the zip directly. The tool is a jar file which can be run directly through a double click. The tool is tested on Windows and Linux. The folder also contains keywords to search for architectural emails, as well as documentation.</p>
Beyond the Code: Mining Self-Admitted Technical Debt in Issue Tracker Systems
<p>Self-admitted technical debt (SATD) is a particular case of Technical Debt (TD) where developers explicitly acknowledge their sub-optimal implementation decisions. Previous studies mine SATD by searching for specific TD-related terms in source code comments.By contrast, in this paper we argue that developers can admit technical debt by other means, e.g., by creating issues in tracking systems and labelling them as referring to TD. We refer to this type of SATD as issue-based SATD or just SATD-I. We study a sample of 286 SATD-I instances collected from five open source projects, including Microsoft Visual Studio and GitLab Community Edition. We show that only 29% of the studied SATD-I instances can be tracked to source code comments. We also show that SATD-I issues take more time to be closed, compared to other issues, although they are not more complex in terms of code churn. Besides, in 45% of the studied issues TD was introduced to ship earlier, and in almost 60%it refers to Design flaws. Finally, we report that most developers pay SATD-I to reduce its costs or interests (66%). Our findings suggest that there is space for designing novel tools to support technical debt management, particularly tools that encourage developers to create and label issues containing TD concerns.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.