Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.7.1
Dataset results
7 results for “Issue Tracking”
Apache Jira Issue Tracking Dataset
<p>This dataset contains the Jira Issue Tracking data of the Apache Software Foundation, enriched with topic modeling information using the BERTopic technique.</p> <p>You can use the dataset with the following steps:</p> <ol> <li>Set up a MongoDB instance.</li> <li>Download the data.</li> <li>Navigate to the download folder and use the mongorestore command (<a href="https://docs.mongodb.com/database-tools/mongorestore/" target="_blank" rel="noopener">https://docs.mongodb.com/database-tools/mongorestore/</a>) with the --gzip flag.</li> </ol> <p>Detailed instructions for setting up/reproducing/updating the dataset are also provided in the website <a href="https://authecesofteng.github.io/semantics-jira-dataset/" target="_blank" rel="noopener">https://authecesofteng.github.io/semantics-jira-dataset/</a> (relevant repos <a href="https://github.com/AuthEceSoftEng/jira-apache-downloader" target="_blank" rel="noopener">https://github.com/AuthEceSoftEng/jira-apache-downloader</a> and <a href="https://github.com/AuthEceSoftEng/jira-topic-extractor" target="_blank" rel="noopener">https://github.com/AuthEceSoftEng/jira-topic-extractor</a>)</p> <p><em>Note: you do not need to download the models .rar files if you do not use them, as the issue-topic distribution (along with probabilities) is already computed and stored in the topics Mongo collection.</em></p>
Replication Package for Identifying Self-Admitted Technical Debt in Issue Tracking Systems using Machine Learning
<p>This dataset includes pre-trained word embeddings and a weighted file that can be used to identify self-admitted technical debt (SATD) from issue tracking systems.</p>
Jbrowse multi-wiggle track issue
<p>SLBP crosslink sites bigwig files</p> <p>Data source: ENCODE consortium <a href="https://www.encodeproject.org/experiments/ENCSR483NOP/">SLBP eCILP datasets</a></p> <p>Bam files were downloaded and crosslink sites at the start sites were extracted and converted to bigwig files</p>
Developing Deep Learning Approaches to Find and Classify Architectural Design Decisions in Issue Tracking Systems
<p>This upload contains three files:</p> <ol> <li>mongodump-JiraRepos_2023-03-07-16 00.archive: Archive containing the issue data pulled from the Jira API.</li> <li>mongodump-MiningDesignDecisions.archive: Archive containing the data of our deep learning models and the labelled issues.</li> <li>mongodump-MiningDesignDecisions-lite.archive: Similar to the archive above, except this one only contains the best trained model (BERT). Also, it does not contain any embeddings or other files.</li> </ol> <p>This archive contains the data of our deep learning models and the labelled issues (MiningDesignDecisions archive).</p>
Managing Assurance Information: A Solution Based on Issue Tracking Systems
<p>Video presentation for NIER Track article " Managing Assurance Information: A Solution Based on Issue Tracking Systems "</p>
Detection of the Fire Drill anti-pattern: 15 real-world projects with ground truth, issue-tracking data, source code density, models and code
<p>This package contains artifacts for <strong>15</strong> real-world software projects. The data is supposed to aid the detection of the presence of the Fire Drill anti-pattern. We include original data, ground truth, code (experimental setups and models), and notebooks. The data supports two distinct methods of detecting the AP: a) through issue-tracking data, and b) through the underlying source code. This version of the dataset corresponds to <strong>v8</strong> of the <a href="https://arxiv.org/abs/2104.15090v8">technical report</a> and the <a href="https://github.com/MrShoenel/anti-pattern-models/releases/tag/arxiv-v8">GitHub repository</a>. The package includes the following:</p> <p>Original data:</p> <ul> <li>For each project, its <strong>original</strong> artifacts (e.g., wikis, meeting minutes, mentor's notes, etc.)</li> <li>Evaluation of raters' notes by the assessor</li> </ul> <p>Fire Drill in issue-tracking data:</p> <ul> <li><strong>Ground truth</strong> for whether and how strong each project exhibits the Fire Drill AP, on a scale from [0,10]. This was determined by two individual raters, who also reached a consensus.</li> <li>Coefficients for indicators for the first method, per project.</li> <li>Detailed issue-tracing data for each project: what occurred and when.</li> <li>Time logs for each project.</li> </ul> <p>Fire Drill in source-code data:</p> <ul> <li><strong>Four</strong> technical reports that document the developed method of how to translate a description into a detectable pattern, and to use the pattern to detect the presence and to score it (similar to the rating). Also includes a report for how activities were assigned to individual commits.</li> <li>Source code density data (metrics) for each commit in each of the nine projects as a separate dataset.</li> <li>Code: a snapshot of the repository that holds all code, models, notebooks, and pre-computed results, for utmost reproducibility (the code is written in R).</li> </ul>
Materials for Publication: Software Feature Request Detection in Issue Tracking Systems
<p>Additional figures, tables, experimant data, code, and results.</p> <p>See README.md for more information.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.