Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “repository mining”
Data format figures-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY
<p>The original data set included noisy, missing and inconsistent data. Data<br> preprocessing improved the quality of the data and facilitated e±cient data<br> mining tasks.<br> Before the experiment, we prepared data suitable to next operation as<br> following steps:<br> ² Delete or replace missing values;<br> ² Delete redundant properties (columns);<br> ² Data Transformation;<br> ² Data Discretization;<br> ² Export data to a required .ar® or .csv format ¯le [11].<br> The original and modi¯ed formats of data set are shown in Figure 1 and<br> Figure 2.<br> Data visualization is also a very useful technique because it helps to deter-<br> mine the di±culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The ¯gure 3 shows the variation<br> of the temperature in time.</p>
Figure 3. Data visualization-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY
<p>Data visualization is also a very useful technique because it helps to deter-<br> mine the di±culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The ¯gure 3 shows the variation<br> of the temperature in time.</p>
Data mining tool to discover DevOps trends from public repositories: Predicting Release Candidates with gthbmining.rc
<p>Public repositories have been performing an essential role in bring- ing software and services to technical communities and general users. Most of the cases, public repositories have a DevOps tool, with a live and historical database behind it, to support delivering and all steps this software or service should adopt before going to production. This paper introduces gthbmining, a data mining set of tools to discover DevOps trends from public repositories, and presents the module gthbmining.rc. Considering the premise of a GitHub public repository, the main contribution here is pre- dicting release candidates, an important label a software release has. The methodology, architecture, components and interfaces are explained, as well as potential users. The results show a reliable and flexible tool, as classifiers metrics and graphics are provided, along with the possibility to add new data mining algorithms in the open source module presented. Related works are also supplied, and a conclusion shows the outcomes gthbmining.rc can provide.</p>
Data supplementing the conference paper "Who you gonna call? Analyzing web requests in Android applications", 14th International Conference on Mining Software Repositories 2017.
<p>This repository contains the data supplementing the paper:</p> <p>M. Rapoport, P. Suter, E. Wittern, O. Lhótak, J. Dolby, "Who you gonna call? Analyzing web requests in Android applications", MSR 2017.</p> <p>A detailed description of the data is included in the archive in README.md.</p>
Data showcase papers published in the Mining Software Repositories (MSR) conference
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <p>citation-table.csv: SWEBOK areas of citing studies<br> citations.bib: Bibliographic details of citing studies<br> citing_dp_dois_citations.txt: Citations of citing studies<br> data_papers.bib: MSR data papers<br> dp_dois_citations.txt: Citations of data papers<br> false_citations.bib: Citing studies that don't actuall use data papers<br> msr-all: Bibliographic details of all MSR papers<br> ndp_dois_citations.txt: Citations of non-data papers<br> ndp_rand_dois_citations.txt: Citations of a randomly chose non-data paper weighted sample<br> self-citations.csv: Data papers citations by their authors</p> <p> </p>
BEIRUT: Repository Mining for Defect Prediction
<p>This artifact includes two CSV files.</p> <p>The sample_metrics.csv contains a sample of 153 metrics extracted from project <a href="https://ratis.apache.org/">ratis</a>.</p> <p>The sample_prediction.csv contains a sample of the prediction results from the application of defect prediction to the extracted metrics.</p>
Appendix-A: Online Repositories Available for Text Mining
<p>Appendix A is associated with Chapter 2: Text data and where to find them of the book: Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
How are software repositories mined? A systematic literature review of workflows, methodologies, reproducibility, and tools
<p>This is the excel spreadsheet dataset containing our analysis of papers performing mining software repositories research from the conferences ICSE, ESEC/FSE, and MSR from the years 2018 - 2020. The data is broken into columns and can be explained at a high-level as follows:</p> <p>Column Content</p> <p>1 The paper being analyzed</p> <p>2 Does the paper state the data they analyzed is available</p> <p>3 Does the paper perform some sort of data analysis or sampling using data others have compiled in the past</p> <p>4 Does the paper state a timestamp for when they begin their work</p> <p>5 Does the paper state the use of systems pre-built to help with MSR work</p> <p>6 - 18 Forms of sampling researchers may have employed to select their data</p> <p>19 What datasets (if any) were used in the analysis</p> <p>20 What tools (if any) were used in the analysis</p> <p>21 How they performed their data sampling workflow</p> <p>22 How they performed their data filtering workflow</p> <p>23 How they performed their data retrieval workflow</p> <p>24 Did they create any scripts in each of these workflows</p> <p>25 - 33 Did they publish a replication package and what is contained within</p> <p>34 Is the paper describing a tool for research or not</p> <p>35 Short description of the paper read</p> <p>36 A high-level category of the work performed in each paper</p>
Supplementary web page for the paper "SEAL: Integrating Program Analysis and Repository Mining"
<p>This is an archive of the supplementary material for the paper “SEAL: Integrating Program Analysis and Repository Mining” including the website and dataset. The website can also be viewed here: <a href="https://se-sic.github.io/paper-SEAL/">https://se-sic.github.io/paper-SEAL/</a></p>
Data showcase papers published in the Mining Software Repositories (MSR) conference (v2.2)
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <ul> <li>citing_dp_dois_citations.txt: Strong and weak citations of (strong and weak) citation papers</li> <li>data_paper_clustering.csv: The clustering process of MSR data papers</li> <li>data_paper_clusters.csv: Clusters of MSR data papers</li> <li>data_papers.bib: Bibliographic details of MSR data papers, along with their assigned clusters (field 'cluster') and strong citations (field 'usedby')</li> <li>dp_dois_citations.txt: Strong and weak citations of MSR data papers</li> <li>msr-all: Bibliographic details of all MSR (data and non-data) papers</li> <li>ndp_dois_citations.txt: Strong and weak citations of MSR non-data papers</li> <li>ndp_rand_dois_citations.txt: Strong and weak citations of a randomly chosen MSR non-data paper weighted sample</li> <li>self-citations.txt: Strong citations of MSR data papers by their authors</li> <li>strong_citation_classification.csv: The classification process of strong citation papers according to the SWEBOK knowledge areas</li> <li>strong_citation_fields.csv: SWEBOK knowledge areas of strong citation papers</li> <li>strong_citations.bib: Bibliographic details of strong citation papers</li> <li>survey_questionnaire.pdf: The final survey questionnaire</li> <li>survey_responses.csv: Anonymized responses of the final survey questionnaire (Email addresses have been excluded for privacy reasons.)</li> <li>weak_citations_notes.bib: Weak citations of MSR data papers and the use they make</li> </ul>
Mining Software Repositories for the Characterization of Continuous Integration and Delivery
<p>Continuous Integration (CI) and Delivery (CD) are software development practices increasingly adopted in industry, which aiming at ensuring quality and stability, and making the whole process more efficient and less error-prone. Software engineers that incorporate these practices may have difficulties trying to get an overview of the impact of adopting them and tracking progress. Therefore, performing analysis regarding the level of the application of CI/CD practices in a software repository can help to understand how such practices are adopted, which can result in a better characterization of their benefits and limitations. To support this analysis, we developed Garimpeiro, a web application that supports the characterization of open source repositories hosted on GitHub regarding the level of adoption of CI/CD practices. This video describes the tool and demonstrates how it can be used.<br> </p>
Replication Package for 'Lessons Learned from Mining the Hugging Face Repository'
<p>Replication Package attached to the 'Lessons Learned from Mining the Hugging Face Repository' article. Within the README and accompanying scripts, you will find detailed instructions to guide you through the analysis conducted in the article. Please, remember to cite the original paper if the replication package is used.</p>
Repository for Analysis of the Mined Dataset (RE: EMT6 Paper)
<p><strong>Table of Contents</strong></p><ol><li>Main Description</li><li>File Description</li></ol><p> </p><p><strong>1. Main Description</strong></p><p>---------------------------</p><p>This is a repository for the mined data and code used pertaining to the manuscript titled "A TCR β chain-directed antibody-fusion molecule that activates and expands subsets of T cells and promotes antitumor activity.".</p><p>The following libraries are required for script execution:</p><ul><li>Seurat</li><li>scReportoire</li><li>ggplot2</li><li>stringr</li><li>dplyr</li><li>ggridges</li><li>ggrepel</li><li>ComplexHeatmap</li></ul><p> </p><p> </p><p><strong>File Descriptions</strong></p><p>---------------------------</p><ul><li>The "Marengo_Code_Review.R" contains all the code needed to generate figures, and perform differential gene expression analysis for the mined datasets. </li><li>The "Cis_paper_EM_BetterEF.rds" file contains mined dataset sourced from DOI: 10.1038/s41586-022-05257-0. Whereas, the "combined_three_effectors.rds" file contains the mined dataset sourced from DOI: 10.1038/s41586-022-05257-0.</li></ul>
Data mining tool to discover DevOps trends from public repositories: Predicting Release Candidates with gthbmining.rc
<p>Public repositories have been performing an essential role in bring- ing software and services to technical communities and general users. Most of the cases, public repositories have a DevOps tool, with a live and historical database behind it, to support delivering and all steps this software or service should adopt before going to production. This paper introduces gthbmining, a data mining set of tools to discover DevOps trends from public repositories, and presents the module gthbmining.rc. Considering the premise of a GitHub public repository, the main contribution here is pre- dicting release candidates, an important label a software release has. The methodology, architecture, components and interfaces are explained, as well as potential users. The results show a reliable and flexible tool, as classifiers metrics and graphics are provided, along with the possibility to add new data mining algorithms in the open source module presented. Related works are also supplied, and a conclusion shows the outcomes gthbmining.rc can provide.</p>
Replication Package for 'Exploring the Carbon Footprint of Hugging Face's ML Models: A Repository Mining Study'
<p>Replication Package attached to the 'Exploring the Carbon Footprint of Hugging Face's ML Models: A Repository Mining Study' article. Within the README and accompanying scripts, you will find detailed instructions to guide you through the analysis conducted in the article.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.