Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

5 results for “Mining Software Repositories”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data supplementing the conference paper "Who you gonna call? Analyzing web requests in Android applications", 14th International Conference on Mining Software Repositories 2017.

<p>This repository contains the data supplementing the paper:</p> <p>M. Rapoport, P. Suter, E. Wittern, O. Lhótak, J. Dolby, "Who you gonna call? Analyzing web requests in Android applications", MSR 2017.</p> <p>A detailed description of the data is included in the archive in README.md.</p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Data showcase papers published in the Mining Software Repositories (MSR) conference

<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <p>citation-table.csv: SWEBOK areas of citing studies<br> citations.bib: Bibliographic details of citing studies<br> citing_dp_dois_citations.txt: Citations of citing studies<br> data_papers.bib: MSR data papers<br> dp_dois_citations.txt: Citations of data papers<br> false_citations.bib: Citing studies that don&#39;t actuall use data papers<br> msr-all: Bibliographic details of all MSR papers<br> ndp_dois_citations.txt: Citations of non-data papers<br> ndp_rand_dois_citations.txt: Citations of a randomly chose non-data paper weighted sample<br> self-citations.csv: Data papers citations by their authors</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

How are software repositories mined? A systematic literature review of workflows, methodologies, reproducibility, and tools

<p>This is the excel spreadsheet dataset containing our analysis of papers performing mining software repositories research from the conferences ICSE, ESEC/FSE, and MSR from the years 2018 - 2020. The data is broken into columns&nbsp;and can be explained at a high-level as follows:</p> <p>Column&nbsp; &nbsp; &nbsp; Content</p> <p>1&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The paper being analyzed</p> <p>2&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Does the paper state the data they analyzed is available</p> <p>3&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Does the paper perform some sort of data analysis or sampling&nbsp;using data others have compiled in the past</p> <p>4&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Does the paper state a timestamp for when they begin their work</p> <p>5&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Does the paper state the use of systems pre-built to help with MSR&nbsp;work</p> <p>6 - 18&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Forms of sampling researchers may have employed to select their&nbsp;data</p> <p>19&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;What datasets (if any) were used in the analysis</p> <p>20&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;What tools (if any) were used in the analysis</p> <p>21&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;How they performed their data sampling workflow</p> <p>22&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;How they performed their data filtering workflow</p> <p>23&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;How they performed their data retrieval workflow</p> <p>24&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Did they create any scripts in each of these workflows</p> <p>25 - 33&nbsp; &nbsp; &nbsp; &nbsp; Did they publish a replication package and what is contained within</p> <p>34&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Is the paper describing a&nbsp;tool for research or not</p> <p>35&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Short description of the paper read</p> <p>36&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;A high-level category of the work performed in each paper</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Data showcase papers published in the Mining Software Repositories (MSR) conference (v2.2)

<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <ul> <li>citing_dp_dois_citations.txt: Strong and weak citations of (strong and weak) citation papers</li> <li>data_paper_clustering.csv: The clustering process of MSR data papers</li> <li>data_paper_clusters.csv: Clusters of MSR data papers</li> <li>data_papers.bib: Bibliographic details of MSR data papers, along with their assigned clusters (field &#39;cluster&#39;) and strong citations (field &#39;usedby&#39;)</li> <li>dp_dois_citations.txt: Strong and weak citations of MSR data papers</li> <li>msr-all: Bibliographic details of all MSR (data and non-data) papers</li> <li>ndp_dois_citations.txt: Strong and weak citations of MSR non-data papers</li> <li>ndp_rand_dois_citations.txt: Strong and weak citations of a randomly chosen MSR non-data paper weighted sample</li> <li>self-citations.txt: Strong citations of MSR data papers by their authors</li> <li>strong_citation_classification.csv: The classification process of strong citation papers according to the SWEBOK knowledge areas</li> <li>strong_citation_fields.csv: SWEBOK knowledge areas of strong citation papers</li> <li>strong_citations.bib: Bibliographic details of strong citation papers</li> <li>survey_questionnaire.pdf: The final survey questionnaire</li> <li>survey_responses.csv: Anonymized responses of the final survey questionnaire (Email addresses have been excluded for privacy reasons.)</li> <li>weak_citations_notes.bib: Weak citations of MSR data papers and the use they make</li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Mining Software Repositories for the Characterization of Continuous Integration and Delivery

<p>Continuous Integration (CI) and Delivery (CD) are software development practices increasingly adopted in industry, which aiming at ensuring quality and stability, and making the whole process more efficient and less error-prone. Software engineers that incorporate these practices may have difficulties trying to get an overview of the impact of adopting them and tracking progress. Therefore, performing analysis regarding the level of the application of CI/CD practices in a software repository can help to understand how such practices are adopted, which can result in a better characterization of their benefits and limitations. To support this analysis, we developed Garimpeiro, a web application that supports the characterization of open source repositories hosted on GitHub regarding the level of adoption of CI/CD practices. This video describes the tool and demonstrates how it can be used.<br> &nbsp;</p>

opencc-by-4.0Oct 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record