Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
207
datasets available to search
ShareScore release 0.7.1
Dataset results
207 results for “github”
On the Outdatedness of Workflows in the GitHub Actions Ecosystem
<p>This replication package contains all the material required to replicate the analyses we made for the paper entitled "On the Outdatedness of Workflows in the GitHub Actions Ecosystem" for Journal of Systems and Software.</p> <p>Refer to the README file(s) for more details on the content of this replication package.</p>
Replication Package for ESEC/FSE 2023 Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects?"
<p>This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models. The GitHub repository of the paper is available at <a href="https://github.com/mcxwx123/Sustainable_projects">https://github.com/mcxwx123/Sustainable_projects</a>.</p> <p>We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars. The raw data of repositories (collected in their first 1,3,5 months(s)) are stored in `Replication Package/data/prodata_1.pkl`, `Replication Package/data/prodata_3.pkl`, and `Replication Package/data/prodata_5.pkl`. The contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m3_t2_k1.pkl`.<br> `Replication Package/data/X_test_m3_t2_k1.pkl` and `Replication Package/data/y_test_m3_t2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table 3 and 4, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table 5. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section 3.1.<br> </p> <p>The explanations for collecting the variables, the examples of variable effects on project sustainability, and the hyperparameter setting of the machine learning models are provided in the README.md file.</p>
Dataset of the Paper "Demystifying Practices, Challenges and Expected Features of Using GitHub Copilot"
<p>This dataset collected from Stack Overflow (SO) and GitHub was used to conduct an empirical study on investigating the practices, challenges, and expected features of using GitHub Copilot. We provide below a brief description of each file in the dataset:</p> <p><strong>1. Dataset (SO).xlsx</strong></p> <p>contains the IDs and URLs of the labelled posts which are related to Copilot from SO, and the data extracted from these related SO posts.</p> <p><strong>2. Dataset (GitHub).xlsx</strong></p> <p>contains the discussion IDs and URLs in the Copilot category of GitHub Discussions, and the data extracted from the relevant discussions.</p> <p><strong>3. Extracted Data (SO+GitHub).xlsx</strong></p> <p>provides the final results of the data extracted from SO posts and GitHub discussions.</p>
Replication Package for "Are Prompt Engineering and TODO Comments Friends or Foes? An Evaluation on GitHub Copilot"
<p>This is the replication package accompanying the submission of "Are Prompt Engineering and TODO Comments Friends or Foes? An Evaluation on GitHub Copilot"</p>
Data accompanying the GitHub repository bartonlab/paper-HIV-latent-reservoir
<p>This dataset contains data from simulations of the HIV-1 dynamics that accompany the GitHub repository bartonlab/paper-HIV-latent-reservoir. The GitHub repository contains code for reproducing results described in the manuscript 'Clonal heterogeneity and antigenic stimulation shape persistence of the latent reservoir of HIV'. See the GitHub repository for details.</p>
The regulatory landscape of the yeast phosphoproteome - GitHub Data
<p>Data for: https://github.com/Villen-Lab/YeastPhosphoAtlasAnalysis</p>
List of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub
<p>This is the list of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub. The licenses are ranged by the amount of files that they cover, there are a total of 95 different licenses. Where possible, the names are presented as identifiers at https://spdx.org/licenses/. "GitHub" stands for no license in the file or the project.</p>
The GitHub Sponsors Data Set
<p>This is the data set for our submission to MSR 2020. "Contributing by Paying: A Data Set and Data-Driven Study on the GitHub Sponsors Program". Please download and unzip the .zip archive to get the data set used in the paper (folder <em>2019-11-30</em>) and recent updates (folder <em>sponsorship_snapshots</em>) on the snapshots of sponsors.</p>
To What Extent Do Developers Discuss End User Human-Centric Issues of Software on GitHub?
<p>Labelled dataset of a random selection of 1230 issue comments from 7 GitHub repositories.</p>
GitHub Open Source Survey 2017
<p>Data and accompanying documentation for the Open Source Survey fielded by GitHub and collaborators in 2017. Respondents were sourced via random sampling from traffic to licensed open source repositories on GitHub.com and from invitations sent to selected open source communities that work on other platforms. The files here include a README, a full copy of the (English) questionnaire, notes for working with the data, and two CSV data files. A report based on the subset of responses sampled from GitHub.com is available at http://opensourcesurvey.org/2017/. </p> <p>See the README for details on sampling methodology, and notes and questionnaire files for question wordings, response options, branching logic, and recoded variables. </p>
Scripts Extracao Github
Open the record for dataset details and reuse information.
tkeijzer/DamsImpactPersistence: DamsImpactPersistence code V1.0 GitHub
<p>https://github.com/tkeijzer/DamsImpactPersistence/commits/V1.0</p>
Informal diagrams (.drawio) mined from GitHub projects with more than 5000 stars
Open the record for dataset details and reuse information.
Survey design: GitHub Copilot in NAV
Open the record for dataset details and reuse information.
Exploring Activity and Contributors on GitHub: Who, What, When, and Where
<p>Supplementary materials.</p>
Understanding the archived projects on GitHub
<p>The supplementary material of the paper. </p>
Copy of datadir for github repo
Open the record for dataset details and reuse information.
scGFT GitHub use-case data
Open the record for dataset details and reuse information.
JITEC for Github
Open the record for dataset details and reuse information.
A Lot of Talk and a Badge: An Exploratory Analysis of Personal Achievements in GitHub
<p>Replication package for the paper "A Lot of Talk and a Badge: An Exploratory Analysis of Personal Achievements in GitHub" </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.