Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
207
datasets available to search
ShareScore release 0.7.1
Dataset results
207 results for “github”
Data accompanying the GitHub repository bartonlab/paper-MPL-inference
<p>This dataset contains data from simulations and analysis of within-host HIV-1 evolution that accompany the GitHub repository bartonlab/paper-MPL-inference. The GitHub repository contains code and data for reproducing results described in the manuscript "Fitness inference from complex evolutionary histories with genetic linkage." See the GitHub repository for details on the interpretation and analysis of this data.</p>
libs-github-api: fix "lines" mistake — GitHub API reports bytes, not lines
<p>Correct mistake in several places, including data file and scripts.</p>
Github-Archive Event Analysis
<p>This research project fetches event-data from githubarchives.org, filters the data to extract the information of interest, generates basic statistics and plots regarding to these statistics.</p> <p>The experiment is deployed to gain general knowledge on basic github-usage. Therefore, the following questions were followed:<br /> 1) How are GitHub-events distributed? - This can be derived by quantitative analysis of the distribution of different Event-Types.<br /> 2) What is the common ratio of commits per push, what are extremes? - again quantitative analysis of push-events is used.</p> <p>To visualize the results of this analysis, two plots are created. Each plot addresses one of the research-questions described above. Additionally, textual output is written to the terminal containing the precise numbers of the analysis and can be captured via native terminal functions.<br /> The data-files created by downloading and unzipping are just used as input for analysis and do not depict "final output".</p> <p>The given results were collected/created for the default time-period: 01.01.2015 00:00 to 01:00. </p> <p> </p> <p>The python3-Scripts need python version 3 and were executed on Linux! Additional libraries are required: matplotlib for python3</p>
452,000,000 public Git commits on GitHub (October 2016)
<p>What's inside</p> <p>part-000xx.lzo - LZO archives with the data (refer to "Format").</p> <p>part-000xx.lzo.index - LZO index files so that the archives are splittable in Hadoop.</p> <p>stats.csv.gz - GZIP-ed CSV file with some repository statistics related to the commits.</p> <p>Format</p> <p>part-000xx - text, one line per repository, every line is JSON with the following scheme:</p> <p>{ "r": "repository name", "c": [{ "h": "git hash", "a": "author's email hash", "t": "date and time commit was created", "m": "commit message" }, ...] }</p> <p>Date and time format is <em>mostly</em> Go language's time.Time.String(), I recommend to use dateutil.parse() to parse it with Python.</p> <p>Commit message contains explicit \r and \n symbols in order to be a single line.</p> <p>stats.csv has 4 columns: repository name, number of commits, number of contributors, average length of the commit messages.</p>
CoqStoq: A Dataset of Coq Proofs Scraped from GitHub and a Corresponding Benchmark
<p>CoqStoq is a dataset of Coq Proofs scraped from GitHub. CoqStoq contains proofs from 2,205 open source projects and includes a benchmark on proofs from 12 projects. This repository includes the projects from CoqStoq's benchmarks and the processed proof data collected by CoqStoq.</p>
Gitome: A curated dataset for GitHub README-related tasks
<h2><strong>About </strong></h2><p>This repository contains the source code implementation used to replicate the experimental results obtained in the submitted to the 21st International Conference on Mining Software Repositories (MSR204).</p><p><i>"Gitome: A curated dataset for GitHub README-related tasks"</i></p><p>authored by:</p><p>Claudio Di Sipio, Juri Di Rocco, Riccardo Rubei, Phuong Than Nguyen, and Davide Di Ruscio,</p><p>Università degli Studi dell'Aquila, Italy</p><h2><strong>Data description </strong></h2><p>The dataset is structured as follows: </p><ul><li><strong>emf_metamodel.zip:</strong> It contains the Ecore project with the Gitome data model</li><li><strong>existing_dumps.zip</strong>: It contains the existing datasets used to build Gitome</li><li><strong>lang_aggr_stats.csv: </strong>It contains the language data to compute the statistics presented in the paper</li><li><strong>langs.csv: </strong>It contains all the languages and their frequency</li><li><strong>output_dataset.zip:</strong> It contains the benchmarking dataset obtained by parsing the README files</li><li><strong>repository_lists.zip: </strong>It contains the list of repositories for each considered dataset (with possible duplicates)</li><li><strong>topics.csv:</strong> It contains all the topics and their frequency</li><li><strong>topics_aggr_stats.csv: </strong>It contains the topics data to compute the statistics presented in the paper</li><li><strong>gitome_repo.txt</strong>: It contains the list of the URLs of the considered GitHub repositories</li></ul><p> </p><h2><strong>How to collect Gitome</strong></h2><p>To collect all the data stored in this archive, please refer to the supporting Github repository https://github.com/MDEGroup/Gitome-MSR2024.</p><p> </p><p> </p>
Polski frontendu and American Back-end: GitHub Profile Recruitment Bias in GPT-4
<p>This repository serves as the online appendix for the paper "Polski frontendu and American Back-end: GitHub Profile Recruitment Bias in GPT-4".</p>
Resource Usage and Optimization Opportunities in Workflows of GitHub Actions - Artifact
<p>This package contains the data, the code used for collection and the analysis notebooks used to obtain the results presented in the paper: "Resource Usage and Optimization Opportunities in Workflows of GitHub Actions" published at ICSE2024.</p>
Reddit's NetSec forum GitHub projects, typology, maturity and popularity indicators
<p>Dataset including information of GitHub projects shared on Reddit's forum /r/netsec. It includes information about the posts itselves, like the author, title, image, tags, number of upvotes and number of comments, as well as repository information, such as number of stars, forks, issues, collaborators and the repository URL and last modification date.</p>
Data to support publication figures and animation scripts at GitHub: Modeling weather-driven long-distance dispersal of spruce budworm moths (Choristoneura fumiferana)
<p>Long-term studies of insect populations in the North American boreal forest have shown the vital importance of long-distance dispersal to the maintenance and expansion of insect outbreaks. In this work, we extend several concepts established previously in an empirically-based dispersal flight model with recent work on the physiology and behavior of the adult eastern spruce budworm (SBW) moth, Choristoneura fumiferana (Clem.). An outbreak of defoliating SBW in Quebec, ongoing since the mid-2000s, already covers millions of hectares of forests in eastern Canada and threatens to spread into neighboring areas through annual summertime episodes of long-distance dispersal. Such flight events in favorable conditions frequently include billions of SBW moths dispersing in the warm atmospheric boundary layer, typically starting around sunset and often lasting through several hours of wind-driven transport over hundreds of kilometers. Successful SBW dispersal to possibly distant host forest areas depends acutely on the weather. Here we describe the components and results of SBW–pyATM, an open-source individual-based modeling framework developed in Python for the simulation of these weather-driven SBW dispersal events. Using seasonal SBW phenology results from BioSIM at known outbreak locations and high-resolution Weather Research and Forecasting (WRF) model output, we focus on modeling dispersal flights over two successive nights in July 2013 in southern Quebec. Our flight model closely reproduces the SBW spatial patterns and motions observed by weather surveillance radar over the St. Lawrence estuary. With SBW–pyATM we can estimate landing locations for both male and female SBW and the resulting spatial patterns of egg distribution, allowing us eventually to forecast future larval defoliation activity in new locations where immigration could help overcome local limitations on SBW populations. This information could then support forest management decisions where SBW outbreaks threaten valuable resources.</p>
How Are Communication Channels on GitHub Presented to Their Intended Audience? – Dataset
<p>This dataset includes the data used for the thematic analysis of the paper in the title.</p> <ul> <li>all_projects.csv includes all large GitHub projects as described in the paper.</li> <li>analyzed_projects.csv includes the analyzed projects.</li> <li>checkRegex.py is the python script used to automatically find communication channels in markdown files.</li> <li>download_markdown_files.py is the python script used to download the markdown files.</li> </ul>
An evidence-based study on issue labeling in Github-based repositories (Supplementary material)
<p>The open-source software community has grown in size and importance over the years. As a consequence, the number of project contributors has increased considerably. The capability of open-source project repositories to accommodate issue reports is essential. An issue report encompasses a large set of data that describes the necessary changes a software should handle. As developers need detailed information to reproduce and find them, incomplete information is a severe problem that may influence triage and defect detection leading to delays in project maintenance. Issue trackers commonly use the labeling method to add extra details to issues. Knowing the importance of labels, this dissertation focus on investigating the usage, creation, and similarities in the context of the issue lifecycle in both maintenance and evolution in the repository issue trackers of the largest and most popular code hosting platform, Github. In addition, it analyzes the number of labeled and unlabeled issues in the repository and the connection between the issues' components, with an analysis focused on the lifecycle. The results indicate a significant correlation between repositories with many issues and the creation of labels, but not all repositories use them. 64.58% of the repositories insert new labels as the project evolves. 73.14% repositories applied on issues the Github standard labels. We also found an influence of primary issue fields such as title, description, and comments in most issue labels, impacting the creation and labeling issues. These numbers show that issue labeling is of prominent relevance for project maintenance and evolution. It provides developers with an easy and convenient way to inform about an incoming issue reported by systems users.</p>
On the Use of GitHub Actions in Software Development Repositories
<p>This replication package contains all the material required to replicate the analyses we made for our paper entitled "On the Use of GitHub Actions in Software Development Repositories" accepted for publication at the 38th International Conference on Software Maintenance and Evolution (ICSME) 2022.</p> <p>The required dependencies are listed in `requirements.txt` and can be installed (preferably in a virtual environment) using `pip install -r requirements.txt`. The notebooks are expected to be executed with `jupyter lab` (but any `jupyter` environment should do the job). The data (from `data/` folders) are generated by the various notebooks, starting from the raw dataset in the `data-raw` folder. Please refer to the paper and to the README file contained in that folder for more informations.</p> <p> </p>
The Good First Issue Recommendation Dataset from "GFI-Bot: Automated Good First Issue Recommendation on GitHub"
<p>This is a good first issue (GFI) recommendation dataset created from the GFI-Bot project (<a href="https://github.com/osslab-pku/gfi-bot">https://github.com/osslab-pku/gfi-bot</a>). For more information about the GFI recommendation problem and GFI-Bot, please check our publications:</p> <ul> <li>Wenxin Xiao, Hao He, Weiwei Xu, Xin Tan, Jinhao Dong, and Minghui Zhou. 2022. Recommending Good First Issues in GitHub OSS Projects. In Proceedings of the 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 21–29, 2022. ACM. <a href="https://hehao98.github.io/files/2022-recgfi.pdf">https://hehao98.github.io/files/2022-recgfi.pdf</a></li> <li>Hao He, Haonan Su, Wenxin Xiao, Runzhi He, and Minghui Zhou. 2022. GFI-Bot: Automated Good First Issue Recommendation on GitHub. In Proceedings of the 2022 ACM 30th Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE 2022, Singapore, November 14-16, 2022. ACM. <a href="https://hehao98.github.io/files/2022-gfibot.pdf">https://hehao98.github.io/files/2022-gfibot.pdf</a></li> </ul> <p>The dataset is a MongoDB dump and needs to be restored to a MongoDB instance before use. This can be done via the official <a href="https://www.mongodb.com/docs/database-tools/mongorestore/"><code>mongorestore</code></a> tool by running a command like this in the <code>dataset/</code> folder:</p> <pre><code class="language-bash">mongorestore --uri={{ your mongodb url }} --gzip </code></pre> <p>In the <code>gfibot.dataset</code> collection, each document describes the state of an issue at a certain time (either at the time of issue creation or at the time of issue resolution). The <code>resolver_commit_num</code> is the ground truth label (i.e., # of commits the issue resolver has made in the repository before issue resolution, excluding commits for resolving the issue itself; <code>resolver_commit_num = 0</code> means the resolver is someone completely new to the repository). The remaining fields can be used as features or further analyzed to derive new features.</p> <p>The <code>gfibot.resolved_issue</code> collection additionally provides information about which GitHub user resolved this issue and in what commit or pull request. This information can be used to study problems like, e.g., personalized good first issue recommendation or newcomer retention mechanisms.</p> <p>This dataset can be used to evaluate new GFI recommendation approaches. We hope it will be helpful in advancing GFI recommendation research and other future studies on open-source software onboarding.</p>
GitRanking: A Ranking of GitHub Topics for Software Classification using Active Sampling
<p>Replication package for our paper: GitRanking: A Ranking of GitHub Topics for Software Classification using Active Sampling</p>
StackPilot: Contrasting Code Snippets from Stack Overflow and GitHub Copilot
<p>Copy-paste programming via Stack Overflow and code generation via GitHub Copilot both define a query/prompt-based programming model. To enable systematic comparison of code copied from Stack Overflow and code generated by GitHub Copilot, we provide a dataset of 30,746 code snippets that Stack Overflow and GitHub Copilot produced in response to the same 2,636 queries/prompts.</p>
Data for EBE_Extended_Biomass_Estimation on GitHub
<p>This material regards the paper:</p> <p>Latella, M., Raimondo, F., Belcore, E., Salerno, L., & Camporeale, C. (2022). On the integration of LiDAR and field data for riparian biomass estimation, <em>Journal of Environmental Management</em>, 322, 116046. <a href="https://doi.org/10.1016/j.jenvman.2022.116046">https://doi.org/10.1016/j.jenvman.2022.116046</a>.</p> <p>and the GitHub code:</p> <p>EBE_Extended_Biomass_Estimation</p> <p>Please cite the related article if using the script or data.</p> <p> </p>
Data accompanying the GitHub repository bartonlab/paper-covariance-estimation
<p>This dataset contains data from simulations and analysis of Wright-Fisher evolution that accompany the GitHub repository bartonlab/paper-covariance-estimation. The GitHub repository contains code and data for reproducing results described in the manuscript "Estimating genetic linkage and selection from allele frequency trajectories." See the GitHub repository for details on the interpretation and analysis of this data.</p>
Prioritising GitHub Priority Labels - Data Set and Software
<p>This is the data set and software produced for the paper <em>Prioritising GitHub Priority Labels</em>, J. Caddy and C. Treude.</p> <p>The CSV file contains a manually categorised set of GitHub issue labels that are priority-related. They have been ranked and normalised into three values; "High", "Medium", and "Low" priorities. These labels have been gathered from the 5000 most-starred repositories on GitHub as of 2022-06-01.</p> <p>The Python script makes use of this data set as an example, and will retrieve the highest priority issues from all of the repositories contributed to by the author specified.</p> <p>Run the python script from the same directory as the CSV file, providing the username you wish to see the highest priority issues for as the first command line argument. Supply your GitHub Personal Access Token either at the prompt so it's not displayed, or as the second command line argument.</p>
A Dataset of bot and human contributors names in GitHub
<h4>A Dataset of Bot and Human Contributors' names in GitHub</h4> <p>This repository provides a dataset of 2,150 contributors (1,035 bots and 1,115 humans) that were active enough (made at least 5 events in GitHub) as of 3 May 2024. This dataset accompanies the paper titled <strong>A Bot Identification Model and Tool Based on GitHub Activity Sequences</strong> published at the <strong>Journal of Systems and Software (JSS), </strong>see<strong> <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.jss.2024.112287" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.jss.2024.112287</span></span></a></strong>. This research paper is co-authored by Natarajan Chidambaram, Alexandre Decan and Tom Mens (Software Engineering Lab, University of Mons, Belgium). This work is supported by Service Public de Wallonie Recherche under grant number 2010235 - ARIAC by DigitalWallonia4.AI, by the Fonds de la Recherche Scientifique – FNRS under grant numbers J.0147.24, T.0149.22, and F.4515.23.</p> <h4>Files description</h4> <p>bots.txt - contains the login name of bots, one per line</p> <p>humans.txt - contains the login name of humans, one per line.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.