Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
207
datasets available to search
ShareScore release 0.7.1
Dataset results
207 results for “github”
Automated Assessment of Mobile Programming Courses: Leveraging GitHub Classroom and Flutter for Enhanced Student Outcomes - Survey complete results
Open the record for dataset details and reuse information.
Identifying Unmaintained Projects in GitHub
<p>The dataset used in the paper "Identifying Unmaintained Projects in GitHub" accepted at 12th International Symposium on Empirical <em>Software Engineering</em> and Measurement (<em>ESEM),</em> 2018.</p>
Dataset and scripts for "Investigating the Cross-Repository Socially Connected Teams in Github"
<p>Dataset and scripts for "Investigating the Cross-Repository Socially Connected Teams in Github"</p> <p>README included in the files</p>
Avaliando as Interações entre Desenvolvedores e ChatGPT no GitHub: Uma análise de dados de Pull Requests
<p>Com o surgimento dos modelos de Linguagem de Grande Escala (LLMs) como o ChatGPT, introduziu um novo conjunto de ferramentas para apoiar desenvolvedores de software na resolução de tarefas de programação. No entanto, a compreensão das interações (ou seja, prompts) entre desenvolvedores e o ChatGPT que resultam em contribuições para o código permanece limitada. Para explorar essa limitação, foi realizada uma avaliação manual de 155 links válidos do ChatGPT, extraídos de 139 Pull Requests (PRs) mesclados à branch principal, revelando as interações entre desenvolvedores e revisores com o ChatGPT que levaram às integrações na branch principal. Os resultados produziram um catálogo de 14 tipos de solicitações feitas ao ChatGPT, categorizadas em quatro grupos principais. Foi identificado um número significativo de solicitações envolvendo revisão de código e a implementação de trechos de código com base em tarefas específicas. Os desenvolvedores também buscaram esclarecer dúvidas solicitando explicações técnicas ou refinamentos de texto para suas páginas web. Além disso, foi verificado que prompts envolvendo a geração de código geralmente exigiram mais interações para produzir a resposta desejada, em comparação com prompts solicitando revisão de código ou informações técnicas.</p>
Exploratory Data Analysis of SonarCloud and GitHub Data
Open the record for dataset details and reuse information.
Towards Filtering Out Deficient Pull Requests Collected through the GitHub API
<p>This is a replication package for an APSEC 2024 ERA paper.</p>
[DataSet] Avaliação Comparativa do GitHub Copilot e do Amazon CodeWhisperer na Geração Automatica de Código-Fonte
Open the record for dataset details and reuse information.
Nice to Meet You: The Role of Communicative Signaling in GitHub Profile Biographies
<p>This repository serves as the online appendix for the paper "Nice to Meet You: The Role of Communicative Signaling in GitHub Profile Biographies".</p>
Nice to Meet You: The Role of Communicative Signaling in GitHub Profile Biographies
<p>This repository serves as the online appendix for the paper "Nice to Meet You: The Role of Communicative Signaling in GitHub Profile Biographies".</p>
Investigando o Uso de Inteligência Artificial em Projetos Brasileiros do GitHub
<p>Dataset com os dados dos 654 repositórios brasileiros do GitHub utilizados no estudo.</p>
GitHub Pull Request Demonstration by Dhruvil Prajapati
<div> <p>In this demonstration, NLU student Dhruvil Prajapati walks us through pull requests for the TOPS SCHOOL GitHub repository. You can watch the video below or find a link in the 'Additional details' section.</p> </div>
Training data for the GitHub repository "buildingsFromSentinel"
<p>Training and testing data for machine learning models predicting the building height and footprint from satellite data in urban areas.</p> <p>Sentinel-1 and -2 data are retrieved from https://scihub.copernicus.eu/ and the GHS built-up grid (here GHSBuilt10) from https://ghsl.jrc.ec.europa.eu/download.php?ds=buS2. GHSBuilt10 is derived from Sentinel-2 global image composite for the reference year 2018 using Convolutional Neural Networks (GHS-S2Net).</p> <p>The dataset contains the following folders:</p> <ul> <li>footprint: PNG images over urban areas with either three or four features: <ul> <li>XXX_labels.png: true-colour images (TCI) retrieved from Sentinel-2 data</li> <li>XXX_labels4.png: TCIs with the band 8 (i.e., near-infrared = NIR) as the fourth dimension in the image.</li> </ul> </li> <li>height: data for different cities <ul> <li>building_height.tif: real building height (only for the training data)</li> <li>sentinel_cropped: satellite images for the same area. Contains Sentinel-1 and -2 data as well as the GHS-Built data with a 10-m resolution.</li> <li>README.txt: information of the origin of the building height data</li> </ul> </li> </ul>
Identifiers in source code extracted from 13,000,000 public GitHub repositories (October 2016)
<p>101 files, indexed LZO archives for Hadoop/Spark.<br> Each file is text lines with the format:</p> <p>(‘<GitHub repo name>', [(‘<name>', <count>),(‘<name>', <count>),(‘<name>', <count>)])</p> <p> </p>
Copernicus Climate Change Service data for the pypsa-entsoe Github repository
<p>Files needed for the https://github.com/matteodefelice/pypsa-entsoe repository.</p>
Replication Package of "Code Generative Techniques and TODO Comments: Friends or Foes? An Evaluation of GitHub Copilot"
<p>The data and scripts contributed by the efforts in "Code Generative Techniques and TODO Comments: Friends or Foes? An Evaluation of GitHub Copilot".</p>
Replication package for our TOSEM paper entitled "An Empirical Study on GitHub Pull Requests' Reactions"
<p>This package contains our dataset and the source code used to collect data from the the top 10,000 most starred GitHub repositories, and the selected six repositories (i.e., Cataclysm-DDA, Julia, Laravel, Node, RPCS3 and Rust), as well as the source code to analyze the data and generate all the figures in the paper. </p> <p>Please carefully read the README.md file for more details.</p>
Resulting Pseudonymized Classification Data for "Automatic Core-Developer Identification on GitHub: A Validation Study"
<p>Resulting pseudonymized classification data of the study "Automatic Core-Developer Identification on GitHub: A Validation Study". The corresponding input data, from which the output data have been derived, can be found here: https://zenodo.org/record/7775078</p>
Pseudonymized Raw Data for "Automatic Core-Developer Identification on GitHub: A Validation Study"
<p>Pseudonymized raw data (i.e., commit data and issue data for 25 GitHub projects) that has been used as input for the study published as "Automatic Core-Developer Identification on GitHub: A Validation Study".</p> <p>The pseudonymized raw data has been extracted via the tools <a href="https://github.com/se-sic/codeface/">Codeface</a>, <a href="https://github.com/se-sic/GitHubWrapper/">GitHubWrapper</a>, <a href="https://github.com/mehdigolzadeh/BoDeGHa">BoDeGHa</a>, and <a href="https://github.com/se-sic/codeface-extraction/">codeface-extraction</a> (and additional manual corrections after sanity checks).</p>
Datasets of issue-commit and issue-method links extracted from GitHub repositories
<p>Contains issue-commit and issue-method links extracted from GitHub repositories.</p> <p>Available on GitHub: https://github.com/pragma-once/utilizing-bert-for-traceability/releases</p>
Dataset and results for the study on issue prioritization in GitHub
<p>Dataset and results for the study on issue prioritization</p> <p> </p> <p>-feature: the extracted features for the selected 274 projects</p> <p> </p> <p>-training_data: the training data for 60 projects used to evaluate the prioritization methods</p> <p>--dataset1: data with multicollinearity features removed</p> <p>--dataset2: data with both multicollinearity features and features with weak or insignificant correlation with issue priority removed</p> <p> </p> <p>-ndcg: the complete results of NDCG@k (k ranging from 1 to 20)</p> <p>--result_1: results based on data with multicollinearity features removed</p> <p>--result_2: results based on data with both multicollinearity features and features with weak or insignificant correlation with issue priority removed</p> <p>--result_cross_project: results of cross projects</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.