Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
207
datasets available to search
ShareScore release 0.7.1
Dataset results
207 results for “github”
Replication package of the Paper "On the Relationships between the Initial Ecology Indicators of OSS Projects and Their Long-Term Popularity: An Exploratory Study on GitHub"
<p>The dataset is collected from GitHub API and GitHub GHTorrent dataset. A brief description of each folder is provided below:</p> <p><strong>1. "Dataset_and_Code" folder</strong></p> <p>Contains the final dataset and algorithms</p> <p><strong>2. "Test_parameters" folders </strong></p> <p>Includes datasets under different parameters and the corresponding reproduction code, which corresponds to the first experiment of RQ1</p> <p><strong>3. "Compare_baseline"folder</strong></p> <p>Includes the dataset used by our method, the dataset used by the baseline method, and the reproduction code, corresponding to the second experiment of RQ1</p> <p><strong>4. "PLS" folder</strong></p> <p>Includes the dataset used by PLS and the corresponding reproduction code, which corresponds to experiment of RQ2</p> <p><strong>5. " Indicator_Calculation " folder</strong><br>It contains the calculation methods for various metrics in the paper, as well as the corresponding key files. </p> <p><span><strong>6. " Appendix " folder</strong><br></span><span>It includes supplementary materials such as the methodology for metric calculations to address the reviewers' questions.</span> </p> <p><strong>Note</strong></p> <p>We have provided a corresponding README file in each folder to help others reproduce our results</p>
Data accompanying the GitHub repository bartonlab/paper-binary-trait-inference
<p>This dataset contains data from simulations and analysis of within-host HIV-1 evolution that accompany the GitHub repository bartonlab/binary_trait. The GitHub repository contains code and data for reproducing results described in the manuscript "Inferring selection for HIV-1 escape from T cell responses using a binary trait model." See the GitHub repository for details on the interpretation and analysis of this data.</p>
seltmann/taxonomy-darwin-core: A GitHub Approach to Publishing Darwin Core Formatted Occurrence Data for Taxonomic Studies
<p><strong>A GitHub Approach to Publishing Darwin Core Formatted Occurrence Data for Taxonomic Studies</strong></p> <p><strong>Description</strong><br> This repository contains a Darwin Core Archive template and instructions for revisionary taxonomists to use to publish their data as a Darwin Core Archive or Darwin Core Compliant CSV file. This repository contains the Darwin Core Archive for "A taxonomic revision of <em>Gryonoides</em> Dodd, 1920 (Hymenoptera: Scelionidae), with a review of the hosts of Teleasinae." The archive was produced by the authors from data contained in mx (Yoder et al. 2006–present).</p> <p><strong>Summary</strong><br> Darwin core archives have emerged as the accepted data-sharing standard for occurrence data about organisms. These occurrences could be observations or specimens in natural history collections. The standard is applied by natural history collections worldwide to share their data between various repositories, including Global Biodiversity Information Faculty (GBIF) and Integrated Digitized Bio collections (iDigBio). This repository can be repurposed as an example template for publishing material examined as a Darwin Core Archive, accessible for data aggregators, journals, and conforming to community standards. This method can be used for any occurrence dataset, such as material examined, species monitoring observation records, ecological observations, or natural history collection data.</p> <p><strong>How to Create a Darwin Core Archive Using this Repository</strong><br> 1- Fork repository on GitHub or download its contents.</p> <p>2- Use the occurrences.csv file as a template. Delete the *Gryonoides* data and add your own. Do not change the column number or order. It is ok to leave in extra columns that you do not use. The only columns that need to be filled out are <em>occurrenceID</em> and <em>BasisOfRecord</em>. Specific definitions of the field names can be found in the Darwin Core Documentation under the <a href="https://dwc.tdwg.org/terms/#occurrence">Occurrence Core</a>.</p> <p>3- Do not touch the meta.xml. This file describes the columns in the occurrences.csv file.</p> <p>4- Edit the eml.xml files to include information about your institution and project. </p> <p>5- Zip the folder to create the archive. If you are using GitHub you can use the zip function for the repository using Code -> Download ZIP.</p> <p>6- Validate the archive using the <a href="https://www.gbif.org/tools/data-validator">GBIF Data Validator tool</a>.</p> <p><strong>Citations</strong></p> <p>Yoder, M.J., Dole, K., Seltmann, K., and Deans, A. 2006-Present. Mx, a collaborative web based content management for biological systematists. http://mx.phenomix.org/index.php/Main_Page<br> </p>
To Type or Not to Type? A Systematic Comparison of the Software Quality of JavaScript and TypeScript Applications on GitHub
<p>This repository contains the data and script (sampling, data collection, statistical analysis) for a repository mining study on GitHub to compare the software quality of JavaScript and TypeScript applications. Please refer to the README for more information.</p>
On Reporting Performance and Accuracy Bugs for Deep Learning Frameworks: An Exploratory Study from GitHub
<p>This repository aims to store the dataset of performance and accuracy bug reports, which belongs to <em>"On Reporting Performance and Accuracy Bugs for Deep Learning Frameworks: An Exploratory Study from GitHub"</em></p>
BotHunter: An Approach to Detect Software Bots in GitHub
<p>Contains the oracle and BIMAN's output.</p>
Example Lminor dataset for Seq2Fun issue #3 (Github)
<p>This is the example dataset for the Seq2Fun issue #3 posted on github <a href="https://github.com/xia-lab/Seq2Fun/issues/3">here</a>.</p> <p>The compressed archive contains the following files:</p> <ul> <li>Directory (<em>raw_reads</em>) which contains subsampled (250k reads) mRNA-Seq fastq files (single-read 50bp) from full <em>Lemna minor</em> extracts. (For the files a respective multiQC html report from our RNASeq mapping pipeline is provided.)</li> <li>Sample table for batch processing (<em>lminor_sampleTable.tab</em>).</li> <li>Bash script (<em>seq2fun_analysis.sh</em>)<em> </em>to perform Seq2Fun analysis (runs smoothly without an error but is unable to map any reads)</li> <li>Output directory (<em>seq2fun</em>) with the files created with the <em>seq2fun_analysis.sh</em> script.</li> </ul> <p>I hope the provided files will help to solve the issue.</p>
Dataset for paper "What Helps a New GitHub Project Achieve Sustained Activity?"
<p>These data files are data for paper "What Helps a New GitHub Project Achieve Sustained Activity?".</p>
Code Smells in Elixir: Results from a Mining Study on GitHub [DATASET]
<p>Dataset used in research</p>
Como os mantenedores usam GitHub Reactions? Um estudo exploratório
<p>Dataset contendo respostas obtidas no estudo "Como os mantenedores usam GitHub Reactions? Um estudo exploratório" aceito para publicação no VEM 2022 (10th Workshop on Software Visualization, Evolution and Maintenance / CBSoft).</p> <p>Resumo: Plataformas de codificação social modernas têm fomentado a comunicação e colaboração no desenvolvimento de software por meio de funcionalidades inspiradas de redes sociais tradicionais. Estudos anteriores mostraram que reactions é uma funcionalidade cada vez mais utilizada, contudo pouco se sabe sobre seu impacto no desenvolvimento de software da perspectiva dos desenvolvedores. Neste trabalho, é apresentado um estudo com 17 mantenedores de pro- jetos populares na plataforma GitHub com intuito de coletar suas percepções sobre tal funcionalidade. Os resultados mostram que a absoluta maioria dos participantes vêem benefícios nas reações (e.g., praticidade de feedaback e métrica de aceitação) e três em cada quatro consideraram as reações ao tomarem decisões de projeto.</p>
Dataset for "Exploring the Verifiability of Code Generated by GitHub Copilot"
<p>Collection of Python implementations and translations to Dafny with verification attempts.</p>
Replication Package for Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects"
<p>This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models. </p> <p>We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars. The raw data of repositories are stored in `Replication Package/data/prodata.pkl`, and the contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m2_k1.pkl`. We sort items by the order in `Replication Package/data/randind.npy`, which can be used to reproduce the same results as in the paper. <br> `Replication Package/data/X_test_m2_k1.pkl` and `Replication Package/data/y_test_m2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table III and IV, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table II. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section III.B.</p>
Replication Package for Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects"
<p>This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models. </p> <p>We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars. The raw data of repositories are stored in `Replication Package/data/prodata.pkl`, and the contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m2_k1.pkl`. We sort items by the order in `Replication Package/data/randind.npy`, which can be used to reproduce the same results as in the paper. <br> `Replication Package/data/X_test_m2_k1.pkl` and `Replication Package/data/y_test_m2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table III and IV, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table II. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section III.B.</p>
Github Repository for "Social interactions generate complex selection patterns in virtual worlds"
<p>Repository for <br><strong>"Social interactions generate complex selection patterns in virtual worlds"</strong><br><em>Francesca Santostefano, Maxime Fraser Franco, Pierre Olivier Montiglio</em><br>Journal of Evolutionary Biology 2024</p>
Pre-processed functional data for the github repo: gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course
<p>This repository was created for the purposes of a practical demonstration for the May 2024 SPM course at UCL.</p> <p>This upload contains subject-level anonymised functional data that have undergone the following pre-processing:</p> <ul> <li>first 5 volumes discarded</li> <li>spatial realingment</li> <li>unwarping</li> <li>normalisation to MNI space</li> <li>smoothing</li> <li>denoising with ICA-AROMA</li> </ul> <p>There are also pre-processed SPM GLM files.</p> <p>These data are to be used in conjuction with the code in the repositry found <a href="https://github.com/gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course" target="_blank" rel="noopener"><strong>here</strong></a>. They only contain a subset of the participants included in the full analysis and are for demonstrattion purposes only.</p>
UC1 DATA FOLDER GITHUB
<p>This folder contains the necessary data to execute the UC1 GitHub repository. It should be included inside the <a href="https://github.com/ICAERUS-EU/UC1_Crop_Monitoring/tree/main">UC1_Crop_Monitoring</a> folder. It includes the following data: </p> <ul> <li> <p>"features": saves some files with information about the row positions in the global vineyard orthomosaic, the plant positions, their health status and the parcels which divide the vineyards.</p> </li> <li> <p>"images": includes some orthomosaic images of the vineyard in different color formats as well as the DEM and NDVI orthomomsaics. </p> </li> <li> <p>"images_row": it contains some images taken with the drone of the plants from a row point of view. </p> </li> <li> <p>"images_saved": some images generated after executing wth some codes of the UC1 repository. </p> </li> </ul> <p> </p>
Dataset for the study on issue links in GitHub
<p>Dataset for the study on issue links in GitHub</p> <ol> <li>dataset</li> </ol> <ul> <li>issues, pulls and commits: raw data for each project</li> <li>links: extracted links and their types from each project</li> <li>samples: constructed sample of each project</li> </ul> <p> 2. models </p> <ul> <li>Binary-Classifier: Binary classifiers trained to predict the existence of links</li> <li>Multi-Classifier: Multiple classifier trained to predict link types</li> </ul>
Dataset of the Paper "Architecture Decisions in Quantum Software Systems: An Empirical Study on Stack Exchange and GitHub"
<p>This dataset was collected from GitHub and Stack Exchange (including Stack Overflow, Quantum Computing Stack Exchange, and Computer Science Stack Exchange) to conduct an empirical study on architecture decisions in quantum software systems. We provide below a brief description of each file:</p><p><strong>1. Dataset (GitHub).xlsx</strong></p><p>contains selected quantum software projects from GitHub with project names, issue IDs, and issue URLs and the data extracted from the GitHub issues that are related to architecture decisions in quantum software development.</p><p><strong>2. Dataset (SO).xlsx</strong></p><p>contains the IDs and URLs of Stack Overflow (SO) labeled posts and the extracted data from the Stack Overflow posts that are related to architecture decisions in quantum software development.</p><p><strong>3. Dataset (QC).xlsx</strong></p><p>contains the IDs and URLs of Quantum Computing (QC) Stack Exchange labeled posts and the extracted data from the Quantum Computing Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>4. Dataset (CS).xlsx</strong></p><p>contains the IDs and URLs of Computer Science (CS) Stack Exchange labeled posts and the extracted data from the Computer Science Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>5. Extracted Data (GitHub+SO+QC+CS).xlsx</strong></p><p>provides the final results of data extracted from the related GitHub issues, SO posts, QC posts, and CS posts.</p>
Evaluating Test Quality in GitHub Repositories: A Comparative Analysis of CI/CD Practices Using GitHub Actions
Open the record for dataset details and reuse information.
Dataset of the Paper "Exploring the Problems, their Causes and Solutions of AI Pair Programming: A Study on GitHub and Stack Overflow"
<p>The dataset collected from GitHub Discussions, GitHub Issues, and Stack Overflow is used to conduct an empirical study on the problems, causes, and solutions of using GitHub Copilot in practice. A brief description of each document in the dataset is provided below:</p> <p><br><strong>1. Dataset(GitHub_Discussions).xlsx</strong></p> <p>contains the Discussion IDs and URLs in the Copilot category of GitHub Discussions, and the data extracted from the related discussions along with analysis results.</p> <p><strong>2. Dataset(GitHub_Issues).xlsx</strong></p> <p>contains the Issue IDs and URLs of the labelled issues which are related to Copilot from GitHub Issues, and the data extracted from the related issues along with analysis results.</p> <p><strong>3. Dataset(SO_Posts).xlsx</strong></p> <p>contains the SO Post IDs and URLs of the labelled posts which are related to Copilot from Stack Overflow, and the data extracted from the related posts along with analysis results.</p> <p><strong>4. Extracted_Data.xlsx</strong></p> <p>contains the final results of the data extracted from GitHub Discussions, GitHub Issues, and SO posts.</p> <p><strong>5. pilot labelling folder</strong></p> <p>contains three .xlsx files (i.e., Pilot_Labelling(GitHub_Discussions).xlsx, Pilot_Labelling(GitHub_Issues), and Pilot_Labelling(SO)), with each file corresponding to one of the three data sources and containing the pilot data labelling results with the Cohen's kappa value.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.