Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

207

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

207 results for “github”

Learn how ShareScore rates datasets ↗
zenodo32/100

Replication package of the Paper "On the Relationships between the Initial Ecology Indicators of OSS Projects and Their Long-Term Popularity: An Exploratory Study on GitHub"

<p>The dataset is collected from GitHub API and GitHub GHTorrent dataset. A brief description of each folder is provided below:</p> <p><strong>1. "Dataset_and_Code" folder</strong></p> <p>Contains the final dataset and algorithms</p> <p><strong>2. "Test_parameters" folders&nbsp;</strong></p> <p>Includes datasets under different parameters and the corresponding reproduction code, which corresponds to the first experiment of RQ1</p> <p><strong>3. "Compare_baseline"folder</strong></p> <p>Includes the dataset used by our method, the dataset used by the baseline method, and the reproduction code, corresponding to the second experiment of RQ1</p> <p><strong>4. "PLS" folder</strong></p> <p>Includes the dataset used by PLS and the corresponding reproduction code, which corresponds to experiment of RQ2</p> <p><strong>5. " Indicator_Calculation " folder</strong><br>It contains the calculation methods for various metrics in the paper, as well as the corresponding key&nbsp;files.&nbsp;</p> <p><span><strong>6. " Appendix " folder</strong><br></span><span>It includes supplementary materials such as the methodology for metric calculations to address the reviewers' questions.</span>&nbsp;</p> <p><strong>Note</strong></p> <p>We have provided a corresponding README file in each folder to help others reproduce our results</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Data accompanying the GitHub repository bartonlab/paper-binary-trait-inference

<p>This dataset contains data from simulations and analysis of within-host HIV-1 evolution that accompany the GitHub repository&nbsp;bartonlab/binary_trait. The GitHub repository contains code and data for reproducing results described in the manuscript &quot;Inferring selection for HIV-1 escape from T cell responses using a binary trait model.&quot; See the GitHub repository for details on the interpretation and analysis of this data.</p>

opencc-by-4.0Jan 2024View details →
zenodo32/100

seltmann/taxonomy-darwin-core: A GitHub Approach to Publishing Darwin Core Formatted Occurrence Data for Taxonomic Studies

<p><strong>A GitHub Approach to Publishing Darwin Core Formatted Occurrence Data for Taxonomic Studies</strong></p> <p><strong>Description</strong><br> This repository contains a Darwin Core Archive template and instructions for revisionary taxonomists to use to publish their data as a Darwin Core Archive or Darwin Core Compliant CSV file. This repository contains the Darwin Core Archive for &quot;A taxonomic revision of <em>Gryonoides</em> Dodd, 1920 (Hymenoptera: Scelionidae), with a review of the hosts of Teleasinae.&quot; The archive was produced by the authors from data contained in mx (Yoder et al. 2006&ndash;present).</p> <p><strong>Summary</strong><br> Darwin core archives have emerged as the accepted data-sharing standard for occurrence data about organisms. These occurrences could be observations or specimens in natural history collections. The standard is applied by natural history collections worldwide to share their data between various repositories, including Global Biodiversity Information Faculty (GBIF) and Integrated Digitized Bio collections (iDigBio). This repository can be repurposed as an example template for publishing material examined as a Darwin Core Archive, accessible for data aggregators, journals, and conforming to community standards. This method can be used for any occurrence dataset, such as material examined, species monitoring observation records, ecological observations, or natural history collection data.</p> <p><strong>How to Create a Darwin Core Archive Using this Repository</strong><br> 1- Fork repository on GitHub or download its contents.</p> <p>2- Use the occurrences.csv file as a template. Delete the *Gryonoides* data and add your own. Do not change the column number or order. It is ok to leave in extra columns that you do not use. The only columns that need to be filled out are <em>occurrenceID</em>&nbsp;and <em>BasisOfRecord</em>. Specific definitions of the field names can be found in the Darwin Core Documentation under the <a href="https://dwc.tdwg.org/terms/#occurrence">Occurrence Core</a>.</p> <p>3- Do not touch the meta.xml. This file describes the columns in the occurrences.csv file.</p> <p>4- Edit the eml.xml files to include information about your institution and project.&nbsp;</p> <p>5- Zip the folder to create the archive. If you are using GitHub you can use the zip function for the repository using Code &nbsp;-&gt; Download ZIP.</p> <p>6- Validate the archive using the <a href="https://www.gbif.org/tools/data-validator">GBIF Data Validator tool</a>.</p> <p><strong>Citations</strong></p> <p>Yoder, M.J., Dole, K., Seltmann, K., and Deans, A. 2006-Present. Mx, a collaborative web based content management for biological systematists. http://mx.phenomix.org/index.php/Main_Page<br> &nbsp;</p>

openother-openNov 2021View details →
zenodo32/100

To Type or Not to Type? A Systematic Comparison of the Software Quality of JavaScript and TypeScript Applications on GitHub

<p>This repository contains the data and script (sampling, data collection, statistical analysis)&nbsp;for a repository mining study on GitHub to compare the software quality of JavaScript and TypeScript applications. Please refer to the README for more information.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

On Reporting Performance and Accuracy Bugs for Deep Learning Frameworks: An Exploratory Study from GitHub

<p>This repository aims to store the dataset of performance and accuracy bug reports, which belongs to&nbsp;<em>&quot;On Reporting Performance and Accuracy Bugs for Deep Learning Frameworks: An Exploratory Study from GitHub&quot;</em></p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

BotHunter: An Approach to Detect Software Bots in GitHub

<p>Contains the oracle and BIMAN&#39;s output.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Example Lminor dataset for Seq2Fun issue #3 (Github)

<p>This is the example dataset for the Seq2Fun issue #3 posted on github <a href="https://github.com/xia-lab/Seq2Fun/issues/3">here</a>.</p> <p>The compressed archive contains the following files:</p> <ul> <li>Directory (<em>raw_reads</em>) which contains subsampled (250k reads) mRNA-Seq fastq files (single-read 50bp) from full <em>Lemna minor</em> extracts. (For the files a respective multiQC html report from our RNASeq mapping pipeline is provided.)</li> <li>Sample table for batch processing (<em>lminor_sampleTable.tab</em>).</li> <li>Bash script (<em>seq2fun_analysis.sh</em>)<em> </em>to perform Seq2Fun analysis (runs smoothly without an error but is unable to map any reads)</li> <li>Output directory (<em>seq2fun</em>) with the files created with the <em>seq2fun_analysis.sh</em> script.</li> </ul> <p>I hope the provided files will help to solve the issue.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Dataset for paper "What Helps a New GitHub Project Achieve Sustained Activity?"

<p>These data files are data for&nbsp;paper &quot;What Helps a New GitHub Project Achieve Sustained Activity?&quot;.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Code Smells in Elixir: Results from a Mining Study on GitHub [DATASET]

<p>Dataset used in research</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Como os mantenedores usam GitHub Reactions? Um estudo exploratório

<p>Dataset contendo respostas obtidas no estudo &quot;Como os mantenedores usam GitHub Reactions? Um estudo explorat&oacute;rio&quot; aceito para publica&ccedil;&atilde;o no VEM 2022 (10th Workshop on Software Visualization, Evolution and Maintenance /&nbsp;CBSoft).</p> <p>Resumo: Plataformas de codifica&ccedil;&atilde;o social modernas t&ecirc;m fomentado a comunica&ccedil;&atilde;o e colabora&ccedil;&atilde;o no desenvolvimento de software por meio de funcionalidades inspiradas de redes sociais tradicionais. Estudos anteriores mostraram que reactions &eacute; uma funcionalidade cada vez mais utilizada, contudo pouco se sabe sobre seu impacto no desenvolvimento de software da perspectiva dos desenvolvedores. Neste trabalho, &eacute; apresentado um estudo com 17 mantenedores de pro- jetos populares na plataforma GitHub com intuito de coletar suas percep&ccedil;&otilde;es sobre tal funcionalidade. Os resultados mostram que a absoluta maioria dos participantes v&ecirc;em benef&iacute;cios nas rea&ccedil;&otilde;es (e.g., praticidade de feedaback e m&eacute;trica de aceita&ccedil;&atilde;o) e tr&ecirc;s em cada quatro consideraram as rea&ccedil;&otilde;es ao tomarem decis&otilde;es de projeto.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Dataset for "Exploring the Verifiability of Code Generated by GitHub Copilot"

<p>Collection of Python implementations and translations to Dafny with verification attempts.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Replication Package for Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects"

<p>This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models.&nbsp;</p> <p>We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars.&nbsp;The raw data of repositories are stored in `Replication Package/data/prodata.pkl`, and the contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m2_k1.pkl`. We sort items by the order in `Replication Package/data/randind.npy`, which can be used to reproduce the same results as in the paper.&nbsp;<br> `Replication Package/data/X_test_m2_k1.pkl` and `Replication Package/data/y_test_m2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table III and IV, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table II. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section III.B.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Replication Package for Paper "How Early Participation Determines Long-Term Sustained Activity in GitHub Projects"

<p>This replication package can be used for replicating results in the paper. It contains 1) a dataset of 290,255 repositories; and 2) Python scripts for training and interpreting models.&nbsp;</p> <p>We recommend manually setup the required environment in a commodity Linux machine with at least 1 CPU Core, 8GB Memory and 100GB empty storage space. We conduct development and execute all our experiments on a Ubuntu 20.04 server with two Intel Xeon Gold CPUs, 320GB memory, and 36TB RAID 5 Storage.</p> <p>We use GHTorrent to restore historical states of 290,255 repositories with more than 57 commits, 4 PRs, 1 issue, 1 fork and 2 stars.&nbsp;The raw data of repositories are stored in `Replication Package/data/prodata.pkl`, and the contribution of features resulting from LIME model is stored in `Replication Package/data/limeres_m2_k1.pkl`. We sort items by the order in `Replication Package/data/randind.npy`, which can be used to reproduce the same results as in the paper.&nbsp;<br> `Replication Package/data/X_test_m2_k1.pkl` and `Replication Package/data/y_test_m2_k1.pkl` store the test dataset for the LIME model. You can run `Replication Package/fitdata.py` to get the results in Table III and IV, run `Replication Package/draw_compare_variable.py` to get Figure 2 and run `Replication Package/allvari_statistics.py` to get Table II. In `Replication Package/Variable_comparison_with_different_parameter.pdf`, we show the LIME results under different parameters. In `Replication Package/sample_pros.csv`, we also provide the list of randomly selected repositories in Section III.B.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Github Repository for "Social interactions generate complex selection patterns in virtual worlds"

<p>Repository for&nbsp;<br><strong>"Social interactions generate complex selection patterns in virtual worlds"</strong><br><em>Francesca Santostefano, Maxime Fraser Franco, Pierre Olivier Montiglio</em><br>Journal of Evolutionary Biology 2024</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Pre-processed functional data for the github repo: gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course

<p>This repository was created for the purposes of a practical demonstration for the May 2024 SPM course at UCL.</p> <p>This upload contains subject-level anonymised functional data that have undergone the following pre-processing:</p> <ul> <li>first 5 volumes discarded</li> <li>spatial realingment</li> <li>unwarping</li> <li>normalisation to MNI space</li> <li>smoothing</li> <li>denoising with ICA-AROMA</li> </ul> <p>There are also pre-processed SPM GLM files.</p> <p>These data are to be used in conjuction with the code in the repositry found <a href="https://github.com/gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course" target="_blank" rel="noopener"><strong>here</strong></a>. They only contain a subset of the participants included in the full analysis and are for demonstrattion purposes only.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

UC1 DATA FOLDER GITHUB

<p>This folder contains the necessary data to execute the UC1 GitHub repository. It should be included inside the <a href="https://github.com/ICAERUS-EU/UC1_Crop_Monitoring/tree/main">UC1_Crop_Monitoring</a> folder. It includes the following data:&nbsp;</p> <ul> <li> <p>"features": saves some files with information about the row positions in the global vineyard orthomosaic, the plant positions, their health status and the parcels which divide the vineyards.</p> </li> <li> <p>"images": includes some orthomosaic images of the vineyard in different color formats as well as the DEM and NDVI orthomomsaics.&nbsp;</p> </li> <li> <p>"images_row":&nbsp; it contains some images taken with the drone of the plants from a row point of view.&nbsp;</p> </li> <li> <p>"images_saved": some images generated after executing wth some codes of the UC1 repository.&nbsp;</p> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset for the study on issue links in GitHub

<p>Dataset for the study on issue links in GitHub</p> <ol> <li>dataset</li> </ol> <ul> <li>issues, pulls and commits: raw data for each project</li> <li>links: extracted links and their types from each project</li> <li>samples: constructed sample of each project</li> </ul> <p>&nbsp; &nbsp; &nbsp; 2. models&nbsp;</p> <ul> <li>Binary-Classifier: Binary classifiers trained to predict the existence of links</li> <li>Multi-Classifier: Multiple classifier trained to predict link types</li> </ul>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset of the Paper "Architecture Decisions in Quantum Software Systems: An Empirical Study on Stack Exchange and GitHub"

<p>This dataset was collected from GitHub and Stack Exchange (including Stack Overflow, Quantum Computing Stack Exchange, and Computer Science Stack Exchange) to conduct an empirical study on architecture decisions in quantum software systems. We provide below a brief description of each file:</p><p><strong>1. Dataset (GitHub).xlsx</strong></p><p>contains selected quantum software projects from GitHub with project names, issue IDs, and issue URLs and the data extracted from the GitHub issues that are related to architecture decisions in quantum software development.</p><p><strong>2. Dataset (SO).xlsx</strong></p><p>contains the IDs and URLs of Stack Overflow (SO) labeled posts and the extracted data from the Stack Overflow posts that are related to architecture decisions in quantum software development.</p><p><strong>3. Dataset (QC).xlsx</strong></p><p>contains the IDs and URLs of Quantum Computing (QC) Stack Exchange labeled posts and the extracted data from the Quantum Computing Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>4. Dataset (CS).xlsx</strong></p><p>contains the IDs and URLs of Computer Science (CS) Stack Exchange labeled posts and the extracted data from the Computer Science Stack Exchange posts that are related to architecture decisions in quantum software development.</p><p><strong>5. Extracted Data (GitHub+SO+QC+CS).xlsx</strong></p><p>provides the final results of data extracted from the related GitHub issues, SO posts, QC posts, and CS posts.</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Evaluating Test Quality in GitHub Repositories: A Comparative Analysis of CI/CD Practices Using GitHub Actions

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset of the Paper "Exploring the Problems, their Causes and Solutions of AI Pair Programming: A Study on GitHub and Stack Overflow"

<p>The dataset collected from GitHub Discussions, GitHub Issues, and Stack Overflow is used to conduct an empirical study on the problems, causes, and solutions of using GitHub Copilot in practice. A brief description of each document in the dataset is provided below:</p> <p><br><strong>1. Dataset(GitHub_Discussions).xlsx</strong></p> <p>contains the Discussion IDs and URLs in the Copilot category of GitHub Discussions, and the data extracted from the related discussions along with analysis results.</p> <p><strong>2. Dataset(GitHub_Issues).xlsx</strong></p> <p>contains the Issue IDs and URLs of the labelled issues which are related to Copilot from GitHub Issues, and the data extracted from the related issues along with analysis results.</p> <p><strong>3. Dataset(SO_Posts).xlsx</strong></p> <p>contains the SO Post IDs and URLs of the labelled posts which are related to Copilot from Stack Overflow, and the data extracted from the related posts along with analysis results.</p> <p><strong>4. Extracted_Data.xlsx</strong></p> <p>contains the final results of the data extracted from GitHub Discussions, GitHub Issues, and SO posts.</p> <p><strong>5. pilot labelling folder</strong></p> <p>contains three .xlsx files (i.e., Pilot_Labelling(GitHub_Discussions).xlsx, Pilot_Labelling(GitHub_Issues), and Pilot_Labelling(SO)), with each file corresponding to one of the three data sources and containing the pilot data labelling results with the Cohen's kappa value.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record