Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

846

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

846 results for “commit”

Learn how ShareScore rates datasets ↗
zenodo36/100

Metrics by commit for Android

<p>This dataset is a set of csv files. Each file contains the metrics by commit of all android code samples analyzed in the paper.&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

Bug-fix commit script and dataset

<p>This contains the scripts and dataset used in &quot;automatic identification of bug-fixing commits in Python open-source projects&quot; for reproduction purposes.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
dryad36/100

Consideration of food and nutrition in blue economy voluntary commitments

<p>Increasing the production of food from the ocean is seen as a pathway towards more sustainable and healthier human diets. Yet this potential is being overshadowed by competing uses of ocean resources in an accelerating 'blue economy'. The current emphasis on production growth, rather than equitable distribution of benefits, has created three unexamined or flawed assumptions that: growth in the blue economy will lead to growth in blue food production; increased production will inevitably lead to improved food and nutrition security; and mariculture production will replace marine capture fisheries. In this Perspective, we argue that if research and policies are pursued without addressing these 'blind spots', 'blue food' contributions to reducing hunger and malnutrition, and to meeting the Sustainable Development Goals, will be limited. Taking a broader food system approach, beyond production to also consider food access, affordability and consumption, will refocus the 'blue food' agenda on making production and consumption more equitable and sustainable, while increasing access for those who need it most.</p>

opencc-zeroJan 2021View details →
zenodo36/100

452,000,000 public Git commits on GitHub (October 2016)

<p>What's inside</p> <p>part-000xx.lzo - LZO archives with the data (refer to "Format").</p> <p>part-000xx.lzo.index - LZO index files so that the archives are splittable in Hadoop.</p> <p>stats.csv.gz - GZIP-ed CSV file with some repository statistics related to the commits.</p> <p>Format</p> <p>part-000xx - text, one line per repository, every line is JSON with the following scheme:</p> <p>{ "r": "repository name", "c": [{ "h": "git hash", "a": "author's email hash", "t": "date and time commit was created", "m": "commit message" }, ...] }</p> <p>Date and time format is <em>mostly</em> Go language's time.Time.String(), I recommend to use dateutil.parse() to parse it with Python.</p> <p>Commit message contains explicit \r and \n symbols in order to be a single line.</p> <p>stats.csv has 4 columns: repository name, number of commits, number of contributors, average length of the commit messages.</p>

opencc-by-nc-4.0Feb 2017View details →
zenodo36/100

Data and code for "The economic commitment of climate change"

<p>This repository contains data and code necessary for replication of the publication:</p> <p><em>The economic commitment of climate change</em>.<br>M. Kotz, A. Levermann, L. Wenz. Nature. 2024.</p> <p>including updates made to address critiques brought forward by the Matters Arising process.</p> <p>See the README document for detailed instructions in its use.</p> <p>Please feel free to contact maxkotz@pik-potsdam.de in case of any questions.</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Involvement vs Commitment L-shaped matrix diagram

<p>Involvement vs Commitment L-shaped matrix diagram of volunteering work in non-profit organizations.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Negative Complement of a Set of Vulnerability-Fixing Commits: Supplementary Material

<div> <div><span>EASE 2024 - Industry track</span></div> <br> <div><span>This archive contains accompanying materials for the paper: </span></div> <br> <div><span>Roc&iacute;o Cabrera Lozoya, Antonino Sabetta, Tommaso Aiello. "Negative Complement of a Set of Vulnerability-Fixing Commits: Method and Dataset" </span></div> <br> <div><span>submitted to the Industry track at EASE 2024 - https://conf.researchr.org/track/ease-2024/ease-2024-industry</span></div> <br> <div><span>It contains the following folders and files:</span></div> <br> <div><span>-</span><span> data</span></div> <div><span> </span><span>-</span><span> commit_pairs_final.csv : Dataset of 534 commit pairs corresponding to a positiive (security-relevant) commit and a negative sample obtained by the approach described in the paper.</span></div> <div><span> </span><span>-</span><span> single_positive.csv : Contains a single positive instance (taken from the MSR2019 dataset) which can be used to test the generate_negative_complement.py script.</span></div> <div><span>-</span><span> scripts</span></div> <div><span> </span><span>-</span><span> generate_negative_complement.py : Takes in a single .java file and obfuscates its developer-defined identifiers.</span></div> <div><span> </span><span>-</span><span> obfuscate_java_file.py : Generates a negative complement for a security-relevant dataset. The ouput is written to couple_dataset.csv.</span></div> <div><span> </span><span>-</span><span> requirements.txt : Requirements needed in the virtual environment to successfully run the previous scripts.</span></div> </div>

opencc-by-4.0Dec 2023View details →
zenodo36/100

A dataset of Linux Kernel commits

<p>Dataset with metadata about more than 1,200,000 changes (commits) of&nbsp;the Linux kernel, corresponding to a period since 2005 to 2023, which<br>can be easily ingested in data analytics systems.<br><br>It also includes a list of more than 90,000 pairs of bug fixing changes&nbsp;and their corresponding bug introducing changes, labeled by developers&nbsp;of the Linux Kernel.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

A Vulnerability Introducing Commit Dataset for Java: an Improved SZZ Based Approach

<p>This archive file contains the dataset generated from project-KB by applying our two-phase improved SZZ based algorithm for extracting vulnerability introducing commits. The package contains also the two tools (FilterBugIntroder, BugIntroducer) that automates the process.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Corporate Governance, Corporate Social Responsibility and Corporate Sustainability: The Moderating Role of Top Management Commitment

<table align="center"> <tbody> <tr> <td> <p><strong>Design/Methodology/Approach:</strong> We used the non-probability sampling technique by employing convenience sampling for data collection. By employing a survey questionnaire, data were collected from 397 employees of SMEs in Ghana. The IBM Statistical Package for Social Science (SPSS) version 25.0 and IBM&#39;s Analysis of Moments of Structures (AMOS) version 24 software packages were employed as analytical tools in this investigation.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Dataset - COMMIT: Consideration of metabolite leakage and community composition improves microbial community reconstructions

<p>Generated draft and consensus reconstructions for the manuscript &quot;COMMIT: Consideration of metabolite leakage and community composition improves microbial community reconstructions&quot; (Wendering &amp; Nikoloski).</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

What Makes a Good Commit Message?

<p><strong>What-Makes-a-Good-Commit-Message?</strong></p> <p>This repository contains the main data and scripts used in &quot;What Makes a Good Commit Message&quot;</p> <p><strong>Notes:</strong></p> <p><em>The code will be maintained on <a href="https://github.com/WhatMakesAGoodCM/What-Makes-a-Good-Commit-Message.git">GitHub</a>. If you have any questions, please feel free to mention issue.</em></p> <p><strong>Dataset:</strong></p> <p>The folder dataset contains the following files.</p> <ul> <li> <p>literature survery.xlsx</p> <ul> <li>It contains the data of 46 relevant literature reviewed in this study (Section 3.2).</li> </ul> </li> <li> <p>Questionnaire.pdf</p> <ul> <li>It is the questionnaire sent to experienced contributors.</li> <li>It contains three questions.</li> <li>It also contains an example of the initialized questionnaire.</li> </ul> </li> <li>Frequency.pdf <ul> <li>It describes the number and proportion of occurrences of a category/subcategory.</li> </ul> </li> <li> <p>posts list.xlsx</p> <ul> <li>It contains all posts we studied in Sec. 3.2.</li> </ul> </li> <li> <p>sampled messages.csv</p> <ul> <li>It contains meta-information of 1649 labeled commit messages.</li> <li>label = 0 means a commit message contains &quot;Why and What&quot;.</li> <li>label = 1 means a commit message contains &quot;Neither Why nor What&quot;.</li> <li>label = 2 means a commit message contains &quot;No What&quot;.</li> <li>label = 3 means a commit message contains &quot;No Why&quot;.</li> <li>if_mulit_commit = 1 means a commit is non-atomic.</li> <li>new_message1 means a message after preprocessing.</li> </ul> </li> <li> <p>maintenance type and expression way.xlsx</p> <ul> <li>It contains the results of our RQ2: the expression ways of Why and What, as well as links to maintenance types.</li> </ul> </li> </ul> <p><strong>CommitMessage (Scripts):</strong></p> <p>The folder contains the following scripts files.</p> <ul> <li> <p>Preprocessor</p> <ul> <li>It contains the preprocessing of commit messages, including the replacement of token in the message, etc.</li> </ul> </li> <li> <p>ModelTraining</p> <ul> <li>It contains the code for our model training, that is, the implementation of different classification techniques.</li> </ul> </li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"

<p>This document provides the underlying dataset for the bibliometric component for the 2022 study &quot;From intent to impact: Investigating the effects of open sharing commitments&quot; by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study&#39;s technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/&nbsp;</p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p>&nbsp;</p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process&nbsp;and clean the data&nbsp;for its eventual re-use in secondary analysis or&nbsp;text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name&nbsp;</td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server&#39;s unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server&#39;s unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of &quot;10.2139/ssrn.&quot; + &#39;ssrn_id&#39;</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Transmission and Distribution Substation Energy Management Considering Large-Scale Energy Storage, Demand Side Management and Security-Constrained Unit Commitment

<p>Data used in &quot;Transmission and Distribution Substation Energy Management Considering Large-Scale Energy Storage, Demand Side Management and Security-Constrained Unit Commitment&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

What really changes when developers intend to improve their source code: A commit-level study of static metric value and static analysis warning changes

<p>This is the dataset for the publication &quot;What really changes when developers intend to improve their source code: A commit-level study of static metric value and static analysis warning changes&quot;.</p> <p>It contains a random sample of 2533 commits from 54 Java Apache open source projects classified by two researchers into perfective, corrective and other changes (manual_labels.csv).&nbsp; Moreover, we include static source code metrics and static analysis warnings for the 2533 changes in al_changes_gt.csv.gz.</p> <p>In addition, we include the full dataset of 125482 commits in all_changes_sebert.csv.gz with all metrics and automatic labels for every commit that was not manually labeled. The automatic labels were provided by a fine-tuned transformer model (BERT) pre-trained exclusively on software engineering data.</p> <p>We also provide the fine tuned version of the pre-trained model in seBERT_fine_tuned_commit_intent.tar.gz as well as a Snapshot of the SmartSHARK MongoDB database used in gathering the raw data in smartshark_emse.agz.</p> <p>The model can be tested live on the <a href="https://user.informatik.uni-goettingen.de/~trautsch2/emse_2021/commit_intent.html">website</a> accompanying the publication.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Raw Dataset for Ocean Commitments Under the Paris Agreement

<p>This dataset is uploaded to complement the manuscript Gallo et al. Ocean Commitments Under the Paris Agreement. The dataset includes calculations of the Marine Focus Factor (MFF) for each Nationally Determined Contribution (NDC) analyzed, as well as information about which countries including which marine categories in their NDCs. Finally, the dataset provides raw data used for the multiple linear regression analysis reported in the manuscript, which tests a series of explanatory variables for their ability to explain variance in the MFF. Each raw datasheet is followed by a spreadsheet that provides additional details on the datasets. We hope that the availability of this raw data will allow other scholars interested in this analysis to build on our findings. </p>

opencc-by-4.0Aug 2017View details →
zenodo36/100

Curated dataset of bug fix commits from "An Empirical Study on Real Bug Fixes"

<p>To cite it:</p> <p><code>@misc{bfdataset,<br> &nbsp; author = {Martin Monperrus},<br> &nbsp; title = {Curated dataset of bug fix commits from &quot;An Empirical Study on Real Bug Fixes&quot;},<br> &nbsp; year = 2017,<br> &nbsp; doi = {10.5281/zenodo.1004734},<br> &nbsp; url = {https://doi.org/10.5281/zenodo.1004734}<br> }</code></p> <pre> &nbsp;</pre> <p>&nbsp;</p>

openother-openOct 2017View details →
zenodo36/100

Commits modifications from 61 projects hosted in Github's repositories

<p>This dataset contains data from 61 opensource projects that can be found in github repositories.The data contains informations about:</p> <ul> <li>committer name</li> <li>committer email</li> <li>modification type - changed, removed or created</li> <li>modified files</li> <li>commit comments</li> <li>commit date</li> </ul> <p>&nbsp;</p> <p>This dataset references 299 major version software, developed in five different programming languages. These languages are:</p> <ul> <li>C++</li> <li>Java</li> <li>Javascript</li> <li>Python</li> <li>Ruby</li> </ul> <p>The dataset was used in my&nbsp;undergraduate thesis</p>

opencc-by-nc-4.0Nov 2018View details →
zenodo36/100

Code diffs and commit messages from top1000-2000 Java projects in GitHub

<p>This dataset contains pairs &lt;code changes, commit message&gt; from top 1000-2000 Java projects in GitHub via&nbsp;<a href="https://developer.github.com/v3/">GitHub Developer API</a>.</p> <p>The structure of&nbsp;files is as follows:</p> <ul> <li>&lt;repository-id&gt;.&lt;repository-name&gt;/&nbsp; &nbsp; &nbsp; &nbsp;# Each directory contains commits of one project. <ul> <li>&lt;commit-id&gt;.&lt;commit-sha&gt;/&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Each directory contains one commit information, and<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# the commit-sha is the primary key of this commit in GitHub.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# The children in this directory&nbsp;have 3 types: the file<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;#&nbsp;&quot;commit_msg&quot;, multiple directories &quot;&lt;filename&gt;&quot;, and the file<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;#&nbsp;&quot;error&quot;.&nbsp; <ul> <li>commit_msg&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# This file contains one line, representing the commit<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;#&nbsp;message.&nbsp;</li> <li>&lt;filename&gt;/&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # The&nbsp;name of this directory is the changed file name. <ul> <li>patch&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# The code diffs of this &lt;filename&gt;, describing<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;which lines are added and which lines are<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;removed with some same lines context.</li> </ul> </li> <li>&lt;filename&gt;/ <ul> <li>patch</li> </ul> </li> <li>error&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# This file contains&nbsp;the error message when crawling<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;from GitHub. If this file exists, the directories<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;&lt;filename&gt; will not exist.&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</li> </ul> </li> </ul> </li> </ul>

opencc-by-4.0Jan 2019View details →
zenodo36/100

359,569 commits with source code density; 1149 commits of which have software maintenance activity labels (adaptive, corrective, perfective)

<p>This dataset comes as SQL-importable file and is compatible with the widely available MariaDB- and MySQL-databases.</p> <p>It is based on (and incorporates/extends) the dataset &quot;<em>1151 commits with software maintenance activity labels (corrective,perfective,adaptive)</em>&quot; by Levin and Yehudai (<a href="https://doi.org/10.5281/zenodo.835534">https://doi.org/10.5281/zenodo.835534</a>).</p> <p>The extensions to this dataset were obtained using&nbsp;<em>Git-Tools</em>, a tool that is included in the&nbsp;<strong>Git-Density</strong>&nbsp;(<a href="https://doi.org/10.5281/zenodo.2565238">https://doi.org/10.5281/zenodo.2565238</a>) suite. For each of the projects in the original dataset, Git-Tools was run in&nbsp;<em>extended</em>&nbsp;mode.</p> <p>The dataset contains these tables:</p> <ul> <li><strong>x1151</strong>: The original dataset from Levin and Yehudai. <ul> <li>despite its name, this dataset has only 1,149 commits, as two commits were duplicates in the original dataset.</li> <li>This dataset spanned 11 projects, each of which had between 99 and 114 commits</li> <li>This dataset has&nbsp;<strong>71</strong>&nbsp;features and spans the projects&nbsp;<em>RxJava, hbase, elasticsearch, intellij-community, hadoop, drools, Kotlin, restlet-framework-java, orientdb, camel</em>&nbsp;and&nbsp;<em>spring-framework</em>.</li> </ul> </li> <li><strong>gtools_ex</strong>&nbsp;(short for <em>Git-Tools, extended</em>) <ul> <li>Contains&nbsp;<strong>359,569</strong>&nbsp;commits, analyzed using Git-Tools in extended mode</li> <li>It spans all commits and projects from the x1151 dataset as well.</li> <li>All 11 projects were analyzed, from the initial commit until the end of January 2019. For the projects&nbsp;<em>Intellij</em>&nbsp;and&nbsp;<em>Kotlin</em>, the first 35,000 resp. 30,000 commits were analyzed.</li> <li>This dataset introduces&nbsp;<strong>35 new</strong>&nbsp;features (see list below), 22 of which are <em><strong>size</strong></em>- or <em><strong>density</strong></em>-related.</li> </ul> </li> </ul> <p>The dataset contains these views:</p> <ul> <li><strong>geX_L</strong>&nbsp;(short for Git-<em>tools, extended, with labels</em>) <ul> <li>Joins the commits&#39; labels from&nbsp;<em>x1151</em>&nbsp;with the extended attributes from&nbsp;<em>gtools_ex</em>, using the commits&#39; hashes.</li> </ul> </li> <li><strong>jeX_L</strong>&nbsp;(short for&nbsp;<em>joined, extended, with labels</em>) <ul> <li>Joins the datasets&nbsp;<em>x1151</em>&nbsp;and&nbsp;<em>gtools_ex</em>&nbsp;entirely, based on the commits&#39; hashes.</li> </ul> </li> </ul> <p>&nbsp;</p> <p>Features of the&nbsp;<strong>gtools_ex</strong>&nbsp;dataset:</p> <ul> <li><strong>SHA1</strong></li> <li><strong>RepoPathOrUrl</strong></li> <li><strong>AuthorName</strong></li> <li><strong>CommitterName</strong></li> <li><strong>AuthorTime </strong>(UTC)</li> <li><strong>CommitterTime </strong>(UTC)</li> <li><strong>MinutesSincePreviousCommit</strong>: Double, describing the amount of minutes that passed since the previous commit. Previous refers to the <strong>parent</strong> commit, not the previous in time.</li> <li><strong>Message</strong>: The commit&#39;s message/comment</li> <li><strong>AuthorEmail</strong></li> <li><strong>CommitterEmail</strong></li> <li><strong>AuthorNominalLabel</strong>: All authors of a repository are analyzed and merged by Git-Density using some heuristic, even if they do not always use the same email address or name. This label is a unique string that helps identifying the same author across commits, even if the author did not always use the exact same identity.</li> <li><strong>CommitterNominalLabel</strong>: The same as&nbsp;<em>AuthorNominalLabel</em>, but for the committer this time.</li> <li><strong>IsInitialCommit</strong>: A boolean indicating, whether a commit is preceded by a parent or not.</li> <li><strong>IsMergeCommit</strong>: A boolean indicating whether a commit has more than one parent.</li> <li><strong>NumberOfParentCommits</strong></li> <li><strong>ParentCommitSHA1s</strong>: A comma-concatenated string of the parents&#39; SHA1 IDs</li> <li><strong>NumberOfFilesAdded</strong></li> <li><strong>NumberOfFilesAddedNet</strong>: Like the previous property, but if the net-size of all changes of an added file is zero (i.e. when adding a file that is empty/whitespace or does not contain code), then this property does not count the file.</li> <li><strong>NumberOfLinesAddedByAddedFiles</strong></li> <li><strong>NumberOfLinesAddedByAddedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>NumberOfFilesDeleted</strong></li> <li><strong>NumberOfFilesDeletedNet</strong>: Like the previous property, but considers only files that had net-changes</li> <li><strong>NumberOfLinesDeletedByDeletedFiles</strong></li> <li><strong>NumberOfLinesDeletedByDeletedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>NumberOfFilesModified</strong></li> <li><strong>NumberOfFilesModifiedNet</strong>: Like the previous property, but considers only files that had net-changes</li> <li><strong>NumberOfFilesRenamed</strong></li> <li><strong>NumberOfFilesRenamedNet</strong>: Like the previous property, but considers only files that had net-changes</li> <li><strong>NumberOfLinesAddedByModifiedFiles</strong></li> <li><strong>NumberOfLinesAddedByModifiedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>NumberOfLinesDeletedByModifiedFiles</strong></li> <li><strong>NumberOfLinesDeletedByModifiedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>NumberOfLinesAddedByRenamedFiles</strong></li> <li><strong>NumberOfLinesAddedByRenamedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>NumberOfLinesDeletedByRenamedFiles</strong></li> <li><strong>NumberOfLinesDeletedByRenamedFilesNet</strong>: Like the previous property, but counts the net-lines</li> <li><strong>Density</strong>: The ratio between the two sums of all lines added+deleted+modified+renamed and their resp. gross-version. A density of zero means that the sum of net-lines is zero (i.e. all lines changes were just whitespace, comments etc.). A density of of 1 means that all changed net-lines contribute to the gross-size of the commit (i.e. no useless lines with e.g. only comments or whitespace).</li> <li><strong>AffectedFilesRatioNet</strong>: The ratio between the sums of&nbsp;<em>NumberOfFilesXXX</em>&nbsp;and&nbsp;<em>NumberOfFilesXXXNet</em></li> </ul> <p>&nbsp;</p> <p>This dataset is supporting the paper&nbsp;<strong>&quot;<em>Importance and Aptitude of Source code Density for Commit Classification into Maintenance Activities</em></strong><strong>&quot;</strong>, as submitted to the&nbsp;<em>QRS2019</em>&nbsp;conference (The 19th IEEE International Conference on Software Quality, Reliability, and Security). Citation:&nbsp;H&ouml;nel, S., Ericsson, M., L&ouml;we, W. and Wingkvist, A., 2019. Importance and Aptitude of Source code Density for Commit Classification into Maintenance Activities. In&nbsp;<em>The 19th IEEE International Conference on Software Quality, Reliability, and Security</em>.</p>

opencc-by-4.0Mar 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record