Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

846

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

846 results for “commit”

Learn how ShareScore rates datasets ↗
zenodo48/100

Github commit data for the article "Beyond Zipf's law: Exploring the discrete generalized beta distribution in open-source repositories"

<p><span>This dataframe corresponds to the data used in the Nowak's et al. 2024 article "Beyond Zipf&rsquo;s law: Exploring the discrete generalized beta distribution in open-source repositories" (see reference below).</span></p> <p><span>It consists of the distirbutions of number of commits per user across a number of GitHub repositories.&nbsp;<br><br>There are three columns:</span></p> <ul> <li><span>repository: the repository name</span></li> <li><span># of commits: the number of commits of a given individual</span></li> <li><span>rank: the user rank in the repository (by decreasing number of commits)<br><br></span></li> </ul> <p><strong><span>Reference:</span></strong></p> <p><span>Nowak, P., Santolini, M., Singh, C., Siudem, G., &amp; Tupikina, L. (2024). Beyond Zipf&rsquo;s law: Exploring the discrete generalized beta distribution in open-source repositories.&nbsp;<em>Physica A: Statistical Mechanics and Its Applications</em>, <em>649</em>, 129927. <a href="https://doi.org/10.1016/j.physa.2024.129927">https://doi.org/10.1016/j.physa.2024.129927</a></span></p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Data for "Deforestation in the Brazilian Amazon could be halved by scaling up the implementation of zero-deforestation cattle commitments"

<p>The processed data supporting the Global Environmental Change publication &quot;Deforestation in the Brazilian Amazon could be halved by scaling up the implementation of zero-deforestation cattle commitments&quot;.</p> <p>These data can be analyzed and visualized with the code at: <a href="https://github.com/sam-a-levy/Levyetal2023_cattlemarketshare">https://github.com/sam-a-levy/Levyetal2023_cattlemarketshare</a></p> <p>For a description of each file &amp; the variables contained, please look to the README file.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The commitment to global sea level rise over the next 500 years: exploring the threat of the Antarctic Ice Sheet to coastal infrastructure

<p>Within Australia alone, more than A$226 billion of coastal infrastructure is vulnerable to the anticipated rise in sea level by the end of the century. The IPCC Fifth Assessment Report concludes that the likely increase in global mean sea level during the 21st century ranges from 26-55 centimetres (under the low-end RCP2.6 climate scenario) to 45-82 centimetres (under the high-end RCP8.5 climate scenario). However, these projections do not take into account the potential for collapse of the marine-based sectors of the Antarctic Ice Sheet.</p> <p>Recent evidence has indicated that the IPCC projections may be under-estimates, with sea level increases of up to 2.5 metres possible by the end of the 21st century. Modelling studies have also demonstrated the potential for the Antarctic Ice Sheet to undergo irreversible collapse during the coming centuries, leading to dramatic increases in global sea level on time scales relevant to critical coastal infrastructure such as refineries and airports. The most extreme prediction is that Antarctica could contribute 15.65&plusmn;2.00 metres to global sea level by the year 2500.</p> <p>Here, we combine climate modelling and ice sheet modelling to explore the evolution of the Antarctic Ice Sheet over the next 500 years under a range of climate scenarios. We run the models many times to take into account gaps in our understanding of ice sheet dynamics. This allows us to generate robust projections of the Antarctic contribution to global sea level from the present to the year 2500, complete with quantified confidence intervals. We conclude that the sea level contribution during the 21st century will be modest, consistent with the IPCC Fifth Assessment Report, but that melting of the Antarctic Ice Sheet will accelerate thereafter. By the year 2500, we predict that the Antarctic contribution to global sea level will be at least 5 metres.</p>

opencc-by-4.0Jun 2020View details →
zenodo44/100

The datasets for "Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data" paper

<p>Paper title:&nbsp; Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data</p> <p>Conference:&nbsp; ICSME 2021</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Studying Bug-Fixing Commits in the WoC Dataset: Replication Package

<p>A replication package for the MSR 2023 Challenge submission titled &quot;Studying Bug-Fixing Commits in the WoC Dataset&quot;.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Replication Package: Assessing time-based and range-based strategies for commit assignment to releases

<p><strong>Abstract:</strong></p> <p>Release is a ubiquitous concept in software development, referring to grouping multiple independent changes into a deliverable piece of software. Mining releases can help developers understand the software evolution at coarse grain, identify which features were delivered or bugs were fixed, and pinpoint who contributed on a given release. A typical initial step of release mining consists of identifying which commits compose a given release. We could find two main strategies used in the literature to perform this task: time-based and range-based. Some release mining works recognize that those strategies are subject to misclassifications but do not quantify the impact of such a threat. This paper analyzed 13,419 releases and 1,414,997 commits from 100 relevant open source projects hosted at GitHub to assess both strategies in terms of precision and recall. We observed that, in general, the range-based strategy has superior results than the time-based strategy. Nevertheless, even when the range-based strategy is in place, some releases still show misclassifications. Thus, our paper also discusses some situations in which each strategy degrades, potentially leading to bias on the mining results if not adequately known and avoided.</p> <p><strong>Instructions:</strong></p> <p>Visit <a href="https://github.com/gems-uff/release-mining">https://github.com/gems-uff/release-mining</a> for instructions about how to use this dataset.</p> <p><strong>Files:</strong></p> <ul> <li>The <em>repos.tgz</em> contains our project <em>corpus </em>comprising 1,414,997&nbsp;releases&nbsp;from&nbsp;100&nbsp;relevant&nbsp;open&nbsp;source&nbsp;projects.</li> <li>The <em>repos.sha1</em> contains the sha1 checksum of <em>repos.tgz</em></li> </ul> <p><strong>Disclaimer:</strong></p> <p>This replication package contains the source code of 100 relevant open source projects. Its purpose is to enable the replication of the study conducted in the paper &quot;Assessing time-based and range-based strategies for commit assignment to releases.&quot;&nbsp;&nbsp;</p> <p>It is essential to check each project license before using the source code or any attached file for any other purposes besides replicating the study.</p> <p>&nbsp;</p>

openmit-licenseJan 2021View details →
zenodo40/100

Commit metadata for projects in TravisTorrent

<p>This dataset contains git metadata for 1,262 Java and Ruby projects hosted at GitHub up to January 27, 2017, 19:18:08 +0000. These are the projects in the TravisTorrent dataset released on January 11, 2017.</p> <p>It consists of an SQLite table containing the following columns:</p> <ul> <li><strong>project</strong>: project name on GitHub, in the form "owner/project"</li> <li><strong>sha</strong>: the commit id</li> <li><strong>message</strong>: the commit message</li> <li><strong>date</strong>: commit date</li> <li><strong>author_name</strong>: name of the commit author (not the committer)</li> <li><strong>author_email</strong>: email of the commit author (not the committer)</li> </ul>

opencc-by-sa-4.0Jul 2017View details →
zenodo40/100

1151 commits with software maintenance activity labels (corrective,perfective,adaptive)

<p>Data format: CSV</p> <p>Separator character: '#'</p> <p><strong>This dataset contains 1151 commits manually labeled with maintenance activities ("c" for corrective, "p" for perfective, "a" for adaptive)</strong> according to the definition by Mockus et al. in  <em>"Mockus, A. and Votta, L.G., 2000, October. Identifying Reasons for Software Changes using Historic Databases. In icsm (pp. 120-130)"</em>.</p> <p>In addition, this dataset also contains <strong>further information (features) extracted from the commits</strong>:</p> <ol> <li>The <strong>source code changes</strong> performed by the commit author as part of a given commit (statement added, statement removed, etc.) <ul> <li>The source code change taxonomy is detailed in <em>"Fluri, B. and Gall, H.C., 2006, June. Classifying change types for qualifying change couplings. In Program Comprehension, 2006. ICPC 2006. 14th IEEE International Conference on (pp. 35-45). IEEE."</em></li> </ul> </li> <li>A binary indication (1/0) whether a given commit contains any of the <strong>keywords from a pre-computed </strong>(according to a word frequency analysis)<strong> set of keywords</strong> <strong>indicative of each maintenance activity</strong>.</li> </ol> <p>The dataset consists of commits sampled from the following open source projects:</p> <ol> <li>RxJava</li> <li>hbase</li> <li>elasticsearch</li> <li>intellij-community</li> <li>hadoop</li> <li>drools</li> <li>kotlin</li> <li>restlet-framework-java</li> <li>orientdb</li> <li>camel</li> <li>spring-framework    </li> </ol> <p>This dataset is a supporting material for the paper <strong>"Boosting Automatic Commit Classification Into Maintenance Activities By Utilizing Source Code Changes", to appear in PROMISE 2017.</strong></p>

opencc-by-sa-4.0Jul 2017View details →
zenodo40/100

Understanding Organizational Commitment and its Factors Influencing the Nurse's Job Satisfaction in Hospitals- A Systematic Literature Review and Further Research Agendas

<div> <p><em><span>Organizational commitment is a crucial concept when it comes to human resource management and Organizational behavior. It has to do with how much a worker commits to and identifies with the goals, values, and purposes of their company. Elevated levels of Organizational commitment are associated with enhanced job satisfaction, less attrition, and better performance. The expected ideal condition, status, and research deficit are all included in this article. A research agenda is determined by applying the ABCD framework to qualitatively analyze the identified research gap.The paper documents the topic and provides helpful information about it, which will aid future scholars. </span></em></p> </div>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Data from: Transcriptome analysis of apical meristem enriched bud samples for size dependent flowering commitment in Crocus sativus reveal role of sugar and auxin signalling

<p><strong>Background</strong></p> <p>Cultivation of <em>Crocus sativus</em> (saffron) faces challenges due to inconsistent flowering patterns and variations in yield. Flowering takes place in a graded way with smaller corms unable to produce flowers. Enhancing the productivity requires a comprehensive understanding of the underlying genetic mechanisms that govern this size based flowering initiation and commitment. Therefore, samples enriched with non-flowering and flowering apical buds from small (&lt;6g) and large (&gt;14g) corms were sequenced.&nbsp;</p> <p><strong>Methods and Results</strong></p> <p>Apical bud enriched samples from small and large corms were collected immediately after break of dormancy in July. RNA sequencing was performed using Illumina Novaseq 6000. <em>De-novo</em> transcriptome assembly and analysis using flowering committed buds from large corms at post-dormancy and their comparison with vegetative shoot primordia from small corms pointed out the major role of Auxin and ABA hormonal regulation. Many genes with known dual responses in flowering development and circadian rhythm like Flowering locus T and Cryptochrome 1 along with a transcript showing homology with small auxin upregulated RNA (SAUR) exhibited induced expression in flowering buds. Thorough prediction of&nbsp;<em>Crocus sativus</em> non-coding RNA repertoire has been carried out for the first time. Enolase was found to be acting as a major hub with protein-protein interaction analysis using Arabidopsis counterparts.</p> <p><strong>Conclusion</strong></p> <p>Transcripts belong to key pathways including phenylpropanoid biosynthesis, hormone signaling and carbon metabolism were found significantly modulated. KEGG assessment and protein-protein interaction analysis confirm the expression data. Findings unravel the genetic determinants driving the size-dependent&nbsp;flowering in <em>Crocus sativus</em>.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

COMMIT Scenario Explorer

<p><strong>Release Notes</strong></p> <p>Version 1.1</p> <ul> <li>&nbsp;Error correction in the data for the <strong>model</strong>: &quot;IMAGE&quot;, <strong>scenario</strong>: &quot;2Deg2020_V4&quot;, <strong>variables</strong>: &quot;Price|Carbon&quot; and &quot;Policy Cost|Area under MAC Curve&quot;</li> </ul> <p>Version 1.0</p> <ul> <li>Initial release of the COMMIT data set.</li> </ul> <p>This scenario explorer presents a set of scenarios developed in the&nbsp;<a href="https://themasites.pbl.nl/commit/">COMMIT</a>&nbsp;project. <a href="https://themasites.pbl.nl/commit/">COMMIT</a>&nbsp;stands for Climate pOlicy assessment and Mitigation Modeling to Integrate national and global Transition pathways. The COMMIT project&rsquo;s main aim has been to improve modelling of national low-carbon emission pathways and analysis of country contributions to the global ambition of the Paris Agreement. The main motivation for this aim was that a common understanding and coherent message from the research community in different parts of the world is crucial to support the international negotiation process on climate policy. COMMIT consisted of a consortium of a large number of national teams, who regularly support domestic climate policy-making in their respective countries (Australia, Brazil, Canada, China, EU, India, Indonesia, Japan, Russia, South Korea, USA) and leading global integrated assessment modelling teams (PBL, PIK, IIASA, RFFCMCC EIEE). The consortium was led by Detlef van Vuuren and Heleen van Soest (PBL Netherlands Environmental Assessment Agency).</p> <p>The set of scenarios presented here revolves around the Bridge scenario, which was developed to study how the emissions gap between Nationally Determined Contributions (NDCs) and the global emissions levels needed to achieve the Paris Agreement&#39;s climate goals can be closed based on good practice policies (GPP). Next to these GPP and Bridge scenarios, the set contains different reference scenarios (a no new policies baseline, a current policies scenario, and an NDC scenario), as well as different scenarios limiting global warming to 2 degrees Celsius. The Bridge scenario builds upon the current policies scenario and assumes that specific good practice policies, which have shown to be effective in some countries, will be implemented globally from 2020 until 2030. After 2030, the bridge scenario transitions to a 2 &deg;C scenario following a cost-effective pathway. A distinction is made between low/medium and high-income countries in terms of timing and stringency of good practice policy targets. The set of policies was defined in dialogue with national model teams, granting a more realistic scenario narrative.</p> <p>Another set of scenarios included here consists of the Reference (here: Baseline) and Low-carbon scenarios as presented by Fragkos et al. in Energy.</p> <p>The data is also available for download and interactive viewing at the&nbsp;<a href="https://data.ene.iiasa.ac.at/commit">COMMIT Scenario Explorer hosted by IIASA</a>. The advantage of getting the data through the scenario explorer is that by registering with your email you will receive updates anytime there is a new version of this data set.</p> <p>The scenario ensemble is licensed under <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>. Please refer to the creative commons site for more information.</p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

Core bibliometric Covid19 and comparable research dataset and code for the study "From intent to impact: Investigating the effects of open sharing commitments"

<p>This document provides the underlying dataset for the bibliometric component for the 2022 study &quot;From intent to impact: Investigating the effects of open sharing commitments&quot; by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study&#39;s technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/&nbsp;</p> <p>Particularly, note that there is an error rate in attribution of signatory status to journal publications and preprints; in their location within specific thematic disease-based areas; or computing of dimension such as identification of data availability statement sections; identification of data depisition mentions within data availability statement sections; or matching of preprints and journal publications.</p> <p>These error rates are expected and have been estimated, please consult the technical report for full details.</p> <p>&nbsp;</p> <p>Definition of data fields is provided is the table below:</p> <table> <tbody> <tr> <td>Column name&nbsp;</td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server&#39;s unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server&#39;s unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of &quot;10.2139/ssrn.&quot; + &#39;ssrn_id&#39;</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>preprint_server</td> <td>Preprint platform on which a preprint has been published, restricted to arXiv, bioRxiv, medRxiv and SSRN for this study.</td> </tr> <tr> <td>journal_title</td> <td>Publishing journal name in the case of a journal publication.</td> </tr> <tr> <td>year</td> <td>The set is restricted to 2020 and 2021 for Covid19 preprints and journal publications. HVRD journal publications restricted to 2018-2019. HVRD preprints were restricted to 2020-2021 instead, to compensate for the lac of year-normalization for preprints, and generally better control findings against the launch of medRxiv in 2019.</td> </tr> <tr> <td>publication_title</td> <td>Title of the individual journal publication or preprint, not that of the publishing journal or preprint server.</td> </tr> <tr> <td>authors</td> <td>First 100 researchers that appear as authors of a preprint or journal publication. These are not parsed and provided for qualitative validation or&nbsp; assessments rather than for further quantitative treatment.</td> </tr> <tr> <td>Covid19</td> <td>Journal publications or preprints are coded 1 if they has been identified as falling into this thematic area through our queries (see the technical annex), 0 otherwise</td> </tr> <tr> <td>HVRD</td> <td>Human viral respiratory disease, the thematic area considered to be the closest to Covid19. Journal publications or preprints are coded 1 if they has been identified as falling into this thematic area through our queries (see the technical annex), 0 otherwise</td> </tr> <tr> <td>Journal_sig</td> <td>Journal publications where the publishing journal and/or its publishing house are Joint Statement signatories. Coded as 1 if they are signatories, 0 if not signatory, null if status could not be determined due to insufficient metadata. Not that all preprint servers included in this study are Joint Statement signatories. This category was fully removed from the models for preprints, rather than all preprints being assigned automatic signatory status.</td> </tr> <tr> <td>RPO_sig</td> <td>Journal publications and preprints where at least one author is affiliated with at least one research performing organization that is a Joint Statement signatory. Coded as 1 ifor signatory, 0 if not signatory, null if status could not be determined due to insufficient metadata.</td> </tr> <tr> <td>Funder_sig</td> <td>Journal publications and preprints where at least one funder supporting the research is a Joint Statement signatory. Coded as 1 ifor signatory, 0 if not signatory, null if status could not be determined due to insufficient metadata. Although funding is attributed to researchers rather than publications, funding metadata is more readily available at the second level. This approach also captures the flexible usage of financial resources that researchers may make accross mulitple concurrently ongoing research projects.</td> </tr> <tr> <td>overton_norm</td> <td>Year and subfield-normalized binary score of whether the journal publications has been cited by one or more policy-related documents from the Overton database. Null scores for journal publications not covered by the database.</td> </tr> <tr> <td>overton</td> <td>Normalizations being unable for preprints, binary score of whether the preprint has been cited by one or more policy-ralated documents from the Overton database. Null scores for preprints not covered by the database.</td> </tr> <tr> <td>daswriting_binary</td> <td>Binary score capturing identification of a data availability statement in the journal publication or preprint using the queries presented in the technical annex. Null scores are for publications and preprints where records of full texts were unavailable for text mining, or were this analysis could not be performed due to licensing restrictions.&nbsp;</td> </tr> <tr> <td>deposition_binary</td> <td>Binary score capturing identification of a data availability statement and data deposition mention therein in the journal publication or preprint using the queries presented in the technical annex. Null scores are for publications and preprints where records of full texts were unavailable for text mining, or were this analysis could not be performed due to licensing restrictions.&nbsp;</td> </tr> <tr> <td>is_oa</td> <td>Binary score capturing OA or free-to-read (also so-calleod &quot;bronze OA&quot; and &quot;green OA&quot;) status of journal publications. Unpaywall categories have been used in a mutually exclusive implementation, with the best (gold &gt; hybrid&gt;bronze&gt;green) possible applicable category being retained. Null scores for journal publications not covered in our Unpaywall dataset. Scores of 0 denote journal publications not available under an OA or free-to-read category.</td> </tr> <tr> <td>is_gold</td> <td>as above</td> </tr> <tr> <td>is_hybrid</td> <td>as above</td> </tr> <tr> <td>is_bronze</td> <td>as above</td> </tr> <tr> <td>is_green</td> <td>as above</td> </tr> <tr> <td>matched_journal_binary</td> <td>For preprints, whether one or more matching journal publications could be identified using the queries identified in the technical, or preprint servers&#39; own lists of preprint-journal publication matches. Null scores for preprints with insufficient metadata information to perform the matching operation.</td> </tr> <tr> <td>matched_journal_doi</td> <td>For those preprints with or more matching journal publications, the DOI(s) of the matching journal publication(s). Note that some of the maching journal publications identified do not have DOIs.</td> </tr> <tr> <td>matched_preprint_binary</td> <td>For journal publications, whether one or more matching preceding preprints could be identified using the queries identified in the technical annex, or preprint servers&#39; own lists of preprint-journal publication matches. Null scores for journal publications without sufficient metadata to run the analysis.</td> </tr> <tr> <td>matched_preprint_id</td> <td>For those journal publications preceded with one or more arXiv, bioRxiv, medRxiv or SSRN preprints, the DOI(s), arXiv ID and/or SSRN ID of the matching preprint(s).&nbsp;</td> </tr> <tr> <td>hasdoi</td> <td>Only journal publications with DOIs were retained in the core quantitative analyses.</td> </tr> <tr> <td>hasacknowledgements</td> <td>Only journal publications with funding acknowledgements (to determine funding-based signatory status) were retained in the core quantitative analyses.</td> </tr> <tr> <td>funder_array</td> <td>Array (but cast as string) of names of the funders on the basis of whose idenitification signatory status has been attributed, where relevant. Null if non-signatory or unknown signatory status.</td> </tr> <tr> <td>RPO_array</td> <td>Array (but cast as string) of names of the research performing organizations on the basis of whose idenitification signatory status has been attributed, where relevant. Null if non-signatory or unknown signatory status.</td> </tr> <tr> <td>DAS_excerpt</td> <td>Journal publication or preprint text excerpt on which succesful identifcation of data availability statements and/or data deposition mentions have been made. Null both where the query could not be run at all, or where the query was negative.</td> </tr> <tr> <td>big5</td> <td>Journal publication published in a journal owned by one of the following five publishing houses: Elsevier, Sage, Springer Nature, Taylor-Francis, Wiley.</td> </tr> <tr> <td>LMIC</td> <td>Journal publication whose authors include at least one researcher affiliated with at least one institution located in a lower-middle income country as defined by the World Bank</td> </tr> <tr> <td>LIC</td> <td>Journal publication whose authors include at least one researcher affiliated with at least one institution located in a low income country as defined by the World Bank</td> </tr> <tr> <td>SouthNorth</td> <td>Journal publication whose authors include at least one researcher affiliated with at least one institution located in a upper-middle income country, a lower-middle income country, or a low income country as defined by the World Bank; as well as at least one researcher affiliated with at least one institution located in a high income country. For the purpose of this indicator, Sicnece-Metrix exceptionally includes China and Bulgaria in the list of high income countries.</td> </tr> <tr> <td>DID_allauthors_OR</td> <td>Journal publication is included in the difference-in-difference model defining signatory publication as EITHER holding journal-based signatory status OR funding-based signatory status, and where no filter has been applied to control for author-level biases.</td> </tr> <tr> <td>DID_authorcontrol_OR</td> <td>Journal publication is included in the difference-in-difference model defining signatory publication as EITHER holding journal-based signatory status OR funding-based signatory status, and where a filter has been applied to control for author-level biases.</td> </tr> <tr> <td>DID_authorcontrol_AND</td> <td>Journal publication is included in the difference-in-difference model defining signatory publication as holding journal-based signatory status AND funding-based signatory status, and where a filter has been applied to control for author-level biases.</td> </tr> <tr> <td>DID_allauthors_AND</td> <td>Journal publication is included in the difference-in-difference model defining signatory publication as holding journal-based signatory status AND funding-based signatory status, and where no filter has been applied to control for author-level biases.</td> </tr> <tr> <td>Preprint_authorcontrol</td> <td>Preprint is included in the the analytical breakdowns where a filter has been applied to control for author-level biases. Note that authors have been kept constant in preprints on the basis of their belonging to all analytical breakdowns in journal publications rather than in preprint-based groups.</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Data for publication "Committed sea-level rise under the Paris Agreement and the legacy of delayed mitigation action".

<p>Data underlying the publication &quot;Committed sea-level rise under the Paris Agreement and the legacy of delayed mitigation action&quot;.</p> <p>Journal: Nature Communications</p> <p>Authors: <em>Matthias Mengel<sup>1*</sup></em><em>, Alexander Nauels</em><sup><em>2</em></sup><em>, Joeri Rogelj</em><sup><em>3,4</em></sup><em>, Carl-Friedrich Schleussner<sup>1,5</sup></em></p> <p>(1) Potsdam Institute for Climate Impact Research (PIK), Member of the Leibniz Association, P.O. Box 60 12 03, D-14412 Potsdam, Germany</p> <p>(2) Australian-German College of Climate &amp; Energy Transitions, The University of Melbourne, Parkville, Victoria 3010, Australia</p> <p>(3) ENE Program, International Institute for Applied Systems Analysis (IIASA), Schlossplatz 1, Laxenburg A-2361, Austria</p> <p>(4) Institute for Atmospheric and Climate Science, ETH Zurich, Universit&auml;tstrasse 16, Zurich 8006, Switzerland</p> <p>(5) Climate Analytics, Ritterstr. 3, 10969 Berlin, Germany</p> <p>(*) email matthias.mengel@pik-potsdam.de</p> <p>Abstract:</p> <p>Sea-level rise is a major consequence of climate change that will continue long after emissions of greenhouse gases have stopped. The 2015 Paris Agreement aims at reducing climate-related risks by reducing greenhouse gas emissions to net zero and limiting global-mean temperature increase. Here we quantify the effect of these constraints on global sea-level rise until 2300 including Antarctic ice-sheet instabilities. We estimate median sea-level rise between 0.7 and 1.2m if net zero greenhouse gas emissions are sustained until 2300, varying with the pathway of emissions during this century. Temperature stabilization below 2&deg;C is insufficient to hold median sea-level rise until 2300 below 1.5m. We find that each 5-year delay in near-term peaking of CO2 emissions increases median year-2300 sea-level rise estimates by ca. 0.2m, and extreme sea-level rise estimates at the 95th percentile by up to 1m. Our results underline the importance of near-term mitigation action for limiting long-term sea-level rise risks.</p> <p>&nbsp;</p> <p>Large zip files provides data. Small zip file python code for plotting and writing supplementary data.</p>

opencc-by-4.0Dec 2017View details →
zenodo40/100

Code and data for "Current fossil fuel infrastructure does not yet commit us to 1.5°C warming"

<p>This package generates all of the model runs and plotting code for &quot;Current infrastructure does not yet commit us to 1.5&deg;C warming&quot;.</p> <p>See enclosed README file for dependencies and how to run.</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

District heating modelling data for the publication "Integration of feed flow temperatures in unit commitment models of future district heating systems"

<p>Modelling data for a district heating system model which has been used for the publication &quot;Integration of feed flow temperatures in unit commitment models of future district heating systems&quot; on the 4th Generation District Heating (4GDH) conference 2018.</p>

opencc-by-sa-4.0Jan 2019View details →
zenodo40/100

Discrete regulatory modules instruct hematopoietic lineage commitment and differentiation

<p>bioRxiv preprint: https://www.biorxiv.org/content/10.1101/2020.04.02.022566v4</p> <p><strong>Contact:</strong> Grigorios Georgolopoulos (<a href="mailto:ggeorgol@altius.org">ggeorgol@altius.org</a>); Jeff Vierstra (<a href="mailto:jvierstra@altius.org?subject=Consensus%20DNase%20I%20footprints">jvierstra@altius.org</a>)</p> <p>Lineage commitment and differentiation is driven by the concerted action of master transcriptional regulators at their target chromatin sites. Multiple efforts have characterized the key transcription factors (TFs) that determine the various hematopoietic lineages. However, the temporal interactions between individual TFs and their chromatin targets during differentiation and how these interactions dictate lineage commitment remains poorly understood. Here we delineate the temporal interplay between the <em>cis</em>- and the <em>trans</em>-regulatory landscape in establishing lineage commitment and differentiation in human hematopoiesis by performing a dense timecourse of chromatin accessibility (DNase I-seq), and gene expression (total and single cell RNA-seq).</p> <p>All data uploaded correspond to human genome build version GRCh38.</p> <p><strong>Contents</strong></p> <ol> <li><strong>DNase I Hotspot (DHS) metadata: </strong>Supplementary_Data_1.txt</li> <li><strong>DNase I Hotspot quantile-normalized counts:</strong> A tab-separated matrix with quantile-normalized DNase I density counts from 79,085 FDR 5% hotspots, across 12 erythroid differentiation timepoints from 3 donors, present in at least n=2 samples. Rows correspond to DHS information in&nbsp;Supplementary_Data_1.txt (hotspots.fdr.0.05.qnorm.counts.tsv.gz)</li> <li><strong>Column information for DNase I Hotspot quantile-normalized counts: </strong>hotspots.fdr.0.05.qnorm.counts.info.tsv</li> <li><strong>Developmentally regulated gene metadata (erythroid): </strong>Supplementary_Data_2.csv</li> <li><strong>Gene matrix of quantile-normalized FPKM values (erythroid): </strong>A tab-separated matrix with the quantile-normalized FPKM values of all detected genes, across 13 erythroid differentiation timepoints from 3 donors. (fpkm_erythroid_qnorm.tsv.gz)</li> <li><strong>Column information for the quantile-normalized FPKM gene matrix (erythroid): </strong>A tab-separated table (fpkm_erythroid_qnorm.info.tsv)</li> <li><strong>CD34+ HSPC TADs at 10kb resolution: </strong>Supplementary_Data_3.bed</li> <li><strong>Day 11 <em>ex vivo</em> erythroid progenitor TADs at 10kb resolution: </strong>Supplementary_Data_4.bed</li> <li><strong>Transcription factor motif enrichment per DHS cluster: </strong>Supplementary_Data_5.csv</li> <li><strong>Correlation information (links) between developmentally regulated DHS and target genes: </strong>Supplementary_Data_6.csv</li> <li><strong>Chromatin anchor loops called from 10kb resolution Hi-C data: </strong>Supplementary_Data_7.bedgraph</li> <li><strong>Developmentally regulated gene metadata (megakaryocytic): </strong>Supplementary_Data_8.csv</li> <li><strong>Gene matrix of quantile-normalized FPKM values (megakaryocytic): </strong>A tab-separated matrix with the quantile-normalized FPKM values of all detected genes, across 13 megakaryocytic differentiation timepoints from 3 donors. (fpkm_megakaryocyte_qnorm.tsv.gz)</li> <li><strong>Column information for the quantile-normalized FPKM gene matrix (megakaryocytic): </strong>A tab-separated table (fpkm_megakaryocyte_qnorm.info.tsv)</li> <li><strong>Marker (differentially expressed) genes per single cell population: </strong>Supplementary_Data_9.csv</li> <li><strong>A SCANPY h5ad Annotated DataFrame object: </strong>Annotated Data frame `anndata` in h5ad format including the gene-by-cell count matrix, Velocyto splicing kinetics (RNA velocity) information layer, along with obs, obsm, var, varm, and uns layers. (SCANPY_anndata_object.h5ad)</li> </ol>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Evaluating Institutional Commitments to Open Scholarly Infrastructure: A Review of Open Access Collection Development Policies

<p>Data prepared for the publication &quot;Evaluating Institutional Commitments to Open Scholarly Infrastructure: A Review of Open Access Collection Development Policies.&quot;</p> <p><strong>oa-cd-policies.csv</strong></p> <p>Scope: This data represents collection development policies that contain substantial mention of open access.</p> <p>Data collection: The policies were sourced using an Advanced Google Search for &quot;open access&quot; AND &quot;collection development policy&quot; at &quot;.edu&quot; domains.</p> <p>Variables:</p> <ul> <li>institution: Free text, name of the institution.</li> <li>carnegie_class: One of <a href="https://carnegieclassifications.acenet.edu/carnegie-classification/classification-methodology/basic-classification/">these options</a>;&nbsp;the Carnegie classification of the institution.</li> <li>institution_type: One of public or private; the funding source of the institution.</li> <li>policy_name: Free text; the title of the policy.</li> <li>supplemental_policy: Link to a supplemental open access policy if linked in the collection development policy.</li> <li>cd_policy_link: Link to the policy.</li> <li>infrastructure: TRUE or FALSE; whether the policy includes a commitment to open access scholarly&nbsp;infrastructure development, including open source platforms, locally hosted platforms, consortia, or an institutional repository.</li> <li>excerpt: Free text; text from the policy that mentions infrastructure.</li> </ul> <p><strong>principles-policies.csv</strong></p> <p>Scope:&nbsp;This data represents those collection development policies from oa-cd-policies.csv&nbsp;that contain commitments in line with the <a href="https://openscholarlyinfrastructure.org/">Principles for Open Scholarly Infrastructure</a>.</p> <p>Variables:</p> <ul> <li>institution: Free text, name of the institution.</li> <li>carnegie_class: One of <a href="https://carnegieclassifications.acenet.edu/carnegie-classification/classification-methodology/basic-classification/">these options</a>;&nbsp;the Carnegie classification of the institution.</li> <li>institution_type: One of public or private; the funding source of the institution.</li> <li>policy_name: Free text; the title of the policy.</li> <li>supplemental_policy: Link to a supplemental open access policy if linked in the collection development policy.</li> <li>cd_policy_link: Link to the policy.</li> <li>principle: One of the three main <a href="https://openscholarlyinfrastructure.org/">Principles</a>.</li> <li>sub_principle: One of the <a href="https://openscholarlyinfrastructure.org/">Sub-Principles</a>.</li> <li>excerpt: Free text; text from the policy that illustrates the sub_principle.</li> </ul>

opencc-by-4.0May 2023View details →
zenodo40/100

IMPETUS - Turning Climate Commitments into Action

<p>&nbsp;</p> <p>A call to action video to&nbsp;launch the climate IMPETUS project and to promote it online and at events.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Dataset of Commit Classification via Diff-Code GCN based on System Dependency Graph

<p>Commit Classification via Diff-Code GCN based on System Dependency Graph</p> <p>The&nbsp;dataset&nbsp;is based on Lobna Ghadhab et al. [1]. Levin et al.[2]&#39;s dataset, and we extract all commits with pure java codes of two versions.&nbsp;</p> <p>In the dataset, evert commit folder have two sub-folder called before and after, they contains two version of codes. we extracted it by pydriller.</p> <p>The dataset have 1213 commits with two version java codes,and it contains three categories:</p> <p>(1) The first category is Corrective, which involves rectifying errors and faults identified during software usage.</p> <p>(2)The second category is Perfective, which entails enhancing software quality attributes, such as performance, maintainability, and usability.</p> <p>(3) Lastly, is Adaptive, which&nbsp;encompasses adapting the software to new environments (e.g., software or hardware) or introducing new functionalities.</p> <p>The dataset have 450 labels of&nbsp;Corrective. 441 for Perfective the rest for&nbsp;Adaptive.</p> <p>&nbsp;</p> <p>[1]L. Ghadhab, I. Jenhani, M. W. Mkaouer, and M.Ben Messaoud, &rdquo;Augmenting commit classification by using fine-grained source code changes and a pretrained deep neural language model,&rdquo; Information and Software Technology, vol. 135, p. 106566, 2021/07/01/2021.</p> <p>[2]S. Levin and A. Yehudai, &rdquo;Using Temporal and Semantic Developer-Level Information to Predict Main</p> <p>tenance Activity Profiles,&rdquo; in 2016 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2016.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Angular GitHub Commits Object-centric Event Log

<p><strong>Overview</strong></p> <p>This real-world object-centric event log in the OCEL 2.0 standard contains an extraction of the commit information from the <a href="https://github.com/angular/angular">GitHub repository</a> used to developed the <a href="https://www.angular.io/">Angular</a> platform. A single code commit in the repository is abstracted to one event in the log. The dataset contains essential information for each commit, such as the timestamp and the contributor&#39;s details. Crucially, commit information is connected to two classes of objects: the file(s) affected by the commit, and the branch(es) in the repository containing the commit.</p> <p><strong>Description</strong></p> <p>GitHub, a popular platform for developers offering the functionalities of the Git versioning system, allows to record single modifications to software projects by contributors; such modifications are grouped in units called <strong>commits</strong>. Commits contain all details of the edits operated on a group of files in the projects. Therefore, all commits of a project constitute a ledger, that allows to rewind or fast-forward all contributions in the project.</p> <p>Commits in a project are arranged in <strong>branches</strong>, which form a tree-like structure. A contributor may create a new branch, essentially a copy of the project, in order to commit modifications safely. Once the contributor is satisfied with the edits, they may <strong>merge</strong> their new branch back into the pre-existing branch (realized by applying the modifications of all the new commits sequentially, and then solving the conflicts that may arise).</p> <p>This log contains an extraction of the commit information of the <a href="https://www.angular.io/">Angular</a> project on <a href="https://github.com/angular/angular">GitHub</a>. The abstraction level is such that every commit corresponds to an event in the log.</p> <p>For each event, the following information is recorded:</p> <ul> <li>a unique identifier (<strong>hash</strong>)</li> <li>the author&#39;s timestamp of the commit (includes timezone information)</li> <li>an <strong>activity label</strong>: the Angular project conforms to the <a href="https://www.conventionalcommits.org/">Conventional Commits</a> initiative, which mandates commit messages containing an initial identifier. This helps to reconstruct a clean activity notion. Some of the labels have been cleaned by hand (for instance, in case of typos)</li> <li>the message of the commit</li> <li>the contributor&#39;s name</li> <li>the contributor&#39;s email (<strong>resource</strong>)</li> <li>a <strong>merge</strong> flag; <strong>True</strong> if the commit is a merge, <strong>False</strong> otherwise</li> <li>information related to the <strong>files</strong> edited by the commit (in case of renames, we track the new name)</li> <li>information related to the <strong>branches</strong> in which the commit appears</li> </ul> <p>Files and branches are two distinct object types in this log. Note that a commit might not be associated to any file. Conversely, a commit always appears in at least one branch.</p> <p>This event log has been extracted with the help of <a href="https://github.com/ishepard/pydriller">PyDriller</a>.</p> <p><strong>Properties</strong></p> <p>This event log has the following properties:</p> <table> <tbody> <tr> <td><strong>Property</strong></td> <td><strong>Value</strong></td> </tr> <tr> <td>Events</td> <td>27847</td> </tr> <tr> <td>Activity Labels</td> <td>67</td> </tr> <tr> <td>Object Types</td> <td>2</td> </tr> <tr> <td>Objects (files)</td> <td>35392</td> </tr> <tr> <td>Objects (branches)</td> <td>119</td> </tr> </tbody> </table> <p><strong>Get started</strong></p> <p>Download the dataset, and position it in the folder of your Python script or console.</p> <p><em>pip install pm4py</em></p> <p>To manipulate object-centric logs programmatically, use the functionality of the <em>ocel</em> package <a href="https://pm4py.fit.fraunhofer.de/static/assets/api/2.7.5.1/api.html#object-centric-process-mining-pm4py-ocel">in the PM4Py library</a>. Additionally, check out the <a href="https://www.ocel-standard.org/beta/tool-support/overview/">tool support</a> for object-centric event logs!</p> <p><em>from pm4py import ocel</em></p> <p><strong>Acknowledgements</strong></p> <p>We thank the Alexander von Humboldt (AvH) Stiftung for supporting our research.</p>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record