Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “Dependency graph”
Goblin: Neo4J Maven Central dependency graph
<p>This repository contains a Neo4j dump of Maven Central dependency graph generated using <a href="https://github.com/Goblin-Ecosystem/goblinDependencyMiner">goblinDependencyMiner</a>.<br>To import this graph into neo4j, <strong>please use a version 4.x</strong>.</p> <p>Our dependency graph structure and metamodel are shown in images "goblin_dg_structure" and "metamodel".</p> <p>The latest available version dates from April 20, 2025, contains <span>16,939,391</span> nodes (712,509 libraries and 16,226,882 releases) and <span>152,434,085</span> edges (136,207,203 dependencies and 16,226,882 versioning edges).</p> <p>This repository contains two dump of the database:</p> <ul> <li><strong>goblin_maven_20_04_25.dump: </strong>This dataset contains the entire Maven Central dependency graph.</li> <li><strong>with_metrics_goblin_maven_20_04_25.dump</strong>: This dataset is the same as the previous one, but enriched with new “AddedValue” nodes (49,393,155 new nodes) representing the following metrics: CVE (dated may 13, 2025), freshness, popularity and speed. More information in this <a href="https://github.com/Goblin-Ecosystem/goblinTutorial">tutorial</a>.</li> </ul> <p>More details in the dedicated paper: <strong>Goblin: A Framework For Enriching And Querying the Maven Central Dependency Graph </strong>(https://doi.org/10.1145/3643991.3644879)<strong> </strong>- 21st International Conference on Mining Software Repositories (MSR'24).<br>If you use it, please <strong>cite</strong> this paper: <a href="https://dl.acm.org/doi/10.1145/3643991.3644879">https://dl.acm.org/doi/10.1145/3643991.3644879</a></p> <p>⚠️ This dataset is the subject of the <strong>Mining Challenge at the MSR 2025 conference</strong>, more information <a href="https://2025.msrconf.org/track/msr-2025-mining-challenge">here</a>.</p>
Maven central dependency graph
<p>The Maven dependency graph is an open dataset of Maven Central artifacts, their dependencies, as well as other relationships. Its main intent is to domesticate the wild within and around the Maven central ecosystem, in particular, and JVM-based libraries at large, making it more harnessable to both academics and industry. It is intended to answer high-level research questions concerning artifacts releases, evolution, and usage trends over time. It can also be used to assist researchers in selecting relevant datasets, among the mass of existing software artifact, for assessing particular empirical software engineering challenges. The complexity of these questions can range from simple pattern matching to advanced big data analysis and machine learning techniques.<br> <br> The accompanying paper to this dataset is has been accepted for publication in the proceedings of the International Conference on Mining Software Repositories 2019 and has received the MSR 2019 Data Showcase Award. This paper is available for download on <a href="https://arxiv.org/abs/1901.05392">arXiv</a>.</p>
The Updated Maven Central Dependency Graph
<p><strong>Maven Central Dependency Graph</strong></p> <p>This is an updated version of the artifact at https://zenodo.org/record/1489120</p> <p>The Maven dependency graph is an open dataset of Maven Central artifacts, their dependencies, as well as other relationships. Its main intent is to domesticate the wild within and around the Maven central ecosystem, in particular, and JVM-based libraries at large, making it more harnessable to both academics and industry. It is intended to answer high-level research questions concerning artifacts releases, evolution, and usage trends over time. It can also be used to assist researchers in selecting relevant datasets, among the mass of existing software artifact, for assessing particular empirical software engineering challenges. The complexity of these questions can range from simple pattern matching to advanced big data analysis and machine learning techniques.</p> <p>The accompanying paper to this dataset is has been accepted for publication in the proceedings of the International Conference on Mining Software Repositories 2019 and has received the MSR 2019 Data Showcase Award. This paper is available for download on <a href="https://arxiv.org/abs/1901.05392">arXiv</a>.</p> <p><strong>What is new?</strong></p> <p>The previous version included artifacts until September 6, 2018.<br> This version includes artifacts until September 10, 2019.</p> <p>This version includes license information as well as information about associated code repository.</p> <p>This version contains 4 201 392 artifacts (version) of 308116 distinct libraries from 47481 distinct group IDs.</p> <p>Note 33 638 artifacts represents version ranges and note actual versions. They can be filtered out by excluding version containing ','.</p> <p><br> <strong>Usage</strong></p> <p>Usage:</p> <ul> <li> Download the archive from zenodo</li> <li> Decompress the archive</li> </ul> <pre><code class="language-bash"># Pull the image and start the container docker run -d --name mm-neo4j -p 7474:7474 -p 7687:7687 -v /path/to/neo4j-data:/data --env=NEO4J_dbms_memory_heap_max__size=8g lyadis/mm-neo4j:latest</code></pre> <ul> <li> Open http://127.0.0.1:7474/browser/ in your browser</li> </ul>
LabelGit: A dataset for software repositories classification using attributed dependency graphs
<p>A dataset for software repositories classification using attributed dependency graphs</p>
Results from "Binary Reduction of Dependency Graphs"
<p>The raw data of the reduction of the 238 bugs as extracted from the runs of the four algorithms ddmin, verify, closure, binary. An extra column also exists that describes the run of Binary Reduction on the list of classes directly.</p> <p>The columns in `deliverable.csv` are predicate (the name of the decompiler), name (the name of the project we ran on), size (number of classes), scc (number of strongly connected components). Then for each tool we have size, scc (final size and scc), iters (iterations to last success), total-iters (iterations before finishing), time (time to last success (s)), total-time (time before finishing), timeout (did the predicate timeout, or succeed), check (did the bug still exist after reduction)</p> <p>The `benchmarks.csv` covers basic statistics about the programs used in the results. Lib is the number of library classes, LOC is lines of source code in the program, classes are the number of classes, edges are the edges in the dependency graph, degree is the average in and out degree in the graph, scc is the number of strongly connected components, out_degree and in_degree is the median in and out degree in the graph.</p>
РИС. 3. Графики Зависимости: А – Ширины (D) и В – фронтальной кривиЗны (D/L) раковины Corbicula fluminea от длины (L). В качестве Зависимых переменных (y) в формулах укаЗаны Ширина (D) и фронтальнаЯ кривиЗна раковины (D/L). FIG. 3. Graphs of correlations between (A) width (D), (B) frontal curvature (H/L) and the shell length (L) of Corbicula fluminea. The formulas include shell width (D) and frontal curvature (D/L) as dependent variables (y). in Особенности аллометрического роста двустворчатого моллюска-вселенца Corbicula fluminea (Bivalvia: Cyrenidae) иЗ бассейна реки Дон
РИС. 3. Графики Зависимости: А – Ширины (D) и В – фронтальной кривиЗны (D/L) раковины Corbicula fluminea от длины (L). В качестве Зависимых переменных (y) в формулах укаЗаны Ширина (D) и фронтальнаЯ кривиЗна раковины (D/L). FIG. 3. Graphs of correlations between (A) width (D), (B) frontal curvature (H/L) and the shell length (L) of Corbicula fluminea. The formulas include shell width (D) and frontal curvature (D/L) as dependent variables (y).
Neo4J Maven Central dependency graph
<p>⚠️ Updated archive here: <a href="../records/11104819">https://zenodo.org/records/11104819</a> ⚠️</p> <p>Neo4j dump of the dependency graph of Maven Central as of April 04, 2023.<br> </p>
Dataset of Commit Classification via Diff-Code GCN based on System Dependency Graph
<p>Commit Classification via Diff-Code GCN based on System Dependency Graph</p> <p>The dataset is based on Lobna Ghadhab et al. [1]. Levin et al.[2]'s dataset, and we extract all commits with pure java codes of two versions. </p> <p>In the dataset, evert commit folder have two sub-folder called before and after, they contains two version of codes. we extracted it by pydriller.</p> <p>The dataset have 1213 commits with two version java codes,and it contains three categories:</p> <p>(1) The first category is Corrective, which involves rectifying errors and faults identified during software usage.</p> <p>(2)The second category is Perfective, which entails enhancing software quality attributes, such as performance, maintainability, and usability.</p> <p>(3) Lastly, is Adaptive, which encompasses adapting the software to new environments (e.g., software or hardware) or introducing new functionalities.</p> <p>The dataset have 450 labels of Corrective. 441 for Perfective the rest for Adaptive.</p> <p> </p> <p>[1]L. Ghadhab, I. Jenhani, M. W. Mkaouer, and M.Ben Messaoud, ”Augmenting commit classification by using fine-grained source code changes and a pretrained deep neural language model,” Information and Software Technology, vol. 135, p. 106566, 2021/07/01/2021.</p> <p>[2]S. Levin and A. Yehudai, ”Using Temporal and Semantic Developer-Level Information to Predict Main</p> <p>tenance Activity Profiles,” in 2016 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2016.</p> <p> </p>
CODE-RADE dependency graph (14/09/2015)
<p>The dependency graph of the CODE-RADE Continuous Integration service at http://ci.sagrid.ac.za</p> <p> </p> <p>Published 14/09/2015. Incomplete.</p>
A Dependency Graph for 460,000 Papers and Their Software Mentions from the CZI Software Mentions Dataset
<p>Using the CZI Software Mentions Dataset and <a href="https://ecosyste.ms">ecosyste.ms</a> we create a graph of papers, their mentioned software, and recursive dependencies of each piece of software across 466,000 papers and three software registries (PyPI, CRAN, and BioConductor).</p>
SoftwareThatMatters - Analyzing the effect of introducing time as a component in Python dependency graphs
<p>This is a data set produced as a result of the research conducted during my bachelor thesis project "Analyzing the effect of introducing time as a component in Python dependency graphs". This set contains 3 parts:</p> <ul> <li>The raw data that was gathered in order to be processed.</li> <li>The processed data that was used for the construction of the graph.</li> <li>The results that were produced during this research.</li> </ul> <p>More information can be found on the repository of this research project: <a href="https://github.com/AndreiPurcaru/SoftwareThatMatters">SoftwareThatMatters</a></p> <p> </p>
Dataset for Generative Model of Software Dependency Graphs
<p>Data set for the paper entitled "A Generative Model of Software Dependency Graphs to Better Understand Software Evolution".</p> <p>Available files are:</p> <ul> <li>Sources archives (102 MB),</li> <li>Extracted dependencies (3.5 MB) and</li> <li>Generated graphs (15 MB).</li> </ul>
Graph of the dependence of the pH of the aqueous extract of the leather tissue of sheepskins in the canning process
<p>Graph of the dependence of the pH of the aqueous extract of the leather tissue of sheepskins in the canning process</p>
Graph of dependence of moisture content of sheepskin leather tissue in the canning process
<p>Graph of dependence of moisture content of sheepskin leather tissue in the canning process</p>
Graph of the dependence of the temperature of welding of the leather tissue of sheepskins during the canning process
<p>Graph of the dependence of the temperature of welding of the leather tissue of sheepskins during the canning process</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.