Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “Software Dependencies”
Fedora and Debian software package dependency networks along with description text associated with nodes
<p>Fedora (version 28) and Debian (version 9.5) software package dependency networks along with description text associated with nodes. Also includes learned vectors by using PCTADW-* as in "Kexuan Sun, Shudan Zhong, and Hong Xu. 2020. Learning Embeddings of Directed Networks with Text-Associated Nodes---with Application in Software Package Dependency Networks. 2020 BigGraphs Workshop at IEEE BigData 2020."</p>
Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'
<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions. </p> <p> </p>
Cataloging Dependency Injection Anti-Patterns in Software Systems
<p><strong>Background</strong> Dependency Injection (DI) is a commonly applied mechanism to decouple classes from their dependencies in order to provide better modularization of software. In the context of Java, the availability of a DI specification and popular frameworks, such as Spring, facilitate DI usage in software projects. However, bad DI implementation practices can have negative consequences, such as increasing coupling, hindering the achievement of DI's main goal. Even though the literature suggests the existence of DI anti-patterns, there is no detailed documentation of such bad practices. Moreover, there is no evidence on their occurrence and perceived usefulness from the developer's point of view. </p> <p><strong>Aims</strong> Our goal is to review the reported DI anti-patterns in order to analyze their completeness and to propose and evaluate a novel catalog of Java DI anti-patterns. </p> <p><strong>Method</strong> We propose a catalog containing twelve Java DI anti-patterns. We selected four open-source and two closed-source software projects that adopt a DI framework and developed a tool to statically analyze the occurrence of the candidate DI anti-patterns within their source code. Also, we conducted a survey through face to face interviews with three experienced developers that regularly apply DI. We extended the survey in order to gather the perception of a set of fifteen expert and novice developers through an online questionnaire. </p> <p><strong>Results</strong> At least nine different DI anti-patterns appeared frequently in the analyzed projects. In addition, the feedback received from the developers confirmed the relevance of the catalog. Besides, the respondents expressed their willingness to refactor instances of anti-patterns from source code.</p> <p><strong>Conclusions</strong> The catalog contains Java DI anti-patterns that occur in practice and are useful. Sharing it with practitioners may help them to avoid such anti-patterns.</p>
LabelGit: A dataset for software repositories classification using attributed dependency graphs
<p>A dataset for software repositories classification using attributed dependency graphs</p>
Package and Dependency Metadata for CZI Hackathon: Mapping the Impact of Research Software in Science
<p>A collection of useful datasets extracted from <a href="https://packages.ecosyste.ms">https://packages.ecosyste.ms</a> and <a href="https://repos.ecosyste.ms/">https://repos.ecosyste.ms</a> for use at the CZI Hackathon: Mapping the Impact of Research Software in Science.</p><p>All data is provided as NDJSON (new line delimited JSON), each line represents a valid JSON object, and they are separated by newline characters. There are <a href="https://pypi.org/project/ndjson/">python</a> and <a href="https://www.rdocumentation.org/packages/ndjson/versions/0.9.0/topics/stream_in">R</a> libraries for reading these files, or you can maually read each line and parse each line as a single JSON object.</p><p>Each ndjson file has been compressed with gzip (actual command: `tar -czvf`) to reduce download size, they expand to significantly bigger files after extraction.</p><h4>Package Data</h4><p>Package names from cran, bioconductor and pypi that have been parsed by the <a href="https://github.com/chanzuckerberg/software-mentions">software-mentions</a> project (data: <a href="https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c">https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c</a>) are collected together with their latest release at time of publishing along with the names of their dependencies, those dependency names have then also been recursively fetched with latest release and dependencies until the full list of transitive dependencies is included. </p><p>Note: This approach uses a simplified method of dependency resolution, always picking the latest version of each package rather than taking into account each dependencies specific version range requirements, this is primarily due to time constraints and allows all software ecosystems to be processed in the same way. A future improvement would be to use each package ecosystem's specific dependency resolution algorithm to compute the full transitive dependency tree for each mentioned software package.</p><h4>GitHub Data</h4><p>Two different approaches were taken for collecting data for referenced GitHub mentions:</p><p>1. `github.ndjson` is metadata for each repository from GitHub, including "manifest" files which are known files that contain dependency information for a project such as requirements.txt, DESCRIPTION and package.json, parsed using <a href="https://github.com/ecosyste-ms/bibliothecary">https://github.com/ecosyste-ms/bibliothecary</a>, which may include transitive dependencies that have been discovered in a `lockfile` within the repository.</p><p>2. `github_packages.ndjson` is metadata for each package that was found on any package manager that references the GitHub url as it's repository url/source/homepage, these packages, like the cran and pypi data above, include the latest release and their direct dependencies. There may be more than one package for each GitHub URL as it is a one to many relationship. `github_packages_with_transitive.ndjson` follows the same format but also includes the extra resolved transitive dependencies of all packages using the same approach as with cran and pypi data above with the same caveats. </p><p>There are also many more ecosystems referenced in these files than just cran, bioconductor and pypi, https://packages.ecosyste.ms provides a standardized metadata format for all of them to enable comparison and simplification of automation.</p><h4>Contact</h4><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: https://github.com/ecosyste-ms/packages/issues</p>
Figure 9. Experimental Page Rank dependency on Markov Chain length with balanced distribution-Study of a Random Navigation on the Web Using Software Simulation
<p>This paper explored different implementation choices for analyzing the most important<br> parameters about a web. Many researchers explored the use of new search engines for studying the<br> evolution of the web (Ntoulas, Cho and Olston, 2004). Another important research is realized about<br> the Link Structure Graph (LSG). The LSG captures a complete hyperlink structure from the web<br> and models link associations reflected in the page layout (Rodrigues, Milic-Frayling and Fortuna,<br> 2007). For further works ideas like extrapolation methods for accelerating page rank calculation can<br> be developed (Kamvar et al., 2003).</p>
Figure 7. Experimental Page Rank dependency on Markov Chain length-Study of a Random Navigation on the Web Using Software Simulation
<p>The next diagram proves that the values for Experimental Page Rank depend on the length<br> of the Markov Chain, while Algorithmic Page Rank remains constant.</p>
A Dependency Graph for 460,000 Papers and Their Software Mentions from the CZI Software Mentions Dataset
<p>Using the CZI Software Mentions Dataset and <a href="https://ecosyste.ms">ecosyste.ms</a> we create a graph of papers, their mentioned software, and recursive dependencies of each piece of software across 466,000 papers and three software registries (PyPI, CRAN, and BioConductor).</p>
Replication Package for "An Empirical Comparison of Dependency Network Evolution in Seven Software Packaging Ecosystems"
<p>This is the replication package for the article "An Empirical Comparison of Dependency Network Evolution in Seven Software Packaging Ecosystems" published in the Empirical Software Engineering journal.</p> <p>This package requires Python 3.5 and all the dependencies that are listed in "requirements.txt".<br> The notebooks (in "notebooks" folder) should be opened and executed with Jupyter.</p> <p>The notebooks require the graphs (in "graphs" folder) to be computed first. To do so, execute "helpers.py" with Python.<br> The graphs are built using the data provided by https://libraries.io under CC BY-SA<br> https://creativecommons.org/licenses/by-sa/4.0/<br> Those data can be found in the "data" folder.</p> <p> </p>
Open Source Software Package, Version and Dependency Metadata
<p>This dataset contains metadata about 7 million open source packages, versions and dependencies from 34 different software ecosystems in 4 different csv files exported from <a href="https://packages.ecosyste.ms/">https://packages.ecosyste.ms</a>. This can be seen as a spiritual successor to https://zenodo.org/records/3626071, up to date for 2023, including more packages and an expanded array of metadata fields.</p><p>Full copies of the postgresql database this data was extracted from are also available: <a href="https://packages.ecosyste.ms/open-data">https://packages.ecosyste.ms/open-data</a> </p><h4>Files</h4><p>Expanded csv sizes:</p><ul><li>packages-1.0.0-2023-10-17.csv - 2.6GB - 7,015,326 rows</li><li>packages_with_repository_fields-1.0.0-2023-10-17.csv - 4.3GB - 7,019,492 rows</li><li>versions-1.0.0-2023-10-17.csv - 14GB - 84,953,382 rows</li><li>dependencies2-1.0.0-2023-10-17.csv - 124GB - 976,944,543 rows</li></ul><p>Contact</p><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: <a href="https://github.com/ecosyste-ms/packages/issues">https://github.com/ecosyste-ms/packages/issues</a></p>
Artifacts of Paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems"
<p>This package contains all the artifacts (i.e. codes & datasets) we use in our paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems" accepted to FSE 2020.</p>
Dataset for Generative Model of Software Dependency Graphs
<p>Data set for the paper entitled "A Generative Model of Software Dependency Graphs to Better Understand Software Evolution".</p> <p>Available files are:</p> <ul> <li>Sources archives (102 MB),</li> <li>Extracted dependencies (3.5 MB) and</li> <li>Generated graphs (15 MB).</li> </ul>
Giving Back: Contributions Congruent to Library Dependency Changes in a Software Ecosystem
<p>Abstract: </p> <p>Widespread adoption of third-party libraries for contemporary software development has led to the creation of large inter-dependency networks, where sustainability issues of a single library can have widespread network effects. Maintainers of these libraries are often overworked, relying on the contributions of volunteers to sustain these libraries. To understand these contributions, in this work, we leverage socio-technical techniques to introduce and formalise dependency-contribution congruence (DC congruence) at both ecosystem and library level, i.e., to understand the degree and origins of contributions congruent to dependency changes, analyze whether they contribute to library dormancy (i.e., a lack of activity), and investigate similarities between these congruent contributions compared to typical contributions. We conduct a large-scale empirical study to measure the DC congruence for the NPM ecosystem using 1.7 million issues, 970 thousand pull requests (PR), and over 5.3 million commits belonging to 107,242 NPM libraries. We find that the most congruent contributions originate from contributors who can only submit (not commit) to both a client and library.<br> At the project level, we find that DC congruence shares an inverse relationship with the likelihood that a library becomes dormant, i.e., the lower the DC congruence, the more likely the project becomes dormant. Finally, by comparing source code of contributions, we find statistical differences in file path and added lines in source code of congruent contributions when compared to typical contributions. Our work has implications to encourage and sustain dependency change congruent contributions, especially to support library maintainers in sustaining their projects.</p> <p>Data Description:</p> <p>Datasets and source code related to (a) the ecosystem-level and package-level DC congruence results, (b) the metric data for our survival model analysis, and (c) the file path and source code similarities between contribution types.</p> <p>For each folder, it includes a text file to describe file details.</p>
Open Source Software Dependency Dataset
<p>Open Source Software Dependency Dataset</p>
[Dataset and Software] Altitude-dependent plasma parameter variations of synthetic EISCAT UHF and VHF incoherent scatter spectra calculated from TIE-GCM results
Open the record for dataset details and reuse information.
Artifacts of Paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems"
<p>This package contains all the artifacts (i.e. codes & datasets) we use in our paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems" accepted to FSE 2020.</p>
Artifacts of Paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems"
<p>This package contains all the artifacts (i.e. codes & datasets) we use in our paper "Understanding and Discovering Software Configuration Dependencies in Cloud and Datacenter Systems" accepted to FSE 2020.</p>
Dataset of the paper "Software Supply Chain Meets Large Language Models: Can Dependency-related Problems Be Solved?"
<p>This is the dataset of the paper "Software Supply Chain Meets Large Language Models: Can Dependency-related Problems Be Solved?"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.