Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

61

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

61 results for “Open Source Software”

Learn how ShareScore rates datasets ↗
zenodo36/100

Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Matching articles with categories)

<p>This document includes which primary study falls into which category with respect to the RQs in the following study: &ldquo;Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review&rdquo;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Matching articles with categories)

<p>This document includes which primary study falls into which category with respect to the RQs in the following study: &ldquo;Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review&rdquo;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Open Source Software Sustainability

<p>The reproducible code and dataset along with the paper submission.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

SoK: Taxonomy of Attacks on Open-Source Software Supply Chains - Visualization Tool Screenshots & Selected Papers

<p>This artifact complements the paper &quot;SoK: Taxonomy of Attacks on Open-Source Software Supply Chains&quot;, submitted at IEEE S&amp;P 2023.</p> <p>The papers selected during the Systematic Literature Review (SLR) are presented in the CSV file.</p> <p>This screenshots display the main features of the visualization tool that allows to explore the taxonomy of attacks on OSS supply chains, as well as the related safeguards and the selected references.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Open Source Software in Data Science

<p>This upload includes an anonymized data set of a survey first launched in 2022. The survey has been revised since. The data set. however, contains answers of the first launch.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Underrepresented Groups in Open Source Software Development

<p>This repository contains the artifacts generated when researching minority groups in open source software development, which aimed to analyze knowledge about minority groups in OSS projects. This set of artifacts consists of two main components:</p> <p><strong>Research Protocol (PDF):</strong></p> <p>The protocol in PDF format provides a detailed overview of the research design, methodology, and objectives of the study. It describes the scope of the research, research questions, and data collection approach. This protocol is a fundamental reference for researchers interested in understanding how the study was designed.</p> <p><strong>Search Data (CSV):</strong></p> <p>The CSV file contains the raw data collected during the research. The data includes results returned by each database, articles considered in each phase of the study, and the set of seed articles. The XLSX spreadsheet is a source for analyzing the data used in the research.</p> <p>This repository aims to promote a more transparent understanding of the process that guided this research. Researchers and the open source developer community can utilize this dataset for academic studies, replications, and informed decision-making to promote equity and representation in open source projects.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Open-Source Software Product Line Extraction Processes: the ArgoUML-SPL and Phaser Cases

<p>Collection of datasets and analysis scripts supporting the information provided in the text.</p> <p>There are two compressed files, one for the ArgoUML data and one for the Phaser data. Each compressed file contains a README describing important information.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Open-source Software Governance Documentation Dataset on GitHub

<p>This dataset contains 710 GitHub-hosted OSS projects, which contain a governance file in the root directory of the project. It also contains commits, issues, and comments on each project.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Proof of Concept database with inputs and outputs of the Master thesis: Analyzing Software Delivery Performance behavior in popular Open Source Software Projects on a Release timeline basis through delivery metrics

<p>The software has become one of the main assets to deliver services today. Thus, software delivery has been dealing with a competitive and dynamic environment where the demand for faster and more assertive deliverables, called here Releases, only increases. Agile development methods emerged helping to accelerate software delivery, embracing industry and open source community. Since then, the software delivery frequency has expanded and improved bringing more adopters of rapid release cycles to reduce their time-to-market. However, using only rapid releases can not be enough as measuring software delivery can answer essential questions, like how software delivery is happening and how it should be. Some approaches for measuring software delivery appeared such as Software Delivery Performance (SDP) where software delivery is measured as a consequence of capabilities evolution. Popularity in Open Source Software Projects (OSSP) means that a project is mature enough in the community to fit the software demand and, therefore, is likely to be ready to be measured through a software delivery approach like SDP. In light of it, this work offers means to analyze SDP behavior in popular Open Source Software Projects on a Release timeline basis through delivery metrics. The results demonstrated that popularity is efficient filtering, as it improves the OSSP delivery, supporting the work&#39;s reliability and accuracy. The source code and methodology are published as a replication package to encourage reproducibility and future research.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Free Open Source Communities Sustainability: Does It Make a Difference in Software Quality?

<p>&nbsp;Free and Open Source Software (FOSS) communities&#39; sustainability is important to maintain and promote access to technology and innovation. While considerable effort has been invested on the topic of FOSS sustainability, limited attention has been given to its impact on community outcomes, e.g., software quality. We conducted an empirical study to examine the influence of FOSS sustainability on software quality. We used projects&#39; data sourced from the Apache Software Foundation Incubator. A total of 236 projects were originally selected. Our final list included a total of 217 projects after applying our exclusion criteria. We used Bayesian data analysis, which incorporates probability distributions to represent the regression coefficients and intercepts. Our findings suggest that our selected sustainability metrics do not significantly affect defect density or code coverage. However, we observed a positive impact of community age on specific code quality metrics, such as risk complexity, number of very large files, and code duplication percentage. Interestingly, our findings show that even when communities are experiencing sustainability, certain code quality metrics are negatively impacted. This implies that code quality practices are not consistently linked to sustainability, and defect management and prevention may be prioritized over the former.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Open Source Software Package, Version and Dependency Metadata

<p>This dataset contains metadata about 7 million open source packages, versions and dependencies from 34 different software ecosystems in 4 different csv files exported from <a href="https://packages.ecosyste.ms/">https://packages.ecosyste.ms</a>. This can be seen as a spiritual successor to https://zenodo.org/records/3626071, up to date for 2023, including more packages and an expanded array of metadata fields.</p><p>Full copies of the postgresql database this data was extracted from are also available: <a href="https://packages.ecosyste.ms/open-data">https://packages.ecosyste.ms/open-data</a>&nbsp;</p><h4>Files</h4><p>Expanded csv sizes:</p><ul><li>packages-1.0.0-2023-10-17.csv - 2.6GB - 7,015,326 rows</li><li>packages_with_repository_fields-1.0.0-2023-10-17.csv - 4.3GB - 7,019,492 rows</li><li>versions-1.0.0-2023-10-17.csv - 14GB - 84,953,382 rows</li><li>dependencies2-1.0.0-2023-10-17.csv - 124GB - 976,944,543 rows</li></ul><p>Contact</p><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: <a href="https://github.com/ecosyste-ms/packages/issues">https://github.com/ecosyste-ms/packages/issues</a></p>

opencc-by-sa-4.0Oct 2023View details →
zenodo32/100

Dataset and scripts for "Sentiment Analysis over Collaborative Relationships in Open Source Software Projects"

<p>Dataset and scripts for &quot;Sentiment Analysis over Collaborative Relationships in Open Source Software Projects&quot;.</p> <p>README is included in the files</p>

opencc-by-4.0Jul 2019View details →
zenodo32/100

Dataset and Code for "Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects"

<p>Dataset and Code for &quot;Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects&quot; with README included</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Refactoring Code Smells in Open Source Projects: A Hands-on Approach to Teaching Software Maintenance

<p>Code smells are suboptimal code structures that can undermine software quality and maintainability. On the one&nbsp;hand, software engineers commonly apply refactoring techniques to address these deficiencies and improve internal quality attributes. On the other hand, when performed manually and without discipline, refactoring can lead to&nbsp;code degradation. Despite its importance, refactoring and code smells are rarely explored in depth in undergraduate&nbsp;computing courses, which can be reflected in industry practices. To address this gap, this paper presents a hands-on approach to teaching code smell refactoring through contributions to Open Source Software (OSS) projects, an&nbsp;environment where developers with diverse skill levels collaborate, and maintaining code quality is particularly&nbsp;challenging. Code smells accumulate over time in such scenarios, hindering software evolution and collaboration.&nbsp;Our study in two undergraduate Software Quality and Software Maintenance courses expands on previous findings&nbsp;by incorporating an in-depth analysis of students&rsquo; learning experiences. The results indicate that: (i) students rec-&nbsp;ognized improvements in code quality after refactoring; (ii) they identified strong connections between refactoring,&nbsp;testing, and debugging; (iii) their confidence decreased when refactoring required changes across multiple files; (iv)&nbsp;code complexity posed a significant challenge to refactoring; (v) students&rsquo; choices of refactoring techniques were&nbsp;influenced by project structure and personal preferences, often combining multiple techniques to address a single&nbsp;smell; (vi) in some cases, refactoring introduced new code smells; (vii) the longest refactoring efforts were also&nbsp;the most likely to reintroduce code smells; (viii) contributing to OSS projects improved students&rsquo; programming&nbsp;skills and fostered a sense of professional growth; (ix) students faced challenges in understanding OSS contribution&nbsp;processes, particularly regarding issue resolution, adherence to contribution guidelines, and responding to maintainer feedback; (x) automated checks and review workflows varied across projects, affecting students&rsquo; ability to&nbsp;submit successful contributions; and (xi) despite these challenges, engagement with OSS enabled students to gain&nbsp;practical experience in collaborative software development. Our findings offer valuable insights for software engineering educators seeking to integrate refactoring practices into coursework while leveraging OSS contributions as&nbsp;an educational tool.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Revisiting reopened bugs in open source software systems

<p>Reopened bugs can degrade the overall quality of a software system since they require unnecessary rework by developers. Moreover, reopened bugs also lead to a loss of trust in the end-users regarding the quality of the software. Thus, predicting bugs that might be reopened could be extremely helpful for software developers to avoid rework. Prior studies on reopened bug prediction focus only on three open source projects (i.e., Apache, Eclipse, and OpenOffice) to generate insights. We observe that one out of the three projects (i.e., Apache) has a data leak issue -- the bug status of \textit{reopened} was included as training data to predict reopened bugs. In addition, prior studies used an outdated prediction model pipeline (i.e., with old techniques for constructing a prediction model) to predict reopened bugs. Therefore, we revisit the reopened bugs study on a large scale dataset consisting of 47 projects tracked by JIRA using the modern techniques such as SMOTE, permutation importance together with 7 different machine learning models. We study the reopened bugs using a mixed methods approach (i.e., both quantitative and qualitative study). We find that: 1) After using an updated reopened bug prediction model pipeline, only 34\% projects give an acceptable performance with AUC $\geqslant$ 0.7. 2) There are four major reasons for a bug getting reopened, that is, technical (i.e., patch/integration issues), documentation, human (i.e., due to incorrect bug assessment), and reasons not shown in the bug reports. 3) In projects with an acceptable AUC, 94\% of the reopened bugs are due to patch issues (i.e., the usage of an incorrect patch) identified before bug reopening. Our study revisits reopened bugs and provides new insights into developer&#39;s bug reopening activities.<br> &nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Replication package for "Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study"

<p>Replication package for the paper&nbsp;<em>Blended Modeling in Commercial and Open-source Model-Driven Software Engineering Tools: A Systematic Study</em>.</p> <p>Protocol</p> <ul> <li><code>/01-protocol/protocol.pdf</code></li> </ul> <p>Data &amp; analysis scripts</p> <p>This replication package is structured as follows:</p> <ul> <li><code>/02-search</code>&nbsp;- Detailed data on the&nbsp;<code>/academic</code>&nbsp;and&nbsp;<code>/grey literature</code>&nbsp;search.</li> <li><code>/03-tools</code>&nbsp;- Identified tools and inclusion/exclusion decisions.</li> <li><code>/04-classification_schema</code>&nbsp;- Classification framework and the corresponding data extraction form.</li> <li><code>/05-data</code>&nbsp;- Clean data in a processable form.</li> <li><code>/06-analysis</code>&nbsp;- Analysis scripts and results.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Open-source release of tensor-network software

<p>This is a package<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-SBmodel-85d3b1b42ae4f41239f8975ec68008b2">1</a></sup>&nbsp;for calculating FLUCTUATIONS of heat transfer in the Spin-Boson model<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-PRX2020-85d3b1b42ae4f41239f8975ec68008b2">2</a></sup>&nbsp;using the&nbsp;<strong>Time Evolving Density matrices using Orthogonal Polynomial Algorithm (<em>TEDOPA</em>)</strong><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Prior2010-85d3b1b42ae4f41239f8975ec68008b2">3</a></sup><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Chin2010-85d3b1b42ae4f41239f8975ec68008b2">4</a></sup>.</p> <p>We employ the&nbsp;<strong>Thermofield-based chain-mapping approach for open quantum systems</strong><sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-PRA2015-85d3b1b42ae4f41239f8975ec68008b2">5</a></sup>&nbsp;that enables us to use a vacuum initial matrix product state (pure) for the environment instead of a thermal state (mixed), thereby speeding up the computation greatly.</p> <p>In this package, we use the ITensor library<sup><a href="https://github.com/aspects-quantum/TFlucn_tedopa#user-content-fn-Itensor-85d3b1b42ae4f41239f8975ec68008b2">6</a></sup>&nbsp;in Julia for tensor network manipulations. &nbsp;</p> <p>This package uses&nbsp;<strong>julia = "1.8.2"</strong> version.</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries in Software Package Ecosystems

<p>Replication package for the paper "Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries in Software Package Ecosystems "</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Example StoX projects used in the publication "StoX - an open source software for marine survey analyses"

<p><strong>Datasets for the paper &quot;StoX - an open source software for marine survey analyses&quot;</strong></p> <p>Two example StoX projects&nbsp;provided as zip files, including input data and all user settings. These data were used to produce Figure 3 of the paper. The R-script used to produce the figure is included in the supporting information&nbsp;S5 of the paper:</p> <p>1.&nbsp;Example_Barents_Sea_cod_swept_area_survey_1999.zip</p> <p>2.&nbsp;Example_Nordic_Seas_herring_acoustic-trawl_survey_2016.zip</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo32/100

Dataset and Code for "Mining Micro-Patterns of Issue Resolution Processes for Open Source Software Projects"

<p>Dataset and Code for &quot;Mining Micro-Patterns of Issue Resolution Processes for Open Source Software Projects&quot; with README included</p>

opencc-by-4.0Oct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record