Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “Software Knowledge”

Learn how ShareScore rates datasets ↗
zenodo44/100

Software Engineering Education Knowledge versus Industrial Needs

<p>Dataset of the research paper:&nbsp;<strong>Software Engineering Education Knowledge versus Industrial&nbsp;Needs</strong></p> <p><em>Contribution</em>: Determine and analyze the gap between software practitioners&rsquo; education outlined in the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and industrial needs pointed by Wikipedia articles referenced in Stack Overflow (SO) posts.<br> <em>Background</em>: Previous work has uncovered deficiencies in the coverage of computer fundamentals, people skills, software processes, and human-computer interaction, suggesting rebalancing.<br> <em>Research Questions</em>: 1) To what extent are developers&rsquo; needs, in terms of Wikipedia articles referenced in SO posts, covered by the SEEK knowledge units? 2) How does the popularity of Wikipedia articles relate to their SEEK coverage? 3) What areas of computing knowledge can be better covered by the SEEK knowledge units? 4) Why are Wikipedia articles covered by the SEEK knowledge units cited on SO?<br> <em>Methodology</em>: Wikipedia articles were systematically collected from SO posts. The most cited were manually mapped to the SEEK knowledge units, assessed according to their degree of coverage. Articles insufficiently covered by the SEEK were classified by hand using the 2012 ACM Computing Classification System. A sample of posts referencing sufficiently covered articles was manually analyzed. A survey was conducted on software practitioners to validate the study findings.<br> <em>Findings</em>: SEEK appears to cover sufficiently computer science fundamentals, software design and mathematical concepts, but less so areas like the World Wide Web, software engineering components, and computer graphics. Developers seek advice, best practices and explanations about software topics, and code review assistance. Future SEEK models and the computing education could dive deeper in information systems, design, testing, security, and soft skills.</p> <p>The following data files are included.</p> <ul> <li><strong>wikipedia_articles.csv</strong>: Wikipedia articles mapped to the knowledge units of the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and the first and second level categories of the 2012 ACM Computing Classification System (CCS).</li> <li> <p><strong>posts_analysis.csv</strong>: Stack Overflow post data and metadata.</p> </li> <li> <p><strong>posts_aggregated_codes.csv</strong>: The aggregated codes that resulted from the manual analysis of the Stack Overflow posts by grouping individual keywords assigned to the posts.</p> </li> <li> <p><strong>survey_questionnaire.csv</strong>:&nbsp;The final survey questionnaire.</p> </li> <li> <p><strong>survey_responses.csv</strong>:&nbsp;Anonymized responses of the final survey questionnaire. (E-mail addresses have been excluded for privacy reasons.)</p> </li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Analysis of the DLR Knowledge Exchange Workshop Series on Software Engineering

<p>This repository is used to analyze the workshops of the DLR internal workshop series on software<br> engineering. These workshops are two-day events of the DLR software engineering community and<br> focus on different main topics every year.</p>

openmit-licenseJun 2018View details →
zenodo36/100

Supporting Software Engineers in IT Security and Privacy through Automated Knowledge Discovery - Dataset

<h1>Supporting Software Engineers in IT Security and Privacy through Automated Knowledge Discovery - Dataset</h1> <h2>&nbsp;</h2> <h2>Dataset corresponding to&nbsp;<a href="https://doi.org/10.1145/3672608.3707798">https://doi.org/10.1145/3672608.3707798</a></h2> <div>This is the dataset corresponding to&nbsp;<a href="https://doi.org/10.1145/3672608.3707798">https://doi.org/10.1145/3672608.3707798</a> "Supporting Software Engineers in IT Security and Privacy through Automated Knowledge Discovery" with the goal of providing a systematic method to discover state-of-the-art security knowledge, focused on threats, measures, and properties from science and project literature with minimal manual effort, and providing software engineers with the awareness and means they need to apply the knowledge to their projects.</div> <p>&nbsp;</p> <h2>Structure of the files</h2> <div>The dataset contains the extraction prompt, the extraction results, and metadata.</div> <div>- <strong>ACM_IEEE_EU_Results.zip</strong> is an archive file with the metadata and extraction results with the following folder structure:&nbsp;</div> <div>&nbsp; &nbsp; - <strong>Meta/ </strong>contains the DOI, title, publication date, and pages (omitted for IEEE since it may contain intellectual property).</div> <div>&nbsp; &nbsp; - <strong>Extract/</strong> contains extraction results generated by the LLM in response to the extraction_prompt.txt file.</div> <div>&nbsp; &nbsp; - <strong>ExtractMeta/</strong> contains the metadata for the extraction, such as execution time.</div> <div>- <strong>extraction_prompt.txt</strong> is the prompt used to extract information from the science and project publications.</div>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.

<p>Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees &nbsp;among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Software Mentions Knowledge Graph

<p>Software plays an essential role in modern society. Its growth and evolution have given rise to a wide variety of pseudonyms to refer to the same software. This multiplicity of names can be confusing, complicating the accurate identification of software and its relationship to other programs and applications.<br> It is necessary to identify mentions of software in texts because they provide key information about its use and application in different contexts. Such identification makes it possible to establish a clearer link between programs, applications and their usefulness in various spheres. In addition, being able to group the different pseudonyms of a software under a single name facilitates its search and study, simplifying the acquisition of knowledge and the exchange of information between users and developers.<br> The term &quot;Alias&quot; refers to the name chosen to represent a group of pseudonyms that identify the same software tool. This alias acts as a common denominator, unifying the different ways in which a software tool can be mentioned in the scientific articles under analysis. From now on, we will refer to this unifying term as &quot;alias&quot; or &quot;group&quot;, while &quot;pseudonym&quot; will designate the various forms of mentions of the same software that have appeared in the scientific literature.<br> Therefore, the main objective of this project is the construction of a knowledge graph [1] that groups the pseudonyms of scientific software tools into a common alias or group. This grouping will be done by analyzing and classifying the mentions of such software in academic publications provided by the article &quot;CZ Software Mentions&quot;, published by the Chan Zuckerberg Initiative.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

SoSEN-KG: Knowledge Graph Dump v0 for SoSEn: Software Search Engine.

<p>SoSEn is a semantic search engine for scientific software. We index scientific software in a knowledge graph, which is used for search and understanding of the software. The graph `graph.ttl` contains information about the software, and `keywords.ttl` contains keyword information about the software. `graph.ttl` can be used alone, or it can be combined with `keywords.ttl` to facilitate tf-idf based keyword search.</p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Relevant Datasets and Software Used for Paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description"

<p>This repository contains relevant datasets and software&nbsp;used in a paper&nbsp;&quot;KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description&quot;. They are used to run the code of <em>KGML-xDTD&nbsp;</em>stored on <a href="https://github.com/chunyuma/KGML-xDTD">Github</a>&nbsp;and support the results of this paper.</p> <p><strong>About the datasets</strong></p> <p>1. <em>bkg_rtxkg2c_v2.7.3.tar.gz</em></p> <p>This tar.gz file contains three sub-folders: tsv_files, scripts, and relevant_dbs. The &quot;tsv_files&quot; sub-folder has the input files that the neo4j software uses. The &quot;scripts&quot; sub-folder contains a shell script with a relevant python script to construct&nbsp;the&nbsp;biomedical knowledge graph. The &quot;relevant_dbs&quot; sub-folder stores two auxiliary databases that <em>KGML-xDTD</em> needs to use.&nbsp;</p> <p>2. <em>indication_paths.yaml</em></p> <p>This file contains the <a href="https://sulab.github.io/DrugMechDB">DrugMechDB</a>&nbsp;MOA paths that we used to evaluate the predicted MOA paths by <em>KGML-xDTD.&nbsp;</em>It is downloaded from the official <a href="https://github.com/SuLab/DrugMechDB">GitHub repository</a> of DrugMechDB.</p> <p>3.&nbsp;<em>training_data.tar.gz</em></p> <p>This tar.gz file contains the processed training data of four data sources (e.g., <a href="https://mychem.info">MyChem</a>, <a href="https://lhncbc.nlm.nih.gov/ii/tools/SemRep_SemMedDB_SKR/SemMedDB_download.html">SemMedDB</a>, <a href="https://bioportal.bioontology.org/ontologies/NDFRT">NDF-RT</a>, <a href="https://unmtid-shinyapps.net/shiny/repodb/">RepoDB</a>) mentioned in the paper. These processed drug-disease pairs have been matched to the identifiers of biological entities used in our biomedical knowledge graph and respectively split into true positive (tp) sets and true negative (tn) sets. We also provide the names of these drug identifiers and disease identifiers under a sub-folder &quot;translated _to_name&quot;.</p> <p><strong>About the software</strong></p> <p><em>neo4j-community-3.5.26.tar.gz</em></p> <p>This tar.gz is the Neo4j community version 3.5.26 downloaded from <a href="https://neo4j.com/download-center/#community">Neo4j Download Center</a>. Although the&nbsp;newer versions are&nbsp;available, due to their big&nbsp;changes in the Neo4j setting that are not compatible with our scripts on Github, we provide the version that we used in our research. If you would like to use the newer version, modifications to our script will be required to import the biomedical knowledge graph into your local Neo4j database with the new setting.</p>

opencc-zeroJan 2023View details →
zenodo28/100

Supplementary Material Requirements Study - Identifying Necessary Green Coding Knowledge for Young Professionals Starting their Careers in the Software Industry Full Audio Transcript with transcribers Notes

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo28/100

Replication package for the paper - "Balanced Knowledge Distribution among Software Development Teams - Observations from Open-Source and Closed-Source Software Development"

<p>Replication package for the&nbsp;paper - &quot;Balanced Knowledge Distribution among Software Development Teams - Observations from Open-Source and Closed-Source Software Development&quot;.</p>

opencc-by-4.0Mar 2022View details →
zenodo20/100

Why should software companies use models to improve knowledge management? A survey in Brazil

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo20/100

Updating a systematic literature review on knowledge management diagnostics in software development organizations

<p>Material to enable future replications.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record