Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Software Sharing”
The Role of Informal Communication in Building Shared Understanding of Non-Functional Requirements in Remote Continuous Software Engineering
<p><strong>Study Information</strong></p> <p>We conducted an ethnography-informed case study of a remote software organization that adopts CSE practices to explore how the organization builds a shared understanding of NFRs. Our study uses semi-structured interviews with a period of observations to answer the following research questions:</p> <p> </p> <ol> <li> <p>How does a remote software organization that adopts CSE practices reach a shared understanding of NFRs?</p> </li> <li> <p>What are the limitations to the shared understanding of NFRs in a remote software organization that adopts CSE practices?</p> </li> <li> <p>What organizational practices for remote collaboration supported a shared understanding of NFRs?</p> </li> </ol> <p> </p> <p>In our study, we refer to our partner organization as Alpha. We used ethnography-informed methods to study Alpha's practices and processes and how they approach a shared understanding of NFRs in their product development. </p> <p> </p> <p><strong>Data Analysis</strong></p> <p>We performed a qualitative study through semi-structured interviews and observations. We use the open, axial and selective coding approach from grounded theory [1] to create our codebook, which informed the results and discussion of our study. Two independent coders held agreement sessions to discuss the codes, consolidate the codes and calculate the inter-rater reliability using the Cohen Kappa's coefficient for measuring observer agreement for categorical data [2]. </p> <p> </p> <p><strong>Artifact Descriptions</strong></p> <p>Our replication package contains three artifacts:</p> <p>1. Codebook.csv: The codebook contains rows for the list of codes used, including the code name and the description of the codes. The codes are the final set of themes derived during the thematic analysis of the interview responses. For example, 'Gaps in communication' means when interview participants describe miscommunications due to team members making assumptions about a project/process or having unclear expectations for a project.</p> <p>2. kappa-scores.csv: This contains the associated kappa values for each round of inter-rater agreement sessions. For each agreement session, the Cohen Kappa's coefficient was calculated from the number of agreements and disagreements of codes within one or two interview transcripts. The Kappa values represent the level of agreement ranging from 0 to 1, where > 0.6 represents substantial agreement. </p> <p>3. Interview-questions.csv: This contains the interview questions used in the semi-structured interviews. Some of the interview questions varied depending on the interviewee’s role, experience and the flow of the interviews.</p> <p><strong> </strong></p> <p><strong>Usefulness</strong></p> <p>We recognize that the value and usefulness of our replication package are yet-to-be-determined. In the interest of transparency of open science, we published our artifacts. We hope that these artifacts are useful to either replicate our findings or to further analyze them to produce other enlightening results.</p> <p><strong> </strong></p> <p><strong>References</strong></p> <p>1. Rashina Hoda, James Noble, and Stuart Marshall. "Grounded theory for geeks". In: Proceedings of the 18th conference on pattern languages of programs. 2011, pp. 1–17.</p> <p>2. J Richard Landis and Gary G Koch. "The measurement of observer agreement for categorical data". In: biometrics (1977), pp. 159–174.</p> <p><strong> </strong></p> <p> </p>
Credit for software creators: An adapted illustration from The Turing Way: Shared under CC-BY 4.0 for reuse
<p>Illustration adapted from <strong>The Turing Way Community, & Scriberia. (2023). Illustrations from The Turing Way: Shared under CC-BY 4.0 for reuse. Zenodo. <a href="https://doi.org/10.5281/zenodo.8169292">https://doi.org/10.5281/zenodo.8169292</a></strong>.</p> <p>The illustration has been adapted by Stephan Druskat (subsumed as co-author in <em>The Turing Way Community</em>):</p> <ul> <li>The speech bubble has been changed to read "More credit" instead of "More credit<strong>s</strong>". (One character removed and alignment adapted.)</li> <li>The inscription on the purse has been changed to read "Credit" instead of "Credit<strong>s</strong>". (One character removed and alignment adapted.)</li> </ul>
Manual classification of dataset and software mention contexts as used, created or shared
<p>This dataset consists of two JSON files containing:</p> <p>- software-contexts.json: 49195 sentences from the SoMeSci and Softcite datasets with 3162 software mentions and with added manual indication if the mentioned software is used, created and/or shared,</p> <p>- dataset-software-extra-contexts.json: an additional set of 481 sentences with dataset or software mentions and with added manual indication if the mentioned software is used, created and/or shared.</p> <p>Softcite dataset: <a href="https://zenodo.org/record/7995565">https://zenodo.org/record/7995565</a></p> <p>SoMeSci dataset: <a href="https://zenodo.org/record/4968738">https://zenodo.org/record/4968738</a></p>
Data: Researcher Perspectives on the Use and Sharing of Software
Open the record for dataset details and reuse information.
Artfiact for the Proc. ACM Softw. Eng. article "Sharing Software-Evolution Datasets: Practices, Challenges, and Recommendations."
<p>This dataset is a collection of all notes taken for the article:</p> <p> David Broneske, Sebastian Kittan, and Jacob Krüger:<br> Sharing Software-Evolution Datasets: Practices, Challenges, and Recommendations. <br> Proc. ACM Softw. Eng. 1, FSE, 2024.<br> https://doi.org/10.1145/3660798</p> <p>Please refer to the readme for a description of the files involved in the zip file.</p>
The Lack of Shared Understanding of Non-Functional Requirements in Continuous Software Engineering: Accidental or Essential?
<p>Study Information</p> <p>We conducted a case study of three small organizations scaling up continuous software engineering to further understand and identify factors that contribute to lack of shared understanding of non-functional requirements (NFRs), and its relationship to rework. To conceptualize lack of shared understanding of NFRs we traced it to rework development tasks. Seeking to shed light on the complex relationship between shared understanding, rework, and CSE, our study examined forty-one NFR-related development tasks identified as rework and was driven by the following research questions:</p> <ol> <li>What contributes to lack of shared understanding of NFRs?</li> <li>Which NFRs are most associated with a lack of shared understanding?</li> <li>What amount of a lack of shared understanding of NFRs is accidental versus essential?</li> </ol> <p>To answer our research questions, we conducted a multi-case study using a mixed-methods approach in collaboration with three independent organizations using qualitative and immersive techniques. For this study our organizations are referred to as Alpha, Beta, and Gamma. We performed a preliminary study of each organization to build a context for each organization. We then collected data from project task management repositories and analyzed the tasks to uncover 41 NFR-related software development tasks as rework due to a lack of shared understanding of NFRs.</p> <p>Data Analysis</p> <p>We performed qualitative analysis through focus groups at each organization. We used an open-coding [1] approach to develop a purely inductive codebook, which minimizes a coder's ability to force a bias of any particular hypothesis. In the initial coding phase, a transcript was independently coded by two coders, after which an agreement session was held to discuss the codes, consolidate the codebook, and to calculate Cohen's kappa coefficient (using sklearn.metrics cohen_kappa_score). We continued coding in pairs until our inter-rater reliability met substantial agreement. After which we individually coded the remaining transcripts followed by an expert rater reviewer.</p> <p>This data represents two artifacts from our research in evaluating the complex relationship between shared understanding, non-functional requirements, continuous software engineering, and rework. It includes the resulting thematic synthesis codebook and inter-rater kappa values.</p> <p>Artifact Descriptions</p> <p>Our replication package contains two artifacts:<br> 1. Codebook.csv: The codebook itself contains a row for each code used (48 in total), including the code name, a brief description of the code, which round of coding that code was introduced, the total number of tasks that code appeared in (at least once), the number of tasks that code appeared in for each organization (Alpha, Beta, and Gamma), the total number of occurrences across all tasks, the number of occurrences across across each organization (Alpha, Beta, and Gamma), and the number of occurrences for each task.</p> <p>For example, the code 'BusinessContext' was used when 'Talking about information from the business side of the organization' and appeared in the first transcript we coded (1). The 'BusinessContext' code appeared in 19/41 tasks (10 at Alpha, 8 at Beta, and 1 at Gamma). Furthermore, the 'BusinessContext' code was used a total of 72 times (43 at Alpha, 22 at Beta, and 7 at Gamma). The remaining 41 columns represent the number of occurrences for each task, e.g. it appeared 3 times in task A-2 (Alpha's second task).</p> <p>2. kappa-values.csv: Contains the interview rounds (in order of coding) and the associated kappa values calculated for our inter-rater coding agreement.</p> <p>Usefulness</p> <p>While we recognize that the value and usefulness of our replication package is yet to-be-determined, in the interest of transparency of open science we published our artifacts. In light of this, we hope that these artifacts are useful to either replicate our findings or to further analyze to produce other enlightening results.</p> <p>References</p> <ol> <li>J. M. Corbin and A. Strauss, “Grounded theory research: Procedures, canons, and evaluative criteria,”Qualitative sociology, vol. 13, no. 1,pp. 3–21, 1990.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.