Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,916

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,916 results for “software,”

Learn how ShareScore rates datasets ↗
zenodo48/100

The potential of low-cost UAVs and open-source photogrammetry software for high-resolution monitoring of alpine glaciers: A case study from the Kanderfirn (Swiss Alps)

<p>This dataset contains high-resolution orthophotos (5 x 5 cm) and digital surface models (25 x 25 cm) of the Kandernfirn Glacier located in the Swiss Alps. Aerial images were aquired with a self-developed fixed-wing Unmanned Aerial Vehicle during ten surveys&nbsp;on five different days in 2017 and 2018. The open-source photogrammetry software OpenDroneMap (version 0.4.1) was used for image processing.</p> <p>The orthophotos and digital surface models were validated through dGNSS point measurements of ground control points. Please refer to the corresponding paper for information on the horizontal and vertical accuracy of the files.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Imalys: ESIS-Software-Tools – Tutorial Data

<p>Imalys is a <a href="https://codebase.helmholtz.cloud/esis/Imalys">software library</a> in the <a href="https://doi.org/10.3390/rs16071139">ESIS project</a>. The Imalys library includes a <a href="https://codebase.helmholtz.cloud/esis/Imalys/tutorial">tutorial</a> that shows how image and vector data can be analyzed and transformed into new products. The examples in the tutorial refer to the sample data in <strong>Imalys_tutorial_data.zip</strong>.</p> <p>The software repository with source code, binaries, documents and the tutorials is available on <a href="https://codebase.helmholtz.cloud/esis/Imalys">GitLab</a> &nbsp;and in an earlier version on <a href="https://github.com/c7sepe2/Imalys_ESIS-Software-Tools">GitHub</a>.</p> <p>In the <a href="https://doi.org/10.3390/rs16071139">ESIS project</a> we are trying to put environmental indicators on a well-defined and reproducible basis. The ESIS software library "Imalys" is supposed to generate the remote sensing products defined for ESIS. Landscape diversity, change and different landuse types can be analyzed in time and space. Landuse borders can be delineated and typical landscape structures can be characterized by a self adjusting process.</p> <p><strong>Acknowledgements:</strong></p> <p>We thank the Helmholtz Association and the Federal Ministry of Education and Research (BMBF) for supporting the DataHub Initiative of the Research Field Earth and Environment. The DataHub enables an overarching and comprehensive research data management, following FAIR principles, for all topics in the Program Changing Earth &ndash; Sustaining our Future.</p> <p>Parts of the work was funded by the Federal Ministry for Digital and Transport of Germany (BMDV, grant no. 45KI19D041, Artificial Intelligence and Mobility - AIAMO project).</p>

opengpl-3.0-or-laterMay 2024View details →
zenodo48/100

Software Contributor Roles Crosswalk

<p>Initial version of a crosswalk of software contributor roles taxonomies, models, schemes, etc.</p> <p>Created during a <a href="https://software.ac.uk/cw23">Collaborations Workshop 2023</a> hack activity.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

Software Architecture Assessment for Sustainability: A Case Study

<h3>Replication Package: Software Architecture Assessment for Sustainability: A Case Study</h3> <p><br>This repository contains the supplementary material to support the paper published at the International Conference on Software Architecture (ECSA) 2024 titled, "Software Architecture Assessment for Sustainability: A Case Study". This repository can be used to replicate the study and carry out a Software Architecture Evaluation of other software systems.<br><br>The online version can be browsed on the linked <a href="https://github.com/S2-group/rep-pkg-ecsa-2024-SA-assessment-for-sustainability-a-case-study/tree/v1.0.0" target="_blank" rel="noopener">Github Repository</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Data and Software for: 'The Radius of the High-mass Pulsar PSR J0740+6620 with 3.6 yr of NICER Data'

<p>Posterior sample files associated with the publication "The Radius of the High-mass Pulsar PSR J0740+6620 with 3.6 yr of NICER Data" by Salmi et al. (2024; <a href="https://doi.org/10.48550/arXiv.2406.14466">arXiv.2406.14466</a>; <a href="https://doi.org/10.3847/1538-4357/ad5f1f">https://doi.org/10.3847/1538-4357/ad5f1f</a>).</p> <p>Also included are: the data products; the numeric model files including the telescope calibration products; model modules in the Python language using the X-PSI framework; and Jupyter analysis notebooks.</p> <p>Please refer to the README for detailed information.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Dataset for: Evaluating phylogenetic methods for quantifying risks and opportunities presented by forks in open source software (master dissertation).

<p>This is the data for my master dissertation [1]. If you wish to get a copy, download it from Zenodo and open docs/master.pdf.</p> <p>Data acquisition and encoding techniques are described in paragraph 3.1.1 (table 3.1).</p> <p>The data is described in more detail in paragraph 4.1 (table 4.2).</p> <p>* fork1_all.csv: MySQL server / MariaDB server<br> * fork2_all.csv: Linux kernel / Android kernel<br> * fork3_all.csv: Apache OpenOffice / LibreOffice</p> <p>==Cite==<br> [1] A. Ortiz-Troncoso. Evaluating phylogenetic methods for quantifying risks and opportunities presented<br> by forks in open source software (master dissertation). Zenodo, 2018. doi: http://doi.org/10.5281/zenodo.1158292</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Survey of participant experience in workshop for testing IGP software setup: supplemental dataset for SimAUD 2018

<p>This is a supplemental dataset for a&nbsp;SimAUD 2018 paper.&nbsp;For the context of the&nbsp;dataset, plots, and description text given here, please refer to the paper:</p> <blockquote> <p><strong>Heinrich, M.K., Zahadat, P., Harding, J., et al.&nbsp;Using interactive evolution to design behaviors for non-deterministic self-organized construction. In <em>Proc. of SimAUD</em> (2018). <em>In print</em>.</strong></p> </blockquote> <p>These survey results <em><strong>(see attached file dataset_survey-responses)</strong></em>, are regarding the experience of participants in a workshop testing&nbsp;the <em>Integrated Growth Projection</em>&nbsp;software setup, including an implementation of the <em>Vascular Morphogenesis Controller</em>, and the Interactive Evolution software <em>Biomorpher</em>.</p> <p>The full-time one-week workshop was held as part of the normal coursework of the Master&#39;s degree program&nbsp;<em>CITAstudio: Computation in Architecture</em>, in&nbsp;the Institute of Architecture and Technology, at [KADK] The Royal Danish Academy, School of Architecture, Copenhagen, Denmark. It was part of the first semester of the 2017-2018 school year. Workshop participants were current Master&#39;s students in the&nbsp;<em>CITAstudio&nbsp;</em>program. The workshop teaching was led by Mary Katherine Heinrich and Phil Ayres, with guest teaching by Payam Zahadat and John Harding, overall program teaching supervision by Paul Nicholas, and teaching assistance by&nbsp;Sebastian Gatz.</p> <p><strong>Survey method:</strong></p> <p>The workshop participants gave survey responses anonymously.</p> <p>Survey&nbsp;responses were collected via Google Forms (https://www.google.com/forms/about/).&nbsp;At the start of the survey, participants gave permissions for use and publication, and verified that they participated in the workshop and had not previously taken the survey.&nbsp;The platform discourages duplicate responses by requiring an email sign-in (which is not visible to the surveyor).</p> <p>Although workshop participants gave permission for survey results to be published before taking the survey, the&nbsp;participants were unaware of the specific intended context and purpose of publishing, prior to taking the survey. Authors of the related&nbsp;paper who were workshop participants had no contact with the process of survey preparation, analysis of its results, or writing of related paper sections.&nbsp;</p> <p>There were 26 workshop participants. Participants were architects or architectural designers.&nbsp;</p> <p>Participants were asked about 1) their prior experience, 2) their understanding of topics before and after the workshop, 3) the helpfulness of specific software aspects for their understanding and their project work, and 4) their likelihood to use specific software aspects in the future.</p> <p>In addition to looking at the full surveyed group, we compare experience sub-groups. Participants select relevant tasks that they have previously completed, from a provided list. They are placed in the <em>Less Experience</em>&nbsp;sub-group if they select one or no tasks, and in the <em>More Experience</em>&nbsp;sub-group if they select two or more.&nbsp;</p> <p><strong>Survey results:</strong></p> <p>Close to two-thirds of workshop participants submitted survey responses (16 of 26, or 61.5%), with at least two respondents per group.&nbsp;One respondent indicated workshop absence; their responses were removed. One respondent indicated that they did not understand two questions, so those two responses were removed. All responses were submitted within 18 days of workshop end.</p> <p>Attached file:<em><strong> Plot_1</strong></em>, caption:</p> <blockquote> <p>Plot 1: <em>Participants&#39; scoring of their understanding of the topics &quot;self-organization&quot; and &quot;Interactive Evolution&quot; respectively, comparing scores before and after the workshop.</em></p> </blockquote> <p>Attached file:<em><strong>&nbsp;Plot_2</strong></em>, caption:</p> <blockquote> <p>Plot 2: <em>(Left) Participants&#39; scoring of their likelihood to use certain aspects of the software setup again, if they were to design a non-deterministic self-organizing behavior, and (right) participants&#39; indications of the helpfulness of those same software aspects.</em></p> </blockquote> <p>Responses regarding understanding <em><strong>(see attached file,&nbsp;Plot_1)</strong></em>&nbsp;give evidence that the <em>Integrated Growth Projection</em> software&nbsp;setup helped participants of both experience levels improve their understanding of related topics. Those with less prior experience improved their understanding more than others, and understanding of &quot;Interactive Evolution&quot; improved slightly more than understanding of &quot;self-organization.&quot;&nbsp;</p> <p>Responses regarding the usefulness of certain software aspects <em><strong>(see attached file,&nbsp;Plot_2)</strong></em> give evidence that: 1) Interactive Evolution helped participants to understand and design a non-deterministic self-organizing behavior <em><strong>(see Plot_2, a)</strong></em>; 2) visualization of the environment and simultaneous viewing of multiple results helped them to understand and design such behaviors <em><strong>(see Plot_2, b and c)</strong></em>; and 3) the <em>Integrated Growth Projection</em>&#39;s features of environment visualization and simultaneous results <em>inside</em>&nbsp;the artificial selection preview windows of the IE setup helped them to evolve behaviors to solve their chosen tasks<em> <strong>(see Plot_2, d&nbsp;and e)</strong></em>.&nbsp;</p> <p>____________________________________</p> <p>The research&nbsp;work involved here is part of EU project<em> flora robotica</em>.<br> <a href="http://www.florarobotica.eu/">http://www.florarobotica.eu/</a><br> Project<em> flora robotica</em>&nbsp;has received funding from the European Union&#39;s Horizon 2020 research and innovation program under the FET grant agreement, no. 640959.</p>

opencc-by-sa-4.0Mar 2018View details →
zenodo48/100

Lookup table data files for MOABS software

<p>These three lookup table data files are required for MOABS software to speed up computation. See more at&nbsp;<a href="https://github.com/sunnyisgalaxy/moabs">https://github.com/sunnyisgalaxy/moabs</a></p>

opencc-by-4.0Jul 2019View details →
Figshare48/100

Checklists for Software Citation: what you need to know

<p>The FORCE11 Software Citation Implementation working group has been developing practical guidance in the form of checklists, which are aimed at authors, reviewers, developers, editors and publishers. The checklists help these audiences to implement software citation principles in their workflows. This talk will give a quick introduction to these checklists and explain how they can be used to improve the practice of open research by giving credit for software.</p>

opencc-by-4.0Oct 2019View details →
zenodo48/100

GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.

<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of&nbsp;<a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username &nbsp;</td> </tr> <tr> <td>html_url &nbsp;</td> <td>URL of the GitHub profile &nbsp;</td> </tr> <tr> <td>type &nbsp; &nbsp;</td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization &nbsp; &nbsp;</td> </tr> </tbody> </table> <p><strong>GitHub repositories&nbsp;<em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name&nbsp;</strong></td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>html_url &nbsp;</td> <td>URL link to the repository &nbsp;</td> </tr> <tr> <td>description</td> <td>GitHub project description &nbsp;</td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project &nbsp;</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated&nbsp;<em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name&nbsp;</strong></td> <td><strong>Description&nbsp;</strong></td> <td><strong>Collection Method&nbsp;</strong></td> </tr> <tr> <td>html_url &nbsp;</td> <td>Repository URL &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) &nbsp;</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license &nbsp;</td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) &nbsp;</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis.&nbsp;</td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) &nbsp;</td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages.&nbsp;</td> </tr> <tr> <td>project_information &nbsp;</td> <td>Presence of basic project information in README (True/False) &nbsp;</td> <td>(Manual) Checked if the readme have basic information about the project.&nbsp;</td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. &nbsp;</td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit &nbsp;</td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory.&nbsp;</td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool &nbsp;</td> <td>Name of the continuous integration tool used</td> <td>(Script-&nbsp;<a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule &nbsp;</td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the&nbsp;<br>.github/workflows directory to detect the presence of (linters)&nbsp;Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) &nbsp;</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the&nbsp;<br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language &nbsp;</td> <td>Programming language used in the repository &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>type &nbsp;</td> <td>Specifies if the profile is a user or organization &nbsp;</td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation &nbsp;</td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group &nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Software vulnerability detection datasets - function/method level

<p>This dataset is for software vulnerability detection and includes source code in eight programming languages (C, C++, Java, JavaScript, Go, PHP, Ruby, Python). All data is collected from GitHub.</p><p>data<i>{programming language}_vul.json: a set of vulnerable code samples in a certain programming language.</i></p><p>data<i>{programming language}_patch.json: a set of patching code samples in a certain programming language.</i></p><p>&nbsp;</p><p>Each source code sample includes the following 16 properties:&nbsp;</p><p><strong>index</strong>: index of code. If is_vulnerable==False, this index indicates that this code is a patch of the indexing vulnerable code.</p><p><strong>code</strong>: raw source code (may include comments).</p><p><strong>is_vulnerable</strong>: the code is vulnerable (<strong>True</strong>) or a patch (<strong>False</strong>).</p><p><strong>programming_language</strong>: programming language of the code.</p><p><strong>method_name</strong>: name of the method.</p><p><strong>file_name</strong>: name of the file where the source code is extracted.</p><p><strong>repo_url</strong>: url of the project repository.</p><p><strong>repo_owner</strong>: owner of the repository.</p><p><strong>committer</strong>: developer who pushed the commit.</p><p><strong>committer_date</strong>: date when the commit was pushed.</p><p><strong>commit_msg</strong>: the commit message.</p><p><strong>cwe_id</strong>: If is_vulnerable==True, the CWE id; otherwise None.</p><p><strong>cwe_name</strong>: If is_vulnerable==True, the name of corresponding CWE; otherwise None.</p><p><strong>cwe_description</strong>: If is_vulnerable==True, the description of corresponding CWE; otherwise None.</p><p><strong>cwe_url</strong>: If is_vulnerable==True, the url to obtain more details of corresponding CWE; otherwise None.</p><p><strong>cve_id</strong>: If is_vulnerable==True, the CVE id; otherwise None.</p>

openmit-licenseSep 2024View details →
zenodo48/100

Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'

<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Software developers are users of the Heureka microservice platform

<p>This data set contains the qualitative analysis of the &quot;software developers are users&quot; study, conducted to investigate the fitting of the Heureka microservice platform (http://doc.soteto.net) to support software developers in participating the&nbsp;change of socio-technical evolutionary-teal organizations.</p> <p>The study is part of the SOTETO project (http://soteto.net)&nbsp;and its results will be published and contextualized by the dissertation thesis of Johann Sell.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

TomoBreast randomized clinical trial's lung-heart outcomes and mortality through the 2020 COVID-19 pandemic: data and software

<p>Dataset and R script to reproduce the analyses of the manuscript:</p> <p>Vinh-Hung V, Gorobets O, Adriaenssens N, Van Parijs H, Storme G, Verellen D, Nguyen NP, Magne N, De Ridder M.</p> <p><strong>Lung-heart outcomes and mortality through the 2020 COVID-19 pandemic in a prospective cohort of breast cancer radiotherapy patients.</strong></p> <p>Cancers 2022;&nbsp;14(24):6241. https:// doi.org/10.3390/cancers14246241</p> <p>https://www.mdpi.com/2072-6694/14/24/6241</p> <p>PubMed:&nbsp;PMID:&nbsp;36551726</p> <p>PMCID:&nbsp;PMC9777311</p> <p>Info on the variables in file&nbsp;"aelq6_public.R"</p> <p>reproduced in "aelq_2_3_readme.txt":</p> <p>"aelq2_base2.txt" = baseline characteristics.</p> <p>"aelq3.txt" = longitudinal maesurements.</p> <p>Variables in "aelq2_base2.txt":</p> <p>"<strong>aelq2_base2.txt</strong>" = baseline characteristics.&nbsp;<br># Age at randomization, years.&nbsp;<br># RTdose: cf TomoBreast papers.&nbsp;<br># 51 Gy = hypofractionated, simultaneous integrated boost<br># 42 Gy = hypofractionated, no boost, mastectomy cases only<br># 50 Gy = conventional, no boost, mastectomy cases only<br># 66 Gy = conventional, sequential boost<br># Weight kg, Height cm,&nbsp;<br># Detection 1=found by screening (senology follow-up/controle)<br># &nbsp;&nbsp; &nbsp;2=found by symptoms (pain, palpable)<br># &nbsp;&nbsp; &nbsp;9=unknown<br># Smoker &nbsp;&nbsp; &nbsp;0= Not smoker<br># &nbsp;&nbsp; &nbsp;1= Smoker<br># &nbsp;&nbsp; &nbsp;2=ex-smoker<br># Mastectomy (and other binary coded) 1= yes<br># chemosched 0=none<br># &nbsp;&nbsp; &nbsp;1= planned after RT (sequential)<br># &nbsp;&nbsp; &nbsp;2= prior to RT and is finished (sequential)<br># &nbsp;&nbsp; &nbsp;3= chemo is on-going or is planned to start with RT (concomitant)<br># hormonetherapy &nbsp;&nbsp; &nbsp;0=no<br># &nbsp;&nbsp; &nbsp;1=tamoxifen (nolvadex)<br># &nbsp;&nbsp; &nbsp;2=Femara (Letrozole)<br># &nbsp;&nbsp; &nbsp;3=zoladex<br># &nbsp;&nbsp; &nbsp;4=tamoxifen + zoladex<br># Laterality 1,=Right, 2=Left, 3=Bilateral<br># LengthFU: length of follow-up, days from randomization</p> <p>"<strong>aelq3.txt</strong>" = longitudinal maesurements.<br># "Nr" = Case ID<br># "Time" in days from origin (origin =date of randomization),&nbsp;<br># if negative =before randomization<br># &nbsp; &nbsp;"KPS" &nbsp; &nbsp; &nbsp; "Weight" &nbsp; &nbsp;<br># "Died" &nbsp; &nbsp; &nbsp;"LocalRec" &nbsp;"Metast" &nbsp; &nbsp;"NewPrim" &nbsp; = binary code, 0=no, 1=yes<br># "fAEBreast" "fAEHeart" &nbsp;"fAELung" &nbsp; "fAEOther"&nbsp;<br># fAE = freedom from breast, heart, lung, other adverse event score<br># "LVEF2" = ejection fraction, %<br># "MacIver" = estimated cardiac strain</p> <p># the following are pulmonary function tests, untransformed units<br># "FVC", "FEV1", "PEF", "VC", "TLC", "RV", "FRC", "Raw", "sRaw", "DLCO",<br># "VA", "PF"</p> <p># "fDY", "fFA", "fPA" = freedom from dyspnea, from fatigue, from pain<br># range 0 to 100 (best)<br># see papers:</p> <p># Van Parijs, H.; Vinh-Hung, V.; Fontaine, C.; Storme, G.; Verschraegen, C.;<br># Nguyen, D.M.; Adriaenssens, N.; Nguyen, N.P.; Gorobets, O.; De Ridder, M.<br># Cardiopulmonary-related patient-reported outcomes in a randomized clinical<br># trial of radiation therapy for breast cancer. BMC Cancer 2021, 21, 1177,<br># doi:10.1186/s12885-021-08916-z.</p> <p># preprint:<br># Van Parijs, H.; Cecilia-Joseph, E.; Gorobets, O.; Storme, G.;&nbsp;<br># Adriaenssens, N.; Heyndrickx, B.; Verschraegen, C.; Nguyen, N.P.;<br># De Ridder, M.; Vinh-Hung, V. Lung-heart toxicity in a randomized&nbsp;<br># clinical trial of hypofractionated image guided radiation therapy for<br># breast cancer. Preprints 2022, 202212, 0214.<br># https://doi.org/10.20944/preprints202212.0214.v1</p> <p>#&nbsp;<br># "Year" = year of the observation<br># example: randomized 1/1/2011, measurement done 1/31/2011, time = 30 days,<br># Year =2011<br>#<br>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

The Role of Informal Communication in Building Shared Understanding of Non-Functional Requirements in Remote Continuous Software Engineering

<p><strong>Study Information</strong></p> <p>We conducted an ethnography-informed case study of a remote software organization that adopts CSE practices to explore how the organization builds a shared understanding of NFRs. Our study uses semi-structured interviews with a period of observations to answer the following research questions:</p> <p>&nbsp;</p> <ol> <li> <p>How does a remote software organization that adopts CSE practices reach a shared understanding of NFRs?</p> </li> <li> <p>What are the limitations to the shared understanding of NFRs in a remote software organization that adopts CSE practices?</p> </li> <li> <p>What organizational practices for remote collaboration supported a shared understanding of NFRs?</p> </li> </ol> <p>&nbsp;</p> <p>In our study, we refer to our partner organization as Alpha. We used ethnography-informed methods to study Alpha&#39;s practices and processes and how they approach a shared understanding of NFRs in their product development.&nbsp;</p> <p>&nbsp;</p> <p><strong>Data Analysis</strong></p> <p>We performed a qualitative study through semi-structured interviews and observations. We use the open, axial and selective coding approach from grounded theory [1] to create our codebook, which informed the results and discussion of our study. Two independent coders held agreement sessions to discuss the codes, consolidate the codes and calculate the inter-rater reliability using the Cohen Kappa&#39;s coefficient for measuring observer agreement for categorical data [2].&nbsp;</p> <p>&nbsp;</p> <p><strong>Artifact Descriptions</strong></p> <p>Our replication package contains three artifacts:</p> <p>1. Codebook.csv: The codebook contains rows for the list of codes used, including the code name and the description of the codes. The codes are&nbsp;the final set of themes derived during the thematic analysis of the interview responses. For example, &#39;Gaps in communication&#39; means when interview participants describe&nbsp;miscommunications due to team members making&nbsp;assumptions about a project/process or&nbsp;having unclear expectations for a project.</p> <p>2. kappa-scores.csv: This contains the associated kappa values for each round of inter-rater agreement sessions. For each agreement session, the Cohen Kappa&#39;s coefficient was calculated from the number of agreements and disagreements of codes within one or two interview transcripts. The Kappa values represent the level of agreement ranging from 0 to 1, where &gt; 0.6 represents substantial agreement.&nbsp;</p> <p>3. Interview-questions.csv: This contains the interview questions used in the semi-structured interviews. Some of the interview questions varied depending on the interviewee&rsquo;s role,&nbsp;experience and the flow of the interviews.</p> <p><strong>&nbsp;</strong></p> <p><strong>Usefulness</strong></p> <p>We recognize that the value and usefulness of our replication package are yet-to-be-determined.&nbsp; In the interest of transparency of open science, we published our artifacts. We hope that these artifacts are useful to either replicate our findings or to further analyze them to produce other enlightening results.</p> <p><strong>&nbsp;</strong></p> <p><strong>References</strong></p> <p>1. Rashina Hoda, James Noble, and Stuart Marshall. &quot;Grounded theory for geeks&quot;. In: Proceedings of the 18th conference on pattern languages of programs. 2011, pp. 1&ndash;17.</p> <p>2. J Richard Landis and Gary G Koch. &quot;The measurement of observer agreement for categorical data&quot;. In: biometrics (1977), pp. 159&ndash;174.</p> <p><strong>&nbsp;</strong></p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Testing and Demonstration Data for DataRig Software

<p>This repository holds the testing and demonstration data&nbsp;for&nbsp;<a href="https://github.com/mscaudill/datarig">DataRig</a>, an opensource software program for downloading datasets from data repositories utilizing RESTful APIs. This repository contains 5 sample datasets.</p> <p>&nbsp;</p> <p><strong>annotations_001.txt</strong></p> <p>This data set is a tab-separated text file containing 6 columns that start on line number 7. The column headers are;&nbsp;</p> <p>&nbsp;&#39;Number&#39;&nbsp;&nbsp; &#39;Start Time&#39;&nbsp; &nbsp; &#39;End Time&#39;&nbsp; &nbsp; &#39;Time From Start&#39;&nbsp; &nbsp; &#39;Channel&#39;&nbsp; &nbsp; &#39;Annotation&#39;</p> <p>There are 13 rows of data under each of these column headers representing the start and end times of annotated events from an eeg recording file in this repository called recording_001.edf. The events describe the behavior of a mouse in 5 sec increments with each behavior being one of &#39;exploring&#39;, &#39;grooming&#39; or &#39;rest&#39;.</p> <p>&nbsp;</p> <p><strong>recording_001.edf</strong></p> <p>A European Data Format file consisting of 4 channels of EEG data lasting approximately 1 hour. The times in the annotations_001.txt file are referenced against this file.</p> <p>&nbsp;</p> <p><strong>sample_arr.npy</strong></p> <p>A numpy array of shape (4, 250) with values sequentially running from 0 to 1000.</p> <p>&nbsp;</p> <p><strong>sample_excel.xls</strong></p> <p>An excel file with a single column of 10 numbers from 0-9 sequentially.</p> <p>&nbsp;</p> <p><strong>sample_text.txt</strong></p> <p>A text file with 4 rows containing 250 values per row. The values in the file run from 0 to 1000 sequentially.</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Bots in Software Development: A Systematic Literature Review [Data Set]

<p>This repository contains al the artifacts of the research: Bots in Software Development: &nbsp;A Systematic Literature Review&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification

<p>We present&nbsp;a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability&#39;s root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle&#39;s continuous integration and deployment (CI/CD) process.&nbsp;</p> <p>&nbsp;</p> <p>The uploaded&nbsp;dataset contains 1000 codes,&nbsp; each comprising 20&nbsp;to 50&nbsp;lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Simulation results of adaptive multicast streaming for videoconferences in software-defined networks

<p>Real-time applications, such as video conferences, have strong Quality of Service requirements for ensuring a decent Quality of Experience. Nowadays, most of these conferences are performed over wireless devices. Thus, an appropriate management of both heterogeneous mobile devices and network dynamics is necessary. Software Defined Networking enables the use of multicasting and stream layering inside the network nodes, two techniques able to enhance the quality of live video streams. In this paper, we propose two algorithms for building and maintaining multicast sessions in a software-defined network. The first algorithm sets up the initial multicast trees for a given call. It optimally places the stream layer adaptation function inside the core network in order to minimize the bandwidth consumption. This algorithm has two versions: the first one, based on shortest path trees is minimizing the latency, while the second one, based on spanning trees is minimizing the bandwidth consumption. The second algorithm adapts the multicast trees according to the network changes occurring during a call. It does not recompute the trees, but only relocates the stream layer adaptation functions. It requires very low computation at the controller, thus making our proposal fast and highly reactive. Extensive simulation results confirm the efficiency of our solution in terms of processing time and bandwidth savings compared to existing solutions such as multiple unicast connections, Multipoint Control Unit solutions and application layer multicast.</p>

opencc-by-4.0Apr 2018View details →
edi48/100

Software for processing data from a fast-responding RINKO EC oxygen/temperature sensor (JFE Advantech Co, Ltd)

This dataset describes how data from a fast-responding JFE Advantech RINKO EC ARO-EC-CM sensor connected to a Nortek Vector is processed to obtain accurate aquatic eddy covariance measurements. The code and documentation are stored in a .zip file. It consists of a manual, Fortran source code, a definition file and a complied executable suitable for running on Microsoft Windows. The software development was supported by NSF funding to PI Berg (OCE-1824144, OCE-2223204).

openCustomJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record