Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

260

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

260 results for “empirical study”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"

<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>

opencc-by-4.0Mar 2021View details →
zenodo48/100

Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study

<p>Efficiency is essential to support responsiveness w.r.t. ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code that supports symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development tends to produce DL code that is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, less error-prone imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. While hybrid approaches aim for the &quot;best of both worlds,&quot; the challenges in applying them in the real world are largely unknown. We conduct a data-driven analysis of challenges&mdash;and resultant bugs&mdash;involved in writing reliable yet performant imperative DL code by studying 250 open-source projects, consisting of 19.7 MLOC, along with 470 and 446 manually examined code patches and bug reports, respectively. The results indicate that hybridization: (i) is prone to API misuse, (ii) can result in performance degradation&mdash;the opposite of its intention, and (iii) has limited application due to execution mode incompatibility. We put forth several recommendations, best practices, and anti-patterns for effectively hybridizing imperative DL code, potentially benefiting DL practitioners, API designers, tool developers, and educators.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Dataset: Ethical Issues in Empirical Studies using Student Subjects: Re-visiting Practices and Perceptions

<p># Dataset for Paper &quot;Ethical Issues in Empirical Studies using Student Subjects: Re-visiting Practices and Perceptions&quot; - Rev 1#</p> <p>This is the dataset for the paper titled &quot;Ethical Issues in Empirical Studies using Student Subjects: Re-visiting Practices and Perceptions&quot;. All mapping study data has the prefix *MAP*, while all survey data the prefix *SUR*. It has been updated for a major revision at Springer Empirical Software Engineering (Rev 1).</p> <p>In case of questions, feel free to contact the author, Grischa Liebel, ORCID: https://orcid.org/0000-0002-3884-815X, current affiliation and email: Reykjavik University, Iceland, grischal@ru.is</p> <p>## Systematic Mapping Study ##<br> The mapping study data is mainly contained in the *MAPmappingStudy.xlsx* file. Different tabs are used for the two phases: exclusion by title and abstract (tab *title_abs*), and exclusion by fulltext (tab *fulltextScreening*). The final set of papers is obtained by filtering the *fulltextScreening* tab by included papers (Column V).</p> <p>The tab *fulltextScreening* contains a number of columns named \*range (e.g., *noStudentsRange*). These columns contain the unified/categorised values for the verbatim values listed in the column with the same name without *range*. For instance, *noStudentsRange* contains the range of students in the primary study, while *noStudents* contains the actual number obtained from the studies.</p> <p>The file *MAPvenues.txt* contains the included venues in the mapping study.</p> <p>Finally, the raw search results are provided as BIB/RIS files with the prefix *MAPRAW*.</p> <p>## Survey ##<br> The survey folder contains the survey pages (named *surveyPageN.pdf*), as well as the raw data in *surveyDataAnon.xlsx*. Note that free-text answers have been aggregated, anonymised, and sorted alphabetically in individual tabs. Similarly, countries that only occur once have been changed to Do Not Disclose answers, and all answers have been sorted randomly. All questions are listed by their acronym. The corresponding questions, as well as possible answers, are described in the *QuestionKey* tab.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

ASVspoof2019LA-Sim: Augmented Dataset for An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems

<p>This is the dataset we augmented to study the channel effects for anti-spoofing. For more details, please refer to our Interspeech 2021 paper: &quot;An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems&quot;.</p> <p>Proceeding: <a href="https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html">https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html</a></p> <p>Arxiv: <a href="https://arxiv.org/pdf/2104.01320.pdf">https://arxiv.org/pdf/2104.01320.pdf</a></p> <p>Code:&nbsp;<a href="https://github.com/yzyouzhang/Empirical-Channel-CM">https://github.com/yzyouzhang/Empirical-Channel-CM</a></p> <p>Contact: you.zhang@rochester.edu</p> <p><strong>Version 1.0</strong> contains the <strong>training</strong> and the <strong>development</strong> set. We have added the <strong>evaluation</strong> set in <strong>version 1.1 </strong>but deleted the training set due to the size limitation, but you can still access the training set in version 1.0.</p> <p>Please check it out.</p> <p>To extract the files, please use the following commands:</p> <pre><code class="language-bash">cat eval.tar.gz-part* &gt; eval.tar.gz tar -xvzf *.tar.gz</code></pre> <p>After concatenation, to make sure the download is complete, you can check with the following:</p> <pre><code>md5sum *.tar.gz 15dea7d28b126994bb6b159778f706af dev.tar.gz 0615052b34ca6c7f58505eaa8647844f eval.tar.gz 3058dd9d407f3c9ae697acca8c34a6c3 train.tar.gz</code></pre> <p>Thanks.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Dataset: A Fly in the Ointment: An Empirical Study on the Characteristics of Ethereum Smart Contracts Code Weaknesses and Vulnerabilities

<p>Dataset: A Fly in the Ointment: An Empirical Study on the Characteristics of Ethereum Smart Contracts Code Weaknesses and Vulnerabilities</p> <p>Majd Soud, Grischa Liebel, Mohammad Hamdaqa<br> majd18@ru.is, grischal@ru.is, mhamdaqa@polymtl.ca.</p> <p>This Dataset includes the following:&nbsp;</p> <p>1. &quot;labeling.xml&quot; files that represents the data for Categories of vulnerabilities in Smart Contracts for four data sources (i.e., Common Vulnerability and Exposure (CVE), Smart Contract Weakness Classification Registry (SWC), Stack Overflow, and GitHub)<br> XML files structure:<br> The XML files can be opened used any editor or any code editor (e.g. Visual Studio Code).</p> <p>1. Each file has a root that is &lt;Card_Table&gt;&lt;/Card_Table&gt; which contains all the cards we labeled.&nbsp;<br> 2. Each card is represented by the &lt;Card&gt;&lt;/Card&gt; and contains the following:<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- The tag marked by &lt;Tag&gt; represents the keyword that was used to search and collect the card from StackOverflow.&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- The URL marked by &lt;URL&gt; of the URL link&nbsp;which contains all the information of the labeled vulnerability.&nbsp;&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- The other tages marked by &lt;OtherTags&gt; that shows all the tags used in the post on Stack Overflow.&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- The&nbsp;expert labeling for the categories of vulnerabilities in each card is represented by &nbsp;&lt;CategoryExpertLabel&gt;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- In more details, some records has the &lt;SecondExpertCategoryLabel&gt; that represents the second expert labeling for the categories of vulnerabilities.<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;- The &lt;CategoryAgreement&gt; used to calculate the inter-rater agreement between the two labelers.&nbsp;<br> &nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Artifacts for the ISSTA 2022 Paper: An Empirical Study on the Effectiveness of Static C Code Analyzers for Vulnerability Detection

<p>This repository contains the evaluation script and the corresponding data of the ISSTA&#39;22 paper &quot;An Empirical Study on the Effectiveness of Static C Code Analyzers for Vulnerability Detection&quot;.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A Corpus of Publicly Available Simulink Models for Model-based Empirical Studies

<p>Abstract: Recent years have seen many empirical studies of model-based cyber-physical systems and commercial CPS development tool chains such as Matlab/Simulink. To benefit such research, this paper presents the by-far largest corpus of freely available Simulink models to date, containing over 1,000 models.</p> <p>Surprising findings based on this corpus include that (a) tool support for metric collection is not adequate and (b) users do not reuse model components as they would in object-oriented programs.</p> <p>The paper both confirms and contradicts earlier findings that are based on significantly fewer models, suggesting the utility of the corpus for future research. While others have not yet leveraged this model corpus, we hope that our freely available corpus and infrastructure will benefit future model-based empirical research and tool development efforts, by reducing the model-collection overhead and thus easing evaluation.</p> <p>Please see our paper &quot;A curated corpus of Simulink models for model-based empirical studies&quot;&nbsp; In Proc. 4th International Workshop on Software Engineering for Smart Cyber-Physical Systems (SEsCPS)</p> <p>See https://github.com/verivital/slsf_randgen/wiki for more details including information related to downloading the Simulink models.</p> <p>http://ranger.uta.edu/~csallner/csallner_bib.html#Chowdhury18Curated</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

An Empirical Study of Container Image Configurations and Their Impact on Start Times (Container Image Data)

<p>Dataset with the container image metadata used for our IEEE/ACM CCGRID 2023 paper &quot;An Empirical Study of Container Image Configurations and Their Impact on Start Times&quot;.</p> <p>Abstract of the paper: A core selling point of application containers is their fast start times compared to other virtualization approaches like virtual machines. Predictable and fast container start times are crucial for improving and guaranteeing the performance of containerized cloud, serverless, and edge applications. While previous work has investigated container starts, there remains a lack of understanding of how start times may vary across container configurations. We address this shortcoming by presenting and analyzing a dataset of approximately 200,000 open-source Docker Hub images featuring different image configurations (e.g., image size and exposed ports). Leveraging this dataset, we investigate the start times of containers in two environments and identify the most influential features. Our experiments show that container start times can vary between hundreds of milliseconds and tens of seconds in the same environment. Moreover, we conclude that no single dominant configuration feature determines a container&#39;s start time and that hardware and software parameters must be considered together for an accurate assessment.</p> <p>Dataset description: Our images dataset contains 200,986 entries with 21 features associated to each container image. In the following, we describe the meaning of each feature. Further information is available in <a href="https://github.com/opencontainers/image-spec">OCI Image Specification</a> and the <a href="https://docs.docker.com/engine/reference/run/">Docker Run Documentation</a>. Besides the 20 features grouped in the five categories below, each dataset entry has a image_id, which is used to uniquely identify the dataset entry.</p> <p>Features</p> <p>Metadata features (prefix: meta)</p> <ul> <li><strong>meta_repo_digest</strong> : The repo digest is a SHA-256 hash which is used to uniquely identify and pull the image from Docker Hub</li> <li><strong>meta_architecture</strong> : The CPU architecture which the binaries in the image are built to run on</li> <li><strong>meta_os</strong> : The name of the operating system which the image is built to run on</li> <li><strong>meta_docker_version</strong> : The Docker version used to built this image</li> </ul> <p>I/O stream features (prefix: io)</p> <ul> <li><strong>io_attach_stdin</strong> : boolean setting to determine whether the console should be attached to the process stdin stream</li> <li><strong>io_attach_stdout</strong> : boolean setting to determine whether the console should be attached to the process stdout stream</li> <li><strong>io_attach_stderr</strong> : boolean setting to determine whether the console should be attached to the process stderr stream</li> <li><strong>io_tty</strong> : boolean setting to determine whether the console should pretend to be a TTY when attached</li> <li><strong>io_open_std_in</strong> : boolean setting to determine whether the process stdin stream should be kept open even if console not attached</li> <li><strong>io_std_in_once</strong> : boolean setting to determine whether the process retrieved input from the stdin stream at least once</li> </ul> <p>Start command features (prefix: cmd)</p> <ul> <li><strong>cmd_args</strong> : Length of list of arguments to use as the command to execute when the container starts</li> <li><strong>cmd_envvars</strong> : Environment variables set per default when the container starts</li> <li><strong>cmd_additional_args</strong> : Length of list for additional arguments to the containers entrypoint</li> </ul> <p>File system features (prefix: fs)</p> <ul> <li><strong>fs_volumes</strong> : Number of volumes to create/use by default</li> <li><strong>fs_size</strong> : Size of this image in bytes</li> <li><strong>fs_virtual_size</strong> : Virtual size of this image in bytes (equals size)</li> <li><strong>fs_graph_driver_name</strong> : Name of the image&#39;s graph driver</li> <li><strong>fs_root_fs_type</strong> : Name of the file system type used in the image</li> <li><strong>fs_layers</strong> : Number of root file system layers</li> </ul> <p>Networking features (prefix: net)</p> <ul> <li><strong>net_ports</strong> : Number of ports to expose per default</li> </ul> <p>&nbsp;</p> <p>Dataset acquisition: The dataset has been acquired from Docker Hub using a web crawler. We used substring matches with the <a href="https://hub.docker.com/explore">Docker Hub Explore function</a>. As search strings, we used all letter combination with sizes 1 to 3, meaning that our first search string was &#39;a&#39; and our last was &#39;zzz&#39;. We included both results from the &#39;recently updated&#39; and the &#39;most popular&#39; selection. We came up with an initial list of 286,294 image names. We then tested we could pull and start these images once. These tests have been conducted from April to June 2022. We sorted out all images that were either not pullable or startable and retrieved all total of 200,986 valid images. In the following, we describe the error types that we encountered and that let to the removal of the causing image from the dataset:</p> <ul> <li>The image manifest was unknown when we tried to download it meaning that is has been renamed or deleted from the time when our web crawler was running</li> <li>The entrypoint command required a dependency that was missing in the image and therefore the container could not be started</li> <li>The image did not specify an entrypoint command and could therefore not be started</li> <li>The image declared an invalid root file system type</li> <li>The image had a malformed root file system</li> <li>The image configuration was incomplete and therefore not all required data could be obtained</li> </ul> <p>See also our CodeOcean capsule with the processing scripts for our paper: https://doi.org/10.24433/CO.4595026.v2</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Model outputs for the study "Guidance in Radiology Report Summarization: An Empirical Evaluation and Error Analysis"

<p>This resources provides pre-processed input data, model checkpoints and model outputs for&nbsp;experiments on the&nbsp;OpenI dataset in below&nbsp;study.&nbsp;</p> <blockquote> <p>Jan Trienes, Paul Youssef, J&ouml;rg Schl&ouml;tterer, and Christin Seifert. 2023. <a href="https://arxiv.org/abs/2307.12803">Guidance in Radiology Report Summarization: An Empirical Evaluation and Error Analysis</a>. In Proceedings of the 16th International Natural Language Generation Conference (INLG), Prague, Czech Republic. Association for Computational Linguistics.</p> </blockquote> <p>For more information please refer to the accompanying paper and code repository (<a href="https://github.com/jantrienes/inlg2023-radsum">https://github.com/jantrienes/inlg2023-radsum</a>).</p> <p><strong>The data is structured as follows:</strong></p> <ul> <li><code>data/preprocessed/</code>&nbsp;includes the dataset(s)&nbsp;for each model</li> <li><code>output/</code>&nbsp;includes one folder for each experiment/model run. The first part&nbsp;of each output path&nbsp;indicates the dataset that was used at inference.</li> <li>For a mapping between model IDs and results in the paper, see below table. All models were also trained <em>with the background section as input.&nbsp;</em>These are available in directories&nbsp;with the&nbsp;<code>-bg-</code> qualifier.&nbsp;</li> </ul> <table> <thead> <tr> <th>Model name in paper</th> <th>Output directory</th> </tr> </thead> <tbody> <tr> <td><em>Results from Table 2</em></td> </tr> <tr> <td>OracleExt</td> <td>openi-unguided/oracle</td> </tr> <tr> <td>BertExt (Liu and Lapata, 2019)</td> <td>openi-unguided/bertext-default</td> </tr> <tr> <td>BertAbs (Liu and Lapata, 2019)</td> <td>openi-unguided/bertabs-default</td> </tr> <tr> <td>GSum (Dou et al., 2021)</td> <td>openi-bertext-default-clip-k1/gsum-default</td> </tr> <tr> <td>GSum w/ LR-Approx</td> <td>openi-bertext-default-clip-lrapprox/gsum-default</td> </tr> <tr> <td>GSum w/ BERT-Approx</td> <td>openi-bertext-default-clip-bertapprox/gsum-default</td> </tr> <tr> <td>GSum w/ Thresholding</td> <td>openi-bertext-default-clip-threshold/gsum-default</td> </tr> <tr> <td>WGSum (Hu et al., 2021)</td> <td>openi-wgsum/wgsum-default</td> </tr> <tr> <td>WGSum+CL (Hu et al., 2022)</td> <td>openi-wgsum-cl/wgsum-cl-default</td> </tr> <tr> <td><em>Results from Table 3</em></td> </tr> <tr> <td>Fixed (k=1)</td> <td>openi-unguided/bertext-default-clip-k1</td> </tr> <tr> <td>LR-Approx</td> <td>openi-unguided/bertext-default-clip-lrapprox</td> </tr> <tr> <td>BERT-Approx</td> <td>openi-unguided/bertext-default-clip-bertapprox</td> </tr> <tr> <td>Thresholding</td> <td>openi-unguided/bertext-default-clip-threshold</td> </tr> <tr> <td>k = |OracleExt|</td> <td>openi-unguided/bertext-default-clip-oracle</td> </tr> <tr> <td><em>Results from Table 4</em></td> </tr> <tr> <td>Fixed (Dou et al., 2021)</td> <td>openi-bertext-default-clip-k1/gsum-default</td> </tr> <tr> <td>Oracle Length</td> <td>openi-bertext-default-clip-oracle/gsum-default</td> </tr> <tr> <td>Oracle Length + Content</td> <td>openi-oracle/gsum-default</td> </tr> <tr> <td><em>Results from Table 5</em></td> </tr> <tr> <td>BertExt w/ k=[1,5]</td> <td>openi-unguided/bertext-default-clip-k{1,2,3,4,5}</td> </tr> <tr> <td>GSum w/ k=[1,5]</td> <td>openi-bertext-default-clip-k{1,2,3,4,5}/gsum-default</td> </tr> </tbody> </table>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Dataset for "Architectural Security Weaknesses in Industrial Control Systems: An Empirical Study Based on Security Advisories' Vulnerability Reports"

<p>Supplementary artifacts to &quot;<em>Architectural Security Weaknesses in Industrial Control Systems (ICS): An Empirical Study based on Disclosed Software Vulnerabilities</em>&quot;</p> <p>Published in the Proceedings from the 2019 IEEE International Conference on Software Architecture (ICSA)</p> <p>Package Contains:</p> <p>- Raw output showing Components, CAWEs, and CVEs per report</p> <p>- Frequency Data (# of reports) for those concerns</p> <p>- ICS Component - Term Dictionary&nbsp;</p> <p>- HTML versions of reports studied in paper</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Dataset: Breaking Type-Safety in Go: An Empirical Study on the Usage of the unsafe Package

<p>This dataset contains all script used in the study, as well as the raw data extracted from the repositories and the&nbsp;processed data used to analyze our RQs in the manuscript &quot;Breaking Type-Safety in Go: An Empirical Study on the Usage of the unsafe Package&quot;.</p> <p>For more information on how to understand the folder structure, scripts, and dataset, please read the README.md.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Software Product Line Traceability and Product Configuration in Class and Sequence Diagrams: an Empirical Study

<p>Software Product Line Traceability and Product Configuration in Class and Sequence Diagrams: an Empirical Study</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Software Product Line Configuration and Traceability: an Empirical Study on SMarty Class and Component Diagrams

<p>Software Product Line Configuration and Traceability: an Empirical Study on SMarty Class and Component Diagrams</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study

<p>Online appendix of the paper entitled: &quot;The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study&quot;. It contains all scripts and data required to replicate the four research questions of the&nbsp;study.</p> <p>Abstract:&nbsp;Software vulnerabilities are weaknesses in source code that can be potentially exploited to cause loss or harm. While researchers have been devising a number of methods to deal with vulnerabilities, there is still a noticeable lack of knowledge on their software engineering life cycle, for example how vulnerabilities are introduced and removed by developers. This information can be exploited to design more effective methods for vulnerability prevention and detection, as well as to understand the granularity that these methods should aim at. To investigate the life cycle of software vulnerabilities, we focus on how, when, and under which circumstances vulnerabilities&nbsp;are introduced&nbsp;in software projects, as well as whether, after how long, and how they&nbsp;are removed. We consider 4,097 vulnerabilities with public patches from the National Vulnerability Database&mdash;pertaining to 1,163 open-source software projects on GITHUB&mdash;and define a six-step process that involves both automated parts (e.g., using the SZZ algorithm to find the vulnerability-inducing commits) and manual analyses (e.g., how vulnerabilities were fixed). The investigated vulnerabilities can be classified in 148 categories, take on average 4.19 commits before being introduced, and remain unfixed for a median of 1,506.50 commits and 691.50 days. Most of them are&nbsp;introduced&nbsp;by developers with high workload, often when doing maintenance activities, and&nbsp;removed&nbsp;with mostly with the addition of new source code aiming at implementing further checks on inputs. We conclude by distilling practical implications on when and how vulnerability detectors should work to better assist developers in early detecting these issues.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

On the Co-evolution of ML Pipelines and Source Code - Empirical Study of DVC Projects

<p>This is a replication package of our paper submission to the Saner 2021 entitled:</p> <p>On the Co-evolution of ML Pipelines and Source Code - Empirical Study of DVC Projects</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

An Empirical Study of Flaky Tests in Python

<p><a href="https://arxiv.org/pdf/2101.09077.pdf">Paper</a></p> <p><a href="https://github.com/se2p/FlaPy">Flapy (our test execution tool)</a></p> <p>TestsOverview.csv<br> - Project Columns: Project_Name, Project_URL, Project_Hash<br> - Test Columns: Test_filename, Test_classname, Test_funcname, Test_parametrization<br> - Flaky_sameOrder_withinIteration: the test is non-order-dependent flaky<br> - Order-dependent: the test is order-dependent flaky<br> - Flaky_Infrastructure: the test shows infrastructure flakiness</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

USE-ME Empirical evaluation pilot study

<p>We report on the pilot assessment of the feasibility of the <strong>USE-ME tool </strong>with master students in computer science which were involved into a DSL course. In total there were four groups consisted of two or three participants which were developing the following DSLs:</p> <ul> <li><strong>DSL Spreadsheets </strong>- a DSL transforms an activity graph into a Gantt chart. The target users of this DSL are to be project managers.</li> <li><strong>Gestures Kinect</strong> - a DSL which supports specification of communication between navy users using a Kinect device. The target users of this DSL are to be a navy operators.</li> <li><strong>Peddy Paper</strong> - a DSL which creates several instances of personalised paddy papers. The target users are paddy paper builders.</li> <li><strong>Smart House </strong>- a DSL which supports a design of the elements and operations that the house can have. It is meant to be used by house owners. </li> </ul> <p>The evaluation consisted of three learning session. The first one took place after four weeks of the DSL development. Students were introduced to the usability evaluation during a 2h theoretical lecture. For next 2h were introduced with the USE-ME tool and were guided to perform installation and set up working environment. Also, the students were given a participation questionnaire to fill in and describe a purpose of their DSL. In the end of a session, they were given the background questionnaire to fill till the following session which was happening 1 week after. Following two weekly 4h sessions were consisting of USE-ME hands on. The students were introduced to the modelling activities followed by Visualino example. For each activity, the same student from the group was using the tool to create USE-ME models for their DSL, while other students from the group were helping in deciding what would be a right specification. Finally, students were asked to try to finish the learned models and deliver them as a part of DSL course. After delivery, the students which were using the tool to model were asked to fill in a feedback questionnaire. </p> <p>The evaluation deliveries related to each DSL (project reports, USE-ME models), and ones related to participants (background and feedback questionnaire), as well as evaluation results, are attached in this data set. </p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

An Empirical Study of Activity, Popularity, Size, Testing, and Stability in Continuous Integration

<p>A good understanding of the practices followed by software development projects can positively impact their success --- particularly for attracting talent and on-boarding new members. In this paper, we perform a cluster analysis to classify software projects that follow continuous integration in terms of their activity, popularity, size, testing, and stability. Based on this analysis, we identify and discuss four different groups of repositories that have distinct characteristics that separates them from the other groups.  With this new understanding, we encourage open source projects to acknowledge and advertise their preferences according to these defining characteristics, so that they can recruit developers who share similar values.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Mark loss can strongly bias estimates of demographic rates in multi-state models: a case study with simulated and empirical datasets

<p>This archive contains the empirical data analysed in the paper 'Mark loss can strongly bias estimates of demographic rates in multi-state models: a case study with simulated and empirical datasets' by Touzalin et al. (https://doi.org/10.24072/pci.ecology.100416). &nbsp;The dataset is provided as a .Rdata file ('TLoss_GMdata.Rdata'), and full description of the content is provided in the file 'Readme_TLdata.csv'. All additional details are available in the main text (https://doi.org/10.24072/pci.ecology.100416) or in the supporting information (https://doi.org/10.5281/zenodo.10204538).</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Dataset: an empirical study on architectural smells through a pipeline for continuous technical debt assessment

<h2><strong>Dataset of the study "An empirical study on architectural smells through a pipeline for continuous technical debt assessment"</strong></h2> <h3><strong>Abstract</strong></h3> <p>In recent years, researchers spent an increasing amount of effort investigating technical debt, with quantitative methods, and in particular static analysis, being the most common approach to investigate such a topic.</p> <p>However, quantitative studies are susceptible, to varying degrees, to external validity threats, which hinder the generalisation of their findings.<br>In response to this concern, researchers strive to expand the scope of their studies by incorporating a larger number of projects into their analyses. This practice is typically executed on a case-by-case basis, necessitating substantial data collection efforts that have to be repeated for each new study.</p> <p>To address this issue, this paper presents an approach for tackling this problem and enabling researchers to study architectural smells, a well-known indicator of architectural technical debt, &nbsp;at a large scale. Specifically, we introduce a novel approach to a data collection pipeline that leverages Apache Airflow to continuously generate up-to-date, large-scale datasets with any static analysis tool.</p> <p>Finally, we use the data collected through the pipeline to study the correlation between architectural smells and logical coupling in order to understand how smells influence maintenance efforts.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record