Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “Code Comprehension”
Compiled database, code and raw data for the article "A Comprehensive Database of Leaf Temperature, Water, and CO2 Fluxes in Young Oil Palm Plants Across Diverse Climate Scenarios for the Evaluation of Functional-Structural Models"
<p>This dataset results from an experiment on young oil palm plants (<em>Elaeis guineensis</em>) in the Ecotron facility from CNRS in Montpellier. Four plants were put in a microcosm one by one with varying climatic conditions to investigate the effect of climate on leaf temperature, CO2, and H2O fluxes at the plant scale. The conditions were defined based on typical daily conditions from a location where it is grown (Libo, Indonesia), <em>i.e.</em>, a day with no rainfall and near-average air temperature and humidity. This base condition was then modified by adding more CO2 (400, 600 and 800ppm), less radiation (typical cloudy sky), and more or less temperature and vapour pressure deficit (± 30%).</p> <p>Find more details from the <code>README.md</code> file in the repository or from the associated <a href="https://github.com/PalmStudio/Biophysics_database_palm" target="_blank" rel="noopener">Github repository</a>.</p>
VulnMiner: A Comprehensive Framework for Vulnerability Collection from C/C++ Source Code Projects
<p>In this repository, we present an initial release of the VulnMiner vulnerability dataset, curated from prevalent projects and annotated with vulnerable and benign instances. This dataset incorporates projects with vulnerabilities labeled as Common Weakness Enumeration (CWE) categories. The developed open-source extraction tool collects vulnerability data utilizing static security analyzers. The study also fosters the machine learning (ML) and natural language processing (NLP) model's effectiveness in accurately classifying vulnerabilities, evidenced by its identification of numerous weaknesses in open-source projects.</p>
Data, code and software to reproduce the article entitled "Modeling soil-plant functioning of intercrops using comprehensive and generic formalisms implemented in the STICS model"
<p>This is the data, code and software to reproduce the article entitled " Modeling soil-plant functioning of intercrops using comprehensive and generic formalisms implemented in the STICS model". Here is a summary of the paper:</p> <p>The growing demand for sustainable agriculture is raising interest in intercropping for its multiple potential benefits to avoid or limit the use of chemical inputs or increase the production per surface unit. Predicting the existence and magnitude of those benefits remains a challenge given the numerous interactions between interspecific plant-plant relationships, their environment and the agricultural practices. Soil-crop models are critical in understanding these interactions in dynamics during the whole growing season, but few models are capable of accurately simulating intercropping systems.</p> <p>In this study, we propose a set of simple and generic formalisms for simulating key interactions in intercropping systems that can be readily included into existing dynamic crop models. This requires simulating important processes such as development, light interception, plant growth, N and water balance, and yield formation in response to management practices, soil conditions, and climate. These formalisms were integrated into the STICS soil-crop model and evaluated using observed data of intercropping systems of cereal and legumes mixtures, including Faba bean-Wheat, Pea-Barley, Sunflower-Soybean, and Wheat-Pea mixtures. We demonstrate that the proposed formalisms provide a comprehensive simulation of soil-plant interactions in various types of bispecific intercrops. The model was found consistent and generic under a range of spring and winter intercrops (nRMSE = 25% for maximum leaf area index, 23% for shoot biomass at harvest, and 18% for yield).</p> <p>This is the first time a complete set of formalisms has been developed and published for simulating intercropping systems and integrated into a soil-crop model. With its emphasis on being generic, sufficiently accurate, simple, and easy to parameterize, STICS is well-suited to help researchers designing <em>in silico</em> the agroecological transition by virtually pre-screening sustainable, manageable intercrop systems adapted to local conditions.</p> <p> </p> <p> </p> <p> </p>
Dataset, Survey, and R Notebooks for "Does Surprisal Predict Code Comprehension Difficulty?"
<p>(Version 1.1 Update)</p> <p>- Removed anonymization after reviewing period ended and paper was accepted at Cogsci 2020.</p> <p>- Added Qualtrics Survey in exported form</p> <p>Dataset and R analysis scripts for the paper "Does Surprisal Predict Code Comprehension Difficulty?". For more details, see "ComprehensionREADME.md" in the included zip file.</p>
Code to generate figures 3 and 4 of: "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."
<p>Code to generate figures 3 and 4 of the manuscript titled "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."</p> <p> </p>
Program Comprehension Challenges in Software Code Review
<p>Software engineers spend more time understanding code than writing it (with up to 70% of their time being devoted to this). A key activity where developers spend a lot of time reading and understanding code is software code review. Yet, little is known on the challenges concerning program comprehension during software code review. This study provides insight into the types of comprehension challenges that occur during code reviews and their causes. We find that missing design rationale is the most common reason for comprehension challenges in software code review. Comprehension challenges occur most commonly around five topics: “program logic,” “code design,” “defensive coding,” “condition checking,” and “concurrency”. We also show that machine learning (ML) can be used to automatically detect comprehension challenges with 74.3% precision and a 66.7% recall. We discuss potential improvements to code review support tools based on our findings.</p> <p>These files contain the trained ML algorithms and all the data collected and used for this study.</p> <p> </p> <p>Data was collected in 2018 with analysis performed in 2018/2019, but completion was delayed due to impacts of COVID19.</p>
Replication package: 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study
<p>Replication package | 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study</p>
Model source code, modified model code, and all scripts of the paper "A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity".
<p>Model source code, modified model code, and all scripts of the paper "A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity".</p>
Model source code, simulation results, and all scripts for the paper "A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity".
<p>FigureTable_Scripts.rar is NCL scripts used for figures and tables. <br> Model_Scripts.rar is the modified model code and scripts used for running model.<br> models.rar is the GAMIL model source code.<br> ModelResults.rar is the model results. Note that only the variables used for making plots.</p>
DATASET - On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code
<p>Study package containing raw and analyzed data from the work entitled "On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code".</p> <p>The package comprises: 1) a summary of the information extracted from all papers mentioned in the Background; 2) the source code snippets used in the three studies; 3) the consent and characterization forms distributed to the participants; 4) the raw data, the aggregated data and other material generated with the collected data.</p>
Replication package - The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding
<p>Replication package for:</p> <p>Marvin Wyrich, Andreas Preikschat, Daniel Graziotin and Stefan Wagner. The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding. To appear in: Proceedings of the 43rd International Conference on Software Engineering (ICSE ’21), Madrid, Spain.</p> <p>- The `data` folder contains dataset, R analysis script, and figures for the paper<br> - the `material` folder contains the snippets used for the experimental tasks (`code snippets` subfolder), the environment in which participants had to understand the code (`environment` subfolder) and the task sheet form to evaluate code comprehension.</p> <p>The R analysis script allows to reproduce all statistical analysis steps of the paper as well as its figures. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.</p> <p>For ethics and privacy reasons, we had to drop the following potentially identifying fields from the openly released dataset, because the fields might enable participants to identify their peers:</p> <p>- Age<br> - Gender<br> - Subject (study plan)<br> - Program<br> - Semester no.</p> <p>These missed details will still allow full reproducibility of the study.</p>
Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"
<div> <div># Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"<br><br>containing the dataset and scripts used to create it.</div> </div>
Replication package: Code Comprehension Confounders: A Study of Intelligence and Personality
<p>Replication package for:</p> <p><em>S. Wagner and M. Wyrich, "Code Comprehension Confounders: A Study of Intelligence and Personality," in IEEE Transactions on Software Engineering, vol. 48, no. 12, pp. 4789-4801, 1 Dec. 2022, doi: 10.1109/TSE.2021.3127131.</em></p> <p>- The `data` folder contains dataset and R analysis script. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.<br> - the `materials` folder contains the experimental code snippets and the translated task sheets to evaluate code comprehension performance.</p> <p>Please note that the raw data does not contain the complete data set as we only make the data of those participants public that explicitly agreed to it (which applies to 130 of the 135 participants).</p>
A Large Scale Empirical Study of the Impact of Spaghetti Code and Blob Anti-patterns on Program Comprehension
<p>Dataset and scripts used for the paper.</p>
Impermanent Identifiers: Enhanced Source Code Comprehension and Refactoring
Open the record for dataset details and reuse information.
Predicting Code Comprehension: A Novel Approach to Align Human Gaze with Code Using Deep Neural Networks
<p><strong>Checkout our Github-Repo for more information, issues, and pull requests: </strong></p> <p><a href="https://github.com/Taremeh/predicting-code-comprehension-eye-tracking/">https://github.com/Taremeh/predicting-code-comprehension-eye-tracking/</a></p> <p> </p> <p>Dataset and Replication Package for our paper "Predicting Code Comprehension: A Novel Approach to Align Human Gaze with Code Using Deep Neural Networks"</p>
Comprehensive Analysis of Recurrence-Associated Small Non-Coding RNAs in Esophageal Cancer [Illumina]
GEO Series GSE55854. Homo sapiens. 18 samples. Type: Expression profiling by array.
Inference of pathways, non-coding RNAs and regulatory elements during iron deprivation of Synechocystis based on comprehensive expression profiling
GEO Series GSE39804. Synechocystis sp. PCC 6803. 12 samples. Type: Expression profiling by array.
Comprehensive analysis of Long non-coding RNA expression in dorsal root ganglion reveals cell type specificity and dysregulation following nerve injury [rodent DRG]
GEO Series GSE107180. Mus musculus; Rattus norvegicus. 30 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.
Comprehensive analyses of function and molecular interaction of differentially expressed non-coding RNAs and mRNA in Hantaan virus infection
GEO Series GSE133751. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.