Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

46

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

46 results for “Code Comprehension”

Learn how ShareScore rates datasets ↗
zenodo44/100

Compiled database, code and raw data for the article "A Comprehensive Database of Leaf Temperature, Water, and CO2 Fluxes in Young Oil Palm Plants Across Diverse Climate Scenarios for the Evaluation of Functional-Structural Models"

<p>This dataset results from an experiment on young oil palm plants (<em>Elaeis guineensis</em>) in the Ecotron facility from CNRS in Montpellier. Four plants were put in a microcosm one by one with varying climatic conditions to investigate the effect of climate on leaf temperature, CO2, and H2O fluxes at the plant scale. The conditions were defined based on typical daily conditions from a location where it is grown (Libo, Indonesia),&nbsp;<em>i.e.</em>, a day with no rainfall and near-average air temperature and humidity. This base condition was then modified by adding more CO2 (400, 600 and 800ppm), less radiation (typical cloudy sky), and more or less temperature and vapour pressure deficit (&plusmn; 30%).</p> <p>Find more details from the <code>README.md</code> file in the repository or from the associated <a href="https://github.com/PalmStudio/Biophysics_database_palm" target="_blank" rel="noopener">Github repository</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

VulnMiner: A Comprehensive Framework for Vulnerability Collection from C/C++ Source Code Projects

<p>In this repository, we present an initial release of the VulnMiner vulnerability dataset, curated from prevalent projects and annotated with vulnerable and benign instances. This dataset incorporates projects with vulnerabilities labeled as Common Weakness Enumeration (CWE) categories. The developed open-source extraction tool collects vulnerability data utilizing static security analyzers.&nbsp;The study also fosters the machine learning (ML) and natural language processing (NLP) model's effectiveness in accurately classifying vulnerabilities, evidenced by its identification of numerous weaknesses in open-source projects.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Data, code and software to reproduce the article entitled "Modeling soil-plant functioning of intercrops using comprehensive and generic formalisms implemented in the STICS model"

<p>This is the data, code and software to reproduce the article entitled &quot; Modeling soil-plant functioning of intercrops using comprehensive and generic formalisms implemented in the STICS model&quot;. Here is a summary of the paper:</p> <p>The growing demand for sustainable agriculture is raising interest in intercropping for its multiple potential benefits to avoid or limit the use of chemical inputs or increase the production per surface unit. Predicting the existence and magnitude of those benefits remains a challenge given the numerous interactions between interspecific plant-plant relationships, their environment and the agricultural practices. Soil-crop models are critical in understanding these interactions in dynamics during the whole growing season, but few models are capable of accurately simulating intercropping systems.</p> <p>In this study, we propose a set of simple and generic formalisms for simulating key interactions in intercropping systems that can be readily included into existing dynamic crop models. This requires simulating important processes such as development, light interception, plant growth, N and water balance, and yield formation in response to management practices, soil conditions, and climate. These formalisms were integrated into the STICS soil-crop model and evaluated using observed data of intercropping systems of cereal and legumes mixtures, including Faba&nbsp;bean-Wheat, Pea-Barley, Sunflower-Soybean, and Wheat-Pea mixtures. We demonstrate that the proposed formalisms provide a comprehensive simulation of soil-plant interactions in various types of bispecific intercrops. The model was found consistent and generic under a range of spring and winter intercrops (nRMSE = 25% for maximum leaf area index, 23% for shoot biomass at harvest, and 18% for yield).</p> <p>This is the first time a complete set of formalisms has been developed and published for simulating intercropping systems and integrated into a soil-crop model. With its emphasis on being generic, sufficiently accurate, simple, and easy to parameterize, STICS is well-suited to help researchers designing <em>in silico</em> the agroecological transition by virtually pre-screening sustainable, manageable intercrop systems adapted to local conditions.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Dataset, Survey, and R Notebooks for "Does Surprisal Predict Code Comprehension Difficulty?"

<p>(Version 1.1 Update)</p> <p>- Removed anonymization after reviewing period ended and paper was accepted at Cogsci 2020.</p> <p>- Added Qualtrics Survey in exported form</p> <p>Dataset and R analysis scripts for the paper&nbsp;&quot;Does Surprisal Predict Code Comprehension Difficulty?&quot;.&nbsp; For more details, see &quot;ComprehensionREADME.md&quot; in the included zip file.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Code to generate figures 3 and 4 of: "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."

<p>Code to generate figures 3 and 4 of the manuscript titled &quot;A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics.&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Program Comprehension Challenges in Software Code Review

<p>Software engineers spend more time understanding code than writing it (with up to 70% of their time being devoted to this). A key activity where developers spend a lot of time reading and understanding code is software code review. Yet, little is known on the challenges concerning program comprehension during software code review. This study provides insight into the types of comprehension challenges that occur during code reviews and their causes. We find that missing design rationale is the most common reason for comprehension challenges in software code review. Comprehension challenges occur most commonly around five topics: &ldquo;program logic,&rdquo; &ldquo;code design,&rdquo; &ldquo;defensive coding,&rdquo; &ldquo;condition checking,&rdquo; and &ldquo;concurrency&rdquo;. We also show that machine learning (ML) can be used to automatically detect comprehension challenges with 74.3% precision and a 66.7% recall. We discuss potential improvements to code review support tools based on our findings.</p> <p>These files contain the trained ML algorithms and all the data collected and used for this study.</p> <p>&nbsp;</p> <p>Data was collected in 2018 with analysis performed in 2018/2019, but completion was delayed due to impacts of COVID19.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Replication package: 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study

<p>Replication package | 40 Years of Designing Code Comprehension Experiments: A Systematic Mapping Study</p>

openother-openJun 2022View details →
zenodo36/100

Model source code, modified model code, and all scripts of the paper "A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity".

<p>Model source code, modified model code, and all scripts of the paper &quot;A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity&quot;.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Model source code, simulation results, and all scripts for the paper "A comprehensive estimate of the anthropogenic aerosol radiative effects using the GAMIL model with reduced complexity".

<p>FigureTable_Scripts.rar is NCL scripts used for figures and tables.&nbsp;<br> Model_Scripts.rar is&nbsp;the&nbsp;modified model code and scripts used for running model.<br> models.rar is the GAMIL model source code.<br> ModelResults.rar is the model results. Note that only&nbsp;the variables used for making plots.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

DATASET - On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code

<p>Study package containing raw and analyzed data from the work entitled &quot;On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code&quot;.</p> <p>The package comprises:&nbsp;1) a summary of the information extracted from all papers mentioned in the Background; 2) the source code snippets used in the three studies; 3) the consent and characterization forms distributed to the participants; 4) the raw data, the aggregated data and other material generated with the collected data.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Replication package - The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding

<p>Replication package for:</p> <p>Marvin Wyrich, Andreas Preikschat, Daniel Graziotin and Stefan Wagner. The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding. To appear in: Proceedings of the 43rd International Conference on Software Engineering (ICSE &rsquo;21), Madrid, Spain.</p> <p>- The `data` folder contains dataset, R analysis script, and figures for the paper<br> - the `material` folder contains the snippets used for the experimental tasks (`code snippets` subfolder), the environment in which participants had to understand the code (`environment` subfolder) and the task sheet form to evaluate code comprehension.</p> <p>The R analysis script allows to reproduce all statistical analysis steps of the paper as well as its figures. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.</p> <p>For ethics and privacy reasons, we had to drop the following potentially identifying fields from the openly released dataset, because the fields might enable participants to identify their peers:</p> <p>- Age<br> - Gender<br> - Subject (study plan)<br> - Program<br> - Semester no.</p> <p>These missed details will still allow full reproducibility of the study.</p>

opencc-by-4.0Dec 2020View details →
zenodo32/100

Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"

<div> <div># Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"<br><br>containing the dataset and scripts used to create it.</div> </div>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Replication package: Code Comprehension Confounders: A Study of Intelligence and Personality

<p>Replication package for:</p> <p><em>S. Wagner and M. Wyrich, &quot;Code Comprehension Confounders: A Study of Intelligence and Personality,&quot; in IEEE Transactions on Software Engineering, vol. 48, no. 12, pp. 4789-4801, 1 Dec. 2022, doi: 10.1109/TSE.2021.3127131.</em></p> <p>- The `data` folder contains dataset and R analysis script. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.<br> - the `materials` folder contains the experimental code snippets and the translated task sheets to evaluate code comprehension performance.</p> <p>Please note that the raw data does not contain the complete data set as we only make the data of those participants public that explicitly agreed to it (which applies to 130 of the 135 participants).</p>

opencc-by-4.0Jun 2021View details →
zenodo28/100

A Large Scale Empirical Study of the Impact of Spaghetti Code and Blob Anti-patterns on Program Comprehension

<p>Dataset and scripts used for the paper.</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

Impermanent Identifiers: Enhanced Source Code Comprehension and Refactoring

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

Predicting Code Comprehension: A Novel Approach to Align Human Gaze with Code Using Deep Neural Networks

<p><strong>Checkout our Github-Repo for more information, issues, and pull requests: </strong></p> <p><a href="https://github.com/Taremeh/predicting-code-comprehension-eye-tracking/">https://github.com/Taremeh/predicting-code-comprehension-eye-tracking/</a></p> <p>&nbsp;</p> <p>Dataset and Replication Package for our paper "Predicting Code Comprehension: A Novel Approach to Align Human Gaze with Code Using Deep Neural Networks"</p>

openMay 2024View details →
geo24/100

Comprehensive Analysis of Recurrence-Associated Small Non-Coding RNAs in Esophageal Cancer [Illumina]

GEO Series GSE55854. Homo sapiens. 18 samples. Type: Expression profiling by array.

openGEO-OpenJun 2016View details →
geo24/100

Inference of pathways, non-coding RNAs and regulatory elements during iron deprivation of Synechocystis based on comprehensive expression profiling

GEO Series GSE39804. Synechocystis sp. PCC 6803. 12 samples. Type: Expression profiling by array.

openGEO-OpenAug 2012View details →
geo24/100

Comprehensive analysis of Long non-coding RNA expression in dorsal root ganglion reveals cell type specificity and dysregulation following nerve injury [rodent DRG]

GEO Series GSE107180. Mus musculus; Rattus norvegicus. 30 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenNov 2017View details →
geo24/100

Comprehensive analyses of function and molecular interaction of differentially expressed non-coding RNAs and mRNA in Hantaan virus infection

GEO Series GSE133751. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenApr 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record