Natural Language-Guided Programming User Study
<p>In this dataset you find the user study data that was used in the <strong><em>Natural Language-Guided Programming</em></strong> paper, which is accepted for Onward! 2021. A preprint can be found here <a href="https://arxiv.org/pdf/2108.05198.pdf">https://arxiv.org/pdf/2108.05198.pdf</a>. The dataset consists of the following files:</p> <ul> <li> <p>benchmark.json contains 201 test cases. Each test case consists of context, a natural language intent and target code. The test cases are intended to evaluate a model that can predict code giving a piece of context code and a natural language intent. The test cases were derived from Jupyter notebooks that were crawled from Github projects with permissive licenses. In the project_metadata field you find information about the original project such as its git url and license.</p> </li> <li> <p>predictions-annotated.json contains predictions of the three models used in the paper for 100 test cases in benchmark.json. Each prediction is accompanied with qualitive assesments from three annotators.</p> </li> <li> <p>train-index.jsonl is the list of github projects that were used for training the models.</p> </li> <li> <p>eval-index.jsonl is a list of github projects that we kept separate for evaluation. The benchmark.json was created from a random subset of the projects in this list.</p> </li> </ul> <p>For more details we refer to the paper.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4