Language Models and Simple, Stupid Bugs

Kevin Jesse; Toufique Ahmed; Premkumar Devanbu; Emily Morgan

doi:10.5281/zenodo.7743263

Published February 25, 2023 | Version v2

Software documentation Open

Language Models and Simple, Stupid Bugs

1. University of California, Davis

Abstract—With the advent of powerful neural language models, AI-based systems to assist developers in coding tasks are becoming widely available: Copilot is one such system. Copilot uses Codex, a large language model (LLM), to complete code conditioned on a preceding “prompt". Codex, however, is trained on public GitHub repositories, viz., on code that may include bugs and vulnerabilities. Previous studies [1], [2] show Codex reproduces vulnerabilities seen in training. In this study, we examine how prone Codex is to generate an interesting bug category, single statement bugs, commonly referred to as simple, stupid bugs or SStuBs in the MSR community. We find that Codex and similar LLMs do help avoid some SStuBs, but do produce known, verbatim SStuBs as much as 2x as likely thanknown, verbatim correct code. We explore the consequences of the Codex generated SStuBs and propose avoidance strategies that suggest the possibility of reducing the production of known, verbatim SStubs, and increase the possibility of producing known, verbatim fixes.

Files

Files (3.0 GB)

Name	Size	Download all
msr_sstubs_llms.tar.gz md5:a1ea6a5c49c9ae58735591fdd160ad5f	3.0 GB	Download

Additional details

SHF:Medium: Studying and Exploiting the Bimodality of Software 2107592: U.S. National Science Foundation

	All versions	This version
Views	286	245
Downloads	62	56
Data volume	208.8 GB	188.7 GB

Language Models and Simple, Stupid Bugs

Creators

Description

Files

Files (3.0 GB)

Additional details

Funding