A Single LLM Is an Incomplete Code Reviewer: Evidence that Independent Review by Multiple LLM Families Recovers Code Defects Any One Model Misses
Authors/Creators
Description
Teams increasingly route code review through a single large language model (LLM). We test whether one model's code review is complete relative to independent reviews by several different model families, using a live software team's corpus with a human-reconciled answer key: 33 code/mixed artifacts, 294 confirmed issues, reviewed by fifteen model versions across eight providers (April to July 2026). For each artifact we measure per-model recall against the reconciled confirmed-issue set (a deliberately generous denominator), with Wilson 95% confidence intervals, pairwise cross-family overlap (Jaccard), per-model unique contributions, and a within-artifact coverage curve. No large-sample model exceeded about 61% recall on code; a typical model caught roughly half of confirmed defects (one newest model scored higher, but on only 12 issues, too few to rank). 56.8% of confirmed defects (167 of 294) were found by exactly one model, cross-family overlap was low (median Jaccard about 0.29), and a permutation null confirms this disjointness is not an artifact of the denominator. Computed within each artifact, among only the reviewers that actually reviewed it, coverage of an artifact's confirmed defects rises from about 47% with one reviewer to about 72% with a second (the largest single step), with diminishing returns after; what the data supports is adding a second independent reviewer, not a specific second provider. Six of the eight providers contributed defects no other model caught; the two newest contributed no unique defects in their limited samples. We could not establish that repeated passes vary, that newer versions detect more, or any fine ranking among the non-weakest versions, and we decline to assert them. Within this single-organization case study, a single LLM pass is an incomplete code review, and independent, different-family review recovers the gap. We recommend running a small panel of two to three independent reviewers (different providers are a sensible default this corpus cannot prove beats re-runs), reconciling with a human who verifies findings against source, expecting roughly half to two-thirds single-model code recall, and not adding further models once a few are in place. This version 2 expands the version 1 corpus (18 artifacts, 154 issues, five providers) and incorporates corrections surfaced by additional independent adversarial review, including a within-artifact recomputation of the coverage curve. Conclusions are directional, not a universal benchmark.
Files
LLM_Code_Reviews_v2.pdf
Files
(119.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:426a20603671e4eddcc24e8089897422
|
119.5 kB | Preview Download |