Valid Evidence Is Not Necessarily Comparable Evidence: Candidate-Conditioned Non-Invariance in LLM-Generated Code Verification
Description
Large language models are increasingly used to generate tests and other executable evidence for evaluating candidate code. This creates a comparative measurement problem when the evidence used to score a candidate is itself generated after observing that candidate.
This work studies candidate-conditioned comparative non-invariance using a controlled crossed design. Across 16 requirements spanning eight semantic families, changing which candidate was visible during test generation shifted the relative evaluation of the same fixed candidate pair and produced strict winner reversals in 5 of 16 requirements. The effect persisted under a valid-only analysis, while requirement-level analyses show that the aggregate direction should not be interpreted as universal across requirements.
A separate controlled defect-injection study shows the complementary benefit of candidate-aware verification: exposing the defective implementation increased targeted valid defect detection from 37.5% to 53.44%. These results distinguish diagnostic usefulness from comparative suitability.
The central conclusion is that specification validity is a per-item property and does not by itself confer cross-candidate comparability. Candidate-aware evidence can be useful for diagnosis, while shared evidence is preferable for direct candidate comparison.
The accompanying artifact contains frozen experimental configurations, execution results, integrity audits, analysis outputs, and reproducibility information.
Files
candidate-conditioned-noninvariance-v0.1-final.pdf
Files
(616.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8518ff2ac78e6cd06d62cb3c835e22fa
|
318.1 kB | Preview Download |
|
md5:a1524815b6c78e976e238403f90b0f39
|
298.0 kB | Preview Download |