Published September 1, 2026 | Version v1

Auditable patch validity for lightweight language-model software repair

Authors/Creators

  • 1. Independent Researcher

Description

Terminal resolve rate obscured where repository-level language-model patches lost utility between generation and execution. This study addressed that mea- surement gap by proposing an auditable patch-validity framework whose con- ceptual novelty was to represent each artifact as monotonic survival through six evidence-linked checkpoints: non-empty generation, syntactic validity, reposi- tory applicability, harness completion, behavioral resolution, and review readi- ness. The first five were measured from archived generations, normalization records, and Docker-harness outcomes; review readiness remained a proposed deployment check. In a 300-artifact diagnostic slice, 64 failed syntax, 77 failed applicability, 151 reached execution but remained unresolved, and 8 resolved. Across 1,600 larger-run artifacts, 776 (48.5%) were invalid or inapplicable and 22 resolved. Staged prompting improved completion in some configurations, but its strongest diagnostic resolution advantage, 4/50 versus 1/50, was exploratory (p = 0.25) and reversed at larger scale. By converting terminal outcomes into syntactic, repository-state, and semantic failure tiers, the framework enabled low-cost pre-execution filtering, targeted remediation, and auditable continuous integration (CI) resource allocation. It was a measurement and triage instrument, not a new repair algorithm.

Files

42714 IJECE 6% latex yekti.pdf

Files (740.4 kB)

Name Size Download all
md5:f09caecdf564e8f4a9674ac8eff13144
740.4 kB Preview Download