Published August 29, 2026 | Version 1.0

A Step Score Is Not a Step Verdict: Three Quantities Under the Name Process Supervision, and the Outcome-Only Results That Make the Distinction Load-Bearing

Authors/Creators

  • 1. SONYTECH

Description

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.

Process reward models are introduced, almost without exception, as models that score whether an individual reasoning step is correct. This paper argues that the phrase "process supervision" is currently attached to three different target quantities, that the dominant scalable label defines the second of them rather than the first, and that the field's step-level benchmarks score the first while its downstream gains are argued for in terms of the third. The three are step validity, a property of a trajectory prefix; prefix value, the probability that some completion policy reaches the correct final answer from that prefix; and step advantage, the change in that probability across a step. Monte Carlo estimation defines prefix value, and one 2026 paper states the consequence directly: the resulting rewards are policy-dependent where step correctness should not be. A leading published argument for the process-reward paradigm makes that policy-relativity explicit rather than accidental, holding that progress should be measured under a prover policy distinct from the base policy and that weak provers can improve stronger ones. The empirical fact that forces the distinction into the open is an anomaly on the validity benchmark itself: on ProcessBench, models trained with no step-level labels at all repeatedly match or beat models trained with them, and one paper reports that adding step labels to an outcome-trained model brings no further improvement. Three readings of that anomaly are set out, together with the dissociation analysis that would separate them, which requires no new annotation. What a validity label would have to supply that a value label does not is stated, along with the label sources whose semantics is validity. No experiments are reported here. The strongest case against this paper's position, including a theorem that cuts against its premise, is stated in full.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Files

a-step-score-is-not-a-step-verdict.pdf

Files (452.2 kB)

Name Size Download all
md5:5512fa2b9b41865699fec1fe31ca46ee
452.2 kB Preview Download