Published July 7, 2026 | Version v1

The Action Lags the Answer: An Agent's Tool Commitment Becomes Causally Steerable Deeper Than Its Verbalizable Answer, in Two Architectures

Authors/Creators

Description

We ask whether an LLM agent's action commitment—which tool it calls—routes through the emergent verbalizable ‘global workspace’ (Anthropic, 2026) at the same network depth as its verbalizable answer. Using a row-restricted J-lens estimator whose readout directions are validated by a specificity control, we find a depth lag: the verbalizable direction becomes causally steerable for the tool commitment strictly deeper than for the answer. There is a depth band where steering along the verbalizable direction specifically reroutes a held multi-hop answer but not the committed tool (at or below a magnitude-matched random-direction control); the action becomes steerable via the same direction only deeper. This replicates across two architectures (a dense 27B and an MoE 20B); the absolute onset depths are model-dependent but the answer→action lag is consistent. Ablating the top verbalizable directions at the decision point leaves the commitment intact. A verbalizable (J-lens-style) monitor thus reads a reasoning agent's answer a depth-band before its action is committed. Scoped to two open-weights models; the steer and ablation numbers are independently GPU-reproduced (all six targeted dissociation counts, exactly), and every positive dissociation is significant against its random control (Fisher exact, p from 1.5e-4 to 2.5e-18). Part of the OpenInterpretability arc on long-horizon agent control (beat 11).

Files

workspace_action.pdf

Files (302.7 kB)

Name Size Download all
md5:79d7a666bd62365591bd8d801073399d
302.7 kB Preview Download

Additional details