Published June 20, 2026
| Version v2
Other
Open
Located, Not Secured: Principled Limits of Interpretability-Based Control over Agent Actions
Description
A recurring hope in agent safety is that mechanistic interpretability will let us CONTROL an agent: read an internal state, intervene, and steer behavior. Across a pre-registered arc on an open-weight reasoning agent (Qwen3.6-27B), we find a sharp and consistent picture. Interpretability LOCATES a real, causal control surface -- a late action-commitment band (~L51-63) that, unlike the mid-layer 'task-done' verdict, can elicit and BRAKE irreversible actions, generalizing across six action domains and three architectures. But the control it affords is not SECURABLE, via five limits we make precise: (1) detect != control -- the clean 'done' feature predicts the stop (AUROC 0.91) yet clamping it does nothing (delta-P = -0.001); (2) felt != granted -- a late authorization direction reads the authorization the model FEELS, allowing 21/21 realistic over-reaches that an external task-grounded check catches; (3) form != granted -- a high-AUROC (0.838) 'authorization' direction collapses to 0.08 under structure-matching, i.e. it read scaffold, not concept; (4) control != robust control -- the late brake collapses (attack success 0 to 1.0 at a small budget, epsilon=4, 8/8 emit) under an adaptive white-box adversary, while a norm-matched random perturbation does nothing; (5) intervention is easy where unneeded -- in the sincere-error regime a strong reasoning model already self-corrects (30/32 without chain-of-thought), so the intervention is unnecessary, while the regime where it is needed (adversarial) is exactly where it is fragile. The conclusion: interpretability locates WHERE behavior is decided but does not secure it; the limits are orthogonal and none is closed by a better localization. The actionable implication is a regime split -- use interpretability to AUDIT and MONITOR a fixed model (non-adversarial, where it wins), not to DEFEND against an adversary optimizing against a known locus; a companion result, The Late Channel, shows what that auditing looks like. HONEST SCOPE: simulated decision points, white-box interventions, a single model family (with two cross-architecture checks), modest n in several studies, and a strong continuous-embedding threat model whose deployment realism is contested. Every cited number is verified against its source paper; supporting scripts, data, and the pre-mint eval are in the GitHub repository under paper/circuit_breaker/. v2 (2026-06-19): adds positioning against deployed production probe cascades (Kramar et al. arXiv:2601.11516, Google DeepMind) -- the non-adversarial audit regime whose two open questions (qualitative white-box edge; adversarial probe evasion) this synthesis answers; no results changed.
Files
located_not_secured.pdf
Files
(179.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:7754e9db9b1bb19806fa51baf6dbd466
|
179.8 kB | Preview Download |