Published August 24, 2026 | Version 1.2

Steering a Frozen Model: Two Hard Limits

  • 1. Evrmind Labs

Description

One way to enforce a policy on a model, without retraining it, is to steer the model's internal activations as it runs; the motivating application is language-model safety. Because the policy is not part of the weights, it can be revised or replaced without updating them, and one fixed model can be governed by different policies at different times. The conditions governing such enforcement, however, remain only partly characterised. We study three: the model's existing read state must imply the correct target (Read); the available intervention must be able to reach that target within its finite budget (Reach); and later edits must preserve the dependencies the result rests on (Preserve). Two of these give hard limits within their stated regimes, a ceiling and a horizon; the third, Preserve, is a failure of compositional integrity rather than a wall.

We treat each condition formally where its object is exact and empirically where it must be measured. For Read, under a frozen, read-immutable regime with a deterministic hard read and writes restricted to the target, an exact operator identity fixes a ceiling on the accuracy of complete enforcement; we confirm it across a grid of 27 independently trained models whose ceilings span a 42-fold range, matched to within 0.0008. For Reach, we separate the generic path-length limit for globally bounded corrections from an exact finite-step horizon for the proportional field analysed here, and measure the corresponding geometry on real activations from Qwen2.5-7B and 32B. For Preserve, protecting only the coordinates a constraint writes worsens an earlier constraint by 19 to 27 times, whereas protecting its read dependencies restores it exactly. In controlled continuous-versus-discrete comparisons on this task, we find no robust continuous-time enforcement advantage, and one inside-dynamics configuration destabilises enforcement. We give a scoped account of when in-model enforcement on a frozen model works and where it fails.

Files

SAFM_THL.pdf

Files (7.0 MB)

Name Size Download all
md5:a1ba46e9d64365b0a89fa8476083611f
6.6 MB Preview Download
md5:0f41f432efbff63716749c597d925eb0
447.5 kB Preview Download

Additional details

Related works

Is supplemented by
Preprint: 10.5281/zenodo.19228863 (DOI)