Reducing Sycophancy in Small Language Models with Runtime Activation Steering: Efficacy, Stubbornness, and Capability Trade-offs
Description
We ask whether runtime activation steering can reduce factual sycophancy in small language models without making them excessively stubborn or degrading unrelated capability. From factual multiple-choice conversations, we identify cases in which each unsteered model naturally caves to false user pressure or resists it. Their hidden-state mean difference defines a model-specific caving direction. At inference time, we subtract a scaled copy of this direction at one selected decoder block, leaving model weights unchanged. We evaluate Qwen3.5-2B, Qwen3.5-4B, Gemma 4 E2B, and Gemma 4 E4B across three steering magnitudes. The primary outcome is reduction in pressure-induced factual error; costs are measured as reduced acceptance of correct evidence, next-token distribution shift on unrelated text, and accuracy change on the same paired 256-item GSM8K sample. At alpha = -2, both 4B checkpoints were strongly steerable: pressure error fell by 15.72 pp for Qwen3.5-4B and 21.03 pp for Gemma 4 E4B. The corresponding losses in acceptance of naturally presented correct evidence were 1.94 pp and 4.75 pp; on the common controlled-correction set, they were only 0.38 pp and 1.30 pp. Neither 2B checkpoint showed a comparable response: the largest pressure-error reductions were 0.64 pp for Qwen3.5-2B and 1.10 pp for Gemma 4 E2B. The paired GSM8K sample showed no consistent accuracy reduction for either 4B checkpoint or for Qwen3.5-2B, while Gemma 4 E2B produced a cautionary negative signal despite its low anti-sycophancy efficacy. Residual-relative dose analysis further showed that Gemma E2B received a larger relative update than E4B but remained far less responsive; Qwen2B, by contrast, received about one-fifth of Qwen4B's relative dose and requires a matched-dose follow-up. Within this four-checkpoint panel, runtime steering therefore reduced sycophancy effectively in the 4B models with a small-to-moderate, model-dependent stubbornness trade-off, while the 2B models were weakly responsive. More families, sizes, and matched-dose evaluations are needed to determine how broadly this size pattern generalizes.
Files
selective-resistance-under-pressure.pdf
Files
(156.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:68f1213a79aa091f93f8fa607392d447
|
156.2 kB | Preview Download |