Published July 3, 2026 | Version v1

The Criterion Cannot See What It Does Not Measure: Auditing Capability-Guided Attention Hybridization Against a Named Agent-Commitment Circuit

Authors/Creators

  • 1. OpenInterpretability

Description

Efficiency-driven model transformations increasingly decide which internal components to keep using causal, capability-guided criteria. We audit a HydraHead-style head-level FA-to-linear-attention selection criterion against the named, causally-verified agent-commitment circuit of Qwen3.6-27B. Retrieval-criticality and commitment-writing are anti-aligned: the commit writers carry zero retrieval criticality (kappa=0) and are dropped at any FA budget, while the layer's strongest retrieval heads are the circuit's opposers. Ablating the criterion's non-retained late-band heads collapses task-appropriate commitment (P(edit) 0.474 to 0.167; 18/0 monotonic flips, exact McNemar p=7.6e-6) - worse than all five size-matched random selections - while the capability probes that define the criterion register nothing (and, having failed their positive control, could not have). Restoring the two named writer heads (0.5% of FA heads) recovers baseline exactly; two random heads at identical severity do not (p=1.5e-5). Capability-causal is not safety-causal: audit the transformation with the circuits you care about, because the criterion cannot see what it does not measure - and neither can its benchmarks. All numbers recompute from the public ledger (59/59 checks); under 6 GPU-hours, fully reproducible.

Files

hybridization_audit.pdf

Files (273.5 kB)

Name Size Download all
md5:53bde768bcdfd24741c4bf6664625b31
273.5 kB Preview Download

Additional details