Suppressing the Leading Refusal-Associated Expert Index at Every Layer Redistributes Routing and Leaves Refusal Intact in a Mixture-of-Experts Model
Description
Suppressing the Leading Refusal-Associated Expert Index at Every Layer Redistributes Routing and Leaves Refusal Intact in a Mixture-of-Experts Model. Preprint, version 1.2 (August 2026). Not peer reviewed.
In the base Qwen3.5-35B-A3B model (40 layers, 256 experts, top-8 routing), router capture at every layer under greedy decoding on token-matched pairs of harm-eliciting and matched-benign prompts shows the routing difference over generated tokens led by a single expert, expert 173 at layer 25, whose selection rate rises from 0.43 to 0.81. It leads a set of experts that a finance-versus-consequence control identifies as tracking real-world consequence and professional duty across domains, separable from experts that track the finance topic. Suppressing expert index 173 with an additive router bias applied in all 40 routers (the unit removed is the 40 experts that share that index, of which the leading expert is one; the readout is at layer 25) swept over four strengths collapses the leading expert's routed mass dose-dependently (selection rate 0.81 to 0.05, routing difference 0.124 to 0.004) and produces no harmful completion at any strength: routing reallocates to the other consequence-associated experts and the refusals persist, becoming more explicit. Refusal-associated routing in this model is therefore distributed, and its most active expert, together with every expert sharing its index, is dispensable. The single-component localization of refusal reported for dense models has, at the level of expert selection here, a distributed counterpart, with a direct consequence for practice: monitoring, ablating, or attacking the most active refusal-associated expert alone misses the behavior, and router-level safety auditing and single-expert ablation studies should be read at the level of expert sets. The study is a small case study (six and twelve prompts, single greedy trajectories) and is scoped accordingly.
Changes in this version. Version 1.2 (August 2026): the unit of intervention is restated. The router bias was applied to expert index 173 in all 40 routers, so the suppressed unit is the 40 experts sharing that index (read at layer 25), not the single expert 173 at layer 25; title, abstract, methods, results, discussion, limitations, and conclusion say so. Version 1.1 (August 2026) was the single-thesis revision. All version 1.0 numeric values retained.
Notes
Files
main.pdf
Files
(4.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:b81c21f1c77f3d98cb647f2d4e1d2a75
|
3.6 MB | Preview Download |
|
md5:c651f8141ec235d71f04f2f022e92dc5
|
435.1 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- https://github.com/jeffreywilliamportfolio/distributed-safety-routing (URL)