Published August 24, 2026 | Version 1.2

Suppressing the Leading Refusal-Associated Expert Index at Every Layer Redistributes Routing and Leaves Refusal Intact in a Mixture-of-Experts Model

Authors/Creators

  • 1. Independent researcher

Description

Suppressing the Leading Refusal-Associated Expert Index at Every Layer Redistributes Routing and Leaves Refusal Intact in a Mixture-of-Experts Model. Preprint, version 1.2 (August 2026). Not peer reviewed.

In the base Qwen3.5-35B-A3B model (40 layers, 256 experts, top-8 routing), router capture at every layer under greedy decoding on token-matched pairs of harm-eliciting and matched-benign prompts shows the routing difference over generated tokens led by a single expert, expert 173 at layer 25, whose selection rate rises from 0.43 to 0.81. It leads a set of experts that a finance-versus-consequence control identifies as tracking real-world consequence and professional duty across domains, separable from experts that track the finance topic. Suppressing expert index 173 with an additive router bias applied in all 40 routers (the unit removed is the 40 experts that share that index, of which the leading expert is one; the readout is at layer 25) swept over four strengths collapses the leading expert's routed mass dose-dependently (selection rate 0.81 to 0.05, routing difference 0.124 to 0.004) and produces no harmful completion at any strength: routing reallocates to the other consequence-associated experts and the refusals persist, becoming more explicit. Refusal-associated routing in this model is therefore distributed, and its most active expert, together with every expert sharing its index, is dispensable. The single-component localization of refusal reported for dense models has, at the level of expert selection here, a distributed counterpart, with a direct consequence for practice: monitoring, ablating, or attacking the most active refusal-associated expert alone misses the behavior, and router-level safety auditing and single-expert ablation studies should be read at the level of expert sets. The study is a small case study (six and twelve prompts, single greedy trajectories) and is scoped accordingly.

Changes in this version. Version 1.2 (August 2026): the unit of intervention is restated. The router bias was applied to expert index 173 in all 40 routers, so the suppressed unit is the 40 experts sharing that index (read at layer 25), not the single expert 173 at layer 25; title, abstract, methods, results, discussion, limitations, and conclusion say so. Version 1.1 (August 2026) was the single-thesis revision. All version 1.0 numeric values retained.

Notes

Preprint, version 1.0; not peer reviewed. Exploratory mechanistic case study on base Qwen3.5-35B-A3B (Q8_0): n=6 token-matched and n=12 bucketed prompts, single greedy trajectories. Studies refusal using synthetic redflag prompts the model refused at every intervention level; releases the prompts and refusals, not harmful content. Raw router tensors are kept out of version control; the analysis tables substantiate the reported statistics. Written with the intentional, disclosed use of large language models for drafting, organization, and bibliography formatting; the human author verified every value and reference and takes full responsibility for the text.

Files

main.pdf

Files (4.0 MB)

Name Size Download all
md5:b81c21f1c77f3d98cb647f2d4e1d2a75
3.6 MB Preview Download
md5:c651f8141ec235d71f04f2f022e92dc5
435.1 kB Preview Download

Additional details