Published September 2, 2026 | Version v1

Layer-Scoped Expert-Budget Expansion Discovers Succinct Convergence in Sparse Mixture-of-Experts Reasoning: Reducing Reasoning Tokens at Matched Accuracy on MMLU-Pro

Authors/Creators

Description

This paper introduces layer-scoped expert-budget expansion, a training-free, runtime-only routing modification for sparse Mixture-of-Experts (MoE) language models. By expanding the expert selection budget ($N \ge K$) exclusively within late transformer layers and applying a linear decay factor to additional experts, the method uncovers a phenomenon termed succinct convergence: holding output correctness constant, the expanded model arrives at the same correct final answer via significantly shorter reasoning trajectories.

Evaluated on the complete MMLU-Pro benchmark (714 questions) using Qwen3.6-35B-A3B, a ties-only paired comparison (577 common correct answers) demonstrates an 8.5% reduction in mean reasoning tokens and a 10.9% decrease in latency (p = 6.5 10^-6}, while maintaining overall accuracy statistically indistinguishable from native routing (84.5% vs 84.0%, p=0.77).

Files

Layer-Scoped Expert-Budget Expansion Discovers Succinct Convergence in Sparse Mixture-of-Experts Reasoning_ Reducing Reasoning Tokens at Matched Accuracy on MMLU-Pro.pdf

Additional details

Dates

Submitted
2026-09-02

Software

Repository URL
https://github.com/vagrillo/llama.cpp/blob/moe-expansion/docs/moe-expansion.md
Programming language
C , C++
Development Status
Wip