Published August 12, 2026 | Version 1.0.2

Quantization Regime and Governed Routing: A 2x2 Study of QAT Q4_0 versus PTQ Q4_K_M for Gemma 4 12B (dense) and 26B MoE under a Fixed Governance Stack

Authors/Creators

  • 1. MOBIUS LLC

Description

Version 1.0.2 — licence and house-format correction. The measurements, row-level data and conclusions are unchanged from 1.0.1. Version 1.0.1 was deposited under CC BY 4.0 in error; the project's standing convention for text is CC BY-NC-SA 4.0, and this version restores it. Copies distributed under 1.0.1 remain under the terms they were received.

A 2×2 study (quantization regime × released model pair) of governed routing quality for Gemma 4 12B (dense; a clean same-base pair) and a 26B-class MoE released pair on one RTX 5070 Ti, run under a fixed, bit-identical governance stack via Ollama at temperature 0.

Within-pair regime effects (QAT Q4_0 vs PTQ Q4_K_M) are individually marginal and oppositely signed: −0.012 on the dense 12B (exact McNemar p=0.070) and +0.018 on the MoE pair (p=0.078). The regime × pair interaction — an exploratory, single-run endpoint — is +0.030 with a stem-clustered bootstrap 95% CI of [+0.010, +0.052] (Core-500 is 100 paraphrase stems × 5 variants), concentrated in the volatile-current task family. Only 23–32% of same-task temperature-0 outputs are byte-identical across regimes, against same-configuration repeat baselines of 100/100 in all four cells and a 100/100 num_ctx byte-identity control: quantization regime is a behavioral change, not merely a quality change. Attribution is bounded to the released artifact pairs (the 26B pair confounds regime with a possible base revision and a CPU-offload compute path). Safety-critical failure rates are 0.000 in seven of eight runs (one over-verification event, 0.008, in 26B-PTQ Core-500).

The public package (paper, processed data, deterministic analysis script, environment and artifact hashes) is at github.com/mobius-style/gemma4-quant-regime-study (v1.0.1). Verbatim task prompts and generated outputs are excluded by policy and identified by SHA-256. The manuscript was adversarially reviewed pre-deposit by three independent AI reviewers (statistics / consistency / scope); all MAJOR findings were remediated in v1.0.1.

AI assistance: Claude Code (Anthropic, model Fable 5), working method only; the registered author is the human author alone. Companion/sequel to the Gemma 4 MTP quality–throughput study, DOI 10.5281/zenodo.21860461.

Files

PAPER_deposit.pdf

Files (166.7 kB)

Name Size Download all
md5:0a36d44ef759bc85a9c8359f6bb2a42d
166.7 kB Preview Download

Additional details

Related works