Published April 19, 2026 | Version v12.2

Compound Prompting Vanishes Under Matched Compute: A Twelve-Wrapper Audit of Small Language Models on MMLU-Pro

Description

</p>

<p>Across 36 (technique × model) cells at n=400 per cell (detecting ≥5 pp effects at 80% power), no cell exceeded plain+CoT with BH-significant support (a single cell, Qwen3-8B + Plan-and-Solve, shows a raw paired Δ of +2.25 pp that does not survive BH at k=36: p=.150, q<sub>BH</sub>=.504, and falls within the audit's undetectable-effect band); three cells on Phi-4 significantly <em>under</em>performed after Benjamini–Hochberg correction at q&lt;0.05.</p>

<p>A mechanistic decomposition attributed the full apparent benefit to CoT alone (+2.23 pp), with persona scaffolding and three-sample Self-Consistency contributing null effects. Eight inter-model aggregation strategies (four in-pool-verifier, four with out-of-pool GPT-4o verifier) all failed on held-out data, with deltas from -0.38 pp to -9.50 pp.</p>

<p><strong>Qwen3-30B-A3B with plain CoT achieves 79.98% on the full 12,032-question MMLU-Pro benchmark</strong>, exceeding GPT-4o's published five-shot-CoT leaderboard value (72.60%) by 7.38 pp at approximately 25× lower inference cost (OpenRouter list pricing), and tying DeepSeek-V3 (80.46%) within noise at comparable pricing. The GPT-4o comparison is protocol-asymmetric and not the primary headline. A contamination-matched frontier peer (Claude Sonnet 4.6) led Qwen3-30B-A3B by <strong>+6.42 pp (McNemar p=2.08×10<sup>-10</sup>)</strong> on a 3-seed stratified 1,200-item paired sample, replicating the within-model plain-CoT gain at frontier scale.</p>

<p>We adopt the label <em>Compute Confound</em> for this specific compound-prompting instance of the compute-asymmetric evaluation problem. Deployment recommendation (for MCQ reasoning at 8B–30B small-open-source scale): a single best-in-class small model plus plain CoT. Within the audited wrapper family on MCQ reasoning at this scale, stop adding wrappers unless compute-matched lift on held-out data can be demonstrated under Benjamini–Hochberg correction.</p>

<p><strong>Project page:</strong> <a href="https://bgml.ai">https://bgml.ai</a><br/>

<strong>Source code:</strong> <a href="https://github.com/BGMLAI">https://github.com/BGMLAI</a></p>

Files

BGML_P1_v12_2_zenodo_bundle.zip

Files (522.5 kB)

Name Size Download all
md5:0fcd96c6eb30a5d5b60d5d354c427f5b
522.5 kB Preview Download

Additional details

Dates

Submitted
2026-04-19

Software