Published April 19, 2026
| Version v12.2
Preprint
Open
Compound Prompting Vanishes Under Matched Compute: A Twelve-Wrapper Audit of Small Language Models on MMLU-Pro
Authors/Creators
Description
</p>
<p>Across 36 (technique × model) cells at n=400 per cell (detecting ≥5 pp effects at 80% power), no cell exceeded plain+CoT with BH-significant support (a single cell, Qwen3-8B + Plan-and-Solve, shows a raw paired Δ of +2.25 pp that does not survive BH at k=36: p=.150, q<sub>BH</sub>=.504, and falls within the audit's undetectable-effect band); three cells on Phi-4 significantly <em>under</em>performed after Benjamini–Hochberg correction at q<0.05.</p>
<p>A mechanistic decomposition attributed the full apparent benefit to CoT alone (+2.23 pp), with persona scaffolding and three-sample Self-Consistency contributing null effects. Eight inter-model aggregation strategies (four in-pool-verifier, four with out-of-pool GPT-4o verifier) all failed on held-out data, with deltas from -0.38 pp to -9.50 pp.</p>
<p><strong>Qwen3-30B-A3B with plain CoT achieves 79.98% on the full 12,032-question MMLU-Pro benchmark</strong>, exceeding GPT-4o's published five-shot-CoT leaderboard value (72.60%) by 7.38 pp at approximately 25× lower inference cost (OpenRouter list pricing), and tying DeepSeek-V3 (80.46%) within noise at comparable pricing. The GPT-4o comparison is protocol-asymmetric and not the primary headline. A contamination-matched frontier peer (Claude Sonnet 4.6) led Qwen3-30B-A3B by <strong>+6.42 pp (McNemar p=2.08×10<sup>-10</sup>)</strong> on a 3-seed stratified 1,200-item paired sample, replicating the within-model plain-CoT gain at frontier scale.</p>
<p>We adopt the label <em>Compute Confound</em> for this specific compound-prompting instance of the compute-asymmetric evaluation problem. Deployment recommendation (for MCQ reasoning at 8B–30B small-open-source scale): a single best-in-class small model plus plain CoT. Within the audited wrapper family on MCQ reasoning at this scale, stop adding wrappers unless compute-matched lift on held-out data can be demonstrated under Benjamini–Hochberg correction.</p>
<p><strong>Project page:</strong> <a href="https://bgml.ai">https://bgml.ai</a><br/>
<strong>Source code:</strong> <a href="https://github.com/BGMLAI">https://github.com/BGMLAI</a></p>
Files
BGML_P1_v12_2_zenodo_bundle.zip
Files
(522.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:0fcd96c6eb30a5d5b60d5d354c427f5b
|
522.5 kB | Preview Download |
Additional details
Dates
- Submitted
-
2026-04-19
Software
- Repository URL
- https://github.com/BGMLAI