Published April 24, 2026 | Version v1

Polyhedral-Constrained Reasoning: A Polyhedral Combinatorics Framework for Structurally Optimal Training of Large Reasoning Models

Authors/Creators

Description

Large reasoning models (LRMs) significantly improve the performance of complex tasks by generating long-chain reasoning, but at the same time introduce serious problems of redundant reasoning and computational waste. Existing methods (such as Self-Braking Tuning) mainly control the reasoning length through heuristics or posterior adjustment, lacking strict structural constraints and optimality guarantees. This paper proposes a novel training paradigm—Polyhedral-Constrained Reasoning (PCR)—which introduces polyhedral representation theory from Polyhedral Combinatorics into reasoning trajectory modeling, formalizing the "effective reasoning space" as a high-dimensional polyhedron, and treating each reasoning path as a point in this polyhedron. By constructing a "reasoning feasible polytope" and its facests, the model learns to approximate the optimal reasoning extreme point during training, thereby achieving structural elimination of redundant reasoning. PCR further introduces a dynamic cutting plane generation mechanism to gradually remove unnecessary reasoning patterns from the feasible polytope, achieving an optimal trade-off between reasoning length and correctness. This method theoretically unifies inference efficiency, correctness, and structural optimality, providing a training framework with interpretable geometric structure for the next generation of LRM.

Files

Polyhedral-Constrained Reasoning A Polyhedral Combinatorics Framework for Structurally Optimal T.pdf