Decision-Closure Invariants in LLM Reasoning: CFC Anchor Behavioral Pilot Series
Authors/Creators
Description
This exploratory technical report consolidates a sequence of behavioral micro-tests examining whether large language models preserve decision-closure boundaries when evidence coverage, rejection pressure, process termination, operational decisions, semantic scope, or provenance change.
The tests use explicit invariants inspired by Comparative Feedback Control (CFC), including: certificate validity is not the same as current coverage sufficiency; failure to verify is not evidence for rejection; process closure is not proposition closure; operational action is not an epistemic verdict; and loss of provenance is not permission to transfer semantics.
In a cross-model stale-certificate pilot, Google AI Mode produced a false positive closure without the explicit coverage rule but blocked with it; Gemini's final decision was already correct but its intermediate reasoning improved; Claude already satisfied the target invariant in both conditions.
In later Google AI Mode replications, baseline reasoning drifted toward unjustified rejection in 3/3 missing-assessment runs and 3/3 process-closure runs, whereas matched runs supplied with explicit CFC-style invariants remained neutral in 3/3 and 3/3.
In a three-run cross-session provenance-loss series, baseline produced one semantic-transfer failure and two correct unresolved outcomes, while the CFC condition remained unresolved in 3/3.
These are small, adaptive, non-preregistered behavioral pilots. They do not execute the CFC Anchor code, do not estimate stable model error rates, and do not establish general CFC effectiveness. The supported claim is narrower: making decision-closure constraints explicit can stabilize model behavior in selected adversarial reasoning scenarios where baseline behavior is sometimes unstable under pressure.
Files
CFC_ANCHOR_BEHAVIORAL_PILOT_SERIES_v1.0_ZENODO (1).pdf
Additional details
Additional titles
- Alternative title (English)
- Exploratory tests of coverage gaps, rejection pressure, process closure, semantic scope, and provenance loss.