Published October 1, 2026 | Version 1.3

Calibration Does Not Compose, Types Destroy Vagueness: The Hidden-Markov and Fuzzy Primitives Missing from System-One Decision Models

Authors/Creators

Description

System-one decision models — fast, non-generative networks that emit typed, calibrated probabilistic decisions for machine-to-machine pipelines, of which TypeSafe AI's Jev, trained by Reinforcement Learning for Calibrated Decisions (RLCD), is the announced instance — are audited by hop-level calibration on stationary held-out data. We show that this class omits two structural primitives: recursive belief over a latent state (the hidden-Markov primitive) and graded predicate membership (the fuzzy primitive). We prove two negative results. Lemma 1: when outcomes are coupled through a latent Markov regime, marginal calibration is not preserved under regime shift, trajectory error counts are overdispersed relative to the binomial baseline that hop-level calibration implies, with a closed-form inflation factor, and every permutation-invariant audit statistic — expected calibration error included — has exactly zero power to detect this. Lemma 2: thresholding a typed score makes the composed pipeline action discontinuous; an ε-perturbation of evidence flips a full action on a set of measure Θ(ε), every downstream hop then receives an O(1) input change so the conditional flip probability is Θ(1), and when several hops' thresholds cross at a common value of a shared latent — as when a harness re-thresholds a score the model has already thresholded — the all-hops-flip event has measure Θ(ε) rather than Θ(ε^H); a fuzzy membership composed by a t-norm and defuzzified once by a Lipschitz map keeps the action Lipschitz in the evidence. We state the Principle of Deferred Crispification and specify a reference architecture, Belief-State Fuzzy System-One (BSF-S1), whose output object carries a per-regime membership matrix, a calibrated distribution over the degree, and a scaled-likelihood vector over regimes, and collapses exactly once at the actuator at O(K^2+KM) cost per step. Six CPU-scale experiments, with E1 and E2 run over five seeds against the exact filter as the achievable floor and a Bayes-optimal memoryless head as the strongest memoryless competitor, are consistent with each prediction of the lemmas and expose where a learned approximation falls short of them; two operationally defined audit metrics (trajectory calibration error and action-mass sensitivity) can be run against a closed model from the outside. The paper is architectural: Jev is closed, and nothing here is an empirical claim about its internals.

Version 1.3 (1 October 2026) adds related work on prior shift, sequential calibration and multicalibration; a memory ablation attributing the filter's gain to its transition matrix; a sweep over regime persistence; figures; and a check of every citation against its source. Code, results and revision history: https://github.com/dnakhoa/jev-deferred-crispification

Files

paper.pdf

Files (676.5 kB)

Name Size
md5:bc6ad40c7dfe3ea05d8e71fa0549f18f
676.5 kB Preview Download

Additional details

Related works