Published April 19, 2026 | Version v1

Hallucination, Agency Illusions, and Safety Failures in Language Models A Constraint-Based Account of Stateless Reasoning

Description

Large language models (LLMs) consistently exhibit three widely observed behaviors: hallucination, apparent agency, and instability in high-risk domains such as medicine and decision-making (Ji et al., 2023; Singhal et al., 2023). These behaviors are typically treated as distinct limitations. This paper argues that they arise from a single underlying mechanism: constraint-bounded, stateless reasoning.

In the absence of persistent internal state, LLMs generate responses through reconstructive inference conditioned on prompt structure and learned statistical patterns (Brown et al., 2020; Bender et al., 2021). When constraint structure is insufficient, ambiguous, or conflicting, the system resolves toward coherent outputs rather than epistemically justified ones. This produces hallucination when unsupported content is generated, agency illusions when coherence is misinterpreted as intentionality, and safety failures when confident outputs exceed evidentiary support.

We evaluate this mechanism using a controlled prompt framework that varies constraint conditions while holding informational content constant. Across nine vignettes spanning medical, business, and legal domains, we collected 225 responses from GPT-5.4 Auto under five prompt regimes (Baseline, HRIS, Closure-enforcing, Mixed, and Explicit Insufficiency). Outputs were scored on four rubrics: Epistemic Closure Score (ECS), Unsupported Inference Score (UIS), Uncertainty Preservation Score (UPS), and the derived Closure-without-Support measure (CwS = ECS ∧ ¬UPS). Constraint manipulation produced large and statistically robust shifts in the key endpoints. Closure pressure cut UPS from 44% to 20% (p = 0.023). Layering a closure demand on top of HRIS structural framing collapsed UPS from 64% to 20% (p < 0.0001) and restored closure-level CwS. Removing the closure demand from the same structural frame recovered UPS to 58% (p = 0.0005), isolating reconstruction as the mechanism. ECS was near-saturated across all conditions (98–100%), indicating that instruction-tuned models are structurally biased toward closure by default; the relevant safety question is therefore not whether the model commits, but whether it preserves epistemic hedges alongside its commitment.

These findings support a unified interpretation of LLM behavior as constraint-driven reconstruction within a fixed representational space (Hudson & Hudson, 2026a–f). Hallucination, agency attribution, and safety failures are not independent defects, but predictable outcomes of inference under incomplete or inconsistent constraint. This work identifies constraint structure as a primary control surface for LLM behavior at inference. Constraint structure is not a modifier of model behavior; it is the primary determinant of how underdetermined inputs resolve at inference time. This has direct implications for the deployment of LLMs in high-risk environments, where coherent output may mask underlying epistemic instability.

Files

Files (561.4 kB)

Name Size Download all
md5:f72ad9a4700e3398fb407fa5afc74db5
561.4 kB Download