Hallucination as Semantic Misalignment: A Categorical Approach via Lax Institutions
Description
Large Language Models (LLMs) frequently produce hallucinations—outputs that are syntactically fluent yet semantically unfaithful to facts, sources, or context. While existing research has identified statistical mechanisms, detection methods, and mitigation strategies, a unified semantic foundation remains elusive. This paper reframes hallucination as semantic misalignment, formalized through the lens of Institution theory (Goguen & Burstall, 1992) extended to 2-categories. We model hallucination as a breakdown of the satisfaction condition: the naturality between sentences and models fails under vocabulary morphisms, quantified by lax natural transformations. This formalism captures the dual nature of hallucination—a structural error in fact-sensitive domains (medicine, law, science) yet a creative resource in artistic contexts (fiction, poetry, metaphor). We propose a domain-layered naturality control framework with strict naturality for factual claims, semi-strict for opinions, and lax for creative expression. Combined with external semantic verification (NLI consistency, semantic entropy, retrieval divergence) and evaluation redesign that does not over-penalize uncertainty, our approach provides a principled path toward tuning rather than eliminating hallucinations. We critically examine the pitfalls of guardrail-centric approaches and advocate for transparent, semantics-first alignment.
Files
Kano2025_Hallucination_Semantic_Misalignment.pdf
Additional details
Software
- Development Status
- Active