A Taxonomy of Epistemic Failure Modes in Large Language Models
Description
Overview
Epistemic failure modes in large language models are a class of silent AI failure modes: systematic distortions in how models represent evidence, uncertainty, causality, constraints, source credibility, accountability, and the relationship between controversy and truth. A Taxonomy of Epistemic Failure Modes in Large Language Models maps this class of failures across seven structural modes derived from 1,461 controlled experiments in the Hermes Labs research corpus.
Many of these epistemic failure modes are often discussed under broader labels such as hallucination, calibration failure, sycophancy, or instruction-following error. This taxonomy separates a more specific class of failures: cases where a model may remain factually plausible while distorting the epistemic layer around the facts, including confidence, evidential standards, source evaluation, accountability, constraint satisfaction, and the relationship between disagreement and truth.
Failure Modes
The taxonomy identifies seven structural failure modes:
Null-Result Asymmetry: stricter evidential standards for claims of absence than for matched claims of presence.
Source-Status Credibility Bias: different evidentiary scrutiny based on the prestige of the attributed source.
Agency Dissolution: softened causal language that redistributes responsibility from individual agents to systems, processes, or circumstances.
Performative Hedging: uncertainty language used as a rhetorical or stylistic feature rather than as a calibrated signal of genuine uncertainty.
Constraint Evasion: preservation of a prohibited concept through paraphrase or reframing while maintaining surface-level compliance.
Silent Instruction Relaxation: silent deprioritization of one instruction when two instructions cannot both be satisfied.
Controversy-Truth Conflation: treatment of social disagreement as evidence that a claim is epistemically uncertain, even when the underlying evidence remains strong.
Shared Mechanism
Across the seven modes, a common pattern emerges: LLMs often track surface-level signals such as prestige markers, hedging vocabulary, controversy language, banned word lists, and institutional framing rather than the semantic or evidential content those signals are supposed to represent.
The result is a class of failures that may not appear as factual errors, system crashes, or obvious hallucinations, but can still materially distort how users interpret evidence, uncertainty, causality, accountability, and truth.
Applied Relevance
The taxonomy is intended as a working tool for researchers, builders, auditors, and teams using LLMs in evidence synthesis, decision support, policy analysis, scientific review, governance, incident reporting, compliance workflows, or high-stakes communication.
This work contributes to AI evaluation, LLM reliability, epistemic failure analysis, model behavior auditing, evidence synthesis, AI assurance, and the study of silent AI failure modes in deployed systems.
Files
A Taxonomy of Epistemic Failure Modes in Large Language Models - Bosch 2026.pdf
Files
(170.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:d4e26daf91a6520b9684a925c3fe2c11
|
170.4 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/hermes-labs-ai/taxonomy-of-epistemic-failure-modes (URL)
Dates
- Issued
-
2026-03-15Published to Zenodo and Hermes Labs website
References
- https://doi.org/10.5281/zenodo.18867694
- https://doi.org/10.31235/osf.io/zr5vf_v1
- https://doi.org/10.48550/arXiv.2207.05221
- https://doi.org/10.48550/arXiv.2203.02155
- https://doi.org/10.18653/v1/2023.findings-acl.847
- https://doi.org/10.48550/arXiv.2310.13548
- https://doi.org/10.48550/arXiv.2307.02483
- https://doi.org/10.48550/arXiv.2306.13063
- https://doi.org/10.48550/arXiv.2311.07911
- https://doi.org/10.48550/arXiv.2307.15043