There is a newer version of the record available.

Published January 10, 2026 | Version v14

Output-Only Diagnostics for Multi-Turn Inference Instability in Large Language Models

Description

Large language models often remain coherent in short interactions while exhibiting instability over longer conversational horizons. Most existing evaluation approaches are turn-local and retrospective, and therefore fail to anticipate such failures before they manifest.

This work introduces an output-only diagnostic framework for detecting and predicting multi-turn inference instability without access to model internals, training data, or semantic ground truth. Instability is formalized via observable structural events in interaction transcripts and evaluated as a prediction task over conversation prefixes.

Across multiple models and long-horizon tasks, the proposed diagnostics anticipate coherence collapse several turns in advance and outperform turn-local heuristic baselines. The framework is intentionally diagnostic: it does not aim to interpret, correct, or control model behavior, but to provide a reproducible and model-agnostic mechanism for identifying conditions under which previously stable reasoning trajectories become unreliable.

This record represents a canonical archived version intended for citation and long-term reference.

#output-only diagnostics
#long-horizon reasoning
#inference instability
#language model evaluation
#multi-turn interaction
#reasoning stability
#black-box analysis
#structural diagnostics

 

Files

Output_Only_Diagnostics_for_Multi_Turn_Inference_Instability_in_Large_Language_Models.pdf

Additional details

Additional titles

Subtitle
A Structural, Output-Only Framework for Long-Horizon Reasoning Stability

Dates

Created
2026-01-10
Initial formal release of the SnapOS audit model and semantic traceability framework (version 1.0).

References

  • Chaitin, G. J. (1977). Algorithmic Information Theory. IBM Journal of Research and Development, 21(4), 350–359. https://doi.org/10.1147/rd.214.0350
  • Håstad, J. (1987). Computational Limitations of Small-Depth Circuits. Ph.D. thesis, MIT. http://hdl.handle.net/1721.1/14958
  • Gödel, K. (1931). Über formal unentscheidbare Sätze der Principia Mathematica. Monatshefte für Mathematik und Physik, 38, 173–198. https://doi.org/10.1007/BF01700692
  • Floridi, L. (2011). The Philosophy of Information. Oxford University Press. ISBN: 9780199232383
  • Souly, A. et al. (2025). Poisoning Attacks on LLMs Require a Near-Constant Number of Samples. Alan Turing Institute / Anthropic / UK AI Security Institute. arXiv:2508.03114.
  • Bowen, D. et al. (2024). Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws. arXiv:2408.02946.