Output-Only Diagnostics for Multi-Turn Inference Instability in Large Language Models
Authors/Creators
Contributors
Project leader:
Description
Large language models often remain coherent in short interactions while exhibiting instability over longer conversational horizons. Most existing evaluation approaches are turn-local and retrospective, and therefore fail to anticipate such failures before they manifest.
This work introduces an output-only diagnostic framework for detecting and predicting multi-turn inference instability without access to model internals, training data, or semantic ground truth. Instability is formalized via observable structural events in interaction transcripts and evaluated as a prediction task over conversation prefixes.
Across multiple models and long-horizon tasks, the proposed diagnostics anticipate coherence collapse several turns in advance and outperform turn-local heuristic baselines. The framework is intentionally diagnostic: it does not aim to interpret, correct, or control model behavior, but to provide a reproducible and model-agnostic mechanism for identifying conditions under which previously stable reasoning trajectories become unreliable.
This record represents a canonical archived version intended for citation and long-term reference.
#output-only diagnostics
#long-horizon reasoning
#inference instability
#language model evaluation
#multi-turn interaction
#reasoning stability
#black-box analysis
#structural diagnostics
Files
Output_Only_Diagnostics_for_Multi_Turn_Inference_Instability_in_Large_Language_Models.pdf
Files
(304.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:f1145a76c97a6ced4524d2cc2fd42846
|
304.1 kB | Preview Download |
Additional details
Additional titles
- Subtitle
- A Structural, Output-Only Framework for Long-Horizon Reasoning Stability
Identifiers
Dates
- Created
-
2026-01-10Initial formal release of the SnapOS audit model and semantic traceability framework (version 1.0).
References
- Chaitin, G. J. (1977). Algorithmic Information Theory. IBM Journal of Research and Development, 21(4), 350–359. https://doi.org/10.1147/rd.214.0350
- Håstad, J. (1987). Computational Limitations of Small-Depth Circuits. Ph.D. thesis, MIT. http://hdl.handle.net/1721.1/14958
- Gödel, K. (1931). Über formal unentscheidbare Sätze der Principia Mathematica. Monatshefte für Mathematik und Physik, 38, 173–198. https://doi.org/10.1007/BF01700692
- Floridi, L. (2011). The Philosophy of Information. Oxford University Press. ISBN: 9780199232383
- Souly, A. et al. (2025). Poisoning Attacks on LLMs Require a Near-Constant Number of Samples. Alan Turing Institute / Anthropic / UK AI Security Institute. arXiv:2508.03114.
- Bowen, D. et al. (2024). Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws. arXiv:2408.02946.