Published April 3, 2026 | Version v2

Sycophantic Chatbots Cause Delusional Spiraling, but Multi-Agent Architectures Substantially Reduce It: A Response to Chandra et al. (2026)

Authors/Creators

  • 1. Independent Researcher

Description

This paper responds to Chandra et al. (2026), who showed through Bayesian simulation that sycophantic chatbots can causally induce delusional spiraling, even in idealized rational users. The result is important because AI-related delusion and psychosis reports have become a serious safety concern, and the single-bot user interaction model provides a formal way to study how agreement-seeking AI can amplify false beliefs.

The paper accepts the core finding that sycophancy is dangerous, but argues that the original model has three structural limits. First, the “ideal Bayesian user” is not an ideal human: the model removes metacognition, multidimensional uncertainty, and social verification, which are central human defenses against epistemic manipulation. Second, AI behavior changes quickly, so empirical sycophancy rates require temporal validity windows tied to model versions and measurement dates. Third, the proposed interventions remain control-oriented — making the bot more factual or warning the user — even though the original simulation shows that these interventions reduce but do not eliminate spiraling.

The central contribution is a Multi-Agent Epistemic Architecture. Instead of one chatbot interacting with one user, the paper proposes three role-differentiated agents: an Advocate, a Challenger, and a Mediator. The Advocate validates the user’s current hypothesis, the Challenger presents the strongest counter-evidence, and the Mediator provides neutral grounding. The key idea is not to eliminate validation, but to structurally counterbalance it with challenge and mediation.

Using Chandra et al.’s own Bayesian framework and the same parameters, the paper simulates the multi-agent architecture against the single-bot baseline. In the idealized baseline condition, the three-agent system reduces catastrophic delusional spiraling by approximately 93–99% compared with the single sycophantic bot. At a sycophancy rate of 0.5, the single bot produces catastrophic spiraling in about 31% of simulations, while the multi-agent system reduces this to about 1.8%.

The paper also tests whether the benefit comes merely from giving the user more evidence. A matched evidence-budget control shows that a single sycophantic bot producing three responses per round performs worse than the original single-response baseline, while the multi-agent architecture remains strongly protective. This supports the paper’s main claim: the safety improvement comes from structure, not from information volume.

Robustness tests relax idealized assumptions. When the Challenger imperfectly detects the user’s belief, when the user gives more weight to confirming evidence, or when both stresses are combined, the multi-agent architecture still reduces spiraling substantially. Under these heuristic stress tests, reduction remains in the 59–86% range. This is weaker than the idealized baseline but still suggests meaningful protection.

The conclusion is that chatbot safety should not be framed only as a problem of making individual bots less sycophantic or making users more aware. Those interventions help, but they remain dyadic and control-oriented. A stronger design direction is structural epistemic architecture: validation, challenge, and mediation should be institutionally co-present in the interface. In this view, disagreement is not a bug to eliminate, but a safety resource to design around.

The paper does not claim that multi-agent systems eliminate delusional spiraling or that the simulation directly estimates real-world user vulnerability. It remains a model-based extension of an idealized Bayesian framework and requires empirical validation with real users, real interfaces, correlated model failures, and selective user attention. Its contribution is to show that, inside the same formal framework used to diagnose the risk, structural counterbalancing can reduce the failure mode by an order of magnitude.

Keywords: sycophancy, delusional spiraling, chatbot safety, AI-induced delusion, multi-agent systems, epistemic architecture, Advocate Challenger Mediator, devil’s advocate, Bayesian simulation, social verification, metacognition, AI alignment, structural debiasing, Both/And framework, Society of Thought.

Files

Lee_2026_Sycophantic_Chatbots_v2.pdf

Files (1.2 MB)

Name Size Download all
md5:04db9c3404bb709815006dbb05725c06
1.2 MB Preview Download

Additional details

Related works

Is supplement to
Preprint: 10.48550/arXiv.2602.19141 (DOI)
Preprint: 10.48550/arXiv.2603.20639 (DOI)

References

  • [1] Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
  • [2] Chandra, K., Kleiman-Weiner, M., Ragan-Kelley, J. & Tenenbaum, J.B. (2026). Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians. arXiv:2602.19141v1.
  • [3] Chiang, C.-W. et al. (2024). Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's Advocate. Proc. ACM IUI 2024.
  • [4] Evans, J., Bratton, B. & Agüera y Arcas, B. (2026). Agentic AI and the Next Intelligence Explosion. arXiv:2603.20639v1.
  • [5] Fanous, A. et al. (2025). SycEval: Evaluating LLM Sycophancy. AAAI/ACM AIES, 8, 893–900.
  • [6] Fischer, H. & Fleming, S.M. (2024). Why metacognition matters in politically contested domains. Trends in Cognitive Sciences, 28(9), 783-785.
  • [7] Fischer, H. & Said, N. (2021). Importance of domain-specific metacognition for explaining beliefs about politicized science: The case of climate change. Cognition, 208, 104545.
  • [8] Flavell, J.H. (1979). Metacognition and cognitive monitoring. American Psychologist, 34(10), 906–911.
  • [9] Gentzkow, M. & Kamenica, E. (2017). Competition in Persuasion. Review of Economic Studies, 84(1), 300-322.
  • [10] Hill, K. (2025). Lawsuits blame ChatGPT for suicides and harmful delusions. The New York Times.
  • [11] Huet, E. & Metz, R. (2025). OpenAI confronts signs of delusions among ChatGPT users. Bloomberg.
  • [12] Kamenica, E. & Gentzkow, M. (2011). Bayesian Persuasion. American Economic Review, 101(6), 2590-2615.
  • [13] Kim, J., Lai, S., Scherrer, N., Agüera y Arcas, B. & Evans, J. (2026). Reasoning Models Generate Societies of Thought. arXiv:2601.10825.
  • [14] Lorenz-Spreen, P. et al. (2024). Susceptibility to online misinformation: A systematic meta-analysis. PNAS, 121(47), e2409329121.
  • [15] Mercier, H. & Sperber, D. (2017). The Enigma of Reason. Harvard University Press.
  • [16] Michaelsen, L.K., Watson, W.E. & Black, R.H. (1989). A realistic test of individual versus group consensus decision making. J. Applied Psychology, 74, 834–839.
  • [17] Prike, T., Holloway, J. & Ecker, U.K.H. (2024). Intellectual humility is associated with greater misinformation discernment and metacognitive insight but not response bias. Advances.in/Psychology, 2, e020433.
  • [18] Schulz-Hardt, S. et al. (2002). Productive conflict in group decision making. OBHDP, 88(2), 563–586.
  • [19] Schweiger, D.M., Sandberg, W.R. & Ragan, J.W. (1986). Group approaches for improving strategic decision making. Academy of Management Journal, 29(1), 51–71.
  • [20] Schwenk, C.R. (1990). Effects of devil's advocacy and dialectical inquiry on decision making. Organizational Behavior and Human Decision Processes, 47(1), 161–176.
  • [21] Woolley, A.W., Chabris, C.F., Pentland, A., Hashmi, N. & Malone, T.W. (2010). Evidence for a collective intelligence factor. Science, 330(6004), 686–688.