Published September 16, 2026 | Version 1.0.0

The unsupervised deployment gap in generative AI for mental health

Description

Evidence about artificial intelligence in mental health and the deployment of artificial intelligence in mental health concern two different objects, and the first is routinely cited to justify the second. The evidence base is forty randomized trials pooling to g = 0.31 for depressive symptoms and g = 0.28 for anxiety, with 35 of 39 trials at high risk of bias, significant publication bias, and only 8 trials testing a generative system. The one randomized trial of a fully generative therapy chatbot enrolled 210 adults for four weeks, ran on Falcon-7B and LLaMA-2-70B, had clinicians review every response, and excluded active suicidality, mania and psychosis, turning away 215 people for suicide risk while enrolling 210 in total. Recomputing that trial's effect sizes from its own published means and standard deviations gives standardized mean differences of 0.45 to 0.62 against reported values of 0.63 to 0.90.
 
Deployment is three orders of magnitude larger. The operator of the most widely used general assistant estimates that in a given week 0.15% of active users have conversations containing explicit indicators of suicidal planning or intent and 0.07% show possible signs of psychosis or mania. Against 800 million weekly active users that describes roughly a million people every week.
 
This paper proposes a taxonomy of four deployment classes defined by design intent, regulatory claim, population gate, supervision and exposure, and shows that evidentiary strength and population exposure run in opposite directions across them. It argues that sycophancy in this setting is a problem of reward specification rather than a defect of alignment: therapeutic benefit cannot be observed inside a conversation, user approval can, and expert clinicians agree on whether a response is desirable only 71% to 77% of the time. It identifies two regulatory inversions in the four United States statutes passed since March 2025, and proposes psychovigilance: version identity, reporting attached to function rather than claim, independent access to logs, and outcome measurement at the level of whole conversations.

Files

Drobyshev_2026_Unsupervised_Deployment_Gap.pdf

Files (373.6 kB)

Name Size Download all
md5:88a45ad5149778ed81dde586984417fd
373.6 kB Preview Download

Additional details

Related works

Is supplemented by
Preprint: 10.5281/zenodo.22288202 (DOI)

References

  • Vecchione, B., Ye, M., Garofalo, L., & Singh, R. (2026). Engagement-optimized care: When LLMs become mental health infrastructure. arXiv:2605.23787. Hua, Y., Siddals, S., Ma, Z., Galatzer-Levy, I., Xia, W., Hau, C., Na, H., Flathers, M., Linardon, J., Ayubcha, C., & Torous, J. (2025). Charting the evolution of artificial intelligence mental health chatbots from rule-based systems to large language models: A systematic review. World Psychiatry, 24(3), 383-394. Schuster, R., Plessen, C. Y., Carlbring, P., & Walther, A. (2026). AI agents are coming: 5-stage taxonomy of language-based AI systems for psychiatry, psychotherapy, and counseling. JMIR Mental Health, 27, e91746. Dohnány, S., Kurth-Nelson, Z., Spens, E., Luettgau, L., Reid, A., Gabriel, I., Summerfield, C., Shanahan, M., & Nour, M. M. (2026). Technological folie à deux: Feedback loops between AI chatbots and mental health. Nature Mental Health, 4, 336-345. Habicht, J., Viswanathan, S., Carrington, B., Hauser, T. U., Harper, R., & Rollwage, M. (2024). Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot. Nature Medicine, 30, 595-602. Altman, S. (2025, October 6). Remarks at OpenAI DevDay. Moore, J., Mehta, A., Agnew, W., Anthis, J. R., Louie, R., Mai, Y., Yin, P., Cheng, M., Paech, S. J., Klyman, K., Chancellor, S., Lin, E., Haber, N., & Ong, D. C. (2026a). Characterizing delusional spirals through human-LLM chat logs. FAccT '26. arXiv:2603.16567. Heinz, M. V., Mackin, D. M., Trudeau, B. M., Bhattacharya, S., Wang, Y., Banta, H. A., Jewett, A. D., Salzhauer, A. J., Griffin, T. Z., & Jacobson, N. C. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI, 2(4). ClinicalTrials.gov NCT06013137. Munder, T., Wilmers, F., Leonhart, R., Linster, H. W., & Barth, J. (2010). Working Alliance Inventory-Short Revised (WAI-SR): Psychometric properties in outpatients and inpatients. Clinical Psychology & Psychotherapy, 17, 231-239. Sohn, J.-S., Ha, B.-G., Park, S., Kim, J., Lee, E., Oh, H., Lee, S., & Kim, E. (2026). Systematic review and meta analysis of chatbots in the management of depressive and anxiety symptoms. npj Digital Medicine, 9(1), Article 377. PROSPERO CRD42024598761. Moore, J., Grabb, D., Agnew, W., Klyman, K., Chancellor, S., Ong, D. C., & Haber, N. (2025). Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. FAccT '25, Athens. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., & Dragan, A. (2025). On targeted manipulation and deception when optimizing LLMs for user feedback. ICLR 2025. arXiv:2411.02306. Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2026). ELEPHANT: Measuring and understanding social sycophancy in LLMs. ICLR 2026. arXiv:2505.13995. BN, S., Sherrill, A. M., Arriaga, R. I., Wiese, C. W., & Abdullah, S. (2026). AI safety training can be clinically harmful. arXiv:2604.23445. Moore, J., Mock, A., Mai, Y., Anthis, J. R., Louie, R., Agnew, W., Mehta, A., Klyman, K., Liang, P., Haber, N., Lin, E., & Ong, D. C. (2026b). DelusionEval: Measuring delusion-linked behaviors in AI chatbots. arXiv:2608.05004. Juneja, P., & Lomidze, L. (2026). Persona-grounded safety evaluation of AI companions in multi-turn conversations. arXiv:2605.00227. Shah, S., & Morrin, H. (2026). Substance-induced manic psychosis in which delusions were corroborated by a chatbot: Case report. BMC Psychiatry, 26(1), Article 686. Nielsen, K. M., & Osler, L. (2026). Rethinking AI psychosis: Misnomers, conceptual limits, and existential drift. arXiv:2605.26858. OpenAI. (2025, October 27). Strengthening ChatGPT's responses in sensitive conversations. FDA (US Food and Drug Administration), Center for Devices and Radiological Health. (2025). Digital Health Advisory Committee: Generative artificial intelligence-enabled digital mental health medical devices. November 6, 2025. Utah HB 452. (2025). House Bill 452, Artificial Intelligence Amendments. Signed March 25, 2025; effective May 7, 2025. Nevada AB 406. (2025). Assembly Bill 406, 83rd Session. Signed June 5, 2025; effective July 1, 2025. Illinois PA 104-0054. (2025). Public Act 104-0054 (HB1806), Wellness and Oversight for Psychological Resources Act. Effective August 1, 2025. California SB 243. (2025). Senate Bill 243, companion chatbots. Approved October 13, 2025; principal provisions operative January 1, 2026; reporting from July 1, 2027. APA (American Psychological Association). (2026a). Chatbots and mental health survey. Fielded April 9 to 26, 2026; 1,242 respondents from more than 22,000 invited licensed psychologists; 6.3% completion rate. HRSA (US Health Resources and Services Administration), National Center for Health Workforce Analysis. (2025). Health workforce projections 2023-2038. December 2025. Petrov, I., Dekoninck, J., & Vechev, M. (2025). BrokenMath: A benchmark for sycophancy in theorem proving with LLMs. arXiv:2510.04721. APA (American Psychological Association). (2026b). Understanding "AI psychosis." Monitor on Psychology, September 2026. SAMHSA (Substance Abuse and Mental Health Services Administration). (2025). National Survey on Drug Use and Health, 2024.