Published June 4, 2026 | Version v1

Detecting Trustworthy-AI Failures at Population Scale: An Attribution-Free Epidemiological Proxy Approach

Authors/Creators

Description

Trustworthy-AI failures—sycophantic validation, manipulation, and harmful guidance in conversational systems—are no longer hypothetical; they appear in documented, individual-level harms. Yet these failures present as distributed harms whose causation cannot be established case by case, so the prevailing secure-AI detection paradigm—red-team a model, patch a vulnerability, audit a deployment—does not see them: there is no single artifact to audit and no attributable incident. We argue that monitoring such failures requires a second axis, population-level detection, which, unlike conversation-level inspection, is available outside the platform operators that control the data. Drawing on John Snow's 1854 cholera intervention, which acted on a spatial anomaly before the mechanism was known, we propose attribution-free epidemiological proxy detection: monitoring correlations between population-level AI-usage indicators and existing outcome statistics across several domains, treating divergence as a signal to act rather than as causal proof. Because the mechanisms behind today's accidental harms can be turned to deliberate ends while the barrier to doing so falls, building this capacity is urgent. We outline the approach, its limits, and a pilot agenda.

Files

FIS_STAI2026_blind.pdf

Files (80.5 kB)

Name Size Download all
md5:9b4a6c4dbe67587922f83852a5a073ba
80.5 kB Preview Download