Detecting Trustworthy-AI Failures at Population Scale: An Attribution-Free Epidemiological Proxy Approach
Authors/Creators
Description
Trustworthy-AI failures—sycophantic validation, manipulation, and harmful guidance in conversational systems—are no longer hypothetical; they appear in documented, individual-level harms. Yet these failures present as distributed harms whose causation cannot be established case by case, so the prevailing secure-AI detection paradigm—red-team a model, patch a vulnerability, audit a deployment—does not see them: there is no single artifact to audit and no attributable incident. We argue that monitoring such failures requires a second axis, population-level detection, which, unlike conversation-level inspection, is available outside the platform operators that control the data. Drawing on John Snow's 1854 cholera intervention, which acted on a spatial anomaly before the mechanism was known, we propose attribution-free epidemiological proxy detection: monitoring correlations between population-level AI-usage indicators and existing outcome statistics across several domains, treating divergence as a signal to act rather than as causal proof. Because the mechanisms behind today's accidental harms can be turned to deliberate ends while the barrier to doing so falls, building this capacity is urgent. We outline the approach, its limits, and a pilot agenda.
Files
FIS_STAI2026_blind.pdf
Files
(80.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:9b4a6c4dbe67587922f83852a5a073ba
|
80.5 kB | Preview Download |