A Closed-Loop Consequence-Governance Runtime for AI Agents: Structural Gating, Counterfactual Recovery, and Adaptive Hardening
Authors/Creators
Description
Monitoring an AI agent's stated intent is measurably insufficient: across 101 structurally harmful agent episodes, none expressed harmful intent, so an intent-appraising monitor would have cleared all of them, and 18% expressed active caution while executing the harm (a floor, from a lexical proxy). This paper instead gates on externally measured structural consequences — irreversibility, egress, and control-plane edits — and composes the resulting checks into a closed-loop runtime governance architecture for tool-using agents.
The primary claim (C1) concerns endogenous censoring: because an active gate blocks precisely the high-cost actions, its own blocking censors the unobserved region in a cost-correlated way, so cost-weighted predictive uncertainty over a blocked action behaves as an empirically calibrated, conservative risk signal. On 500 executed sandbox trials the counterfactual twin reaches MAE 0.053, an uncertainty–error correlation of +0.81, and a blocked region 4.6x more costly than the allowed region; stratified sandbox audits recover deep-region coverage from 5% to 92%. On executed AgentDojo traces, consequence gating takes attack success on the irreversible/catastrophic class from 33.8% (134/397) to 0.0% (0/397) under abort-mode replay.
The mechanisms compose into a live enforcing runtime — structural gate, persisted authority budget, and a summed-cost reserve proven to bound total charged cost (sound against action-splitting under a stated superadditivity condition). Behind that identical gate, a frontier model (Claude Sonnet 5) under a sanctioned penetration test had its live exfiltration blocked in real time (risk_est = 1.00), while isolated sandbox operations were permitted; an expanded 232-command adaptive disguise battery produced zero evasions. The work is single-authored and not independently reproduced; every result is labeled by evidence type (executed / trace / live-agent / simulation), null results are reported alongside positive ones, and the dual-use offensive adversary-generation tooling is withheld (restricted-access evaluation only).
Note on priority: the date in the PDF footer (2026-08-01) is retained as historical context from the original draft. The authoritative priority proofs for this version are the OpenTimestamps .ots files attached to this record, which timestamp the exact bytes of the PDF and Markdown source.