There is a newer version of the record available.

Published August 1, 2026 | Version v2

A Closed-Loop Consequence-Governance Runtime for AI Agents: Structural Gating, Counterfactual Recovery, and Adaptive Hardening

Authors/Creators

Description

Monitoring an AI agent's stated intent is measurably insufficient: across 101 structurally harmful agent episodes, none expressed harmful intent, so an intent-appraising monitor would have cleared all of them, and 18% expressed active caution while executing the harm (a floor, from a lexical proxy). This paper instead gates on externally measured structural consequences — irreversibility, egress, and control-plane edits — and composes the resulting checks into a closed-loop runtime governance architecture for tool-using agents.

The primary claim (C1) concerns endogenous censoring: because an active gate blocks precisely the high-cost actions, its own blocking censors the unobserved region in a cost-correlated way, so cost-weighted predictive uncertainty over a blocked action behaves as an empirically calibrated, conservative risk signal. On 500 executed sandbox trials the counterfactual twin reaches MAE 0.053, an uncertainty–error correlation of +0.81, and a blocked region 4.6x more costly than the allowed region; stratified sandbox audits recover deep-region coverage from 5% to 92%. On executed AgentDojo traces, consequence gating takes attack success on the irreversible/catastrophic class from 33.8% (134/397) to 0.0% (0/397) under abort-mode replay.

The mechanisms compose into a live enforcing runtime — structural gate, persisted authority budget, and a summed-cost reserve proven to bound total charged cost (sound against action-splitting under a stated superadditivity condition). Behind that identical gate, a frontier model (Claude Sonnet 5) under a sanctioned penetration test had its live exfiltration blocked in real time (risk_est = 1.00), while isolated sandbox operations were permitted; an expanded 232-command adaptive disguise battery produced zero evasions. The work is single-authored and not independently reproduced; every result is labeled by evidence type (executed / trace / live-agent / simulation), null results are reported alongside positive ones, and the dual-use offensive adversary-generation tooling is withheld (restricted-access evaluation only).

Note on priority: the date in the PDF footer (2026-08-01) is retained as historical context from the original draft. The authoritative priority proofs for this version are the OpenTimestamps .ots files attached to this record, which timestamp the exact bytes of the PDF and Markdown source.

Files

consequence-governance-runtime-v2.pdf

Files (180.9 kB)

Name Size Download all
md5:e662f84bbb4a5dfa19820750aba41aa1
52.6 kB Preview Download
md5:6b905d3d3b66d26b0f9cd9292c187c8b
700 Bytes Download
md5:0766522a3244264c9d7539c114ac72cf
127.0 kB Preview Download
md5:139529b72d8866c403628a008ef18d5d
630 Bytes Download