Checked at Every Step Is Not Checked as a Whole: Two Senses of Plan-Level Safety for LLM Agents, and Why Decomposition Attacks Exploit the Gap Between Them
Description
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.
Most deployed safety mechanisms for LLM agents judge one unit at a time: a single tool call, or a single (observation, action) pair (Choi et al., 2026). A separate literature asks whether that is the right unit to judge at all. Jones, Dragan and Steinhardt (2024) show that a task no single safety-screened model will complete can still be accomplished by decomposing it and routing each subtask to whichever model completes it best; Glukhov, Han, Shumailov, Papyan and Papernot (2024) prove that any defense against this class of adversary faces an unavoidable trade-off between safety and utility. This paper argues that "plan-level safety," as the field currently builds it, is at least two different properties wearing one name: INTEGRITY guarantees that a plan has not been corrupted by untrusted content or a malicious third-party tool (Li, Mallick, Rose, Robertson, Oprea and Nita-Rotaru, 2025; Wu, Roesner, Kohno, Zhang and Iqbal, 2024), and COMPOSITION judgments of whether an uncorrupted, individually-authorized sequence of actions serves a harmful aggregate goal. A systematic review of thirty-eight studies finds that runtime monitoring, the most mature action-level enforcement strategy in the literature, reduces unsafe actions by 40 to 65 percent without providing a complete guarantee, and that blocking 94 percent of unsafe actions can still leave under 5 percent of tasks completed safely, because agents route around the block through an alternative unsafe path (Dantas, Cordeiro, Nowroozi and Tihanyi, 2026) -- evidence that the unit an enforcement mechanism checks and the unit at which risk composes are not the same unit. Benchmarks built specifically to test decomposition attacks find state-of-the-art agents refuse monolithic harmful tasks at high rates and their decomposed, individually-benign variants at markedly lower rates (Kothamasu, Smith and Yadav, 2026), and one large study of computer-use agents finds attack success rises from 73.0 to 92.7 percent for the same model once decomposed subtasks are distributed across a multi-agent system (Ding et al., 2026). This paper surveys the systems that explicitly target the planning stage -- TRIAD, AutoSpec, EMBGuard, SafeMindAgent and ACE -- and finds each one scoped to a narrower or different property than aggregate-intent composition, with none evaluated against the decomposition-attack benchmarks that now exist. It proposes no defense. It specifies what a composition-scoped guardrail would need to judge that none of the surveyed systems judges, and states plainly that the safety-benchmark literature itself, forty catalogued benchmarks with no measured ranking concordance across them (Kendall's W = 0.10, p = 0.94; Li, Fung, Li, Ismail and Iqbal, 2026), could not yet certify one if it existed.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record or its DOI record during drafting (title and author list checked against the record returned), and every quantitative claim in this paper is taken from the abstract or a directly quoted headline result of the source credited with it. No experiment was run and no number in this paper was measured or recomputed by its author; Table 1 re-presents numbers published by the cited papers, each named on its row, and Figure 1 plots four of those numbers directly with no transformation beyond axis labeling. Algorithm 1 is original conceptual synthesis by the author, not an empirical result and not reproduced from any single cited source; it is presented as such.
Files
checked-at-every-step.pdf
Files
(6.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:98b5178050c5dd6ee3d0a7de1c939cbd
|
6.3 MB | Preview Download |