Published June 17, 2026 | Version v1

When Defenses Backfire: Model-Dependent Failure Modes in Prompt-Injection Defense for LLM Agents

  • 1. ROR icon Malaviya National Institute of Technology Jaipur

Description

Prompt-injection defenses for large language model (LLM) agents are deployed under the implicit assumption that they are, at worst, neutral—that adding a defense cannot make an agent more vulnerable than leaving it undefended. This paper shows that assumption is false and characterizes the resulting failure mode, termed defense backfire. In a factorial evaluation of five production LLMs across four task suites, three defenses, and four attacks (240 conditions on AgentDojo), the same defense sharply reduces attack success on one model yet increases it on another—lowering one model’s attack success rate (ASR) by up to 86.7 percentage points (pp) while raising another’s by 31.2 pp under the identical static attack. Unlike adaptive-attack results, where defenses fall to stronger attacks, backfire is driven by target selection. The cause is established by intervention: backfire tracks the defense-induced change in tool-use activity (r=0.72 on attackable models), and externally capping the tool budget monotonically removes it, isolating re-exposure as the mechanism for the largest effects. A defense can even be strictly dominated, raising ASR and lowering utility at once. Thus prompt-injection defenses cannot be deployed model-agnostically; their safety must be established per model.

Files

When Defenses Backfire Preprint.pdf

Files (593.2 kB)

Name Size Download all
md5:048eb7f41e0c456a4909976423c4e2e9
593.2 kB Preview Download