AgentShield: A Zero-Trust Runtime Guardrail Architecture for Autonomous Multi-Agent AI Systems with Bidirectional Context Synchronization
Description
The rapid migration of Large Language Models (LLMs) from conversational interfaces to autonomous multi-agent software engineering systems has exposed profound security vulnerabilities. When autonomous agents operate across heterogeneous topologies—spanning cloud-hosted reasoning engines and local execution environments—they are acutely vulnerable to indirect prompt injection, tool-call hijacking, privileged command escalation, and memory poisoning. Conventional boundary defenses (such as input sanitizers and prompt wrappers) fail to address lateral privilege escalation between collaborating agents.
To resolve this critical architectural vulnerability, this paper introduces AgentShield, a zero-trust runtime verification and guardrail framework designed for decentralized multi-agent systems operating over the Model Context Protocol (MCP). AgentShield enforces continuous, non-bypassable policy verification across all intra-agent communications and system tool dispatches. The framework incorporates: (1) an inline bidirectional semantic interceptor that evaluates agent intents before system execution, (2) a multi-lingual token triage engine capable of detecting adversarial jailbreaks in low-resource and code-switched dialects, (3) a cryptographically signed persistent shared ledger ensuring tamper-evident state continuity, and (4) an automated capability attenuator for operating system and file operations.
We evaluate AgentShield across 1,500 adversarial scenarios covering multi-step tool-injection benchmarks and real-world developer workflows. Empirical results demonstrate that AgentShield mitigates 98.4% of prompt injection and tool-escalation attacks while introducing less than 11.8ms of median runtime latency overhead. The core architecture is validated via two production-grade open-source packages released on the Python Package Index (PyPI): prema-agentshield and gemini-antigravity-bridge.
Files
AgentShield__A_Zero_Trust_Runtime_Guardrail_Architecture_for_Autonomous_Multi_Agent_AI_Systems_with_Bidirectional_Context_Synchronization__2_.pdf
Files
(178.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e85b9547b5ef6fac744286cff989b225
|
178.2 kB | Preview Download |