Intent Is Not a Property of the Record: Agent Memory Contamination, the Two Detector Families That Cover Disjoint Failures, and the Base Rate Nobody Has Measured
Description
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.
Persistent agent memory has acquired a security literature that models contamination as adversarial injection, and detectors are trained and evaluated on attacker-generated distributions. A separate empirical literature reports the same downstream damage with no attacker present, when an agent writes its own erroneous output into the store and experience-following propagates it. This paper argues that the two produce objects a record-level detector cannot separate, and that this is a structural fact rather than a contingent gap in detector quality. Three lines converge on it: the attack literature is explicitly optimising injected records to lie inside the benign distribution, one prominent attack has the agent generate the malicious record itself, and the dependability taxonomy places the malicious/non-malicious distinction on the fault, defined as the adjudged or hypothesised cause of an error, rather than on the error a detector can observe. The paper then partitions memory defences by the variable each one reads, and finds a coverage split that the field's framing hides: content-based and behavioural detectors can in principle fire on the benign case, but content-based ones face a published impossibility result against an adaptive attacker and behavioural ones a published false-positive inversion, while origin-binding and information-flow defences carry machine- checked guarantees against the adversarial case and are blind to the benign one by construction, because a self-generated wrong record has impeccable provenance. Only outcome- and consistency-based methods read a variable defined for both, and they are the least developed of the four. One published evaluation supplies the decisive measurement: a trajectory signature reaching AUC 0.99 against poisoned traffic yields 100 percent false positives on benign memory-grounded sends under its own preregistered follow-up, and its author concludes that the signature is an attack precondition rather than a maliciousness predicate. The paper corrects the assumption that the benign case is known to be the more common one; the benign-to-adversarial ratio in a deployed store has not been measured, and the argument here does not need it. No experiments are reported. Five studies that would settle the open part are named, and the case against the paper's own position is stated in full.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own text. The author is responsible for the final text and for all claims made in it.
Files
intent-is-not-in-the-record.pdf
Files
(472.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e0de27a6a959a1c74d80d80a644465db
|
472.4 kB | Preview Download |