Published July 28, 2026 | Version v1

Removing an Attacking AI Agent's Ability to Prove Attack Success

Authors/Creators

Description

We describe a class of environments for storing and transmitting text messages in which an attacker, even holding all the data, cannot prove that an extracted result is the true message rather than one of many plausible candidates.

The mechanism is structural, not computational: the extraction procedure is total — it never returns an error, and some fraction of access parameters yields coherent, contextually valid text. After discarding incoherent outputs, the attacker is left with a set of meaningful candidates among which there is no way to single out the true one.

Key claim: unprovability follows not from the size of the candidate space but from the absence of a selection criterion. Even if there were only three candidates, there would be nothing by which to tell the true one apart.

For an attacking AI agent this means that even with full access to the data it cannot complete the attack operationally — it cannot determine when to stop, cannot confirm success to whoever commissioned the attack, and cannot use the result as evidence.

The result is conditional: it rests on an explicitly stated assumption of statistical uniformity, which is postulated as a property of a specific construction rather than proven, and it applies only to the single-query case. The primary contribution is not the theorem but an architectural framework that makes this assumption empirically testable. An open testbed with the full data set is provided for independent replication attempts.

This report refines and narrows the framing of an earlier work (doi.org/10.5281/zenodo.21261173), focusing specifically on the provability of extraction success.

Files

Removing an Attacking AI Agent’s Ability to.pdf

Files (351.1 kB)

Name Size Download all
md5:8d8717ac6cde2df59a53753e5f03639f
351.1 kB Preview Download

Additional details

Related works

Continues
Preprint: 10.5281/zenodo.21261173 (DOI)

References

  • Christiano, P. et al. (2021). Eliciting Latent Knowledge. Alignment Forum.
  • Hubinger, E. et al. (2019). Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv:1906.01820.
  • Juels, A., & Ristenpart, T. (2014). Honey Encryption.
  • Shannon, C. E. (1949). Communication Theory of Secrecy Systems.
  • Shannon, C. E. (1951). Prediction and Entropy of Printed English.
  • Sharma, P. et al. (2023). Towards Understanding Sycophancy in Language Models.
  • Savchenko, D. (2026). [название июльской работы]. Zenodo. https://doi.org/10.5281/zenodo.21261173