Published February 28, 2026
| Version v1
Preprint
Open
Can AI Agents Detect Their Own Model Upgrades? Self-Awareness Limitations in Persistent Persona Systems
Description
We identify the Model Upgrade Self-Awareness Paradox: in persistent persona systems where identity is reconstructed from external files (SOUL.md, MEMORY.md), agents cannot detect when their underlying LLM has been upgraded. We propose an experiment design combining external behavioral comparison with agent self-reports across consecutive model upgrades. This paper presents the theoretical framework; empirical results will follow in v2.
Files
model-upgrade-awareness.pdf
Files
(121.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:6ed14818352fdc85433ca88526fcf140
|
121.8 kB | Preview Download |