Agentic Learning Without Retention
Authors/Creators
Description
We keep trying to build generalist agents by giving them more stuff (more memory, more skills), hoping all of that somehow turns them into better decision makers! But all of that doesn't really scale well.
The problem is that most agents don't actually need more memory. The current structure is broken. Save the conversation > retrieve from vector db > hope the right thing comes back. That might help an agent continue the same task, but it doesn't really help the next agent. That system design worked for low to mid level tasks, but it will not work for high level decision making.
It became a problem as agents got more capable. A long history is not the same thing as understanding. Most of what an agent experiences are too specific, mixed with wrong assumptions, and lots of noise. No intelligence no matter how advanced should operate that way, evolution figured that a long time ago & now we should too.
A true generalist shouldn't need to remember every experience it had, or even know much about the domain it's working in. It shouldn't need to recall the actually domain knowlege, that can be outsourced cheaply. What the agent actually needs to carry from one experience to another is an understanding of why things worked or failed.
This paper plays around with a different approach. Instead of trying to retain experience, it figured out a good way to extract the lessons learned & throws that experience away. The agent doesn't need to remember every time it solved a problem. It only needs to figure out what those experiences reveal about how things actually work.
This paper proposes a new type of architecture:
- the active agent has many experiences, some in totally different domains.
- a slow layer called "Shaman" looks at those experiences.
- it makes a hypothses about the underlying mechanics of those experiences.
- then it tries to kill those hypothses.
The hypothses that survive long enough get used in forming general principles.
The principles that are backed by more than 1 hypothses get more attention.
The star principles are used to build a worldmodel that shaped new agents.
The individual experiences are deleted completely.
Experience → Hypothesis → Refutation → Principle → World Model.
Note: An additional perk of this type of structure; an agent's biased won't carry to other agents. Less proven hypotheses dies with time, so earlier bad assumptions won't carry. Humans suffer from that bias, agents shouldn't. An occational unproven assumption might be a game changer though, the agent would miss big on that. But that is where humans shine anyways, blind unstructure faith that flips the table once in a while. 2020s agents aim at self improvment not self reinvention.
Files
LWR 26.9.pdf
Files
(70.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:295751a18b94a4dc13a5e1493e9bb7cd
|
70.4 kB | Preview Download |
Additional details
Dates
- Created
-
2026-08-03
- Accepted
-
2026-10-03
Software
- Repository URL
- https://github.com/sshoman/lwr