Whose Goals? Autotelic Agents Generate Goals Within Spaces They Are Given: The Goal Space and the Referee Stay Outside the Agent, and a Foundation Model in the Loop Relocates Them Rather Than Removing Them
Description
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.
Autotelic agents are described as learning to represent, generate, select and solve their own goals, and a new generation of systems built on foundation models is reported to do so without hand-coded goal representations, without human intervention, or from zero data. This paper separates four things that the phrase "the agent generates its own goals" bundles together: the goal space in which any goal can be expressed, the selector that decides which goal to pursue next, the success test that decides whether a goal was reached, and the referee that decides whether the goals produced were new or worth having. An audit of 28 published systems and system families, from engineered-goal-space robotics to language-model task proposers, records where each component sits. The referee is outside the agent in all of them, but that is structural rather than a finding: a published evaluation is by definition its authors'. The substantive results concern the other three. The selector is computed by the agent in all 28. The goal space is written by the designers in the engineered systems and learned from designer-chosen data or objectives in the learned-space systems. In every foundation-model system it is bounded, on top of what pretraining makes available, by something the designers wrote or chose: a prompt, a seed list, an output format, a task-type menu, a simulator or the environment itself. The success test is supplied in about half of the rows (at least 13 of the 28), and where it is the model's own, the two systems that measured it against an external reference found it less accurate on later or harder self-generated goals. Where goals leave the region occupied by human examples, human raters score them as less understandable and less human-like, though the published analysis cannot separate a referee that fails to credit new goals from a generator that produces worse ones. The published ablations show that the designer-held components are not inert, but they do not rank them above the self-directed ones: in Absolute Zero, removing two designer-chosen task types and replacing the proposer's own earlier tasks with a fixed prompt cost comparable amounts, and the LMA3 comparison is qualitative. A foundation model in the loop moves the space and the referee from the designer's grammar into a pretraining corpus and a prompt; it does not remove them. On present evidence, the foundation-model systems generate goals within a supplied space, which is Sigaud et al.'s first-order open-ended case, and the tests that would support a second-order reading have not been run.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref record, during drafting (title and author list checked against the record returned). The full texts (arXiv HTML or PDF renderings) of the load-bearing sources were read for the passages and numbers attributed to them, including MAGELLAN, Voyager, OMNI, OMNI-EPIC, LMA3, IMAGINE, ACES, Absolute Zero, R-Zero, Automated Capability Discovery, Minimo, Enhanced POET, the IMGEP-UGL paper of Pere et al., DADS, Goals as Reward-Producing Programs, and the definitional papers of Sigaud et al., Hughes et al. and Sheth et al. Every quantitative claim is taken from the abstract, main text or a table of the source credited with it; where two published numbers are subtracted, the text says so. No experiment was run and no number in this paper was measured by its author. Table 2 and Figure 1 re-present published values, each named with its source. The four-component decomposition, the audit classification in Table 1, Algorithm 1 and the proposed tests in Section 13 are conceptual synthesis by the author, not empirical results, and are presented as such.
Files
whose-goals.pdf
Files
(458.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:b1c9e3fe45a4db276a4fe9527b5d95c7
|
458.0 kB | Preview Download |