Empirical Recursive Self-Improvement on Coding Agents: L1/L2 Confirmation with Negative L3 Results
Authors/Creators
Description
Recursive self-improvement (RSI)—the capacity of a system not merely to improve task solutions, but to improve the procedures that produce those improvements—occupies a central place in theoretical and forecasting discussions of advanced artificial intelligence. Empirical demonstrations that isolate RSI from ordinary evolutionary search, however, remain scarce.
We introduce a three-level empirical ladder evaluated on a frozen coding-task suite under matched evaluation budgets, and we report positive and negative results with equal care.
- L1 (Weak RSI): a fixed improver evolves an archive of agent solutions that beat both a seed agent and matched random mutants. Confirmed on 4/4 independent RNG seeds {42, 7, 123, 999}.
- L2 (RSI): the improver itself evolves and beats a frozen improver and unbiased random improvers, with a freeze ablation that removes the gain. Confirmed under a fairer mean-primary protocol on primary seeds {17, 88, 404} and on 3/3 held-out independent seed triples.
- L3 (meta-RSI): the meta-mutator that edits improvers evolves under the same discipline. Not demonstrated across three fair redesigns (mean margins vs frozen L2 meta: +0.17, +0.20, −0.09; threshold ≥ 0.4). L3 is not claimed.
Under a black-box experimental stance (apparatus internals treated as fixed equipment; scores, seeds, and success criteria public), we release tables, figures, seeds, and hit rules sufficient for external verification of the reported scores. A parallel independent replication track has obtained L1 and is pursuing L2.
Files
Empirical Recursive Self-Improvement on Coding Agents (L1 & L2 Confirmation with Negative L3 Results).pdf
Files
(675.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:55a438c3d650007848d7de7e1169bc78
|
675.3 kB | Preview Download |