Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang
cs.AI, cs.CL, cs.CV, cs.LG
2026-10-08
A frozen LLM maintains a rulebook compiled into executable code; it clears all 25 ARC-AGI-3 games at RHAE 100.0 with 44% of human actions, and a learned Pong controller wins 21:0.
The hard part of dropping an LLM agent into an unfamiliar game is rarely inventing good actions; it is not knowing how the world works. Code world models are an established answer: write the environment's transition rules as an executable function and let a planner search inside it. Generalisation is where that breaks. A finite interaction history leaves a set of programs that all reproduce the observed transitions yet disagree on unseen states. The paper's example: if every wall seen so far sits at column 7, “a box stops at a wall” and “a box stops at column 7” explain the same trajectory and only diverge when a wall shows up elsewhere. Replay rejects inconsistent programs; it cannot choose among the consistent ones. Earlier Memento papers worked the policy side, episodic memory and reusable skills. This one takes the model side: how environment rules get learned explicitly, stored, and continually revised while the LLM stays frozen.
The core design, Code as Model, is a pair of artefacts:
The split is deliberate. The rulebook carries rule-level reasoning and stays inspectable; the code fixes how rules are applied. An implementation bug is repaired without touching the rulebook, while a genuine change in understanding must be written into the rulebook first and only then compiled. A new executable passes two gates: the LLM judges it faithful to the rulebook, and cell-exact replay must reproduce every recorded transition; failure at either gate sends the proposal back to reflection. “The LLM thinks it is right” becomes “it runs the entire history”.
At runtime this is a five-stage loop: observe, reflect (retain, revise the rulebook, or repair the code), revise, compile, verify. A verified model goes to a planner that searches a shortest path inside the executable; every executed step is checked against observation, and the first mismatch halts execution and feeds the counterexample back to reflection. Model history lives in a Git repository the agent itself can diff and roll back. When several candidates pass verification, the system prefers the shortest rulebook-plus-code description, an MDL-style simplicity bias.
A single model carries a self-confirmation risk: a wrong but unfalsified hypothesis may keep choosing actions that never expose its error. The population extension keeps N rulebook-executable pairs on one shared interaction history, samples one behavioural equivalence class per planning episode, and repairs falsified members individually; N=1 recovers the single-model setting. The backbone is Claude Opus 5 in Claude Code at extra-high reasoning effort, frozen throughout; the harness is built on the baseline1 codebase, which the paper discloses.
Main experiment: all 25 public games; only real environment actions count against the budget, while planning, replay, and model construction are free.
| Method | Games cleared | Mean RHAE |
| Continual Harness | 3/25 | 20.5 |
| DreamTeam | 6/25 | 38.1 |
| OPINE-World | 20/25 | 78.4 |
| NOOA | 19/25 | 85.1 |
| baseline1 | 25/25 | 99.0 |
| Memento 3 (single model) | 25/25 | 100.0 |
RHAE 100 is the ceiling and requires clearing every level at the efficiency bar. Total actions: 7,518, or 44% of the 17,135-action human baseline; baseline1, the only other system to finish all games, used 8,347 (49%). The same backbone under the ARC Prize standard harness scores 59.3 RHAE points lower, a comparison that isolates what the scaffolding alone is worth.
Removing the rulebook while keeping the code world model and everything else identical, on one game per difficulty band: all runs still reach RHAE 100, but actions rise from 617 to 677 (+9%) and agent turns from 678 to 830 (+18%). The population extension was run once, on wa30 at N=2: actions drop from 899 to 597, with eight of nine levels equal or better.
The Pong case study swaps planning for feedback control. The learned model drives a controller that wins 21:0 in three episodes with different openings, with zero LLM calls at execution. Learning cost is 9,504 emulator frames, 1/42 of EfficientZero's reported 4×10⁵ budget and roughly 1/21,000 of the 2×10⁸ frames behind the model-free RL baselines. The same LLM with no world model scores -19.
For agent builders, the paper backs an engineering judgment: in unfamiliar environments, external memory built as explicit hypotheses plus an executable plus cell-exact verification converges faster than leaving the understanding in context, and the 18% turn reduction in the ablation is the rulebook's net contribution. The 59.3-point harness comparison shows how much a scaffold alone can move the same model. The Pong numbers are the practical ones: the controller runs without the LLM, and its learning cost sits two to four orders of magnitude below RL, which fits settings where actions are cheap and samples are expensive. Git-versioned model memory is a detail worth copying outright, since rollback makes aggressive rule changes non-fatal.
On the main benchmark, honesty is due: the edge over baseline1 is 1.0 RHAE point and about 10% fewer actions, an incremental gain. The reusable part is the verification loop and the action-efficiency accounting.
The authors' own list: the public set is close to saturated by strong frontier models, so completion counts cannot isolate architectural contributions and comparisons need model choice, reasoning budget, repeated-run stability, and held-out evaluation; the code-rulebook fidelity check is an LLM judgement rather than an exact one; the MDL bias does not remove uncertainty, a single model can be self-confirming, and the population is a computational approximation to the version space, not a calibrated posterior.
Close reading adds more. The ablation covers three games and the population a single game at N=2, samples too small for strong claims; the stated reason is that each Claude Code run costs substantial time and money. Pong stops learning at the first 21:0 win and evaluates exactly that checkpoint, with each baseline following its own protocol, so the direction is credible but the precision is not; full-frame replay is also acknowledged to retain residual rendering and opponent-paddle errors. The method leans on Claude Opus 5 at extra-high reasoning effort, and the paper gives no cost figure.