UCL's Memento 3 lets a frozen LLM self-improve via rulebooks, clearing all 25 ARC-AGI-3 games
UniversityCollegeLondon · hf · 2026-10-09
UCL introduces Memento 3, a model-based route to recursive self-improvement where the underlying LLM stays frozen. The agent maintains a natural-language "rulebook" as persistent semantic memory of revisable hypotheses about environment dynamics, compiling it into executable code for prediction and planning.
Key points:
- A continual loop of observation, reflection, rule revision, compilation, and verification uses prediction errors to refine both rulebook and code
- Code updates are accepted only when the LLM judges them faithful to the rulebook and cell-exact replay reproduces observed transitions
- A population extension runs multiple world models in parallel, sharing evidence and using prediction disagreement to guide exploration
Results: on ARC-AGI-3 the single-model agent clears every level of all 25 public games, achieves mean Relative Human Action Efficiency (RHAE) of 100.0, and uses only 44% of human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 across three episodes with different openings, without further LLM calls.
More from coding & agent
- Karpathy: 'Prompting is going away' — Graphs are the end state of the LLM → Agents evolution — VeryWellVersed · 2026-10-09
- AgenticBBO-Bench benchmarks LLM agents for black-box optimization; GPT-6 Astra and DeepSeek-V4.1-Flash on Pareto frontier — Ming Chen · 2026-10-09
- Run a Free Claude Code-Style Coding Agent Locally with Ollama, No API Bills — nikola_mr64990 · 2026-10-09
- Dev open-sources an AI podcast with dual AI hosts covering daily tech news, built on Doubao's podcast API — vista8 · 2026-10-09
- Hugging Face ships tokenizers v0.23.3, fixing long-lasting breaking change from 0.23.1 — LysandreJik · 2026-10-09
- Two years later: the bet on LLMs for decompilation and program understanding paid off — paul_cal · 2026-10-09