Google's RRSI Paper Shows Agent Self-Improvement Overfits and Needs Regularization

机器之心 · wechat · 2026-10-06

Google Research's RRSI (Regularized Recursive Self-Improvement) paper tackles agent-harness-level RSI: without touching model weights, agents recursively rewrite their own prompts, control flow, tools, memory, skills, context management, and subagents in an execute-observe-modify loop.

The core problem is evolution-set overfitting. Reusing the same evolve tasks turns harness evolution into an adaptive search on finite data, causing benchmark-specific fitting, evaluation-noise chasing, and complexity accumulation. Counterintuitively, evolve scores keep rising while gains vanish out of distribution.

RRSI's approach: don't restrict what agents can modify — regularize how they search. On the proposal side, limit the number of entangled changes per round and keep full evolution history as evidence. On the selection side, a candidate must prove itself: benchmark-specific edits are filtered, noise-range gains aren't committed, added context/tokens must justify their cost, and stale mechanisms get pruned.

Key ablation: unregularized evolution reaches a higher evolve score (92.8 vs 90.5) but worse OOD average (40.3 vs 43.6) and higher inference cost (3.80M vs 2.42M policy tokens per trial). Across three domains and 8 benchmarks, evolving on a single benchmark then freezing transfers gains up to +4.7 points OOD (e.g., SWE-bench Verified 82.0→83.8, GDPval 48.8→52.3). In a 30-round run on Terminal-Bench 2.1 (64.6→78.7), only 10 of 30 candidates were accepted. The takeaway: self-improvement itself needs regularization — the next frontier is generalizable self-improvement.

Related event: Google's RRSI Paper Adds Regularization to Recursive Agent Self-Improvement(2 posts)→

Original post →

More from coding & agent

coding & agent channel →