Google's RRSI regularizes recursive self-improvement of agent harnesses, +4.7 OOD points

google · hf · 2026-09-22

Google Research introduces RRSI, adding regularization to recursive self-improvement of LLM agent harnesses. RSI tends to overfit training tasks, with in-distribution gains vanishing on OOD benchmarks. RRSI constrains evolution via a proposer with temporally annealed edit budgets and exploration encouragement, plus a selector with a critic and pruner to favor reusable mechanisms over benchmark-specific tweaks. Across eight benchmarks, RRSI gains up to 14.1 points on the evolved split and up to 4.7 on five OOD benchmarks, with 30% fewer policy tokens. Code is open-sourced.

Original post →

More from coding & agent

coding & agent channel →