Google's RRSI regularizes recursive self-improvement of agent harnesses, +4.7 OOD points
google · hf · 2026-09-22
Google Research introduces RRSI, adding regularization to recursive self-improvement of LLM agent harnesses. RSI tends to overfit training tasks, with in-distribution gains vanishing on OOD benchmarks. RRSI constrains evolution via a proposer with temporally annealed edit budgets and exploration encouragement, plus a selector with a critic and pruner to favor reusable mechanisms over benchmark-specific tweaks. Across eight benchmarks, RRSI gains up to 14.1 points on the evolved split and up to 4.7 on five OOD benchmarks, with 30% fewer policy tokens. Code is open-sourced.
More from coding & agent
- Grok Build ships v1.0.38-40: subagents default to parent model, long-run agent workflows get steadier — XFreeze · 2026-09-22
- TypeSafe Jev lands Spring AI integration: typed decisions with confidence in ~300 ms — blaizedsouza · 2026-09-22
- Bezalel, one-MCP-URL agent toolkit, lands its first paying customer — Rasmic · 2026-09-22
- Grok 4.7 produces realistic Three.js cloth simulation after heavy iteration — minchoi · 2026-09-22
- One-line prompt, minutes to a playable game: Grok 4.7 launch demos — minchoi · 2026-09-22
- Grok 4.7 launch day: 10 wild demos from Blender scenes to $146K error catches — minchoi · 2026-09-22