Google open-sources RRSI: agents self-improve their harness, lifting Terminal-Bench 2.1 to 80.2%
sudoraohacker · x · 2026-09-29
Google Research released RRSI (Regularized Recursive Self-Improvement), which lets an agent automatically evolve its own harness—prompts, control flow, tools, memory and context management—while keeping the model frozen, using evaluation to decide which edits to keep.
To counter harness-evolution overfitting (big in-distribution gains that vanish out of distribution), RRSI regularizes both sides:
- Proposal side: an annealed budget caps bundled edits; the proposer is conditioned on the full edit history so falsified hypotheses aren't redrawn; stalled runs get redirected to unexercised components
- Selection side: a critic screens candidates for suite-specific logic before evaluation; gains must exceed evaluation noise; stale components get pruned
Per the project report, with Claude Opus 4.8 fixed, Terminal-Bench 2.1 rose from 74.2% to 80.2%, and SWE-bench Verified (not used for selection) rose from 82.0% to 83.8%. Paper and project page are linked; real-world gains still need on-task validation.
Related event: Google Open-Sources RRSI for Recursive Agent Self-Improvement(2 posts)→
More from coding & agent
- OpenAI launches always-on Workspace Agents in ChatGPT with Slack deploys and MCP support — testingcatalog · 2026-09-29
- BaRe-Mem: Bayesian Reliability Memory Makes Multi-Agent Consultation Robust to Misleading Advisors — NanyangTechnologicalUniversity · 2026-09-29
- SkillDRE Evolves Malicious Agent Skills via Dual-Stage Feedback, 45.28% Attack Success — Pengyu Zhu · 2026-09-29
- 19-year-old claims $750K profit from Claude Code arbitrage bot built in 2 days — Aiden_Tech_Ai · 2026-09-29
- Agent edits the workflow at build time, plain script at run time: token-saving browser automation — ZennoLab_Guru · 2026-09-29
- Pydantic open-sources Monty: a Rust-based Python sandbox for running AI-generated code (8.4k stars) — samuelcolvin · 2026-09-29