Google open-sources RRSI: agents self-improve their harness, lifting Terminal-Bench 2.1 to 80.2%

sudoraohacker · x · 2026-09-29

Google Research released RRSI (Regularized Recursive Self-Improvement), which lets an agent automatically evolve its own harness—prompts, control flow, tools, memory and context management—while keeping the model frozen, using evaluation to decide which edits to keep.

To counter harness-evolution overfitting (big in-distribution gains that vanish out of distribution), RRSI regularizes both sides:

Per the project report, with Claude Opus 4.8 fixed, Terminal-Bench 2.1 rose from 74.2% to 80.2%, and SWE-bench Verified (not used for selection) rose from 82.0% to 83.8%. Paper and project page are linked; real-world gains still need on-task validation.

Related event: Google Open-Sources RRSI for Recursive Agent Self-Improvement(2 posts)→

Original post →

More from coding & agent

coding & agent channel →