Google's RRSI Paper Adds Regularization to Stop Agent Self-Improvement Overfitting

cihangxie · x · 2026-09-23

Google researchers released RRSI (arXiv:2609.24972, with code and project page), showing that recursive self-improvement of agent harnesses overfits training tasks — in-distribution gains shrink on out-of-distribution benchmarks. Their fix constrains both the proposer (annealed edit budget, encouraging unexplored trajectories) and the selector (a critic to screen benchmark-specific proposals, a pruner to drop small/costly/useless edits), validated across 8 benchmarks spanning coding, agentic workspace and engineering.

Related event: Google Proposes RRSI: Regularizing Recursive Self-Improvement of Agent Harnesses(8 posts)→

Original post →

More from coding & agent

coding & agent channel →