RSIAgent Paper: Training-Free Self-Improvement Lets Kimi-K3 and GLM-5.3 Beat GPT-6

rohanpaul_ai · x · 2026-09-25

The arXiv paper "RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments" (Sibo Zhu et al., 50 pages) introduces a training-free multi-agent framework coordinating curriculum, actor, and verifier agents to continually explore environments, validate outcomes, and retain reusable action-condition-consequence causal knowledge in memory.

Its key design is a broad-then-deep exploration strategy: parallel broad self-exploration discovers diverse environment structures, then focused deep exploration uncovers hard cases, hidden constraints, boundary conditions, and unknown causal dependencies. The resulting memory is frozen and directly reusable without parameter updates.

On OSWorld-v2 and Agent's Last Exam, RSIAgent substantially boosts strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.

Related event: RSIAgent: Training-Free Self-Improvement Beats GPT-6 with Open Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →