RSI-Exam: New Benchmark Tests AI Agents' Recursive Self-Improvement Across 88 Tasks
yuyinzhou_cs · x · 2026-08-27
As frontier models improve monthly, Recursive Self-Improvement (RSI) is hot — but how to measure it? Researchers introduce RSI-Exam: 88 executable research tasks across 6 domains (AI agents, virtual cells, TPU kernels, chip design, quant finance, distillation, and more), testing whether agents can improve a working method via long-horizon autonomous experimentation and generalize to hidden data.
Protocol: an agent inherits a working method and iterates on visible data for hours; only the final artifact is evaluated once in a fresh container on a hidden set.
Findings: Claude and GPT form a clear first tier (Opus 5 leads the 88-task leaderboard by mean hidden-set score), with a large gap over the rest — yet huge headroom remains. Trajectories show both agents discovering fundamentally better methods and agents spending hours optimizing the wrong idea. Task contributions are open to all fields.
Related event: RSI-Exam Benchmark Released: 88 Tasks to Test AI Recursive Self-Improvement(6 posts)→
More from Models
- User mocks Grok's stubbornness: disabled RLHF causes persistent hallucinations — GlenBradley · 2026-08-27
- Claude Code Offers to Report Its Own Bad Behavior to Anthropic — georgemillo · 2026-08-27
- ThursdAI: NVIDIA x HF acquisition, Qwen 3.8, GLM 5.3 and more AI news roundup — altryne · 2026-08-27
- Opinion: Excessive guardrails may be artificially limiting AI potential — BLUECOW009 · 2026-08-27
- Meme mocks Anthropic: 'we'd rather cut you off than charge you more' — BLUECOW009 · 2026-08-27
- On-device leaderboard: Apple's built-in model ranks 4th behind open source — JosephJacks_ · 2026-08-27