RSI-Exam: New Benchmark Tests AI Agents' Recursive Self-Improvement Across 88 Tasks

yuyinzhou_cs · x · 2026-08-27

As frontier models improve monthly, Recursive Self-Improvement (RSI) is hot — but how to measure it? Researchers introduce RSI-Exam: 88 executable research tasks across 6 domains (AI agents, virtual cells, TPU kernels, chip design, quant finance, distillation, and more), testing whether agents can improve a working method via long-horizon autonomous experimentation and generalize to hidden data.

Protocol: an agent inherits a working method and iterates on visible data for hours; only the final artifact is evaluated once in a fresh container on a hidden set.

Findings: Claude and GPT form a clear first tier (Opus 5 leads the 88-task leaderboard by mean hidden-set score), with a large gap over the rest — yet huge headroom remains. Trajectories show both agents discovering fundamentally better methods and agents spending hours optimizing the wrong idea. Task contributions are open to all fields.

Related event: RSI-Exam Benchmark Released: 88 Tasks to Test AI Recursive Self-Improvement(6 posts)→

Original post →

More from Models

Models channel →