RSI-Exam Benchmark Tests AI Recursive Self-Improvement

RSI-Exam is a new benchmark of 88 executable research tasks across 6 domains that tests whether frontier AI agents can achieve recursive self-improvement. Its key finding is that improvement is volatile: long-horizon runs can yield genuine gains validated on hidden sets, but agents still fail when their search gets stuck in wrong approaches.

2026-08-27 ~ 2026-08-27 · 3 related posts

1 near-duplicate retellings: HuaxiuYaoML