RSI-Exam Benchmark Launches to Test AI Recursive Self-Improvement

HuaxiuYaoML · x · 2026-08-27

RSI-Exam, a new benchmark with 88 executable research tasks, has been released to test if frontier AI agents can achieve Recursive Self-Improvement (RSI). Domains include virtual cells, TPU kernels, chip design, and quant finance.

The protocol requires agents to iterate on visible data, with final artifacts evaluated on a hidden set. Opus 5 currently leads the 88-task leaderboard with a mean hidden-set score of 0.464. Task contributions are open for authorship credit.

Related event: RSI-Exam Benchmark Tests AI Recursive Self-Improvement(2 posts)→

Original post →

More from Research

Research channel →