RSI-Exam released: Benchmarking recursive self-improvement in AI agents

HuaxiuYaoML · x · 2026-08-31

RSI-Exam introduces a benchmark with 88 executable research tasks across 6 domains to test if AI agents can achieve Recursive Self-Improvement (RSI). Domains include virtual cells, TPU kernels, chip design, and quantitative finance. 35 tasks are now public, with Opus 5 leading the leaderboard. The project calls for expert contributors to author, review, or audit tasks.

Original post →

More from coding & agent

coding & agent channel →