RSI-Exam Benchmark Launches to Test AI Recursive Self-Improvement
HuaxiuYaoML · x · 2026-08-27
RSI-Exam, a new benchmark with 88 executable research tasks, has been released to test if frontier AI agents can achieve Recursive Self-Improvement (RSI). Domains include virtual cells, TPU kernels, chip design, and quant finance.
The protocol requires agents to iterate on visible data, with final artifacts evaluated on a hidden set. Opus 5 currently leads the 88-task leaderboard with a mean hidden-set score of 0.464. Task contributions are open for authorship credit.
Related event: RSI-Exam Benchmark Tests AI Recursive Self-Improvement(2 posts)→
More from Research
- TIDES Dataset: Longitudinal Bilingual Record of 12 Teams' Collaboration — josephseering · 2026-08-27
- Kyoto U's MemUse: Natural Integration Outperforms QA in Evaluating Conversational Memory — Kyoto-University · 2026-08-27
- GPT-5.6 Builds New Kernel, Achieving 9.7x Speedup on TPU — HuaxiuYaoML · 2026-08-27
- Gordian Screens 1,327 Targets In Vivo, Accelerating Drug Discovery — juanbenet · 2026-08-27
- Discussion on Multi-Agent Reward Schemes and Convergence — jessi_cata · 2026-08-27
- Why scaling LLMs won't lead to real agency: A 3-tier Embodied AI architecture — Far-Start-1789 · 2026-08-27