Introducing RSI-Exam: 88 Tasks Testing Whether Agents Can Recursively Self-Improve

HuaxiuYaoML · x · 2026-08-27

RSI-Exam is a new benchmark testing whether frontier AI agents can achieve recursive self-improvement — turning a weak method into one that performs better on hidden data. The task bank holds 88 executable research tasks across 6 domains, spanning virtual cells, TPU kernels, chip design, quantitative finance, agent harnesses, and model distillation, and is now open for task contributions from every field.

Each task is not a question with a known short answer but an executable research environment with a working starting artifact, a task-native metric, visible data for iterative research, hidden data for final evaluation, and a fixed research budget. The agent may modify or replace the inherited method; what matters is whether the final artifact still performs better when rerun in a fresh container on the hidden set. Opus 5 currently leads at 0.464, ahead of GPT-5.6-sol (0.433) and GLM 5.3 (0.403).

Related event: RSI-Exam Benchmark Tests AI Recursive Self-Improvement(3 posts)→

Original post →

More from coding & agent

coding & agent channel →