Introducing RSI-Exam: 88 Tasks Testing Whether Agents Can Recursively Self-Improve
HuaxiuYaoML · x · 2026-08-27
RSI-Exam is a new benchmark testing whether frontier AI agents can achieve recursive self-improvement — turning a weak method into one that performs better on hidden data. The task bank holds 88 executable research tasks across 6 domains, spanning virtual cells, TPU kernels, chip design, quantitative finance, agent harnesses, and model distillation, and is now open for task contributions from every field.
Each task is not a question with a known short answer but an executable research environment with a working starting artifact, a task-native metric, visible data for iterative research, hidden data for final evaluation, and a fixed research budget. The agent may modify or replace the inherited method; what matters is whether the final artifact still performs better when rerun in a fresh container on the hidden set. Opus 5 currently leads at 0.464, ahead of GPT-5.6-sol (0.433) and GLM 5.3 (0.403).
Related event: RSI-Exam Benchmark Tests AI Recursive Self-Improvement(3 posts)→
More from coding & agent
- Stanford's Self-Verification Boosts DeepSeek Past Claude — 机器之心 · 2026-08-27
- Claude iterates furiously on OpenSCAD code automatically — _Stocko_ · 2026-08-27
- Two vLLM recipes for Blackwell: NVFP4 KV cache buys 262K context and more streams — SeanHighness · 2026-08-27
- Agent Trading Demo: Links X API to iMessage for Inverse Cramer Strategy — MurrLincoln · 2026-08-27
- Paper Studies Agent Communication in Long-Horizon Competitive Settings — xuanalogue · 2026-08-27
- Zoetrope visualizes Claude Code sessions as live flow graphs — Saboo_Shubham_ · 2026-08-27