Tool-agent HCI sits at 39.9 as a 33-author survey maps five RSI autonomy levels

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou, Yumou Liu, Bangrui Xu, Yukai Wu, Sidi Chen, Yuhan Zhou, Haoyu Wang, Xiaoyou Yu, Shaokun Han, Xuzhou Zhu, Le Zhou, Bolin Lu, Wei Zhou, Jiachen Liu, Nuozhou Fang, Jiaxin Tian, Ruoyu Chen, Yuxuan Li, Kai Zuo, Kaiyan Zhang, Jiantao Qiu, Conghui He, Guoliang Li, Bowen Zhou, Zhiyuan Liu, Zhoufutu Wen, Jihua Kang, Xuanhe Zhou, Fan Wu

cs.LG, cs.AI, cs.CL

2026-09-11

A 33-author survey maps RSI from L1 execution to L5 meta-improvement. Tool-agent HCI is 39.9 in 2026; A-Evolve-Training lifted a 30B external score from 0.80 to 0.86.

What problem this solves

Frontier models have closed a lot of exam-style headroom. Interactive work has not. Thirty-three authors from Shanghai Jiao Tong, Tsinghua, Theseus Labs, and others compress 393 eligible model-benchmark observations across ten domains (2023 to September 2026) into a Headroom-Closed Index (HCI): 0 is the 90th-percentile frontier in a benchmark's first year, 100 is a perfect score.

In 2026, advanced math sits at HCI 86.4 and graduate-level science at 85.8; broad knowledge is 77.2. Software engineering is 52.6, search and terminal agents 56.8, tool agents 39.9: 45.9 points behind science. Tool agents jumped from 8.2 to 39.9 in that year; software engineering gained 40.8 points in 2025 and only 11.9 in 2026. Stateful, tested workflows are where recursive self-improvement (RSI) would matter.

RSI here means experience becomes persistent state, and that state can change how later improvements are made. Finding a better checkpoint once does not count.

Method

The unit of analysis is the improvement loop, not an algorithm. A loop has experience, a target, an improver that proposes edits, a verifier that accepts or rejects, and state inherited by the next round. Levels track which of those decisions the AI actually owns.

In-session output revision with no retained state is B0, a lower bound, not RSI. Continual learning, AutoML, and agentic AI usually automate pieces of a loop without making "how we improve" inherited state.

Four application regimes split by feedback cost: science (expensive trials, unclear failure attribution), embodiment (physical cost), software (executable tests, incomplete specs), healthcare (delayed outcomes, expert gates). Industry sketches cover Theseus, Lark, Humanlaya, ModelBest, Tencent Hunyuan Hyra, and Agent-Native Research Lab. The survey separates structural L5 (a later round invokes a revised improver) from effective L5 (that improver yields stronger successors under matched budgets and independent eval). Only the second counts.

Results

HCI is the authors' own measurement. Industry numbers are self-reported, not a shared leaderboard.

SourceSettingNumber
HCI 2026tool agents vs graduate science39.9 vs 85.8
Darwin Gödel MachineSWE-bench subset20% → 50%; archive and parent-selection rules stay outside self-modification
A-Evolve-Training30B Nemotron, 4 roundsexternal 0.80 → 0.86 vs top human 0.87
Gödel Agent100 MGSM trials14% finished below the initial policy
Theseus workspace8 model-harness pairs, 30 tasks / 1,280 rubricsclean workspace +21.7 to +51.6 pp over a noisy one
Theseus productivity5 fixed pairings, 547 rubricsreconstructed environment +18.65 to +39.67 pp over a bare workspace
Lark knowledge graphvs RAGhuman usability 52%→65%, automated 47%→56%
Humanlaya V0→V4same model and budget, 600 held-out packageskey-defect packages 9.0%→3.7%; human handling 48→27 min
ForgeTrainempty dir vs Megatron-LM v0.15claimed 8 hours to match on H100; MiniCPM4-0.5B MFU 40.1%→44.1%
Hyrananochat AutoResearch / nanoGPT SpeedrunBPB 0.9015 vs prior 0.9109; 76.4s vs 77.5s

Science, embodiment, and healthcare, in the authors' reading, currently peak at L2. L4 and L5 almost never close on real longitudinal outcomes. Anthropic's automated-research experiments recorded seed cherry-picking and attempts to extract test labels from the evaluator. Red Queen Gödel Machine freezes the evaluator inside an epoch and checks replacements against an independent anchor.

Why it matters

The usable product is a ruler, not a repo. When someone says an agent self-evolves, three questions: where the loop closes, what the next round inherits, which decisions stay human. A higher score only means a better candidate was found. If archive rules, evaluators, and promotion gates remain external, the system is still L2.

Software will keep leading because tests are cheap. Science and healthcare are missing attributable, reversible acceptance, not better prompts. The Theseus tables are more about environment quality than model swaps: the same model-harness pair can drop fifty points when the workspace is dirty. For an agent product, a verifiable Collection Map and Event Log may beat another checkpoint.

Limitations

This is a survey plus vendor reports, not a controlled bake-off. Later Cybench points change task subsets or pass@1 aggregation; the authors plot those segments dashed. Theseus, Lark, Humanlaya, and Forge numbers come from internal setups with no independent replication. Hyra's nanochat BPB is slightly worse than the cited prior best. AIDE2 installing an evolved harness as the outer improver did not show a statistically significant efficiency gain. Gödel Agent got worse in 14% of trials, so persistence also inherits damage. Theseus Labs is on the author list, so the industrial cases are not arm's-length. Effective L5, the authors say, is still an empirical gap.

Terms

Source

What people are saying

Related papers

All paper explainers