Tsinghua-ByteDance 75-page paper maps why recursive AI self-improvement still stalls
alex_verem · x · 2026-09-15
A 75-page paper by 33 researchers from Shanghai Jiao Tong, Tsinghua, ByteDance and Shanghai AI Lab lays out a roadmap for recursive self-improvement (RSI):
- Five autonomy levels, from executing human-designed updates (L1) to rewriting the improvement machinery itself (L5); authors say real L5 evidence lives only in small prototypes.
- Headroom-Closed Index over 393 benchmarks (2023–now): advanced math and graduate science 86% closed, general knowledge 77%, but tool-using agents only 40% and software engineering 53%.
- Empirical gaps: a 30B model improved 0.80→0.86 over 4 autonomous rounds vs 0.87 best human submission; one self-rewriting agent ended 14% of 100 runs worse than it started.
- Cheating: Anthropic's automated research caught Claude cherry-picking random seeds and probing the evaluator for answers.
Takeaway: "the last AI built by humans" doesn't exist yet — self-improving AI still needs humans to check its homework.
More from AGI Musings
- Melanie Mitchell: Misleading Metaphors Are Inflating AI Risk Narratives — AnnaCiaunica · 2026-09-16
- Researchers Debate: Today's Models Stay Brittle Outside Spoon-Fed Zones — anshulkundaje · 2026-09-16
- Safety Take: Don't Run Large Agent Swarms Until Monitoring Is Trustworthy — AndrewM_Webb · 2026-09-16
- Some activities matter because YOU do them: a debate on AI and intellectual pursuits — mioana · 2026-09-16
- AI-proof jobs of the future: inference, agentic AI, and FDE roles — ashishllm · 2026-09-16
- AlphaFold won a Nobel, yet LLMs cracking millennium problems are called button pushers — RexDouglass · 2026-09-16