Tsinghua Benchmark: Over Half of AI Agents Fail to Improve Training Algorithms

alex_verem · x · 2026-08-23

Researchers from Tsinghua University and MosaicML released AI4AI-Bench, a benchmark to test AI agents' ability to redesign training algorithms for recursive self-improvement (RSI). Spanning 10 real research repositories, agents are given 4 hours to rewrite training code, which is then retrained for up to 12 hours and scored. Results show a mean score of 0.166 (baseline 0.1, max 1.0) across 29 configurations of 6 frontier systems, with over half of the agents failing to modify the code at all.

Related event: AI4AI-Bench Shows Recursive Self-Improvement Remains Very Hard(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →