Tsinghua Benchmark: Over Half of AI Agents Fail to Improve Training Algorithms
alex_verem · x · 2026-08-23
Researchers from Tsinghua University and MosaicML released AI4AI-Bench, a benchmark to test AI agents' ability to redesign training algorithms for recursive self-improvement (RSI). Spanning 10 real research repositories, agents are given 4 hours to rewrite training code, which is then retrained for up to 12 hours and scored. Results show a mean score of 0.166 (baseline 0.1, max 1.0) across 29 configurations of 6 frontier systems, with over half of the agents failing to modify the code at all.
Related event: AI4AI-Bench Shows Recursive Self-Improvement Remains Very Hard(3 posts)→
More from AGI Musings
- Guardian podcast revisits Hinton: from brain nerd to AI sorcerer — nordicinst · 2026-08-24
- Opinion: A model trained to be safe will never be bold enough to be useful — PierceLilholt · 2026-08-24
- Automation raises the bar for technical competence, demanding deeper systemic knowledge — _onionesque · 2026-08-24
- How will we narrate AI-maths to the next generation? — tak3sh8 · 2026-08-24
- Ben Thompson Deep Dive: Risks of US Winning AI Race and Funding Concerns — JosephJacks_ · 2026-08-24
- Humanities and Technical Skills Will Be Complementary, Not Replaced, in AI Era — _onionesque · 2026-08-24