AI4AI-Bench Reveals Current LLMs Struggle with Recursive Self-Improvement

_akhaliq · x · 2026-08-22

Einsia introduced AI4AI-Bench, a benchmark testing if agents can improve AI training algorithms themselves. Across 10 real repositories, the average score was just 0.166, with the best model (Opus 5) reaching only 0.288. The median exploration cost per task also surged from $1.69 to $34.60, highlighting the significant difficulty of effective recursive self-improvement (RSI).

Original post →

More from Research

Research channel →