GPT-6-astra Tops Updated RSI-Exam Leaderboard as Fable 5.1 Debuts Second

The RSI-Exam leaderboard, which tests whether AI agents can autonomously improve methods and generalize to hidden data, added several new models. GPT-6-astra remains first at 0.5126, while Anthropic's Fable 5.1 debuts second at 0.4813.

2026-09-18 ~ 2026-09-19 · 3 related posts

1 near-duplicate retellings: airesearch12