RSI-Exam leaderboard update: Fable 5.1 debuts at #2 as GPT-6-astra holds #1

airesearch12 · x · 2026-09-19

The RSI-Exam benchmark added new frontier models to its board: Anthropic's Fable 5.1 landed directly at #2 with 0.4813, alongside Muse Spark (Meta) and Seed-Evolving-0909 (ByteDance Seed). OpenAI's GPT-6-astra remains #1 at 0.5126. As of this release, no model has yet reached the frontier-calibrated reference score. The poster also asked to add the eval to the benchmarkheaven list.

Related event: GPT-6-astra Tops Updated RSI-Exam Leaderboard as Fable 5.1 Debuts Second(3 posts)→

Original post →

More from Models

Models channel →