RSI-Exam leaderboard update: Fable 5.1 debuts at #2 as GPT-6-astra holds #1
airesearch12 · x · 2026-09-19
The RSI-Exam benchmark added new frontier models to its board: Anthropic's Fable 5.1 landed directly at #2 with 0.4813, alongside Muse Spark (Meta) and Seed-Evolving-0909 (ByteDance Seed). OpenAI's GPT-6-astra remains #1 at 0.5126. As of this release, no model has yet reached the frontier-calibrated reference score. The poster also asked to add the eval to the benchmarkheaven list.
Related event: GPT-6-astra Tops Updated RSI-Exam Leaderboard as Fable 5.1 Debuts Second(3 posts)→
More from Models
- Cognition ships SWE-2 coding model: 1 point behind Fable 5.1 at 64% less cost — AxSaucedo · 2026-09-19
- Jev Hits 36M Views in 2 Days, Community Ships 6 Open Clones — Latent Space · 2026-09-19
- "Best explanation of Jev yet": called an unlock for many use cases — Roger_M_Taylor · 2026-09-19
- Claude Fable 5.1 hits 90% on ARC-AGI-2 semi-private at $4.49 per task — geoffwolfe · 2026-09-19
- After 7B+ tokens, developer says DeepSeek V4.1 Flash is his default workhorse — gaganghotra_ · 2026-09-19
- Surprising result: AI models can refer to and react to their own internal pain states — Laneless_ · 2026-09-19