Opus 5.5 Tops RareBench at 52% as Grok 4.7 and GPT-6 Sol Show Regressions

danielmckinn0n · x · 2026-09-26

GamowLabs' RareBench leaderboard crowns Anthropic's Opus 5.5 at 52%, adding 6 pp over previous champ Fable 5.1. Notable findings: confirmation of a widely reported Grok regression (4.7 loses 5 pp to 4.6), GPT-6 Sol slightly underperforming GPT-5.6 (possibly within eval error), and Luna surprisingly finishing last — hinting at a minimum scale threshold for the task. Pure Fable 5.1 scores are coming via Anthropic's Life Sciences Verification Program.

Original post →

More from Models

Models channel →