Opus 5.5 Tops RareBench at 52% as Grok 4.7 and GPT-6 Sol Show Regressions
danielmckinn0n · x · 2026-09-26
GamowLabs' RareBench leaderboard crowns Anthropic's Opus 5.5 at 52%, adding 6 pp over previous champ Fable 5.1. Notable findings: confirmation of a widely reported Grok regression (4.7 loses 5 pp to 4.6), GPT-6 Sol slightly underperforming GPT-5.6 (possibly within eval error), and Luna surprisingly finishing last — hinting at a minimum scale threshold for the task. Pure Fable 5.1 scores are coming via Anthropic's Life Sciences Verification Program.
More from Models
- China Telecom's Xing4.0 trends on HF: 29B MoE trained entirely on Ascend 910C — AdinaYakup · 2026-09-26
- "Anthropic is killing OpenAI" is just another hype cycle, argues exasperated dev — TheMoonMidas · 2026-09-26
- Claude is getting better at flagging fabricated info in retrieved search results — lilyraynyc · 2026-09-26
- Community tester declares Opus 5.5 no longer loops: "OMG SAVED" — scaling01 · 2026-09-26
- User has Claude write a scholarly history of caste in India, from Harappa to today — soumitrashukla9 · 2026-09-26
- Did Claude quietly remove the effort slider? Users wonder — ThePeterMick · 2026-09-26