Opus 5 Scored 53.4% on FrontierCode at Launch, Just 48% on the 5.5 Card
mathemagic1an · x · 2026-09-23
Andrew Carr flagged that Opus 5 reported 53.4% on FrontierCode at launch but shows 48% on the same eval's 5.5 card — prompting the coinage "Benchflation": benchmark updates silently rewriting historical scores and eroding leaderboard comparability.
Related event: Opus 5 FrontierCode score drops from 53.4% to 48% in new card(2 posts)→
More from Models
- OrcaRouter stress-tests JEV: dropping autoregressive decoding could cut inference cost 10-100x — Dan_Jeffries1 · 2026-09-23
- JEV is just calibrated classification over a label set, not deterministic output — tzmartin · 2026-09-23
- Claude's new model claims pixel-perfect visual understanding, demos it with a raindrop story — bookwormengr · 2026-09-23
- France's t0-beta, a 256M-parameter open time-series foundation model, hits top-3 on GIFT-Eval and fev-bench — AxSaucedo · 2026-09-23
- Pirate Face: a 'Pirate Bay for LLMs' as a fallback if Hugging Face gets censored — Atagor · 2026-09-23
- Developer finds Luna 6 is a 'massive downgrade' from Luna 5.6 despite better benchmarks — skilliard7 · 2026-09-23