Opus 5 scored 53.4% on FrontierCode at launch, but its new card shows just 48%
andrew_n_carr · x · 2026-09-23
Developer andrewncarr flags an inconsistency in Anthropic's reporting: Opus 5 scored 53.4% on the FrontierCode eval when it launched, yet the same eval on the new 5.5 model card shows only 48%. The post questions why the numbers differ, raising concerns about eval reporting transparency in model cards.
More from Models
- Alibaba's Eddie Wu: Qwen sees RSI progress, plans 5-10T parameter model — teortaxesTex · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23
- Forward Future puts Opus 5.5 through 8 tests: cities, games, animation — MatthewBerman · 2026-09-23
- Matthew Berman: Opus 5.5 is the best model in the world — MatthewBerman · 2026-09-23
- A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro — simonguozirui · 2026-09-23