Claude Opus 5 tops FrontierCode at 53.4% and 63.6% with medium test-time compute
andrew_n_carr · x · 2026-07-25
A FrontierCode chart shows Claude Opus 5 scaling strongly with more test-time compute on both the main and extended sets.
- On the main set, Opus 5 peaks at 53.4% at medium effort, ahead of Claude Opus 4.8 and GPT-5.6 Sol in the shown runs.
- On the extended set, it reaches 63.6% at medium effort, with the chart suggesting a non-linear sweet spot rather than monotonic gains.
The post uses the benchmark to argue that Opus 5 is a very strong coding model, especially when given the right reasoning budget.
More from Models
- Claude Opus 5 can misread a document despite knowing the underlying facts — teortaxesTex · 2026-07-25
- Early Opus 5 feedback says Claude’s writing is now “4o-level slop,” despite stronger intelligence — jdjohnson · 2026-07-25
- Users say Anthropic’s new model feels faster and stronger than Fable — emax · 2026-07-25
- Claude Opus 5 scores 30.2% on ARC-AGI-3 public demo environments — GregKamradt · 2026-07-25
- Release blog teaser shows a near-tie on FrontierCode agentic coding benchmark — hardmaru · 2026-07-25
- User Accuses Anthropic of Gaming ARC-AGI-3 by Training Specifically on Benchmark Patterns — VraserX · 2026-07-25