FrontierCode shows Claude Opus 5 peaking at medium effort, not max compute
philhchen · x · 2026-07-25
- The chart compares Claude Opus 5, Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol on FrontierCode v1.1.
- It shows test-time compute scaling across five reasoning-effort settings, with both main-set and extended-set scores.
- On the main set, Claude Opus 5 reaches its best score of 53.4 at medium effort.
- On the extended set, it peaks at 63.6 also at medium effort, while higher effort does not keep improving monotonically.
Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(34 posts)→
More from coding & agent
- Teams strip 80% of Claude Code’s system prompt in a new context-engineering guide — EricBuess · 2026-07-25
- A blunt rebuttal says autonomous agents are being oversold for long-horizon hacking tasks — ctjlewis · 2026-07-25
- Opus 5 model card shows 5-agent coding teams reach 0.6 score 2.2× faster — OfirPress · 2026-07-25
- Devin adds Claude Opus 5 as FrontierCode 1.1 shows near-Fable performance at half cost — _sholtodouglas · 2026-07-25
- A week using Claude Code, Codex, and Gemini CLI showed the same repo-breaking patterns — AIcademy-academy · 2026-07-25
- ChatGPT Work agent can now log into websites and keep sessions across runs — OpenAIDevs · 2026-07-25