FrontierCode shows Claude Opus 5 peaking at medium effort, not max compute
philhchen · x · 2026-07-25
- The chart compares Claude Opus 5, Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol on FrontierCode v1.1.
- It shows test-time compute scaling across five reasoning-effort settings, with both main-set and extended-set scores.
- On the main set, Claude Opus 5 reaches its best score of 53.4 at medium effort.
- On the extended set, it peaks at 63.6 also at medium effort, while higher effort does not keep improving monotonically.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from coding & agent
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11