Opus 5 beats higher effort on FrontierCode at medium effort
kieranklaassen · x · 2026-07-25
A repost of a thread claiming that Opus 5 scores better on FrontierCode at medium effort than at higher effort, even though effort still improves results on some other evals.
The implication is that more compute or longer deliberation does not always translate into better coding performance on every benchmark, and FrontierCode may be exposing a different failure mode than other tests.
Related event: Opus 5 Coding Paradox: Higher Reasoning Leads to Lower Scores(13 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11