Opus 5 beats higher effort on FrontierCode at medium effort
kieranklaassen · x · 2026-07-25
A repost of a thread claiming that Opus 5 scores better on FrontierCode at medium effort than at higher effort, even though effort still improves results on some other evals.
The implication is that more compute or longer deliberation does not always translate into better coding performance on every benchmark, and FrontierCode may be exposing a different failure mode than other tests.
More from Models
- Claude docs say thinking content is encrypted, raising a distillation problem — zainhas · 2026-07-25
- ProCreations grug-27b trends on Hugging Face with Qwen3_5-style agentic tags — ProCreations · 2026-07-25
- User says ChatGPT Image 2 now beats Gemini on image editing and accuracy — dreamwieber · 2026-07-25
- Grok 4.5 posts the biggest week-over-week usage gain in Augment’s model picker — Daniel_Farinax · 2026-07-25
- Can any AI really watch 24 fps video with strong comprehension? — mattshumer_ · 2026-07-25
- Kimi weights could turn the debate into hardware economics versus V4 — teortaxesTex · 2026-07-25