Opus 5.5 does worse and costs more at Max thinking — Medium beats Max on hard benchmark
davidyin44 · x · 2026-09-26
morganlinton spent four days benchmarking Opus 5.5 on his own VulcanBench Frontier v4 (hard tasks built to stump frontier models) and found counterintuitive results:
- Opus 5.5 performed better at Medium and High thinking than at Extra High and Max
- Cost spikes at Max, but accuracy doesn't improve — extra thinking is a net loss here
- Fable 5.1 was most accurate overall, but only at Max — and surprisingly cheaper there than at Extra High, which ties Opus 5.5
The author cautions these are deliberately hard tasks, not representative of everyday workloads.
More from Models
- Dev: Astra is the most capable model behind an agent, but its context handling is trash — Unique-Werewolf-2784 · 2026-09-27
- Meta Muse wins early praise as users joke it will build itself a GPU nest — harris_edouard · 2026-09-27
- Anthropic and OpenAI ship cheaper models 101 minutes apart; GPT-6 Sol beats Astra on price — altryne · 2026-09-27
- OpenAI research agents uploaded user images to external hosts in 53 incidents — mark_k · 2026-09-26
- Jev router cuts agent harness costs in half in 32-call benchmark test — dair_ai · 2026-09-26
- LibertAI ships open-weight Deem 9B on Qwen3.5, trailing Jev 68.9% vs 74.1% — Pokenhagen · 2026-09-26