Anthropic’s Opus 4.8 beats version 5 for non-coding work, despite weaker benchmarks
burkov · x · 2026-07-28
Anthropic’s Opus 4.8 is described as better than version 5 for non-coding work, even though Opus 5 allegedly beats Fable 5 on benchmarks.
The post argues that this looks like a “benchmaxed” model: stronger benchmark scores, but worse real-world usefulness than 4.8 on general tasks. It also says the author still prefers Codex for coding, and is hesitant about Kimi K3 because it appears slower and more expensive, which the post attributes to China lacking enough high-end GPUs.
More from Models
- Kimi K3 may need 1.4 TB for MXFP4 inference, far beyond a 96 GB RTX PRO 6000 — HanchungLee · 2026-07-28
- Moonshot’s Kimi K3 lands in Japan with 2.8T open weights and $3/$13 pricing — DavidBennett__ · 2026-07-28
- Users report DeepSeek’s web search is degrading and mixing languages on every query — teortaxesTex · 2026-07-28
- Joseph Jacks says open-weight models may now be only 1–3 months behind frontier AI — JosephJacks_ · 2026-07-28
- Claude meme turns evals into poetry, joking about 0.03 deceptive alignment — maxsloef · 2026-07-28
- Gemini Flash 3.6 is claimed to match Sol 5.6 quality at 70% lower cost — bindureddy · 2026-07-28