A 12-hour side-by-side test says Opus 5 feels worse than Opus 4.8 on analysis work
mynaame · reddit · 2026-07-28
After about 12 hours of side-by-side testing, the author says Opus 5 was disappointing on non-coding tasks.
Compared with Opus 4.8, Opus 5 felt lazier, more assumption-driven, more likely to skip reading PDFs and Markdown references, and more likely to answer in blocky chunks with poor paragraphing. For coding tasks, they said the two models felt almost the same and did not differ much on routine development work.
Their conclusion is that Opus 4.8 was more coherent and precise for analysis-style work, while Opus 5 did not deliver a meaningful jump for their use case.
More from Models
- Users discuss what they actually use Opus 5 for beyond coding — remilouf · 2026-07-29
- Anthropic Hints at Achieving Recursive Self-Improvement, Calls for Pacing Frontier — daniel_mac8 · 2026-07-29
- Claude Opus 5 tops DeepSWE with a 74% score and a claimed 28% cost edge — daniel_mac8 · 2026-07-29
- LiquidAI’s 230M LFM2.5 encoder trends on Hugging Face — LiquidAI · 2026-07-29
- Anthropic may be 1.5 generations ahead internally, with Fable 5.1 weeks away — haider1 · 2026-07-29
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29