After a full day testing Opus 5, the author sees little improvement
jiayuan_jy · x · 2026-07-26
The author says they spent a full day testing Opus 5 and did not notice any significant improvement.
Their takeaway is that it mostly just gets the job done, with performance feeling similar to GPT 5.6 sol and Opus 4.8.
Related event: Claude Opus 5 Early Tests: Better Efficiency but Overly Proactive(29 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11