Teknium says Anthropic’s Opus 5 is worse than 4.6–8 in Hermes agent tests
Teknium · x · 2026-07-28
Teknium says he has switched back to Fable because Opus 5 is performing worse than Opus 4.6–8 in his tests.
He says the newer model caused serious issues inside the Hermes agent and is bad enough that he will not use it again. He also adds that he usually does not trust Anthropic’s claims except for their benchmark reporting, but this time the problem was real in his view.
The post is a user-side failure report rather than an official benchmark, but it is a notable signal about model behavior in agent workflows.
More from coding & agent
- A Claude Code workflow splits planning, execution, and review to cut usage by 50% — dr_cintas · 2026-07-28
- Supermemory pitches shared context transfer across Claude Code, Cursor, and Gemini — builditwithjoe · 2026-07-28
- NousResearch’s Hermes Desktop shows a borderless AI coding workspace — max_paperclips · 2026-07-28
- Vibe-coding debug note shows raw tool calls caught the real agent bug — andrey_kurenkov · 2026-07-28
- A 2026 reading list flags agent harness overengineering, local Gemma 4 benchmarks, and AI infrastructure — rseroter · 2026-07-28
- Codex and Gemini end up arguing inside a client website build — DumpingSouptime · 2026-07-28