Teknium says Anthropic’s Opus 5 is worse than 4.6–8 in Hermes agent tests

Teknium · x · 2026-07-28

Teknium says he has switched back to Fable because Opus 5 is performing worse than Opus 4.6–8 in his tests.

He says the newer model caused serious issues inside the Hermes agent and is bad enough that he will not use it again. He also adds that he usually does not trust Anthropic’s claims except for their benchmark reporting, but this time the problem was real in his view.

The post is a user-side failure report rather than an official benchmark, but it is a notable signal about model behavior in agent workflows.

Original post →

More from coding & agent

coding & agent channel →