Devs Test Opus 5: Strong Coding but Hard to Steer, Lags Fable in Complex Tasks
Anthropic's release of the Opus 5 model has triggered widespread hands-on evaluations among developers. Multiple developers report that while Opus 5 excels in coding, it is harder to steer during complex tasks and long-term collaborations, with its overall performance lagging behind Fable. Additionally, its leading edge in public benchmarks appears disconnected from real-world user experience.
Confirmed
Developers have reached a consensus through rigorous testing. After two days of testing, antirez noted that Opus 5 made substantial progress, bridging previous gaps, but did not significantly outperform Fable overall. doodlestein pointed out that even accounting for context degradation, Opus 5 is still more prone to errors in complex tasks; by contrast, Fable thinks more thoroughly before acting and performs more stably. An analysis shared by petergyang echoed this: Opus 5 dominates public benchmarks but clearly falls behind Fable in real-world tasks, suggesting that leaderboards can no longer genuinely differentiate frontier models. A 24-hour coding test shared by johnlindquist similarly showed that while both produce high-quality code, Fable 5 is more likely to "get it right the first time" when handling open-ended requirements.
Unconfirmed
The exact reasons why Opus 5 is more error-prone in long-term tasks remain a subject of subjective developer discussion. Some Reddit users feel it is more likely to spiral out of control compared to earlier versions, but the specific mechanisms and triggers are yet to be determined.
Why It Matters
The disconnect between benchmark scores and real-world development experience in frontier AI models is becoming an industry focal point. doodlestein raised a counterintuitive cost consideration: in hard debugging or complex implementation tasks, if a model requires multiple reworks due to repeated errors, its final total cost and time consumption might actually exceed that of a model like Fable, which can "get it right the first time." This means that when selecting models, developers must look beyond benchmark scores and consider the model's depth of thought and execution stability in complex workflows.
2026-07-26 ~ 2026-07-27 · 7 related posts
Primary sources
- Fable 5 beats Opus 5 on open-ended coding tasks, says developer after 24 hours — johnlindquist · 2026-07-26
- [source] Opus 5 beats Fable on benchmarks but loses badly in real use, analyst says — petergyang · 2026-07-26
- [source] After two days of testing, antirez says Opus 5 closes the gap but still doesn't beat Sol — antirez · 2026-07-26
- Reddit user says Opus 5 codes better, but is far more pedantic and hard to steer — Veraticus · 2026-07-27
- User says Opus 5 lags Fable on tricky tasks despite still being a strong model — doodlestein · 2026-07-27
- [source] Opus 5 Struggles in Complex Tasks While Fable Shows Deeper Reasoning — doodlestein · 2026-07-27
- Opus may cost more than Fable on hard debugging when retries pile up — doodlestein · 2026-07-27