Opus 5 Tests: Impressive Single-Shot Generation, But Lags Behind Fable in Complex Tasks

Less than three days after its release, Claude Opus 5 has sparked a wave of hands-on testing within the developer community. The model demonstrates stunning 3D and physics simulation capabilities in generating complex applications via a single prompt (such as FPS games, Minecraft replicas, and snowboard simulators), with overall strength surpassing Opus 4.8. However, in deep evaluations of long-horizon tasks and complex development scenarios, multiple developers believe its comprehensive performance still lags behind Fable, once again highlighting the disconnect between public benchmark scores and real-world experience.

Confirmed

* **Outstanding Single-Shot Generation**: Users like @eyishazyer showcased Opus 5 generating complete shooter games, consulting-grade table slides, and a Minecraft replica with real-time lighting and block physics using just a single prompt. Its 3D performance was described as "stunning" by @TAbrodi.

* **Overall Superior to Predecessor**: @eyishazyer confirmed that Opus 5's overall capabilities exceed Opus 4.8, though it inherits some of Fable's stylistic quirks (like Density and FableSpeak).

* **Complex Task Performance Lags Behind Fable**: @doodlestein and @petergyang pointed out that although Opus 5 has higher benchmark scores, it lacks depth of thought when handling complex tasks, is more prone to errors, and is less stable than Fable. @johnlindquist also believes Fable 5 is stronger at open-ended coding.

* **Long-Horizon Collaboration Pain Points**: @Veraticus reported that Opus 5 is harder to steer in long-horizon tasks; @doodlestein added that due to the need for constant rework, the actual cost of using Opus 5 in hard debugging might be higher than Fable, which "gets it right the first time."

* **Suited as a Sub-Agent**: @iskander relayed the perspective that while Opus 5 is technically competent, it lacks high-level creative divergence, making it better suited for a "sub-agent" role dispatched by another master control model.

Unconfirmed

* @antirez mentioned that while Opus 5 has made substantial progress and bridged the gap, whether it has retaken the overall lead in user experience against competitors (like Sol) still requires more real-world testing.

Why It Matters

The release of Opus 5 reiterates the severe disconnect between "benchmark scores" and "real-world development experience" in current AI models. For developers, a model's stability and probability of "getting it right the first time" during hardcore debugging and long-horizon complex tasks determine actual productivity and usage costs far more than flashy benchmark scores or single-shot demos.

2026-07-26 ~ 2026-07-27 · 17 related posts

Full story(20 episodes)→

Primary sources

1 near-duplicate retellings: eyishazyer