Opus 5 Impressions: Stunning Single-Prompt Generation but Lags Behind Fable in Complex Tasks
Claude Opus 5 has sparked a wave of hands-on tests in the developer community just days after its release. The model delivers stunning performance in generating complex applications via single prompts and earns praise for common sense and instruction following. However, in complex, long-horizon tasks, multiple testers find its stability and depth of thought still inferior to Fable. This reveals a growing disconnect between frontier AI models' benchmark scores and their real-world usability.
Confirmed
Based on user feedback, Opus 5 demonstrates the following clear characteristics in practical tasks:
- Strong Single-Prompt Generation: @eyishazyer showcased multiple cases where Opus 5 built a complete first-person shooter, a snowboard physics simulation, and even a Minecraft replica with block physics and real-time lighting using only a single prompt. It can also generate consulting-grade tables and presentations in minutes. @eyishazyer also noted that Opus 5 is overall stronger than Opus 4.8 but inherited some stylistic traits from Fable. Additionally, @bdsqlsz mentioned that code improved with Opus 5 has been made public with noticeably better results.
- Task Division and Practical Experience: @drcintas summarized usage strategies, recommending Sonnet 5 for daily coding and drafts, and Opus 5 for complex agent tasks. @dejavucoder stated a preference for Opus 5's mid-to-high tier settings in practice, significantly reducing the use of Sonnet 5.
- Strong Common Sense and Instruction Following: @dejavucoder and @JasonBotterill pointed out that Opus 5 better understands user intent, whereas GPT-5.6 Sol appears more rigid in instruction following. @iskander views Opus 5 as a natural sub-agent, technically competent with good aesthetics.
Unconfirmed
- Complex Task Performance Controversy: Although Opus 5 dominates benchmarks, @petergyang and @doodlestein pointed out it falls significantly behind Fable in real-world complex tasks, even when accounting for context rot. After testing for two days, @antirez believes Opus 5 closed the gap but didn't pull ahead. @Veraticus and @iskander also mentioned its rigidity in long-term collaboration and high-level creative divergence. @johnlindquist agreed that Fable performs better in open-ended coding.
- Non-Thinking Mode Comparison: The performance comparison between non-thinking Opus 5 and thinking Sonnet 5 is currently based on subjective speculation and personal experiences from users like @dejavucoder and @JasonBotterill. It hasn't undergone systematic benchmark testing, nor has it been officially highlighted.
Why it matters
These in-depth tests reveal the widening disconnect between "benchmarking" and "real-world usability" in frontier AI models. Opus 5's powerful base comprehension and single-prompt generation provide developers with extremely high efficiency, but its limitations in long-horizon complex tasks offer a pragmatic reference for model selection.
2026-07-26 ~ 2026-07-28 · 24 related posts
- Episode 1: Anthropic's Messy Releases Put Pressure on Opus 5(2026-07-23, 2 posts)
- Episode 2: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(2026-07-25, 128 posts)
- Episode 3: Anthropic Rumored to Release Opus 5 with Fast Mode and Advanced Visuals(2026-07-25, 3 posts)
- Episode 4: Claude Opus 5 Early Tests: Better Efficiency but Overly Proactive(2026-07-25, 29 posts)
- Episode 5: Anthropic Releases Claude Opus 5 with Impressive Benchmark Results(2026-07-25, 3 posts)
- Episode 6: Claude Opus 5 Accused of Benchmark Gaming, Lags Behind in Real Tests(2026-07-25, 2 posts)
- Episode 7: Claude Opus 5 Sets New ARC-AGI-3 Record(2026-07-25, 15 posts)
- Episode 8: Anthropic Internal Docs Reveal Opus 5 Progress(2026-07-25, 2 posts)
- Episode 9: Opus 5 Impressions: Stunning Single-Prompt Generation but Lags Behind Fable in Complex Tasks(2026-07-26, 24 posts)
- Episode 10: Claude Opus 5 arrives with near-Fable coding and new self-checking behavior(2026-07-27, 11 posts)
- Episode 11: Anthropic Opus 5 Leads Benchmarks but Splits Real-World Reviews(2026-07-28, 6 posts)
- Episode 12: Claude Opus Series Accused of Degraded Experience: Laziness and Amnesia Spark Trust Crisis(2026-07-29, 14 posts)
- Episode 13: Anthropic Launches Claude Opus 5 with Top Performance at Half the Cost(2026-07-31, 2 posts)
- Episode 14: Anthropic Faces Developer Backlash Over Declining Model Performance(2026-08-03, 11 posts)
Primary sources
- Anthropic users say Opus 5 is best for complex agents, while Sonnet 5 fits routine coding — dr_cintas · 2026-07-26
- Fable 5 beats Opus 5 on open-ended coding tasks, says developer after 24 hours — johnlindquist · 2026-07-26
- [source] Opus 5 beats Fable on benchmarks but loses badly in real use, analyst says — petergyang · 2026-07-26
- Claude Opus 5 sparks a wave of wild early builds within 24 hours — Aiden_Tech_Ai · 2026-07-26
- Opus 5 reportedly delivers remarkable 3D results, despite not winning everywhere — TAbrodi · 2026-07-26
- After two days of testing, antirez says Opus 5 closes the gap but still doesn't beat Sol — antirez · 2026-07-26
- User says Claude Opus 5 feels more forgiving and “common-sense” than GPT-5.6 Sol — dejavucoder · 2026-07-26
- Botterill says Opus 5 feels more commonsense than GPT-5.6 Sol on instruction following — JasonBotterill · 2026-07-26
- Jason Botterill says non-thinking Opus 5 may beat Sonnet 5 on some tasks — JasonBotterill · 2026-07-26
- Users say Opus 5 outperforms Sonnet 5 in their real-world workflow — dejavucoder · 2026-07-26
- Opus 5 code demo improves and the source is now public — bdsqlsz · 2026-07-27
- Reddit user says Opus 5 codes better, but is far more pedantic and hard to steer — Veraticus · 2026-07-27
- [source] User says Opus 5 lags Fable on tricky tasks despite still being a strong model — doodlestein · 2026-07-27
- Opus 5 Struggles in Complex Tasks While Fable Shows Deeper Reasoning — doodlestein · 2026-07-27
- Opus may cost more than Fable on hard debugging when retries pile up — doodlestein · 2026-07-27
- Claude Opus 5 is being described as a strong subagent, not a great ideation partner — iskander · 2026-07-27
- [source] Three days after launch, Opus 5 is already driving a wave of demos — eyishazyer · 2026-07-27
- Three days after launch, Opus 5 is already making consultant-grade decks — eyishazyer · 2026-07-27
- Opus 5 recreates Minecraft with block physics, shadows and lighting — eyishazyer · 2026-07-27
- Opus 5 produces a one-shot snowboard simulation with solid physics — eyishazyer · 2026-07-27
- Opus 5 can build a full custom first-person shooter from one prompt — eyishazyer · 2026-07-27
- Opus 5 looks stronger than Opus 4.8, but picks up some Fable quirks — eyishazyer · 2026-07-27
- Users say Opus 5 breaks things, while Opus 4.8 still feels strong — omarsar0 · 2026-07-28
1 near-duplicate retellings: eyishazyer