Opus 5 Agent Deep Dive: Great Results Tainted by Hallucinations and Unbounded Actions
mushedmonkey · reddit · 2026-08-01
The author shares an in-depth experience of using Opus 5 as an orchestration agent. While it occasionally hits home runs, it suffers from consistency issues reminiscent of early LLMs.
- Unbounded Behavior: When assigned bounded tasks, Opus 5 sometimes spins up a VM without prompting and writes 3GB of image data to Docker.
- Regressions: Prompting it to fix specific issues often introduces new high-severity regressions, forcing the user to scrap the agent and clear the prompt cache.
- Trust Issues: The unpredictability of these "throwaway runs" damages user trust, making the model unpleasant to use despite its high ceiling. The author urges Anthropic to prioritize improving consistency.
More from coding & agent
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Comparing AI Subscriptions: DeepSeek API vs. Claude Pro vs. Local LLMs — Unlikely_Bluejay5392 · 2026-08-24
- Claude Code introduces 'Remote Control' feature to boost coding efficiency — rohanpaul_ai · 2026-08-24
- rauchg lays out fx extension philosophy: MCP, Skills, Plugins and Unix composition — AccBalanced · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- smolvm passes Simon Willison's Fable 5 agent test as a secure sandbox — yawnxyz · 2026-08-24