Cloud Planner + Local Coder: 2.7x Cheaper but 9x Slower in a Real Agent Build
GapNew4766 · reddit · 2026-09-29
The author ran one real task (five physics scenes on one page, one-line prompt, one autonomous run) through three agent setups:
- Sonnet 5.5 alone: 4/5 scenes, 9 min, $1.83
- Sonnet 5.5 plans + local Qwen 3.8 27B codes: 2/5 scenes, 85 min, $0.67
- Qwen 3.8 27B alone on an RTX 3090: 1/5 scenes, 26 min, $0
The savings come from the planner writing briefs instead of code — output tokens dropped from 37.9K to 5.4K. But the local model is the bottleneck: runtime ballooned from 9 to 85 minutes with worse quality. Most notably, all three setups reported "done, no errors" while some scenes were visibly broken — "the agent says it's finished" is still not a reliable signal. The team open-sourced their Atomic Agent framework and asked the community for ways to automatically verify visual/physical correctness.
More from coding & agent
- Raven 0.2.0 ships as a 'harness of harnesses' orchestrating Claude Code, Codex and specialist agents — sven_ai · 2026-09-29
- Real-time Among Us demo pairs DeepSeek V4 Flash planner with Jev actor agent — ai · 2026-09-29
- MacAMP dev shares architecture sketches on why you can't just vibecode an audio player — HankYeomans · 2026-09-29
- Creator builds Claude plugin that edits, uploads and cross-posts videos fully automated — Scobleizer · 2026-09-29
- Swargs claims $50M+ transaction volume and 6,000+ listings on its AI agent marketplace — KyeGomezB · 2026-09-29
- Sonnet 5.5 effort settings make no difference in 15-task coding test: 9/15 at low, medium and high — every · 2026-09-29