Cloud Planner + Local Worker: 7x Fewer Output Tokens but a Hidden Cache Hit-Rate Cost

GapNew4766 · reddit · 2026-09-29

The Atomic Agent team (MIT-licensed, open source) ran one real build through three setups to measure where money and quality actually go in an orchestrator/worker architecture. Task: five hand-written 3D physics scenes on one canvas, fixed camera, one-line prompt, one autonomous run each. Cloud: Sonnet 5.5 via OpenRouter; local: Qwen 3.8 27B Q4 on a single RTX 3090.

| Setup | Scenes | Time | Cloud cost |

|---|---|---|---|

| Sonnet 5.5 alone | 4/5 | 9 min | $1.83 |

| Sonnet plans, Qwen codes | 2/5 | 85 min | $0.67 |

| Qwen alone | 1/5 | 26 min | $0 |

Architecture details

Key findings

The authors note it's one run per setup — a field report — and are soliciting methods to verify visual/physical correctness beyond "no console errors". Repo: github.com/AtomicBot-ai/atomic-agent

Related event: Open-Source Agent Tested: Cloud Planning Plus Local Coding Cuts Cost 2.7x but Runs 9x Slower(2 posts)→

Original post →

More from coding & agent

coding & agent channel →