Cloud Planner + Local Worker: 7x Fewer Output Tokens but a Hidden Cache Hit-Rate Cost
GapNew4766 · reddit · 2026-09-29
The Atomic Agent team (MIT-licensed, open source) ran one real build through three setups to measure where money and quality actually go in an orchestrator/worker architecture. Task: five hand-written 3D physics scenes on one canvas, fixed camera, one-line prompt, one autonomous run each. Cloud: Sonnet 5.5 via OpenRouter; local: Qwen 3.8 27B Q4 on a single RTX 3090.
| Setup | Scenes | Time | Cloud cost |
|---|---|---|---|
| Sonnet 5.5 alone | 4/5 | 9 min | $1.83 |
| Sonnet plans, Qwen codes | 2/5 | 85 min | $0.67 |
| Qwen alone | 1/5 | 26 min | $0 |
Architecture details
- The orchestrator is read-only for the whole turn: writes are refused inside tool results rather than by removing tools, because changing the tool list would change the prompt prefix and throw away the KV cache every turn.
- Delegation is a contract: each task declares what it provides and requires; name mismatches are rejected before any worker starts. Tasks run in dependency waves.
- Workers are one-shot, memory-less sessions receiving one brief plus a quote of the original request.
- Verification before done: the orchestrator runs the result in a throwaway copy with a headless browser and checks the console.
Key findings
- Savings come almost entirely from output tokens: the planner writes briefs, not code — 7x fewer output tokens, 2.7x lower cost.
- Cache hit rate is the hidden cost: cloud-only input read 92.5% from cache, while an orchestrator waiting on slow local workers got about a third.
- The local worker is the bottleneck: 85 minutes vs 9, and worse quality.
- "Verified" doesn't mean correct: all runs passed the headless check with clean consoles, yet some scenes were visually broken (hollow cubes, an unbreakable wall). Crash checks don't catch wrong physics.
The authors note it's one run per setup — a field report — and are soliciting methods to verify visual/physical correctness beyond "no console errors". Repo: github.com/AtomicBot-ai/atomic-agent
More from coding & agent
- AlignOPSD fixes decision-timestamp mismatch in agent distillation, beating GRPO by 5.5-8.7% — Mingju Chen · 2026-09-29
- CompoWorld scales general agent training by composing environments from reusable services — AllSpark-Research · 2026-09-29
- Adaptive Consistency Graph lifts long-horizon agent success from 44.5% to 50.2% — Beihang · 2026-09-29
- ByteDance's TraceDance auto-builds agent behavior benchmarks from real deployment traces — ByteDance · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29
- ADHD is a superpower for juggling 10 AI agents, dinner, and shitposting at once — enggirlfriend · 2026-09-29