Using GPT-6.1 Sol only as planner cuts coding API bill by 77%, local 27B writes code

GapNew4766 · reddit · 2026-09-30

A developer benchmarked three ways to build three small 3D games (pool, bowling, foosball): letting GPT-6.1 Sol do everything, having Sol only plan while a local Qwen 3.8 27B on a single RTX 3090 writes the code, and going fully local. The planner/worker split cut API spend from $0.75 to $0.17 — a 77% saving — at the cost of runtime rising from 6.6 to 43.4 minutes.

The rationale: most tokens in a coding run are the code itself, while planning and review are a small slice, so pushing the bulk of output onto a per-token-free local model saves the most. The open-source tool Atomic Agent implements this as a "Fusion" mode, but any framework that lets you set separate planner and editor models works.

Caveats: savings varied (87% on pool vs 57% on bowling, one run each); it saves API spend only, GPU and power costs remain; and the planner only checks that promised files exist, not that they work — review has to catch the rest.

Related event: Cloud Planner Plus Local Coder Cuts API Costs by 77%(3 posts)→

Original post →

More from coding & agent

coding & agent channel →