Using GPT-6.1 Sol only as planner cuts coding API bill by 77%, local 27B writes code
GapNew4766 · reddit · 2026-09-30
A developer benchmarked three ways to build three small 3D games (pool, bowling, foosball): letting GPT-6.1 Sol do everything, having Sol only plan while a local Qwen 3.8 27B on a single RTX 3090 writes the code, and going fully local. The planner/worker split cut API spend from $0.75 to $0.17 — a 77% saving — at the cost of runtime rising from 6.6 to 43.4 minutes.
The rationale: most tokens in a coding run are the code itself, while planning and review are a small slice, so pushing the bulk of output onto a per-token-free local model saves the most. The open-source tool Atomic Agent implements this as a "Fusion" mode, but any framework that lets you set separate planner and editor models works.
Caveats: savings varied (87% on pool vs 57% on bowling, one run each); it saves API spend only, GPU and power costs remain; and the planner only checks that promised files exist, not that they work — review has to catch the rest.
Related event: Cloud Planner Plus Local Coder Cuts API Costs by 77%(3 posts)→
More from coding & agent
- Artificial Analysis open-sources agent inference benchmark, first results for DGX Spark, RTX 5090, M5 Pro — ArtificialAnlys · 2026-09-30
- Feed Your Agents Markdown: How a Good API Turned cf CLI Into an Unexpected Agent Tool — irvinebroque · 2026-09-30
- Cursor + data export makes debugging GA4/GSC traffic spikes a 10-minute job — gaganghotra_ · 2026-09-30
- Intern-Decision open models (0.8B–4B) beat Jev on multimodal decision benchmarks — max_paperclips · 2026-09-30
- xiaohu shows the computer assigned to his AI agent: browser preinstalled and it can make calls — xiaohu · 2026-09-30
- Editable artifacts may beat screenshots for testing agents' visual understanding — OliviaYii · 2026-09-30