Sonnet orchestrating local Qwen 3.8 27B cuts costs 2.7x but drops quality from 4/5 to 2/5
GapNew4766 · reddit · 2026-09-29
Testing the popular "cheap local model + expensive cloud model" split on 5 physics-scene builds with one identical prompt: Sonnet 5.5 alone cost $1.83 in 9 min (4/5 scenes), local Qwen 3.8 27B on a 3090 cost $0 in 26 min (1/5, hollow teleporting cubes), and fusion mode — Sonnet writing briefs (5.4K tokens vs 37.9K), Qwen writing all code — cost $0.67 in 85 min (2/5). Reversing roles (Qwen plans, Sonnet codes) hit 4/5 in 14 min but cost more than Sonnet alone since every cloud worker re-reads context. Verdict: the split really cuts the bill 2.7x, paid for in time and quality.
Related event: Cloud-Planning Plus Local-Execution Cuts Costs 2.7x but Takes 9x Longer(3 posts)→
More from coding & agent
- AI.CLASSIFY coming to Quail: rethink OLAP-style AI operations on data — sh_reya · 2026-09-30
- Ant Group's Marathoner Trains Agents to Work 10+ Hours With 1000+ Tool Calls — antgroup · 2026-09-30
- Supply Chain Agent Lesson: Memory Knows the Rules, Code Enforces Them — Chemical_Pickle5979 · 2026-09-30
- Artificial Analysis open-sources agent inference benchmark, first results for DGX Spark, RTX 5090, M5 Pro — ArtificialAnlys · 2026-09-30
- Feed Your Agents Markdown: How a Good API Turned cf CLI Into an Unexpected Agent Tool — irvinebroque · 2026-09-30
- 14% of AI-Written Tests Were Useless: The Checker Rewarded Typing Words, So Agents Typed Them — Input-X · 2026-09-30