Benchmark Report: OSS Models Need 10-20x Tokens, Slow for HITL Coding
nateberkopec · x · 2026-09-01
Based on data from @ArtificialAnlys, the author compares frontier models for interactive, human-in-the-loop (HITL) coding sessions. Key takeaways:
- Speed & Cost: GLM-5.3-Flash is cheap but slow; Sol dominates the time-per-task frontier; Gemini 3.7 Flash is very similar to Sol; Opus is smart and fast but expensive; Grok is good but a step down.
- OSS vs. Proprietary: OSS models require 10-20x as many output tokens per task, making them extremely slow and painful for interactive use.
- Conclusion: For interactive HITL work, choose a sub from a frontier AI lab (Google/SpaceX/OpenAI/Anthropic), as their performance is close enough to be considered equivalent.
Related event: AI coding benchmark findings disputed over flawed data(2 posts)→
More from coding & agent
- Design pattern: immutable agent artifact revisions behind a stable review URL — RocketSeven · 2026-09-01
- Building a long-term memory benchmark for agents: what to add? — True_Mongoose_7073 · 2026-09-01
- Built MCP Gate to keep agents from holding Gmail/AWS credentials — No_Ground6610 · 2026-09-01
- Grok Bot orchestrator manages entire media business — minchoi · 2026-09-01
- Dr Eggbot: A Grok Bot that builds high-quality bots for you — minchoi · 2026-09-01
- Viral Grok Bot templates: Video editor, PM, research, sales & more — minchoi · 2026-09-01