35B Agentic Bakeoff: KAT-Coder Matches Qwen at Half the Token Cost
IvGranite · reddit · 2026-07-28
A developer conducted a rigorous head-to-head comparison of several 35B parameter coding models in agentic tasks. The test included 4 models, 6 tasks, 5 repetitions each, totaling 120 runs.
Results show that KAT-Coder-V2.5-Dev matched the top stock pass rate (29/30, tied with Qwen3.5-35B) using only half the input tokens. It also exhibited the cleanest tool behavior, with zero malformed tool-call leaks in 30 runs. In contrast, stock Qwen3.6 was the strongest analyst but a massive token waster with frequent tool-call format leaks. Ornith underperformed due to mechanical failures and whole-file rewrites.
More from coding & agent
- New tool puts Claude, Codex, Grok in shared sessions with your teammates — sergeykarayev · 2026-09-23
- Dev claims 20k more commits coming: Opus 5.5 and GPT-6 Sol supercharge his output — doodlestein · 2026-09-23
- A JEV-powered Wireshark classifier accidentally uncovered real backdoors on a home network — multiply_matrix · 2026-09-23
- 299 real intents tested: classifier routing trails GLM-4-Flash by 3 points but is 6.5x faster — Sufficient_Flower860 · 2026-09-23
- OpenExecutive: open-source virtual executive team of 8 specialist AI agents hits 5.1k GitHub stars — tom_doerr · 2026-09-23
- Framer launches Skills: teach your design agent reusable workflows, design systems and CMS rules — soleio · 2026-09-23