35B Agentic Bakeoff: KAT-Coder Matches Qwen at Half the Token Cost

IvGranite · reddit · 2026-07-28

A developer conducted a rigorous head-to-head comparison of several 35B parameter coding models in agentic tasks. The test included 4 models, 6 tasks, 5 repetitions each, totaling 120 runs.

Results show that KAT-Coder-V2.5-Dev matched the top stock pass rate (29/30, tied with Qwen3.5-35B) using only half the input tokens. It also exhibited the cleanest tool behavior, with zero malformed tool-call leaks in 30 runs. In contrast, stock Qwen3.6 was the strongest analyst but a massive token waster with frequent tool-call format leaks. Ornith underperformed due to mechanical failures and whole-file rewrites.

Original post →

More from coding & agent

coding & agent channel →