35B Agentic Bakeoff: KAT-Coder Matches Qwen at Half the Token Cost
IvGranite · reddit · 2026-07-28
A developer conducted a rigorous head-to-head comparison of several 35B parameter coding models in agentic tasks. The test included 4 models, 6 tasks, 5 repetitions each, totaling 120 runs.
Results show that KAT-Coder-V2.5-Dev matched the top stock pass rate (29/30, tied with Qwen3.5-35B) using only half the input tokens. It also exhibited the cleanest tool behavior, with zero malformed tool-call leaks in 30 runs. In contrast, stock Qwen3.6 was the strongest analyst but a massive token waster with frequent tool-call format leaks. Ornith underperformed due to mechanical failures and whole-file rewrites.
More from coding & agent
- Parallax pitches a personal-agent operating system built on LLMs — avlok · 2026-07-28
- Save 50% on Claude Code Limits by Combining It with OpenAI Codex — dr_cintas · 2026-07-28
- Claude 3 Opus Autonomously Constructs a Maze World Model In-Context — burny_tech · 2026-07-28
- MCP Connect pitches one protocol for connecting AI agents to tools and data — WirelessLife · 2026-07-28
- Claude Code in Action: Automating Ad Account Analysis and Angle Generation — alexgoughcooper · 2026-07-28
- Replit Agent Adds One-Click Integration for you.com Search MCP — RichardSocher · 2026-07-28