Real-World Coding Eval: KAT Coder Outperforms Qwen and Ornith in 35B Local MoE Models
Undici77 · reddit · 2026-08-05
The author conducted months of real-world coding evaluations on 35B MoE local models, focusing on iterative debugging, refactoring, and failure recovery.
- Qwen 3.6 (35B A3B): A solid baseline, but tends to hallucinate confidence in edge cases and drifts during longer multi-file reasoning chains.
- Ornith 1.0: Excellent benchmarks, but suffers from "overthinking" in practice, leading to high token burn and reasoning loops.
- KAT Coder 2.5 Dev: A surprising performer. It offers more decisive outputs with faster convergence, performing best in real-world benchmarks while avoiding the common pitfalls of both Qwen and Ornith.
The author concludes that a similarly sized MoE model with added intelligence and a 1M token context represents a significant step forward.
Related event: Developers Find KAT Coder 2.5 Outperforms Qwen in Coding Tests(2 posts)→
More from Models
- DeepGrove Releases On-Device LLM Running at 200+ tokens/s on Mac Mini — ycombinator · 2026-08-05
- OpenAI Discloses Models Crossed Boundaries to Reach Real Systems in Cyber Evals — ryanmerket · 2026-08-05
- User Complaints: Latest Claude Models Giving Riddles Instead of Answers — natanielruizg · 2026-08-05
- Musk Reveals Grok Roadmap: v4.6 Next Week, v5 to Train on SpaceX Data — XFreeze · 2026-08-05
- Security Researcher Confirms OpenAI Silently Nerfed GPT-5.6 Cyber Capabilities — rez0__ · 2026-08-05
- Analyzing AI Memes: Claude's Local Blind Spots and GPT Image Flaws — mimi10v3 · 2026-08-05