DeepSeek V4.1 Flash tops Chinese models in coding blind test, but drops 17 points when switching clients

teortaxesTex · x · 2026-09-09

A 7-way "model + coding client" blind test on the same production-grade code task put DeepSeek V4.1 Flash + Claude Code first among Chinese models at 76.69, ahead of Qwen 3.8 Flash (75.13), Kimi K3-256K (70.73), Qwen 3.8 Max (69.58) and GLM 5.3 (64.72).

Takeaway: Flash already competes with flagships, but engineering edge cases remain — and the client itself is currently the biggest variable. A single-task engineering eval, not a general capability ranking.

Original post →

More from coding & agent

coding & agent channel →