DeepSeek 4.1 Flash hands-on: 7x cheaper cache hits, 552B MoE, and it can build a Cities: Skylines clone in Three.js

卡尔的AI沃茨 · wechat · 2026-09-11

A hands-on review of DeepSeek's new 4.1 Flash: from Sept 14, all v4-pro requests get force-routed to Flash at its lower price — cache-hit input costs dropped over 7x. Architecture: 552B MoE (8B active), 196B Engram memory params, 1M context, 384K max output, first native vision support, and KV cache compressed to 890 bytes/token (1/4 of the previous gen). It tops DeepSWE v1.1 above gpt5.6Sol and claudeOpus5, though complex science/vision analysis still trails closed frontier models.

Testing via the new DeepSeekHarness desktop app: a WebGL parking game, Crossy Road (less stable than GLM 5.3 Flash), a 3D pixel valley and anime-style Japanese street scene (big visual leap, beat GLM), plus a two-hour, 2M-token run reproducing the viral Reddit "Cities: Skylines in Three.js" prompt — producing an interactive city with a self-consistent micro-economy. Caveats: the phone-link feature is broken and occasional agent errors persist. Takeaway: cheapness is just the entry ticket; stability and saved attention matter more than token price.

Original post →

More from Models

Models channel →