DeepSeek 4.1 Flash hands-on: 7x cheaper cache hits, 552B MoE, and it can build a Cities: Skylines clone in Three.js
卡尔的AI沃茨 · wechat · 2026-09-11
A hands-on review of DeepSeek's new 4.1 Flash: from Sept 14, all v4-pro requests get force-routed to Flash at its lower price — cache-hit input costs dropped over 7x. Architecture: 552B MoE (8B active), 196B Engram memory params, 1M context, 384K max output, first native vision support, and KV cache compressed to 890 bytes/token (1/4 of the previous gen). It tops DeepSWE v1.1 above gpt5.6Sol and claudeOpus5, though complex science/vision analysis still trails closed frontier models.
Testing via the new DeepSeekHarness desktop app: a WebGL parking game, Crossy Road (less stable than GLM 5.3 Flash), a 3D pixel valley and anime-style Japanese street scene (big visual leap, beat GLM), plus a two-hour, 2M-token run reproducing the viral Reddit "Cities: Skylines in Three.js" prompt — producing an interactive city with a self-consistent micro-economy. Caveats: the phone-link feature is broken and occasional agent errors persist. Takeaway: cheapness is just the entry ticket; stability and saved attention matter more than token price.
More from Models
- GPT-6 Pro produces candidate proof for Erdős problem #488, passing two arithmetic checkers — basedjensen · 2026-09-11
- HighLevel claims early alpha access to rumored OpenAI GPT-Live-1, tests it in voice AI across 4M call insights — OpenAIDevs · 2026-09-11
- CritPt eval reportedly so broken that Ant built a fixed version, per F5.1 system card — xeophon · 2026-09-11
- Polymarket opens betting on whether OpenAI's GPT-6 Astra loses public access — Polymarket · 2026-09-11
- Early users find OpenAI's GPT-6 Astra surprisingly good at generating SVG icons — floguo · 2026-09-11
- Google AI Search Leaks Its 'Hidden Thoughts' on Pet-Toxicity Conflict — real_maximpulse · 2026-09-11