DeepSeek V4.1 Flash beta goes live with new architecture, native multimodality, up to 507 tokens/s
智东西 · wechat · 2026-09-08
DeepSeek opened beta testing for an intermediate version of DeepSeek V4.1 Flash, built on a new architecture with native multimodal support, faster speed and lower cost. Developers can call it via model name deepseek-v4.1-flash-expires-on-0910 without changing baseurl; the beta auto-expires on Sept 10.
- Pricing matches V4 Flash: idle hours at ¥0.05/¥1.5/¥4.5 per million tokens (cache-hit input / cache-miss input / output); peak hours ¥0.1/¥3/¥9
- Rate limited to 20 concurrent requests per account
- Community tests report output speeds up to 507 tokens/s, averaging over 300 tokens/s in the pelican-on-a-bicycle SVG test
DeepSeek has shipped rapidly since July (V4 Flash GA, V4 Pro, open-source DeepSeekHarness, V4 Flash Vision Exp), suggesting a new model update cycle.
More from Models
- CPU-Only LLM Tests: 35B MoE at Q2 Beats a 2B Model Despite Half the Speed — ML-Future · 2026-09-08
- Vision model tier list updated: GPT-6 Astra dethrones Gemini — evilsocket · 2026-09-08
- From IMO Gold to Solving a Millennium Problem in Just ~14 Months — felpix_ · 2026-09-08
- Updated AA Intelligence Index Puts Gemini 3.8 Flash Below GLM 5.3 Flash — Able-Line2683 · 2026-09-08
- V4.1-Flash-Vision hands-on: sharper and much faster, but overthinks and gets cheeky — teortaxesTex · 2026-09-08
- Users With Cyber Access Keep Hitting Claude Guardrails Daily, Sparking Overblocking Complaints — IgorCarron · 2026-09-08