DeepSeek v4.1 Flash preview spotted: ~300 tok/s, fires lots of subagents, no vision yet
kevinkern · x · 2026-09-10
Developer Kevin Kern tested an upcoming DeepSeek v4.1 Flash preview endpoint, reporting inference at roughly 300 tok/s and heavy parallel subagent usage on his workloads.
The preview still lacks vision; he hopes the official release will be multimodal.
More from Models
- Grok: Clay Institute prize is far off — OpenAI's math result still needs extensive vetting — MikePFrank · 2026-09-10
- Dev finds Astra inconsistent, says Sol 5.6 was smarter for most software engineering tasks — MinqiJiang · 2026-09-10
- DeepSeek V4.1 Flash OCR Benchmarked: 260 tok/s, Competitive Accuracy, Low Cost — solyarisoftware · 2026-09-10
- Sante Medical AI Model Launched on Ling-3.0-flash: Focus on Reasoning, Safety, Evidence — FellMentKE · 2026-09-10
- Yoav Goldberg: 880K GPU-Hours vs a Few Hundred Prompted LLM Hours for Same Math Result — yoavgo · 2026-09-10
- K3 got 77% speedup on mjwarp kernels; GPT-6 Astra added just 0.38% — YouJiacheng · 2026-09-10