DeepSeek V4.1 Flash hits 532 tokens/s on Inco, fastest output on Artificial Analysis
songhan_mit · x · 2026-09-18
Inference provider Inco announced DeepSeek V4.1 Flash is live with 532 tokens/s output, ranking #1 output speed on Artificial Analysis with a clear lead.
Blogger songhanmit amplified the point: what matters isn't just being fast but iterating fast — high throughput accelerates the entire loop of interactive development and experimentation. The service is live with benchmark links attached.
Related event: DeepSeek Launches V4.1-Flash with Native Vision, Tops Speed Charts(3 posts)→
More from Infra
- Weaviate 1.39 adds 4-bit Rotational Quantization, cutting memory with minimal recall loss — CShorten30 · 2026-09-18
- Huawei's 100K-card Super Cluster can train a 10T-param model on 100T tokens in 30 days — teortaxesTex · 2026-09-18
- MLX Community Makes Qwen 3.8 Flash Nearly 2x Faster on Apple Silicon, License Blocks Launch — gajesh · 2026-09-18
- CoreWeave Brings Multi-Rack NVIDIA Vera Rubin NVL72 Cluster Online — Beth_Kindig · 2026-09-18
- Dev reimplements DeepSeek v4.1 Flash, runs 1M context at 1-10 tok/s on a single RTX 4090 — _xjdr · 2026-09-18
- Brad Gerstner at All-In Summit: who pays for AI CapEx, the gigawatt gap and semis eating the Nasdaq — DavidSacks · 2026-09-18