Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month
BLUECOW009 · x · 2026-09-21
Tobi Lütke (Shopify CEO) revealed a Dell server running DeepSeek 4.1 Flash locally, at roughly 300 tokens per second with quality somewhere between Opus 4.7 and Opus 5. Not cheap, he says, but a one-time fixed cost now buys about a billion high-quality tokens per month — local inference economics are closing in on frontier APIs.
More from Infra
- Researchers formally verify the Kubernetes control plane with a compositional CORE spec — tianyin_xu · 2026-09-21
- 99.7% cache hits: engineered DeepSeek Harness with self-hosted GLM-5.3 — burny_tech · 2026-09-21
- Solo dev open-sources 4 systems projects, asks engineers to roast them — Accomplished_Row1433 · 2026-09-21
- Andrew Chen: strong LLMs are far from running on phones, on-device AI faces bandwidth, heat and model-size hurdles — andrewchen · 2026-09-21
- Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP — bring_back_the_v10s · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21