DeepSeek 4.1 Flash Hits 200 TPS Full Precision on 4 Max-Qs with Just 64GB RAM
TheZachMueller · x · 2026-09-11
A developer ran DeepSeek 4.1 Flash at full precision with the DSpark engine on 4 Max-Q chips with only 64GB of system RAM, reaching 200 TPS.
The key trick: offloading the 200GB Engram (essentially a hash table) to NVMe instead of keeping it in RAM works surprisingly well and saves cost. The author says there's more headroom and a vLLM recipe is coming.
More from Infra
- Google Commits $15B to AI Infrastructure Buildout in Finland — LinkedInNews · 2026-09-11
- DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression — ollama · 2026-09-11
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11