DeepSeek 4.1 Flash Hits 200 TPS Full Precision on 4 Max-Qs with Just 64GB RAM

TheZachMueller · x · 2026-09-11

A developer ran DeepSeek 4.1 Flash at full precision with the DSpark engine on 4 Max-Q chips with only 64GB of system RAM, reaching 200 TPS.

The key trick: offloading the 200GB Engram (essentially a hash table) to NVMe instead of keeping it in RAM works surprisingly well and saves cost. The author says there's more headroom and a vLLM recipe is coming.

Related event: DeepSeek V4.1 Flash tested across tasks: near-frontier performance at a fraction of the cost(15 posts)→

Original post →

More from Infra

Infra channel →