DeepSeek V4 Hits 1,328 tok/s Prefill on Single RTX 6000 via Krasis

mrstoatey · reddit · 2026-08-04

Using the inference optimization tool Krasis (v1.0.19), a developer achieved significant speedups for DeepSeek-V4-Flash-0731 (INT4) on a single RTX PRO 6000 (96GB).

By streaming the model across VRAM and system RAM, the tool keeps 6,440 of the hottest experts resident while offloading the rest. In tests, prefill speeds reached an impressive 1,328 tok/s for a 23k token prompt, with decode speeds of 29.4 tok/s at a 1K context. This offers a highly cost-effective local deployment setup for coding agents that frequently send massive contexts.

Original post →

More from coding & agent

coding & agent channel →