Running DeepSeek V4 Flash Locally: Demystifying Quantization Naming and Precision

TheZachMueller · x · 2026-08-01

Unsloth AI announced that DeepSeek V4 Flash 0731 can now be run locally, supporting lossless 4-bit quantization on 168GB RAM and 3-bit on 110GB RAM.

Addressing community questions about quantization naming, developer danielhanchen explained that because llama.cpp lacks native FP8 support, it uses names like Q8KXL (MXFP4+BF16, 100% lossless) and Q4KXL (MXFP4+Q80, 96% same top-1%). He emphasized that FP8 is not equivalent to Q80, stating that claiming FP8 is lossless is incorrect.

Related event: Unsloth Releases Quantized DeepSeek V4 Flash 0731 for Local Lossless Running(5 posts)→

Original post →

More from Infra

Infra channel →