Running DeepSeek V4 Flash Locally: Demystifying Quantization Naming and Precision
TheZachMueller · x · 2026-08-01
Unsloth AI announced that DeepSeek V4 Flash 0731 can now be run locally, supporting lossless 4-bit quantization on 168GB RAM and 3-bit on 110GB RAM.
Addressing community questions about quantization naming, developer danielhanchen explained that because llama.cpp lacks native FP8 support, it uses names like Q8KXL (MXFP4+BF16, 100% lossless) and Q4KXL (MXFP4+Q80, 96% same top-1%). He emphasized that FP8 is not equivalent to Q80, stating that claiming FP8 is lossless is incorrect.
More from Infra
- antirez looks into deploying LLMs on DGX Spark — antirez · 2026-08-01
- Modal Releases Comprehensive GPU Glossary Covering Hardware to Software Stack — charles_irl · 2026-08-01
- ARM Introduces FEAT_CSSC: Native Popcount for General-Purpose Registers — lemire · 2026-08-01
- Train Your Own Model When Inference Exceeds $750/Day: Pallet's Playbook — marcbhargava · 2026-08-01
- Local Inference of 91GB Audio Model: 127GB RAM Needed for 1M Context — andimarafioti · 2026-08-01
- Llama 3.1 405B Hits 5.6k t/s on Cerebras for Select Customers — kimmonismus · 2026-08-01