antirez Enables Lossless MXFP4 Local Inference for DeepSeek v4 Flash

antirez · x · 2026-08-01

Developer antirez announced that his DwarfStar branch now supports running the lossless MXFP4 DeepSeek v4 Flash GGUF model he published on Hugging Face.

The setup achieves over 20 tokens/s inference speed on 128GB RAM systems, even with SSD streaming. This provides an efficient way to test the actual DS4F weights locally without any quantization.

Original post →

More from Infra

Infra channel →