antirez Enables Lossless MXFP4 Local Inference for DeepSeek v4 Flash
antirez · x · 2026-08-01
Developer antirez announced that his DwarfStar branch now supports running the lossless MXFP4 DeepSeek v4 Flash GGUF model he published on Hugging Face.
The setup achieves over 20 tokens/s inference speed on 128GB RAM systems, even with SSD streaming. This provides an efficient way to test the actual DS4F weights locally without any quantization.
More from Infra
- DeepSeek V4 Flash local benchmark nearly matches top frontier models from 5 months ago — joorklee · 2026-08-01
- New Method Pre-routes MoE Layers to Optimize I/O for Edge Streaming — dai_app · 2026-08-01
- 5TB of Data Stored on a Tiny Glass Slab Marks Microscopic Storage Breakthrough — TansuYegen · 2026-08-01
- Open-Source Engine 'Waste' Runs Kimi K3 on Just 29GB of RAM — galapag0 · 2026-08-01
- Running Qwen 27B on 3x 2080Ti at 55tps: Optimal Config Shared — AccountGotLocked69 · 2026-08-01
- DeepSeek-V3 Trained With Only 180K GPU-Hours, Slashing MoE Compute Costs — teortaxesTex · 2026-08-01