Running DeepSeek V4 Flash on Mac: M3 Ultra Hits 43 tok/s
Professional-Bear857 · reddit · 2026-08-04
A developer on Reddit shared what they consider the best quantization scheme for running DeepSeek V4 Flash on a Mac (M3 Ultra, 192GB+ RAM).
- Model & Tool: Used the Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX quantization from Hugging Face, which supports dspark/mtp.
- Performance: Generation speed increased as cached tokens grew, starting at 34 tok/s and ending at 43 tok/s. The test included the default chat prompt and a 13k token query.
More from Infra
- Save ~48MB RAM Per Execution Using `node --run` Over `npm run` in Node 22+ — DanielLockyer · 2026-08-04
- CoreWeave Plans First APAC Data Centers in Indonesia with 360MW Capacity — dinabass · 2026-08-04
- DeepSeek V4 Flash Quantization Benchmark: IQ3_XXS 2x Faster with No Quality Loss — Spicy_mch4ggis · 2026-08-04
- Analyzing the Transpose Bottleneck in mxfp8 Quantization and VRAM Optimization — dejavucoder · 2026-08-04
- Bittensor's SayGm Offers Single API Key Access to 38 Major AI Models — bittingthembits · 2026-08-04
- Potential Ban Could Slow Data Center Buildouts by 50%, Spike Optical Component Demand — zephyr_z9 · 2026-08-04