DeepSeek V4 Flash Local Benchmark: MXFP4 Quantization Balances Speed and Top Scores

WonderRico · reddit · 2026-08-05

A Reddit user updated their local LLM benchmark with results for DeepSeek V4 Flash 0731.

Running the MXFP4 quantized version from Bartowski, the model achieved impressive speeds of 1K tokens/s for prefill and 90 tokens/s for generation using Dspark. The data shows this version scores the highest yet while maintaining extreme efficiency, though it notably lacks vision capabilities.

Original post →

More from Infra

Infra channel →