DeepSeek V4 Flash Local Benchmark: MXFP4 Quantization Balances Speed and Top Scores
WonderRico · reddit · 2026-08-05
A Reddit user updated their local LLM benchmark with results for DeepSeek V4 Flash 0731.
Running the MXFP4 quantized version from Bartowski, the model achieved impressive speeds of 1K tokens/s for prefill and 90 tokens/s for generation using Dspark. The data shows this version scores the highest yet while maintaining extreme efficiency, though it notably lacks vision capabilities.
More from Infra
- Inference Farming: Turning Idle Consumer GPUs into Decentralized AI Revenue — 0xJeff · 2026-08-05
- Foxconn's Monthly Revenue Growth Accelerates Again in July Driven by AI Server Demand — firstadopter · 2026-08-05
- AMD MI355X Cluster Hits 1M+ tokens/s on Llama 2 70B in Multi-Node Test — AntDX316 · 2026-08-05
- a16z Talks Radiant: Building Portable Nuclear Microreactors to Power AI Infrastructure — a16z · 2026-08-05
- AI Compute Doubles Every 9 Months: 200M H100-Equivalent Chips by 2028 — TansuYegen · 2026-08-05
- Running Minimax H3 Video Generation Locally on RTX 4070: Stunning Results — Athem · 2026-08-05