DeepSeek-V4-Flash hits 51.5 tok/s on M3 Ultra
antirez · x · 2026-08-27
Developer Ivan Fioravanti continues optimizing DeepSeek-V4-Flash inference on M3 Ultra via DwarfStar, boosting speed from 45.7 to 51.5 tok/s (+12.6%).
Key Details:
- Applies Q80 to Q4K quantization for attention+head (not mathematically 1-1 with mxfp4).
- Quality Gate: 85/92 on full eval (vs 82/92 baseline); AIME score improved to 22→24/25.
- Attempting to recover CompSec using imatrix calibration.
- Code branch is on GitHub; Hugging Face model release pending.
More from coding & agent
- Claude Code 2.1.247 adds cost optimization command — ClaudeCodeLog · 2026-08-27
- Developer runs self-built Pokémon Emerald on original Game Boy Advance hardware — IanArawjo · 2026-08-27
- MARS: Multi-Specialist LLM Relay System Boosts Competitive Programming Pass Rates — Andrei Mikhailov · 2026-08-27
- Today's agents only think when pinged: a case for dissonance-driven cognitive initiative — GlenBradley · 2026-08-27
- Agent Architecture Recap: CoS Coordinating 20+ Agents and High-Volume PR Workflows — RachelVT42 · 2026-08-27
- Microsoft and Google's WebMCP standard adds agent buttons to websites — HankYeomans · 2026-08-27