DeepSeek v4 Flash GGUF quant released for DS4 engine, doubles speed to 30+ tok/s
returnity · reddit · 2026-08-01
Developer nazeshinjite released a GGUF quantized version of DeepSeek v4 Flash (DwarfStar) with a separate DSpark MTP head. Running on the DS4 engine, it achieves over 30 tok/s on MBP M5 Max, double llama.cpp's speed, making it suitable for agentic workflows. The quant uses antirez's Q2-Q4 mixed imatrix recipe, and the author seeks community feedback for performance improvements.
More from coding & agent
- DeepSeek-V4-Flash-High Tops Price-Performance in Frontend Code Arena, Ranks #7 Overall — arena · 2026-08-01
- claude-pulse: Real-Time Status Bar Monitor for Claude Code Usage Limits — tom_doerr · 2026-08-01
- Waterloo's R2L Lab to Recruit PhDs, Focusing on Agents and Reasoning Research — hllo_wrld · 2026-08-01
- Codex Agent Autonomously Files Bug Reports with Support Chatbot — ___Patrice___ · 2026-08-01
- AgentIR: Deep Research Agents That Leverage Reasoning Context for Retrieval — hllo_wrld · 2026-08-01
- Getting Computer-Use Clicks Under 10ms: Cloud Sandbox Architecture Optimization — charles_irl · 2026-08-01