DeepSeek v4.1 Flash on-device test: q2 runs at 16 tok/s but tool calls go off the rails
challis88ocarina · reddit · 2026-09-12
- Running DeepSeek v4.1 Flash (q2) on an M3U 32/80c device: 300 tok/s prefill, 16 tok/s decode holding until 110k tokens — far better sustained speed than llama.cpp's steep drop-off, with GPU at 95% throughout.
- Tool calls work but choices are questionable: in a hermes workflow it ran find twice for a remote file locally, then ls /.ssh/ and grepped the entire remote machine despite the path already being in context; even 18k tokens of hermes prompt didn't rein it in.
- It ignores the code-execution tool (same as 4.0), making it less efficient than GLM 5.3 Flash (21-19 tok/s in comparison).
- Trade-off: big VRAM means waiting for big models; q4 gains look limited — bring on the M5U.
More from Infra
- Glass core substrates show 2x better warpage than organic core without stiffener — jwt0625 · 2026-09-12
- d-Matrix partners with NVIDIA to plug Raptor XPUs into NVLink Fusion rackscale systems — bookwormengr · 2026-09-12
- Running out of context on a large codebase: how to auto-handoff long-running local LLM tasks — Developer-Y · 2026-09-12
- The Economist: Nvidia is the central bank of AI — tolugenius · 2026-09-12
- What MoE/LLM runs well offline on a 24GB M5 MacBook Air? — itis_whatit-is · 2026-09-12
- Draft model hits ~60 tok/s running Qwen3.8-27B at 131k context on a 16GB GPU — pneuny · 2026-09-12