Local Q8 Showdown: DeepSeek-V4-Flash-Vision Finishes Tasks 2x Faster Than Qwen3.8
pabloodiablo · reddit · 2026-09-07
The author compares DeepSeek-V4-Flash-Vision (DSV4FV) vs Qwen3.8-Flash-Next (Q38FN), both at Q8KXL, as local daily drivers on 2x Strix Halo 128GB via USB-C 4 with Llama (RPC).
Key findings:
- Raw token generation: DSV4FV is 40% slower than Q38FN.
- Task completion: DSV4FV is 2x faster — a task taking Q38FN 25 minutes on medium took DSV4FV just 12 minutes, which the author attributes to fewer hallucinations and less getting stuck.
- Q38FN's 'xhigh' mode is unusable: a 25-minute task on medium failed to complete in 3 hours.
- DSV4FV on max finished the same task in 37-44 minutes.
- Behavior: Qwen3.8 over-interprets underspecified instructions, adding too much and bogging down in its own creativity.
Verdict: DSV4FV is the better choice for professional coding, though the author remains a big Qwen fan overall.
More from Infra
- Microsoft open-sources tgrep, a trigram-indexed grep up to 52x faster than ripgrep — jedisct1 · 2026-09-07
- How should billing work when an AI system auto-selects the model? — Colddew-YJ · 2026-09-07
- Can you run Qwen Next on a 3090 + 64GB CMP 170HX? Local deployment help — JustinPooDough · 2026-09-07
- SmolVM: open-source microVM sandbox runs OpenClaw 2.0 in isolation, boots in milliseconds — aniketmaurya · 2026-09-07
- Hesamation recommends the best technical book on training LLMs at scale — free to read — Hesamation · 2026-09-07
- After His OpenAI Key Was Stolen, He Found Stratum: a Docker-Layer Secret Scanner Crunching 700K Layers Daily — Ubunta · 2026-09-07