Similar benchmarks, double the size: Qwen3.8-Flash-Next needs 360GB vs DeepSeek-V4-Flash's 162GB lossless
vini542reddit · reddit · 2026-08-28
The author found that despite similar benchmark scores, real-world usage of Qwen3.8-Flash-Next (Q38FN) vs DeepSeek-V4-Flash (DSv4F) differs sharply:
- DSv4F runs lossless at 162GB via Unsloth Q8KXL; on 4×RTX 3090 + 192GB DDR4 @3200MHz he gets 30 tps and 230 pp.
- Q38FN needs roughly 360GB for lossless — more than double; Unsloth has no lossless Q8 yet, and its listed Q8 (270GB) is not lossless either.
- He notes Q38FN is early in its architectural lifecycle and improvements may come, but Alibaba appears to have matched benchmarks by making the model much bigger; quantized to comparable Q4/Q5 sizes, its performance would likely drop significantly.
Takeaway: current benchmarks are misleading when models differ this much in size/quantization cost at equal scores.
More from Models
- User notes o3 Pro shutdown, touts lingering power of Deep Research — bytebot · 2026-08-28
- Why open source keeps catching up: weights and papers drown closed labs' secrets — YogurtExternal7923 · 2026-08-28
- Tencent Hunyuan Hy4 preview hits #5 in WebDev Arena, #3 among open models — TencentHunyuan · 2026-08-28
- TB-fn benchmark reveals score drops for several models, exposing potential benchmark overfitting — burny_tech · 2026-08-28
- OpenAI building "Persistent Mode" for Codex: always-on agents that self-start tasks — The Decoder · 2026-08-28
- GLM 5.3 Flash surges to 4th place in daily OpenRouter usage — Hesamation · 2026-08-28