Why Qwen3.8 27B costs more than DSv4 Flash? KV cache and quantization explained

ihatebeinganonymous · reddit · 2026-08-20

A Reddit user asks why serving Qwen3.8 27B is more expensive than DeepSeek V4 Flash despite the latter having 10x more parameters. Analysis reveals Qwen's KV cache is much larger and quantization differs (DeepSeek native Q4 vs Qwen BF16/Q8). Model size amortizes but KV cache doesn't, making DSv4 Flash ideal for multi-user serving and Qwen for local single-user. 200k context would require huge RAM on Qwen.

Original post →

More from Infra

Infra channel →