Why Qwen3.8 27B costs more than DSv4 Flash? KV cache and quantization explained
ihatebeinganonymous · reddit · 2026-08-20
A Reddit user asks why serving Qwen3.8 27B is more expensive than DeepSeek V4 Flash despite the latter having 10x more parameters. Analysis reveals Qwen's KV cache is much larger and quantization differs (DeepSeek native Q4 vs Qwen BF16/Q8). Model size amortizes but KV cache doesn't, making DSv4 Flash ideal for multi-user serving and Qwen for local single-user. 200k context would require huge RAM on Qwen.
More from Infra
- LiquidAI releases QAD checkpoints, 4-bit models retain 97% BF16 accuracy — JosephJacks_ · 2026-08-20
- AI data center controversy silly? Footprint less than an almond field — justin_hart · 2026-08-20
- What dense GEMM MFUs are people seeing on Blackwells with mxfp8? — isidentical · 2026-08-20
- Rise of S3-native products as a new developer trend — DanielLockyer · 2026-08-20
- Terminal tool cmux reported for high memory usage — DanielLockyer · 2026-08-20
- Loudoun County Gets Rich From Data Centers, Residents Question the Cost — AndyMasley · 2026-08-20