DeepSeek-V4-Flash Inference Blocked: vLLM Lacks Support for New confidence_head

teortaxesTex · x · 2026-08-01

Developers found that DeepSeek-V4-Flash-0731 includes a DSpark confidencehead, but the current public vLLM NVIDIA loader drops it because the head is not yet wired into the inference process.

Original post →

More from Infra

Infra channel →