DeepSeek-V4-Flash Stuck in Doom Loop on llama.cpp Vulkan

KingCpzombie · reddit · 2026-08-04

A developer encountered a 'Doom Loop' issue while trying to locally deploy the DeepSeek-V4-Flash (Q8KXL quantized version) using llama.cpp Vulkan build (10216). The model gets stuck repeatedly outputting meaningless 'Need maybe' tokens before generating actual content.

Hardware & Config:

The user is unsure if building the absolute latest version directly from GitHub is required to fix this.

Original post →

More from Infra

Infra channel →