llama-server crashes with 'non-consecutive token position' on AMD/Vulkan

Gold-Drag9242 · reddit · 2026-08-17

A user reports frequent crashes with non-consecutive token position warnings when running llama-server on Windows with AMD 7900XTX (Vulkan backend). The issue persists across various models like Gemma and Qwen during multimodal tasks. Splitting prompts used to help but fails with the latest Qwen3.8 build.

Original post →

More from Infra

Infra channel →