llama-server crashes with 'non-consecutive token position' on AMD/Vulkan
Gold-Drag9242 · reddit · 2026-08-17
A user reports frequent crashes with non-consecutive token position warnings when running llama-server on Windows with AMD 7900XTX (Vulkan backend). The issue persists across various models like Gemma and Qwen during multimodal tasks. Splitting prompts used to help but fails with the latest Qwen3.8 build.
More from Infra
- Open-source zxLLM predicts LLM VRAM usage & KV-cache needs with high precision — Capable_Item_5918 · 2026-08-17
- Immense compute barrier blocks new frontier AI labs — 0xsachi · 2026-08-17
- WSJ: Nine Tech Giants Hold $3 Trillion in Off-Balance-Sheet AI Obligations — alvelda · 2026-08-17
- Command example to run Qwen3.8-27B GGUF on DGX Spark — ariG23498 · 2026-08-17
- How to start running Qwen 3.8 locally with a 3090 GPU? — ozymandizz · 2026-08-17
- Qwen MLX Challenge Launches to Benchmark Local Model Inference Speed — corruptbytes · 2026-08-17