Debugging infinite loop error when running Qwen 27B on llama.cpp

Fieser_Fettsack · reddit · 2026-08-29

A user reports a recurring infinite loop of slashes (//) causing session crashes when running the Qwen3.8 27B model via llama.cpp, often triggered after the /compact feature. The setup involves a main server (RTX 3060) and an RPC worker (RTX 3080) with MTP and KV cache enabled. Despite extensive troubleshooting—including adjusting quantization, batch sizes, cache types, and Docker configurations—the issue persists. The post includes detailed docker-compose, Dockerfile, and models.ini for community assistance.

Original post →

More from coding & agent

coding & agent channel →