Debugging infinite loop error when running Qwen 27B on llama.cpp
Fieser_Fettsack · reddit · 2026-08-29
A user reports a recurring infinite loop of slashes (//) causing session crashes when running the Qwen3.8 27B model via llama.cpp, often triggered after the /compact feature. The setup involves a main server (RTX 3060) and an RPC worker (RTX 3080) with MTP and KV cache enabled. Despite extensive troubleshooting—including adjusting quantization, batch sizes, cache types, and Docker configurations—the issue persists. The post includes detailed docker-compose, Dockerfile, and models.ini for community assistance.
More from coding & agent
- Robotics Experiment: Claude Coding Failed Completely, Infrastructure Bugs Hinder Progress — verdakorz · 2026-08-30
- Dev Uses AI to One-Shot an Android Port of His 11-Year-Old Hand-Coded Wedding Canvas Art — steren · 2026-08-30
- Fully Open Source Stack: Qwen and Hermes Create a Self-Modifying PC Experience — ramagetime · 2026-08-30
- AGENTS.md vs SKILL.md: What's the difference in AI development? — _jaydeepkarale · 2026-08-30
- AI Agent workflow evolution: from simple triggers to verified production steps — kashifmanzoor · 2026-08-30
- Spent $380 on a Looping GPT-4 Script, So I Built a Multi-Provider Cost Monitor — Ok_Anything_8323 · 2026-08-30