Fixed Chat Template for Qwen 3.8: Enables Reasoning Effort Control and Tool Calling
ex-arman68 · reddit · 2026-08-14
A developer released a fixed Jinja chat template for Qwen 3.5, 3.6, and 3.8 models. Qwen 3.8 introduces prompt-steered reasoning effort, but the official template has issues like inability to disable thinking, poisoned chat history, and tool calling crashes. The fixed template supports full reasoning effort control, thinking toggle, KV cache hits, native llama.cpp support, and works with various inference engines.
More from coding & agent
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Comparing AI Subscriptions: DeepSeek API vs. Claude Pro vs. Local LLMs — Unlikely_Bluejay5392 · 2026-08-24
- Claude Code introduces 'Remote Control' feature to boost coding efficiency — rohanpaul_ai · 2026-08-24
- rauchg lays out fx extension philosophy: MCP, Skills, Plugins and Unix composition — AccBalanced · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- smolvm passes Simon Willison's Fable 5 agent test as a secure sandbox — yawnxyz · 2026-08-24