llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default
Bulky-Priority6824 · reddit · 2026-09-03
llama.cpp release b10763 deprecates --chat-template-kwargs. The preservereasoning chat template kwarg is now enabled by default after argument processing unless explicitly set via --reasoning-preserve / --no-reasoning-preserve; the server logs the effective state and only warns "has no effect" when explicitly enabled on an unsupported template.
More from Infra
- Reddit: Picking the Best Chat Model for a 3090 Ti Local AI Butler — MarcusAurelius68 · 2026-09-03
- MTP vs MTP+Ngram on Qwen3.8 Flash: 10% Speed but 3x Token Usage — esw123 · 2026-09-03
- Perplexity Computer demos fully local operation on NVIDIA DGX Spark — chrmanning · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03