Call for Dynamic Thinking Mode Support in llama.cpp

TheWaffleKingg · reddit · 2026-08-18

With the rise of long-thinking models like Qwen 2.5 32B, users are requesting llama.cpp to support dynamically adjusting the amount of "thinking" during inference. The ability to switch between high, medium, and low thinking modes on the fly would save significant time for easier tasks.

Original post →

More from Infra

Infra channel →