Discussion on Qwen3.8-27B Reasoning Budget Configuration and llama.cpp Implementation

Thin_Pollution8843 · reddit · 2026-08-24

A user discusses issues with the reasoning budget feature in Qwen3.8-27B, noting that the model sometimes thinks too deeply and hits the 128k output token limit. The post compares the official Qwen thinkingbudget configuration with the llama.cpp --reasoning-budget parameter implementation (which includes a stop message) versus a pi plugin implementation, seeking a more generic solution.

Original post →

More from coding & agent

coding & agent channel →