Discussion on Qwen3.8-27B Reasoning Budget Configuration and llama.cpp Implementation
Thin_Pollution8843 · reddit · 2026-08-24
A user discusses issues with the reasoning budget feature in Qwen3.8-27B, noting that the model sometimes thinks too deeply and hits the 128k output token limit. The post compares the official Qwen thinkingbudget configuration with the llama.cpp --reasoning-budget parameter implementation (which includes a stop message) versus a pi plugin implementation, seeking a more generic solution.
More from coding & agent
- AI fakes memory: why it gets confidently wrong without forgetting — PrajwalTomar_ · 2026-08-24
- Observation suggests Codex continues running tasks long after weekly credits run out — gandamu_ml · 2026-08-24
- AI Eval Experts: Claude Excels at Top-Down Evals, But Bottom-Up Criteria Are All You — HamelHusain · 2026-08-24
- Google's AI research agents discover 66 novel biomarkers in automated biomedical study — imjustnewatai · 2026-08-24
- Using 6 Grok Agents to Build a Fully Automated Company Operation — elonmusk · 2026-08-24
- Parsewave Suggests LLM Evaluation Needs Real-World Environments, Not Just Q&A — trashnash007 · 2026-08-24