Testing Qwen 3.8 27B minimal thinking in a specific harness
rosie254 · reddit · 2026-08-15
A user observed that Qwen 3.8 27B barely performed reasoning tasks in a specific harness. The issue is attributed to potential system prompt design and tool overload (8k tokens), tested at high reasoning effort.
More from Models
- Grok 4.6 demonstrates ability to generate interactive Moon city experience — techartist_ · 2026-08-16
- Qwen3.8-27B-AEON-PURE scores perfect on all God Mode Tier tests — StephanSturges · 2026-08-16
- DeepSeek Models 0731 and 0813 Overfitting Differences Spark Technical Debate — teortaxesTex · 2026-08-16
- User tests Muse Glimmer 30B vs. Qwen 3.8 27B — MacaroonDancer · 2026-08-16
- Anthropic refuses to fill forms while Grok offers to place orders — pswider · 2026-08-16
- Meta open-sources Muse Glimmer but keeps powerful Muse Spark behind API — HaktanSuren · 2026-08-16