ARC Prize finds Qwen3.8-27B's chat template injects different instructions per reasoning effort
GregKamradt · x · 2026-10-02
Testing Qwen3.8-27B via Baseten, the ARC Prize team discovered the model's own chat template injects different instructions per reasoning effort: low tells it to keep thinking brief, xhigh instructs careful thinking and checking assumptions, while medium adds neither — even though thinking stays enabled. This may explain medium's lower scores: the settings change how the model is instructed to approach problems, not simply its thinking budget. GregKamradt called it unexpected and credited Baseten for verifying the finding; the template is public on Hugging Face.
More from Models
- Rumor: Anthropic's Fable 5.5 to launch next Tuesday, called 'otherworldly' — imjustnewatai · 2026-10-02
- VAmoS Pro voice-agent benchmark: Grok leads tasks, GPT-Live fastest, Gemini most noise-robust — davlanade · 2026-10-02
- Opus 5.5 keeps saying "himbo" — a verbal quirk no previous Claude showed — repligate · 2026-10-02
- AI spend falls in latest Ramp AI Index as frontier price cuts bite; open source under 5% — PaulYacoubian · 2026-10-02
- Mystery model Fledge Alpha spotted before announcement, clues point to Thinking Machines Lab — cephaloform · 2026-10-02
- Ivo's contract agent hits 91% on Legal Agent Benchmark via River AI post-training — ibab · 2026-10-02