Qwen 3.8 27B beats Muse 30B on benchmarks but flubs multi-question prompts, user finds

octagoncat23 · reddit · 2026-09-17

A Reddit user reports growing distrust of benchmarks: Qwen 3.8 27B outscores Muse 30B, but in his testing Muse is exponentially better at long-context adherence and multi-step reasoning.

Original post →

More from Models

Models channel →