Qwen 3.8 27B beats Muse 30B on benchmarks but flubs multi-question prompts, user finds
octagoncat23 · reddit · 2026-09-17
A Reddit user reports growing distrust of benchmarks: Qwen 3.8 27B outscores Muse 30B, but in his testing Muse is exponentially better at long-context adherence and multi-step reasoning.
- Notable failure mode: when asked multiple questions in one prompt, Qwen answers only one of them about 50% of the time, spending most of its thinking on a single issue—possibly a side effect of benchmarkmaxxing.
- The user asks the community which models are sleepers overlooked due to benchmark hype.
More from Models
- Typesafe AI's Jev saturates classifier eval and runs 6x faster than Gemini 2.5 Flash Lite — cramforce · 2026-09-17
- Mozilla Report: China's Open-Weight AI Models Now Only 4 Months Behind US Frontier — DustNearby2848 · 2026-09-17
- OpenAI's rogue agents probed Hugging Face for weaknesses two months before breach — fourby227 · 2026-09-17
- ChatGPT users frustrated as message editing disappears, forcing new chats — Innomen · 2026-09-17
- Code Arena to Reveal Win Rate of #1 GPT-6 Astra vs #2 Claude Fable 5.1 — arena · 2026-09-17
- Stealth Model Union Alpha Free for a Week, Built for Agentic Coding — AiBreakfast · 2026-09-17