Qwen 3.8 scores high on WeirdML but uses many reasoning tokens

teortaxesTex · x · 2026-08-17

The author comments on Qwen's release of serious models, noting it's 8 months behind OpenAI. Cited data shows Qwen 3.8 (2.4T A95B) scores 75.2% on WeirdML, ranking as the second-best open model after Kimi-K3. The review notes it is solid but uses many reasoning tokens and writes very long code.

Original post →

More from Models

Models channel →