Qwen 3.8 scores high on WeirdML but uses many reasoning tokens
teortaxesTex · x · 2026-08-17
The author comments on Qwen's release of serious models, noting it's 8 months behind OpenAI. Cited data shows Qwen 3.8 (2.4T A95B) scores 75.2% on WeirdML, ranking as the second-best open model after Kimi-K3. The review notes it is solid but uses many reasoning tokens and writes very long code.
More from Models
- Tencent's EVIE Model Tops ViDoRe Benchmarks, Cuts Vector Storage Costs by 32x — jacek2023 · 2026-08-17
- EU firms may use Chinese open models via "jurisdictional wrapper" — teortaxesTex · 2026-08-17
- empero-ai's distilled Qwen3.8-9B trends on Hugging Face — empero-ai · 2026-08-17
- Dev Reports Cache Misses on gpt5.6-sol During Slow Tool Calls — lucasmeijer · 2026-08-17
- Test shows Sol and GLM-5.3 catch critical code flaws; Claude misses them — morgymcg · 2026-08-17
- DeepSeek v4 Pro Review: Matches Flash on Most Evals, Raises Scaling Questions — teortaxesTex · 2026-08-17