New Qwen and GLM Open Models Benchmark Against Claude Opus
kimmonismus · x · 2026-08-26
A significant day for open-weight and local AI. Qwen3.8-Flash-Next (125B MoE, 6B active) beats Claude Opus 4.6 Max on 8 of 9 benchmarks. GLM-5.3-Flash (320B total, 18B active) scores close to Opus 4.8 on Terminal-Bench and leads on several agentic benchmarks. Both models are MIT-licensed, natively multimodal, and support 1M context. While they require serious workstations to run, the narrowing gap with frontier models serves as a wake-up call.
More from Models
- Anthropic reportedly releasing Fable 5.1 model soon — mark_k · 2026-08-27
- Observers doubt Simile/Aaru claims over lack of datasets and peer review — daveholtz · 2026-08-27
- Navigator n2 released: A frontier 27B computer-use model — DhruvBatra_ · 2026-08-27
- Paper reveals reasoning models' 'thinking' behaviors are often uncorrelated with correct answers — Jeande_d · 2026-08-27
- Alibaba Releases FP8 Quantized Qwen3.8-Flash-Next Model — Qwen · 2026-08-27
- Qwen 3.8 27b coding performance shocks community, rivaling GPT 5.5 on consumer hardware — GrokiniGPT · 2026-08-27