Open-weight Qwen3.8 beats Claude Opus 5 on 4 of 5 enterprise personas just 15 days after launch
ryanshrout · x · 2026-09-16
On Signal65's PINNACLE enterprise benchmark, Claude Opus 5 topped the board at launch on Aug 31 — but 15 days later Alibaba's open-weight Qwen3.8-2.4T-A95B (medium reasoning effort) beats it on four of five personas and ties the fifth, making 15% fewer weighted errors on multi-step enterprise work. The best open model trailed by 1.4x at launch. Ryan Shrout: the frontier moved, but open weights closed the two-week gap — the pace any pause must contend with.
More from Models
- US Government Spotted Using Qwen Embedding Model for RAG Lookup — PsychologicalSoup251 · 2026-09-16
- ChatGPT $100-Tier User: GPT-6 Astra Burns 3x Usage Budget Yet Codes Worse Than GPT-5.6 Sol — F0xy1337 · 2026-09-16
- Microsoft reportedly to limit future AI models, prompting Clippy-meme jokes — matt_slotnick · 2026-09-16
- Gemini 3.8 Live now tryable live in Google AI Studio — _philschmid · 2026-09-16
- Gemini 3.8 Live launches: #1 voice model at $0.005/min with async background thinking — _philschmid · 2026-09-16
- Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking with 97-language real-time voice — GoogleAI · 2026-09-16