Qwen 3.8 Max Preview Benchmarks Leaked
bdsqlsz · x · 2026-07-19
The Qwen 3.8 Max preview reportedly shows strong performance in internal benchmarks: - It first took a "candy test," leading the author to believe its math capabilities have surpassed GPT 5.5 high and Kimi K3. - Cited claims indicate Qwen3.8-Max trails Fable 5 by only 12.6 points in coding; in the Cowork test featuring 400 real-world agent tasks, it has already beaten Opus 4.8 Max. - In specific results, Qwen leads Opus 4.8 Max by 5 points in Cowork tasks, also outperforming Kimi K3, GLM-5.2, and the previous-gen Qwen3.7-Max. - The original post stresses this is merely a preview, with some prepped improvements yet to be deployed, meaning the final release could be even stronger.
Related event: Leaked Alibaba Qwen 3.8 Max Shows Strong Benchmark and Coding Performance(6 posts)→
More from coding & agent
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21
- Axcess adds an MCP connector for WCAG accessibility checks that scanners miss — modelcontextprotocol · 2026-07-21
- X post asks whether Cursor Composer, built on Kimi models, would also be banned — max_paperclips · 2026-07-21
- A developer’s Codex usage is draining pooled enterprise credits at a small company — Distinct_Relation_62 · 2026-07-21
- Qwen Code ships cua-driver-rs 0.7.3 with relative coordinates and MCP filtering — github-actions[bot] · 2026-07-21
- Matt Pocock says every new codebase turns legacy within days — mattpocockuk · 2026-07-21