Qwen 3.8 Max Preview Benchmarks Leaked
bdsqlsz · x · 2026-07-19
The Qwen 3.8 Max preview reportedly shows strong performance in internal benchmarks:
- It first took a "candy test," leading the author to believe its math capabilities have surpassed GPT 5.5 high and Kimi K3.
- Cited claims indicate Qwen3.8-Max trails Fable 5 by only 12.6 points in coding; in the Cowork test featuring 400 real-world agent tasks, it has already beaten Opus 4.8 Max.
- In specific results, Qwen leads Opus 4.8 Max by 5 points in Cowork tasks, also outperforming Kimi K3, GLM-5.2, and the previous-gen Qwen3.7-Max.
- The original post stresses this is merely a preview, with some prepped improvements yet to be deployed, meaning the final release could be even stronger.
Related event: Leaked Alibaba Qwen 3.8 Max Shows Strong Benchmark and Coding Performance(6 posts)→
More from coding & agent
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11