Leaked Alibaba Qwen 3.8 Max Shows Strong Benchmark and Coding Performance

Recently, a preview version of Alibaba's Qwen 3.8 Max model was leaked, sparking heated community discussions. The model allegedly has 2.4T parameters and is officially claimed to be the "strongest model except for Fable 5." Currently, it has demonstrated exceptionally strong performance in multiple benchmark and practical development tests.

Benchmarks and Capability Evaluation

Regarding benchmarks, the leaked KingBench 3 leaderboard shows Qwen 3.8 Max scoring 81.25, closely trailing the top-ranked Fable 5 (82.5) and surpassing several Claude models. @bdsqlsz pointed out that in the "candy test," its mathematical ability exceeded GPT 5.5 high and Kimi K3, with outstanding coding capabilities as well. In practical applications, @赛博禅心 tested the model via Claude Code on a development task involving multiple payment logics, confirming its robust coding ability and extremely fast response speed.

Reactions and Impact

Users on platforms like Reddit have been actively verifying the model's authenticity. Faced with the strong benchmark data, @orange stated that if Qwen 3.8 truly surpasses GPT-5.6, the technological gap between Chinese and US models could shorten to about 3 months, though he plans to make a final judgment after actually trying it. Furthermore, @Scobleizer reposted, highlighting a realistic challenge: while open-source frontier models are getting larger, benchmark scores do not equate to affordable inference costs, which is a practical hurdle for the model's future deployment.

2026-07-19 ~ 2026-07-20 · 6 related posts

Full story(17 episodes)→

Primary sources