Qwen3.8-Max Preview Tested: Strong in Coding, Faces Scrutiny Over Thinking Time

Alibaba has released its new flagship base model, Qwen3.8-Max-Preview, reportedly packing 2.4 trillion parameters. The model is currently available for trial on the official website and the coding app Qoder. Its claimed robust performance has triggered a wave of intensive hands-on testing within the community.

Highlights: Coding and Workflows

The model has received considerable praise for coding and task execution. @johnseach noted it generated 1,500 lines of astrophysics code in one go with zero errors; @APPSO found its web generation speed exceptionally fast, capable of rapidly outputting complex pages like a Three.js 3D sports car. During workflow tests in Qoder, @HeyNayeem pointed out that the model can continuously advance tasks to completion with almost no human intervention. Additionally, @青稞AI integrated it into complex project features like shared wallets, and @Askmasr_mod also affirmed its outstanding creative writing capabilities.

Controversy: Thinking Time and Expectations

Despite its impressive capabilities, the preview version's flaws and polarized reception are quite apparent. @curiousily_ observed in tests that the model sometimes gets stuck in thought loops, and its frontend skills are not as advertised. @Askmasr_mod pointed out that its biggest issue is excessively long "thinking time," where even simple prompts can take a while. Regarding overall reception, @davidtsong noted a clear divide on X (Twitter), with some arguing it falls short of the "Fable level" expectations. @teortaxesTex candidly stated that Qwen preview versions are traditionally rough and serve more as research references, estimating its actual performance to be somewhere around GLM 5.2 or slightly better.

Testing Advice

Addressing the model's inconsistent performance, @terryyuezhuo (retweeting) suggested evaluating the API rather than the web UI, as the web chat is merely a free trial portal, whereas the API reflects its true capabilities. @vista8 also reminded users that the web interface is currently mainly suitable for testing text and simple code, though feedback from peers indicates the model is indeed getting stronger.

2026-07-19 ~ 2026-07-21 · 13 related posts