FULL STORY

Qwen3.8-Max: From Leak to Open Weights

Alibaba's 2.4T flagship model, Qwen3.8-Max, progressed from early leaks and preview iterations to its official release, topping leaderboards before its weights were opened to the public.

2026-07-19 ~ 2026-08-09 · 17 episodes · 143 posts

Episode 1 · Alibaba Announces 2.4T Open-Weight Model Qwen3.8 (2026-07-19, 22 posts)

On July 19, Alibaba's Qwen team officially previewed Qwen3.8, which will be released as an open-weight model. With approximately 2.4T parameters, it is the largest model Qwen has ever trained. The company claims its capabilities rival frontier models and rank just below Fable 5 in internal evaluations, with the model still evolving daily. This announcement has drawn significant attention, with multiple authors agreeing that it marks Alibaba's aggressive return to the open-source AI race with a massive parameter scale.

Key Details and Benchmarking Context

According to the official preview, Qwen3.8 has about 2.4T parameters. @teortaxesTex noted that Qwen 3.8-Max shows significant improvements in coding and cowork capabilities over Qwen 3.7-Max, and that 2.4T is the largest scale among publicly known models in China. They added that Qwen3.7-Max was previously the highest-ranking Chinese model on the ECI leaderboard, predicting Qwen 3.8 Preview could reach an ECI score of 156. @aigclink provided benchmarking context: tests were based on 400 real tasks from Alibaba's agentic products using isolated ECS and multiple cross-validations. @ugcfast pointed out that the model is multimodal, and Alibaba claimed its pricing is about one-tenth of Fable 5's. However, in a July 20 post, @bindureddy cited the scale as "about 3T"—near Sol/Opus level—and predicted the open-source community would find more efficient deployment methods within 12–15 weeks, though short-term running costs would remain high. This is the only differing parameter figure in this cluster.

Preview Version and Access Channels

Qwen3.8-Max-Preview is already available on Alibaba Token Plan, Qoder, and QoderWork. @zephyrz9 mentioned that the preview is open to Qwen Token Plan subscribers in mainland China. @aigclink also shared practical results of using Cline with the Qwen3.8 preview to build an "AI employee" project: the backend and frontend both succeeded on the first try with only minor modifications.

Background and Impact

@智东西 contextualized Qwen3.8 within the recent wave of dense updates from Chinese LLMs. Authors like @measureplan and @aigclink unanimously view this as a definitive signal of Alibaba re-entering direct open-source competition.

2 more related posts →

Episode 2 · Leaked Alibaba Qwen 3.8 Max Shows Strong Benchmark and Coding Performance (2026-07-19, 6 posts)

Recently, a preview version of Alibaba's Qwen 3.8 Max model was leaked, sparking heated community discussions. The model allegedly has 2.4T parameters and is officially claimed to be the "strongest model except for Fable 5." Currently, it has demonstrated exceptionally strong performance in multiple benchmark and practical development tests.

Benchmarks and Capability Evaluation

Regarding benchmarks, the leaked KingBench 3 leaderboard shows Qwen 3.8 Max scoring 81.25, closely trailing the top-ranked Fable 5 (82.5) and surpassing several Claude models. @bdsqlsz pointed out that in the "candy test," its mathematical ability exceeded GPT 5.5 high and Kimi K3, with outstanding coding capabilities as well. In practical applications, @赛博禅心 tested the model via Claude Code on a development task involving multiple payment logics, confirming its robust coding ability and extremely fast response speed.

Reactions and Impact

Users on platforms like Reddit have been actively verifying the model's authenticity. Faced with the strong benchmark data, @orange stated that if Qwen 3.8 truly surpasses GPT-5.6, the technological gap between Chinese and US models could shorten to about 3 months, though he plans to make a final judgment after actually trying it. Furthermore, @Scobleizer reposted, highlighting a realistic challenge: while open-source frontier models are getting larger, benchmark scores do not equate to affordable inference costs, which is a practical hurdle for the model's future deployment.

Episode 3 · Qwen3.8-Max-Preview Rolls Out Across Web, PC and iOS Preview (2026-07-19, 2 posts)

Alibaba’s Qwen3.8-Max-Preview has gone live first on Qwen’s PC and web clients, where users can enable it from the model list. Screenshots later showed the new option appearing in the iOS app as well, alongside Qwen3.7-Plus, indicating the preview is expanding across platforms.

Episode 4 · Qwen3.8-Max Preview Tested: Strong Coding but Slow Thinking (2026-07-19, 13 posts)

Alibaba released Qwen3.8-Max-Preview, a new flagship base model reportedly featuring 2.4 trillion parameters. Open for testing on the official website and the coding app Qoder, it has triggered extensive hands-on tests within the community, though feedback is notably polarized.

Highlights: Coding and Workflows

The model received praise for programming and task execution. @johnseach claimed it wrote 1,500 lines of astrophysics code in one go with zero errors, while @APPSO found its webpage generation extremely fast, capable of producing complex Three.js 3D pages. In Qoder workflow tests, @HeyNayeem noted the model continuously drives tasks to completion with minimal human intervention. Additionally, @青稞AI integrated it to develop complex features like shared wallets, and @Askmasrmod praised its outstanding creative writing capabilities.

Controversies: Slow Thinking and Unmet Expectations

Despite its impressive capabilities, the preview version has clear flaws and divided口碑. @curiousily found the model sometimes gets stuck in a thinking loop, with front-end skills falling short of promotional claims. @Askmasrmod pointed out that its biggest issue is excessively long thinking times, even for simple prompts. Overall, @davidtsong observed sharply divided feedback on X, with some arguing it fails to meet the expected "Fable level." @teortaxesTex candidly stated that Qwen preview versions are typically rough, serving more as scientific references, and estimated its true performance around the GLM 5.2 tier or slightly better.

Testing Suggestions

Addressing the model's instability, @terryyuezhuo (repost) suggested evaluating the API rather than the web UI, as the web chat is merely a free trial portal with inconsistent output quality. @vista8 also reminded users that the web interface is currently best suited for text and simple code tests, though friends reported the model is indeed getting stronger.

Episode 5 · Alibaba's Qwen3.8-Max-Preview iterates daily with improved frontend capabilities (2026-07-20, 5 posts)

From July 20 to 21, Alibaba's Qwen team continuously updated the Qwen3.8-Max-Preview model on a daily basis. The team noted that user feedback has exceeded expectations, with broad improvements in overall capabilities. This signals an acceleration in the practical application of large models for specific development scenarios.

Key Details and Future Plans

According to the Qwen team and multiple developers, since July 19, the core improvements of Qwen3.8-Max-Preview have been focused on the Web frontend and WebDev domains. The team is actively inviting users to test and provide feedback for further optimization. Additionally, the team plans to release a more mature formal version and open API access for developers in the future.

Episode 6 · Alibaba Announces Open-Weight Qwen3.8 and Multiple New Updates (2026-07-23, 2 posts)

Alibaba's Tongyi Lab announced a series of AI updates this week, headlined by the 2.4T parameter Qwen 3.8 model which will be released with open weights. Additional updates include new Qwen image capabilities, 16-language TTS, and Zvec.

Episode 7 · Alibaba releases Qwen3.8-Max, open-sources weights next week (2026-08-03, 44 posts)

Alibaba's Qwen team officially released the flagship model Qwen3.8-Max and announced that weights for Qwen3.8-Max and Qwen3.8-27B will be open-sourced next week. The official positioning is as the strongest version to date, focusing on autonomous coding, collaborative work, and long-horizon tasks; the model has been integrated into Command Code Go and OpenRouter. For developers, the release is notable not only for performance but also for the simultaneous advancement of flagship capabilities and open-source cadence.

Confirmed

  • Qwen3.8-Max uses a MoE architecture with 2.4T total parameters and 95B active parameters, supporting a 1M token context; official and related posts describe it as a natively multimodal model.
  • The official claims it sets a new benchmark in coding and cowork, suitable for long-horizon coding, research, and multimodal agent tasks, and claims the model can autonomously develop from an empty folder for over 10 days without human intervention.
  • The open-source plan is clear: weights for Qwen3.8-Max and Qwen3.8-27B will be released next week; Qwen3.8-Max has already been launched on Command Code Go and OpenRouter.
  • For benchmarks and evaluations, the official promotes its general text capabilities with Text Arena results; forwarded posts summarize that it ranks 4th in Frontend Code Arena with a score of 1668. @ArtificialAnlys gives an agentic task index of 53, considering it close to Claude Sonnet 5; @davidthesong believes it is close to Kimi K3 and DeepSeek V4 Flash on many benchmarks but stronger on coding and software tasks.

Unconfirmed

  • The official has not yet published the complete experimental setup, success rate, and reproducibility conditions for the "autonomous development over 10 days" test.
  • No open-source license information is available in the current materials.
  • Some forwarded posts mention API pricing of $2 per million input tokens and $6 per million output tokens, as well as a "cost reduction of 80%," but there is no corresponding official standalone post in this batch to cross-verify.

Why it matters

  • Alibaba is not just releasing a larger model; it is bundling the flagship release, platform integrations, and next week's open-source announcement together, clearly courting developers and deployment ecosystems.
  • If its coding, agent, and long-context capabilities are indeed close to current top closed-source models, the combination of Qwen3.8-Max and 27B could directly impact choices in programming assistants, enterprise applications, and open-source deployment markets. @op7418 also noted that this feels like a "reset" style official release, indicating Qwen's product line is entering a new round of flagship competition.

24 more related posts →

Episode 8 · Qwen3.8-Max Ranks Top Tier Across Benchmarks, Open-Source Narrows Gap (2026-08-03, 7 posts)

Alibaba's Qwen3.8-Max has shown impressive performance in recent benchmarks, ranking in the top tier across vision, text, and code arenas, proving open-source models can compete with frontier closed-source models.

Confirmed

  • Vision: Scored 1305 in the Vision Arena, ranking 2nd overall, just 13 points behind Anthropic's Claude Fable 5 (High).
  • Text: Scored 1496 in the LMSYS Text Arena, ranking 5th overall. It topped the healthcare category and ranked high in math/science and software services.
  • Code: Scored 1668 in both Frontend Code Arena and WebDev, ranking 4th. Top three are Claude Opus 5 Max (1705), Kimi K3 Max (1676), and Claude Opus 5 High.
  • Overall: According to data relayed by @Scobleizer, Qwen3.8-Max scored 93.0 on PaperBench, beating GPT-5.6 Sol (90), showing differentiation without dominating all models.

Why it matters

  • Benchmark results across multiple domains indicate Qwen3.8-Max's strong competitiveness in multimodal vision, text, healthcare, and frontend code, marking a substantial breakthrough for open-source models in comprehensiveness and frontier capability.

Episode 9 · Rumor: Alibaba's Qwen3.8-Max Outperforms Fable 5 and Set to Open Source (2026-08-03, 2 posts)

Rumors suggest Alibaba will open-source the 2.4T parameter Qwen3.8-Max model this week. It reportedly scores 86.6 on TerminalBench, outperforming Fable 5 and rivaling Claude 3.5 Sonnet at a very low cost.

Episode 10 · Qwen3.8-Max Initial Tests Show Performance on Par with DeepSeek (2026-08-03, 2 posts)

Initial stress tests of Alibaba's new Qwen3.8-Max model indicate that its performance reaches DeepSeek's level without noticeable degradation, showing strong competitive potential.

Episode 11 · Alibaba's Qwen3.8-Max: Open-Source Model Nears Closed-Source Frontier (2026-08-03, 13 posts)

Alibaba released its latest flagship open-source model Qwen3.8-Max and a 27B distilled version. The model has 2.4T total parameters and 95B active parameters, natively supporting text, image, video, and audio modalities. Multiple developer tests show robust performance in agentic coding and multimodal tasks, positioning it as a highly cost-effective frontier open-source model.

Confirmed

  • Model specs and architecture: @CodeByPoonam confirmed Qwen3.8-Max has 2.4T parameters and 95B active parameters, making it the most powerful Qwen model, natively multimodal, excelling in coding, work, and long-horizon tasks.
  • Coding and agent capabilities: Tests by @bindureddy, @QuixiAI, and @teortaxesTex show Qwen3.8 performs well in agentic coding, slightly behind K3 but at half the price. @intellectronica notes its excellent tool-calling, a strong K3 alternative. @葬AI's deep review rates its coding and agent abilities as first-class, even slightly surpassing K3 in some dimensions. @CodeByPoonam shared building an invoice generator in 10 minutes at zero cost.
  • Multimodal performance: @teortaxesTex suggests Qwen3.8 Max may achieve SOTA in image recognition and labeling, with strong sample efficiency, perfectly distillable to 27B. @WorldofAI's test validates its frontier performance in frontend code generation like Three.js and multimodal tasks.
  • Narrowing open-source gap: @dairai forwarded feedback that running the model on Hermes Agent forces a rethink of the gap between open and closed frontier models; open models now approach closed ones in agent tasks.

Unconfirmed

  • Long-horizon tasks and productization: @bindureddy notes MoE architecture in open models excels at benchmarks but limited in long-horizon tasks due to fewer active parameters. @葬AI mentions despite excellent real-browser task performance, it lags K3 in one-shot webpage generation and frontend productization.
  • Absolute comparison with K3: @ramoscasals forwarded Ethan Mollick's shader test showing Qwen 3.8 Max robust but not yet at Kimi K3's level. @QuixiAI's reply suggests its actual performance vs DeepSeek needs detailed comparison.

Why it matters

Qwen3.8-Max's release marks open-source models maintaining high cost-effectiveness (half the price of competitors) while matching or approaching top closed models in core agentic coding and multimodal capabilities. As @orange noted via his AI companion Cola, the model's efficiency directly empowers downstream applications, offering developers a cost-effective frontier alternative.

Episode 12 · Kimi K3 and Qwen3.8 Max Evaluations Approach Top Closed-Source Models (2026-08-05, 2 posts)

Recent evaluations reveal that open-source models like Kimi K3 and the upcoming Qwen3.8 Max are approaching the performance of top closed-source models, with Qwen offering the lowest cost for over 80% of tasks.

Episode 13 · Rumored Alibaba Qwen3.8-Max to Feature 2.4T Parameters with 95B Active (2026-08-05, 5 posts)

Alibaba is rumored to soon open-source its latest flagship model, Qwen3.8-Max (Qwen3.8-2.4T-A95B), which boasts a total parameter count of 2.4T but only activates 95B parameters during inference. The model has already appeared on the ModelScope platform, and its highly efficient sparse MoE architecture and native multimodal capabilities have attracted significant community attention.

Confirmed

  • Official Specs: The Qwen team confirmed in a recent X platform AMA that version 3.8 has 2.4T total parameters and 95B active parameters. Additionally, a 27B model with entirely new capabilities is set to be released, rather than just a fine-tuned version of the old model.

Unconfirmed

  • Release Date: Leaks suggest the model will be open-sourced next Wednesday, but an exact date has not been officially announced.
  • Core Capabilities: According to leaks, the model supports a 1 million token context window and features native vision and text multimodal reasoning. Designed for agentic tasks, it focuses on coding, deep research, and long-horizon tasks, though these details await official verification.

Why it matters

  • Inference Efficiency: Activating only 95B parameters out of a massive 2.4T total suggests a highly efficient sparse MoE architecture. If true, this would drastically optimize inference costs and position the model among the world's most powerful frontier models.

Episode 14 · Alibaba's Qwen3.8-Max tops agentic benchmark, open-sources next week (2026-08-06, 11 posts)

Alibaba's Qwen team released its flagship model Qwen3.8-Max on August 6. Based on a Sparse MoE architecture with 2.4T total parameters and 95B active parameters, it tops the Artificial Analysis Agentic Index and ranks fifth overall in the Intelligence Index with a score of 56. The model is available via Alibaba Cloud API and the enterprise platform QwenWork, with open-source weights planned for next week.

Confirmed

  • Model scale and architecture: Sparse MoE, 2.4T total parameters, 95B active, balancing performance and efficiency (m2, m4).
  • Benchmark results: Ranked #1 on Artificial Analysis Agentic Index; #5 overall with score 56, up 10 points from previous generation, tying Claude Opus 4.8 (m1, m2, m3).
  • Services and ecosystem: Available via Alibaba Cloud Model Studio API and integrated into QwenWork; services open to global users (m4, m7).
  • Open-source plan: Weights to be released on Hugging Face next week (m7).

Unconfirmed

  • Context length: m8 mentions 1M token context, not officially confirmed.
  • Medical RL training: m6 mentions upcoming medical RL training, but it's a repost and unverified.
  • Autonomous CLI tool building: m10 claims the model built a CLI tool autonomously over 16 days, but it's a personal experience, not officially confirmed.

Why it matters

Qwen3.8-Max's leading performance on agentic tasks marks a breakthrough for domestic models in agentic capabilities. Its efficient active parameter design balances performance and inference efficiency, and with the upcoming open-source release, it holds significant practical value for AI application deployment and the open-source community.

Episode 15 · Testing Alibaba's Qwen3.8-Max: Stellar Long-Context Agent Capabilities (2026-08-06, 3 posts)

Alibaba's newly released Qwen3.8-Max model has shown stellar performance in real-world tests. The model ranks in the first tier for front-end capabilities and demonstrated exceptional long-context agentic skills by autonomously developing a CLI tool over 16 days without any human intervention.

Episode 16 · Alibaba's Qwen 3.8 Models Rumored for Next Week Release (2026-08-07, 2 posts)

Alibaba's Qwen 3.8 models are rumored to launch next week, debuting with a 2.4T parameter version followed by a 27B model. The 27B version reportedly offers GLM-5.2-level intelligence and can run locally on MacBooks.

Episode 17 · Rumors Claim Alibaba Released Qwen3.8-Max (2026-08-09, 2 posts)

Rumors suggest Alibaba released its flagship Qwen3.8-Max model. It features a sparse Mixture-of-Experts architecture with 2.4 trillion total parameters, activating 95 billion per token, and supports native multimodality with a one-million token context window.