Qwen4 training now: Alibaba teases Max/Flash/Plus variants, 5-10T params for Qwen4.5-5
ChrisGPT · x · 2026-09-22
Community reports from Alibaba's Apsara Conference clarify the Qwen roadmap and correct earlier mix-ups between Qwen4 and the Qwen4.5-5 plans:
- Alibaba's public August 26 architecture preview already includes qwen4exp / Qwen4ExpForConditionalGeneration, with hybrid linear/sparse attention, gated residual streams, N-gram lookup tables offloadable to system RAM, and Muon-based training
- Qwen4-Max, Qwen4-Flash, Qwen4-Plus and Qwen4-27B are coming; Qwen4 is currently training
- The 5-10 trillion parameter scaling target applies to future Qwen4.5 and Qwen5, not Qwen4
- Alibaba demoed recursive self-improvement on Qwen3.8-Max: 33 automated iterations over a month lifted its Artificial Analysis score from 40 to 45; the broader roadmap targets long-running complex tasks and progress toward ASI
Related event: Qwen4 in Training, Aims for 5-10 Trillion Parameters Next(2 posts)→
More from Models
- Grok 4.7 lands: near Opus 5 on AA-Briefcase at ~50% cost per task — NicoVerderosa · 2026-09-22
- Whittle distills on HF: 27B-A3B MoE quant claimed to run on 8GB VRAM laptops — depressedclassical · 2026-09-22
- Developer paying $500/month says Claude quota cuts mean he can't work a full 8-hour day anymore — TejasKumar_ · 2026-09-22
- Gemini mistakes "someone wants to kill me" for self-harm, spams suicide hotlines — WideImagination8644 · 2026-09-22
- User presses Claude Code lead on whether usage resets will actually be banked — Angaisb_ · 2026-09-22
- Anthropic's 'Constitution' Is Just RLAIF, Argues Researcher — Its Real Benefit Is Cutting Human Labeling — WillRinehart · 2026-09-22