Qwen3.8-Flash slashes prices: flat 2.7 RMB output, cheapest cache hits yet
AlchainHust · x · 2026-08-27
Alibaba released Qwen3.8-Flash along with an open-weights next-gen architecture version. API pricing is a flat 0.8 RMB/M input tokens, 2.7 RMB/M output, 0.1 RMB cache hits — no peak/off-peak tiers, vs DeepSeek-V4-Flash's 4.5–9 RMB output. 100M input tokens cost 80 RMB vs $500 on Claude Opus 4.6.
Official benchmarks compare against Qwen3.8-27B, Qwen3.7-Plus, DeepSeek-V4-Flash and Claude Opus 4.6. The author also cites a V2EX post where a company moved to shift-based scheduling to dodge peak-hour token pricing.
Related event: Alibaba Releases Qwen3.8-Flash Multimodal MoE Model to Wide Acclaim(26 posts)→
More from Models
- ThursdAI Recap: GLM 5.3 and Qwen 27B Drop, OpenAI Pauses RL for Safety — altryne · 2026-08-28
- Qwen3.8-Flash-Next runs 162k context on dual 7900 XTX — BigYoSpeck · 2026-08-28
- Glitch Exposes Gemini's Internal Thoughts — Regular_Preference64 · 2026-08-28
- GLM-5.2 Introduces Monitors to Combat Reward Hacking in RL — burny_tech · 2026-08-28
- Anthropic Luna Max test shows generous limits, high speed — timpera · 2026-08-28
- User calls Grokbot 'terrible', cites missing tasks — krishnan · 2026-08-28