Zhipu releases GLM-5.3-Flash: 1M context on domestic chips
智东西 · wechat · 2026-08-27
Zhipu AI has officially claimed the viral "Niu Lai" model as GLM-5.3-Flash (320B-A18B), a natively multimodal model supporting image, video, file, and text input.
Tech Architecture & Compute:
- Hybrid Attention: Uses a mix of sparse and linear attention to significantly reduce long-context compute costs. Introduces IndexPool to compress cache vectors.
- Domestic Chip Deployment: The model is powered by 100,000 domestic chips, achieving 1M context inference and massive global concurrency, marking the first large-scale overseas deployment of domestic chips.
Performance & Pricing:
- Performance: Ties with Claude Opus 4.8 on the AA Intelligence Index; coding capability matches Opus 4.8 in internal evals.
- Pricing: Input at 0.8 RMB/1M tokens, output at 2.8 RMB/1M tokens. Extreme price-performance ratio, roughly 1/10th of GLM-5.3 and 1/40th of Opus 4.8.
Benchmarking:
- Multimodal: Accurately reads exams, clones websites from screenshots, and edits 48-min speeches into highlight reels.
- Coding & Gaming: Developed a 3D diving game from scratch, replicated 2D MOBA mechanics (e.g., Honor of Kings) from video, and designed a skin plugin for an open-source project.
- Productivity: Summarized a 58-page thesis in 5.5 mins, generated analytical PPTs from charts, and created a 36-page job-seeking report.
More from Models
- NVIDIA Provides Day-0 Support for Alibaba's Qwen3.8-Flash-Next with NeMo — Alibaba_Qwen · 2026-08-27
- OpenRouter leaderboard: Real token consumption data outweighs media hype — sujingshen · 2026-08-27
- OpenAI's token efficiency may stem from training budget awareness — teortaxesTex · 2026-08-27
- Qwen 3.8 UX improvement: Simple operations no longer cause token redundancy anxiety — infieldmitt · 2026-08-27
- GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX — teortaxesTex · 2026-08-27
- Rumor: Her-like ChatGPT Voice Mode with GPT-6 Coming This Year — flowersslop · 2026-08-27