GLM-5.3-Flash Launches: 1M Context for $0.15 with Hybrid Attention
gharik · x · 2026-08-27
DeepInfra launched the GLM-5.3-Flash API, a natively multimodal model with 320B total parameters but only 18B active. It features a hybrid sparse and linear attention architecture to maintain accuracy on a 1M-token window while reducing compute costs. Priced at $0.15 per million input tokens, it supports tool calling and structured output.
More from Models
- AI enters adolescence: Small models beating large ones in specific domains — DhruvBatra_ · 2026-08-27
- TerminalBench Deep Dive: GPT-5.6 Leads, Many Models Drop in Rank — abeirami · 2026-08-27
- Gemini 3.5 Transcribe Live Beats GPT Live Transcribe: 5.8% WER at 0.25s First Partial — ArtificialAnlys · 2026-08-27
- Qwen3.8-Flash-Next Runs on Dual DGX Sparks — NVIDIAAI · 2026-08-27
- New leader emerges on LLM leaderboard, claiming the crown — jonathan_wilke · 2026-08-27
- Yutori's Batra: most of the web will never get agent APIs — pixels in, clicks out — DhruvBatra_ · 2026-08-27