.15… · AGI Hunt

GLM-5.3-Flash Launches: 1M Context for $0.15 with Hybrid Attention

gharik · x · 2026-08-27

DeepInfra launched the GLM-5.3-Flash API, a natively multimodal model with 320B total parameters but only 18B active. It features a hybrid sparse and linear attention architecture to maintain accuracy on a 1M-token window while reducing compute costs. Priced at $0.15 per million input tokens, it supports tool calling and structured output.

Related event: Zhipu Open-Sources GLM-5.3-Flash: Frontier Performance at a Fraction of the Cost(39 posts)→

Original post →

More from Models

Models channel →