OpenRouter coding model share: GLM 5.3 Flash leads at 29.2%, DeepSeek V4.1 Flash at 25.6%
togethercompute · x · 2026-09-29
Together claims the #1 spot on OpenRouter token share across top open coding models:
- Z.ai GLM 5.3 Flash: 29.2%
- DeepSeek V4.1 Flash: 25.6%
- Kimi K3: 18.9%
The linked OpenRouter page adds details:
- DeepSeek V4.1 Flash: a 552B-parameter sparse MoE built on the new Causal Encoder-Decoder (CED) architecture, activating 8B params on input and 16B on output. Native image understanding trained jointly from pre-training. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation. Priced at $0.30/M input and $1.20/M output with 1.05M context; positioned as the cost-efficient tier, reportedly exceeding V4 Pro on speed and task completion. Suited for coding, terminal, and computer-use agents plus long-horizon tasks.
- GLM 5.3 Flash: natively multimodal, hybrid sparse + linear attention for efficient long-context, at $0.15/M input and $0.50/M output with 1.05M context.
- GLM 5.3: a large-scale reasoning model for complex software engineering and long-horizon agent tasks, with a 1M-token context window.
More from coding & agent
- Garry Tan Backs AI Browsers Like AsideAI: Agent Password Management Is the Key Layer — garrytan · 2026-09-29
- Meeting Prep Agent Separates App State in SQLite from Long-Term Memory — Datrika_Nandhini · 2026-09-29
- Chestnut unveils 18-DOF Aero Hand plus matching exoskeleton for zero-gap humanoid data capture — chris_j_paxton · 2026-09-29
- RubyLLM hits 1.1M downloads in a month, 13M total — kieranklaassen · 2026-09-29
- Incident Memory Agent Turns Postmortems into Verified Triage Context via Hindsight — nagakeerthan · 2026-09-29
- InstaCloud launches agent-native serverless cloud that lets Claude Code and Cursor provision infra via one command — testingcatalog · 2026-09-29