GLM-5.3 FlashX spotted: 200 tok/s, 1M context, reportedly on 100k Chinese chips
gaganghotra_ · x · 2026-09-18
Third-party eval account LuminaBench reports that Zhipu has quietly rolled out GLM-5.3 FlashX (unconfirmed by the vendor). From Vercel-side signals it appears to be a much faster variant of GLM-5.3 Flash: up to 200 tok/s generation, 1M context window, and native multimodality. The account also claims it runs on 100,000 Chinese-made chips. Official details are pending.
More from Models
- 'Astra Is OpenAI's Smartest and Dumbest Model': Users Report Wildly Inconsistent Behavior — 0xkarasy · 2026-09-18
- Another Day, Another Fake Benchmark King: 'Astra-Level' Free Models Keep Flopping — bindureddy · 2026-09-18
- User burns through ChatGPT Pro weekly quota in 2 days running Codex — craigbalding · 2026-09-18
- Users speculate Gemini 4 training as Gemini 3.8 Flash slows on Antigravity — haider1 · 2026-09-18
- Dev swapped to Opus 5 after OpenAI credits ran out — 'this is not great' — lucasmeijer · 2026-09-18
- 234 Days Later: Opus-Class Intelligence Runs Locally on a Single 8GB RTX 3060 — JFPuget · 2026-09-18