GLM-5.3 FlashX spotted: 200 tok/s, 1M context, reportedly on 100k Chinese chips

gaganghotra_ · x · 2026-09-18

Third-party eval account LuminaBench reports that Zhipu has quietly rolled out GLM-5.3 FlashX (unconfirmed by the vendor). From Vercel-side signals it appears to be a much faster variant of GLM-5.3 Flash: up to 200 tok/s generation, 1M context window, and native multimodality. The account also claims it runs on 100,000 Chinese-made chips. Official details are pending.

Original post →

More from Models

Models channel →