GLM-5.3-Flash hits 270 tok/s: 10% higher quality than 5.2 at one-tenth the cost

Yuchenj_UW · x · 2026-08-28

A Databricks inference team member reports GLM-5.3-Flash serving at 270 tok/s. On the team's OfficeQA Pro v2 benchmark, GLM-5.3-Flash delivers 10% higher quality than GLM-5.2 at roughly 1/10 the cost — fast, cheap and good at the same time.

Original post →

More from Infra

Infra channel →