GLM-5.3 hits 310 tok/s on Databricks inference, claimed strongest OSS coding model
jeremyphoward · x · 2026-09-03
A Databricks-affiliated post reports GLM-5.3 running at 310 tok/s on Databricks inference, ranking #1 in both speed and latency. On their internal coding benchmark, GLM-5.3 is the strongest open-source coding model, competitive with Fable 5 and Opus 4.8. Jeremy Howard replied asking how to use it.
Note: these are vendor-side claims based on an internal benchmark, not independently verified.
More from Infra
- Perplexity open-sources its Mac inference server optimized for Qwen 3.6 on Apple Silicon — Specter_Origin · 2026-09-03
- GLM-5.3-Flash hits 1,005 tok/s locally on dual RTX PRO 6000 Blackwell cards — BanghuaZ · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03