GLM 5.3 open weights arrive; DFlash 2 speculative decoding hits 4.4x FP8 throughput
gan_chuang · x · 2026-08-29
Zhipu has released GLM 5.3 with open weights. Third-party team incoai shipped Day 0 support: a DFlash 2 speculative decoding drafter, an NVFP4 quantized checkpoint, and a live endpoint on TokenRouter's GB300s — claiming up to 4.4x the throughput of native FP8 under autoregressive decoding. As zhijianliu notes, DFlash 2 ten days ago was the first piece; GLM 5.3 completes the stack.
Related event: GLM 5.3 open weights launch with Day-0 decoding and quantization support(2 posts)→
More from Infra
- Microduck BOM Analysis: Is a $10 Price Point Feasible? — NickPassig · 2026-08-29
- Dual 3090 Qwen3.8-27B: DFlash2 Speed Test and Full Config Guide — Old_Ad_6033 · 2026-08-29
- Essay: Data centers are next-generation ports — communities will build around the flow of intelligence — McDonaghMatthew · 2026-08-29
- Video explainer: how data centers actually drive electricity prices lower — aronchick · 2026-08-29
- Mac Studio M5 Ultra Dilemma: Dual 96GB Linked vs Single 256GB? — anonmt57 · 2026-08-29
- Polymarket prices 68% chance a US state enacts a data center moratorium in 2026 — Polymarket · 2026-08-29