GLM 5.3 open weights released with NVFP4 checkpoint achieving 4.4x throughput
songhan_mit · x · 2026-08-29
Zai has released the GLM 5.3 open-weight model. The IncoAI team, in collaboration, announced Day 0 support featuring several technical optimizations:
- DFlash 2 Drafter: Speculative decoding implementation.
- NVFP4 Checkpoint: Utilizing NVFP4 format for efficiency.
- Live Endpoint: Powered by TokenRouter's GB300s.
Benchmarks indicate up to 4.4× throughput improvement over native FP8 with autoregressive decoding.
Related event: GLM 5.3 open weights launch with Day-0 decoding and quantization support(2 posts)→
More from Infra
- Microduck BOM Analysis: Is a $10 Price Point Feasible? — NickPassig · 2026-08-29
- Dual 3090 Qwen3.8-27B: DFlash2 Speed Test and Full Config Guide — Old_Ad_6033 · 2026-08-29
- Essay: Data centers are next-generation ports — communities will build around the flow of intelligence — McDonaghMatthew · 2026-08-29
- Video explainer: how data centers actually drive electricity prices lower — aronchick · 2026-08-29
- Mac Studio M5 Ultra Dilemma: Dual 96GB Linked vs Single 256GB? — anonmt57 · 2026-08-29
- Polymarket prices 68% chance a US state enacts a data center moratorium in 2026 — Polymarket · 2026-08-29