GLM-5.3-Flash Q4 Hits 37.4 t/s at 300k Context on M3 Ultra via Custom Kernels
IngeniousIdiocy · reddit · 2026-09-09
A developer deep-tuned GLM-5.3-Flash (Q4) local inference on M3 Ultra, open-sourcing the work in the ds4 repo's glm53-m3ultra branch:
- Decode: fused dozens of small kernels into larger dispatches, raising memory-bandwidth utilization from 59% to 81%; 29→40 t/s short-context, 24→38 t/s at 62k.
- Long-context attention: replaced staged sort/merge of candidate lists with parallel scans, byte-identical outputs, end-to-end 300k context 21.6→37.4 t/s.
- Prefill: batching weight reuse and wider loads took 366→550 t/s (+50%); cold 62k processing dropped from 170s to 113s.
- Speculative drafter: windowed admission controller backs off when not paying; +4% on a 32-request agent session, +20–50% on SQL/JSON, 1% cost on prose.
- Quality: NLL essentially unchanged (0.300766 vs 0.300804) and 90/100 first-token matches on the 100-prompt QA set.
The optimizations are M3 Ultra-specific, based on detailed measurements of its two-die memory behavior, system cache, and Metal scheduling.
More from Infra
- 27B 1-bit model runs in the browser at 25-30 tok/s on a 6GB RTX 3060 laptop (WebGPU) — mentria-ai · 2026-09-09
- The expensive part of coding agents isn't the agents—it's the silent retries ($900 for one task) — mrtrly · 2026-09-09
- AI infra is unbundling: model, harness, inference and compute become four separate choices — _changxu · 2026-09-09
- Running a 27B Q4 Model at 131K Context on a Single RTX 3090: ~25.6 tok/s Measured — bjivanovich · 2026-09-09
- Polymarket puts 73% odds on a US state enacting a data center moratorium by 2026 — Polymarket · 2026-09-09
- Musk's empire goes all-in on US manufacturing: Terafab phase one tops $16.8B — XFreeze · 2026-09-09