GLM 5.3 Flash Inference Extremely Slow on Apple Silicon
CentrifugalMalaise · reddit · 2026-08-30
User tested the Unsloth quantized GLM 5.3 Flash model on an M2 Ultra Mac Pro, finding the time to first token intolerably slow at 3-4 minutes. In contrast, Qwen3.5 397B runs decently on the same hardware at around 23 tps. The user suspects inference engines have not yet optimized for GLM 5.3 Flash and seeks successful implementation examples on Apple Silicon.
More from Infra
- OpenAI's Jalapeño Chip Beats Nvidia GB200 in Efficiency — Beth_Kindig · 2026-08-30
- Developer Dumps Qualcomm for Rockchip to Ship Edge AI Devices Faster — kscottz · 2026-08-30
- Empire of AI's datacenter water figure is off by ~4,500x: liters vs cubic meters — altryne · 2026-08-30
- Community fork enables MiniMax H3 on dual GPUs with live preview support — karma3u · 2026-08-30
- Own your harness, and if possible, own the model layer too — omarsar0 · 2026-08-30
- M4 Max Benchmarks: oMLX Wins at Long Context, Prefix Caching is Key — vitordeas · 2026-08-30