GLM 5.3 Flash Inference Extremely Slow on Apple Silicon

CentrifugalMalaise · reddit · 2026-08-30

User tested the Unsloth quantized GLM 5.3 Flash model on an M2 Ultra Mac Pro, finding the time to first token intolerably slow at 3-4 minutes. In contrast, Qwen3.5 397B runs decently on the same hardware at around 23 tps. The user suspects inference engines have not yet optimized for GLM 5.3 Flash and seeks successful implementation examples on Apple Silicon.

Original post →

More from Infra

Infra channel →