GLM-5.3-Flash Q4 Hits 37.4 t/s at 300k Context on M3 Ultra via Custom Kernels

IngeniousIdiocy · reddit · 2026-09-09

A developer deep-tuned GLM-5.3-Flash (Q4) local inference on M3 Ultra, open-sourcing the work in the ds4 repo's glm53-m3ultra branch:

The optimizations are M3 Ultra-specific, based on detailed measurements of its two-die memory behavior, system cache, and Metal scheduling.

Original post →

More from Infra

Infra channel →