320B Model Runs on Mac: OrcaSAQ Quantization Brings GLM-5.3 to Apple Silicon

alejandroll10 · x · 2026-08-27

OrcaRouter has natively brought the 320B-parameter GLM-5.3-Flash to Apple Silicon using a new method called OrcaSAQ (Orca Sensitivity-Aware Quantization). It supports 2/3/4/6-bit native MLX quantization without calibration and is architecture-aware. At 6-bit, it achieves 97.76% Top-1 agreement with FP8, delivering near-FP8 fidelity at a fraction of the memory.

Original post →

More from Infra

Infra channel →