320B Model Runs on Mac: OrcaSAQ Quantization Brings GLM-5.3 to Apple Silicon
alejandroll10 · x · 2026-08-27
OrcaRouter has natively brought the 320B-parameter GLM-5.3-Flash to Apple Silicon using a new method called OrcaSAQ (Orca Sensitivity-Aware Quantization). It supports 2/3/4/6-bit native MLX quantization without calibration and is architecture-aware. At 6-bit, it achieves 97.76% Top-1 agreement with FP8, delivering near-FP8 fidelity at a fraction of the memory.
More from Infra
- Prediction: Consumer Desktops Will Soon Run k3-Quality Models — AaronBergman18 · 2026-08-27
- Anthropic Locks in 460MW Compute for $45B, Revealing GPU Economics — zephyr_z9 · 2026-08-27
- Kioxia Plans $6.27B Investment for Third Fab in Iwate — zephyr_z9 · 2026-08-27
- DeepSeek-V4-Flash hits 51.5 tok/s on M3 Ultra — antirez · 2026-08-27
- Colibrì Engine Update: Runs 2.8T Param Models, Boosts Speed via Expert Caching — solyarisoftware · 2026-08-27
- Jensen Huang: Data Center Investment Payback Period Under One Year — zephyr_z9 · 2026-08-27