Preparing for M5 Ultra 512GB: which model quants fit and perform best locally

Ok_Warning2146 · reddit · 2026-08-31

A user preparing for a 512GB M5 Ultra (with TB5 enclosure + 4TB SSD) shares a planned local model lineup: GLM-5.3 MLX mixed 4/8bit (427.8GB), GLM-5.3-Flash (181.9GB), Qwen3.8-Flash-Next (106.2GB), DeepSeek-V4-Flash (165GB), and Kimi-K3 q8 (451GB). They ask M3 Ultra 512GB owners whether these quants are optimal under the RAM constraint, what max context GLM-5.3 and Kimi-K3 can reach, pp/tg numbers, and whether 8-bit variants are worth it.

Original post →

More from Infra

Infra channel →