Preparing for M5 Ultra 512GB: which model quants fit and perform best locally
Ok_Warning2146 · reddit · 2026-08-31
A user preparing for a 512GB M5 Ultra (with TB5 enclosure + 4TB SSD) shares a planned local model lineup: GLM-5.3 MLX mixed 4/8bit (427.8GB), GLM-5.3-Flash (181.9GB), Qwen3.8-Flash-Next (106.2GB), DeepSeek-V4-Flash (165GB), and Kimi-K3 q8 (451GB). They ask M3 Ultra 512GB owners whether these quants are optimal under the RAM constraint, what max context GLM-5.3 and Kimi-K3 can reach, pp/tg numbers, and whether 8-bit variants are worth it.
More from Infra
- Jensen Huang mobilizes over $500B for AI infrastructure financing with major asset managers — RachelVT42 · 2026-08-31
- Devs compare Mujoco and Isaac: Single-purpose sim vs. full-stack features — KyleMorgenstein · 2026-08-31
- Nvidia bets big on physical AI with China as key customer amid ban paradox — pstAsiatech · 2026-08-31
- Huawei's Kirin 2026 processor with LogicFolding architecture coming this fall — pstAsiatech · 2026-08-31
- Huawei revenue up 9.55% as it pours 25% into R&D for self-reliance — pstAsiatech · 2026-08-31
- Ask: does generating every token re-execute the full parameter set in an LLM? — MarinatedPickachu · 2026-08-31