Nativ adds GLM-5.3 support: 505 tok/s on M3 Ultra
lllucas · x · 2026-08-27
Nativ announces day 0 support for GLM-5.3-Flash. On an M3 Ultra (512GB), the 320B MoE model achieves up to 505 tok/s prefill and 32 tok/s decode at 4-bit MLX, with peak memory under 380GB. Nativ is an open-source, SwiftUI-based macOS client for local AI.
More from Infra
- NVIDIA FLARE Cuts Federated VLM Training Traffic by 99% — dl_weekly · 2026-08-27
- Open Source AI Share Hits 62% on Vercel, Eclipsing Closed Source Models — gajesh · 2026-08-27
- Same Budget: 256GB Mac or Two DGX Sparks for 70B Inference? — Whyme-__- · 2026-08-27
- Edviro Builds World Model to Unify Data Center Operations — ycombinator · 2026-08-27
- Chinese Models Top US in Token Usage on OpenRouter; Efficiency Becomes Advantage — AccBalanced · 2026-08-27
- SandboxAQ Open-Sources Switch for Shared AI-Agent Workspaces — Codeblix_Ltd · 2026-08-27