M3 Ultra local benchmark: GLM5.2 and Qwen speeds limited
natesiggard · x · 2026-08-26
Testing GLM5.2 and Qwen 3.8 27B on an M3 Ultra with 512GB RAM, the user found GLM5.2 runs at 16 t/s (fine for background tasks) and Qwen 3.8 27B is capped under 30 t/s. Even if M5 doubles this speed, it won't match SOTA fast tokens for daily driving.
Related event: M3 Ultra Local LLM Tests Show Coding Still Out of Reach(2 posts)→
More from Infra
- Together AI: Fine-tune and deploy Qwen3.8 27B on dedicated infra — togethercompute · 2026-08-26
- Software folks don't get chip making: AlphaChip is just the tip of the iceberg — saranormous · 2026-08-26
- Nvidia BlueField 4 CPU Architecture Enables Tenant Network Isolation — bookwormengr · 2026-08-26
- NVIDIA Vera Rubin NVL72 racks enter production with 100% automated assembly — nvidia · 2026-08-26
- Hardware Prices Won't Recover: Choosing LLM Inference Hardware — nemuro87 · 2026-08-26
- For the price of one DGX Spark, you can build 2x 3090s with 128GB RAM — QuixiAI · 2026-08-26