User Benchmarks Qwen3.8 27b on M5 Max: 8t/s (bf16), 17t/s (8bit)
julianharris · x · 2026-08-16
A user benchmarked the Qwen 3.8 27B model (bf16, 262k context) on an M5 Max with 128GB RAM. Using the oMLX + opencode stack, the speed achieved was 8t/s at bf16 precision and 17t/s at 8bit quantization. The test revealed significant CPU throttling, with expectations that 4bit quantization and speculative decoding will improve performance in the future.
More from Infra
- Colibri: Run 2.8T-parameter MoE models on a 25GB laptop with pure C inference engine — alex_verem · 2026-08-16
- Local AI doesn't need to replace frontier cloud models; hybrid is the destination — ingliguori · 2026-08-16
- Polymarket: 69% Chance a US State Enacts Data Center Moratorium by End of 2026 — Polymarket · 2026-08-16
- New AI Agents Costing More Than Planned: The Silent Budget Killer — rvp · 2026-08-16
- Missouri farmer offers land for data center after county halts 500-acre project — Polymarket · 2026-08-16
- ArchAgent v2 achieves multi-level data prefetching via evolutionary search — dair_ai · 2026-08-16