User Benchmarks Qwen3.8 27b on M5 Max: 8t/s (bf16), 17t/s (8bit)

julianharris · x · 2026-08-16

A user benchmarked the Qwen 3.8 27B model (bf16, 262k context) on an M5 Max with 128GB RAM. Using the oMLX + opencode stack, the speed achieved was 8t/s at bf16 precision and 17t/s at 8bit quantization. The test revealed significant CPU throttling, with expectations that 4bit quantization and speculative decoding will improve performance in the future.

Original post →

More from Infra

Infra channel →