Qwen3.8-27B launches on Mac with 933 tok/s prefill on M5 Max

max_paperclips · x · 2026-08-15

Qwen3.8-27B is now supported on MLX-VLM and NativAI upon release. Initial benchmarks on M5 Max 128GB show a Prefill speed of 933 tok/s and a Decode speed of 33 tok/s. The model maintains coherence and reasonable performance up to a 256k context length.

Original post →

More from Infra

Infra channel →