DFlash 2: Qwen3.8-27B hits 70 tok/s on Apple M5 Max

Lianhuiq · x · 2026-08-19

Inco AI released DFlash 2, an inference optimization technique. A demo shows Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro, which is 4.6× faster than autoregressive decoding.

Key Improvements:

Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→

Original post →

More from Infra

Infra channel →