DFlash 2: Qwen3.8-27B hits 70 tok/s on MacBook with 4.6x speedup

songhan_mit · x · 2026-08-19

Inco AI released DFlash 2, an inference acceleration technology. Benchmarks show Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro, achieving up to 4.6× the speed of autoregressive decoding with identical output quality. The technology, seeded at Z Lab, gains speed by accepting one extra token per pass for free.

Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→

Original post →

More from Infra

Infra channel →