Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x
Inco AI has released DFlash 2, an inference acceleration technology that delivers up to 4.6x speedup over standard autoregressive decoding. Tests show Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro.
2026-08-19 ~ 2026-08-19 · 3 related posts
- DFlash 2: Qwen3.8-27B hits 70 tok/s on Apple M5 Max — Lianhuiq · 2026-08-19
- DFlash 2 released: up to 4.6× speedup for AI inference — igilitschenski · 2026-08-19
1 near-duplicate retellings: songhan_mit