DFlash 2 released: up to 4.6× speedup for AI inference

igilitschenski · x · 2026-08-19

Inco AI announced DFlash 2, an inference stack innovation that delivers up to a 4.6× speedup over standard autoregressive decoding. Benchmarks show Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro. Originally seeded at Z Lab and upgraded at Inco AI, it unlocks performance gains for both frontier-scale models and smaller models on personal devices.

Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→

Original post →

More from Infra

Infra channel →