DFlash 2: Qwen3.8-27B hits 70 tok/s on Apple M5 Max
Lianhuiq · x · 2026-08-19
Inco AI released DFlash 2, an inference optimization technique. A demo shows Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro, which is 4.6× faster than autoregressive decoding.
Key Improvements:
- Over 20% more output per verification pass with 1% added latency.
- Predicts the entire token block in parallel instead of autoregressive drafting.
- Widely adopted in the ecosystem by NVIDIA, Google, and CoreWeave.
Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→
More from Infra
- NVIDIA monopoly hard to break? Paradigm shifts are key — 2C_ornot2C · 2026-08-19
- Anthropic's Multi-Level Monitoring for Astra Inference Revealed — AccBalanced · 2026-08-19
- Brex Report: Infrastructure Wins Over Apps in AI Hype — simonguozirui · 2026-08-19
- Alchemy Author: Not Just for Complex Projects — Simplest Way to Build Any Infra — samgoodwin89 · 2026-08-19
- Test: DeepSeek Harness achieves 99% cache hit rate with GLM and Kimi — sandyyevans · 2026-08-19
- 51WORLD Launches Embodied Data Infrastructure, Boosting Efficiency 10x — 量子位 · 2026-08-19