DFlash 2: Qwen3.8-27B hits 70 tok/s on MacBook with 4.6x speedup
songhan_mit · x · 2026-08-19
Inco AI released DFlash 2, an inference acceleration technology. Benchmarks show Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro, achieving up to 4.6× the speed of autoregressive decoding with identical output quality. The technology, seeded at Z Lab, gains speed by accepting one extra token per pass for free.
Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→
More from Infra
- DFlash 2 released: up to 4.6× speedup for AI inference — igilitschenski · 2026-08-19
- Will we run 30B+ parameter models fast on small GPUs in the future? — absurdother · 2026-08-19
- Periodic Labs trains trillion-parameter models on Miles framework, 3x throughput boost — hsu_byron · 2026-08-19
- Local LLM Speed Bottlenecks: RTX 4090 vs. 5090 Performance Analysis — Viktri1 · 2026-08-19
- LLM Inference Engineering: From KV Cache to vLLM and SGLang — techNmak · 2026-08-19
- NVIDIA H100 Concurrency Response of Plain Global Loads Analyzed — ssh4net · 2026-08-19