DFlash 2 released: up to 4.6× speedup for AI inference
igilitschenski · x · 2026-08-19
Inco AI announced DFlash 2, an inference stack innovation that delivers up to a 4.6× speedup over standard autoregressive decoding. Benchmarks show Qwen3.8-27B running at 70 tok/s on an M5 Max MacBook Pro. Originally seeded at Z Lab and upgraded at Inco AI, it unlocks performance gains for both frontier-scale models and smaller models on personal devices.
Related event: Inco AI Launches DFlash 2, Boosting On-Device LLM Inference Up to 4.6x(3 posts)→
More from Infra
- AMD posts async RL walkthrough on MI355X and benchmarks vs B300 — AnushElangovan · 2026-08-19
- Andrej Karpathy releases llm.c: Train LLMs in raw C/CUDA — goyalshaliniuk · 2026-08-19
- Anthropic's $50B Buildout Shows Financing Is Not the Short-Term Compute Bottleneck — FinanceYF5 · 2026-08-19
- Epoch AI: funding won't bottleneck frontier compute; model to scale past 20GW — FinanceYF5 · 2026-08-19
- Five project companies issued $15.18B in debt for 1.43GW of data centers — FinanceYF5 · 2026-08-19
- Anthropic leveraged under $9B revenue into nearly $50B AI infrastructure — FinanceYF5 · 2026-08-19