DFlash2 on Qwen3.8 27B hits ~200tk/s for code, requires more VRAM
Hefty_Wolverine_553 · reddit · 2026-08-19
A Reddit user tested Inco AI's DFlash2 speculation decoding with Qwen3.8 27B on an RTX 5090. Results show significant speedups for code generation, peaking at 200 tokens/s and averaging 120 tokens/s per request, compared to previous MTP performance. However, 'thinking' tasks saw a drop to 80-90 tk/s. The trade-off is higher memory usage, forcing a context size reduction from 220k to 160k. Detailed configuration flags were provided.
More from Infra
- PA Governor Enacts Nation's Strictest AI Data Center Standards via Executive Order — TinfoilTricorn · 2026-08-19
- Miles v0.1 Open Source RL Framework Launches for LLMs — ying11231 · 2026-08-19
- OpenAI Codex Lead Reveals Tokenizer Inefficiency Can Spike API Bills by 34% — 新智元 · 2026-08-19
- GitHub traffic explosion driven by agents: paid-only hosting isn't the solution — steipete · 2026-08-19
- Input tokens consume 54% of usage in agentic coding, revealing costly 'communication tax' — rseroter · 2026-08-19
- Fixing Model Agnosticism in Vercel AI Gateway — jasonkneen · 2026-08-19