Qwen3.8-27B achieves 2x decode speedup with DFlash2 on 256k context
maddie-lovelace · reddit · 2026-08-21
Tests on Qwen3.8-27B-UD-Q4KXL with DFlash2 on a single RTX 5090 show roughly 2x decode speedup (40 to 75 tps) at 256k context with only a 15% prefill slowdown. The post includes specific reproduction commands and configuration notes.
More from coding & agent
- Grok Bot cleans 100k emails, unsubscribes via autonomous browser use — elonmusk · 2026-08-21
- First OpenAI Codex Ambassador appointed in Salt Lake City — paw_lean · 2026-08-21
- Human-Agent Collaboration: How Modern Teams Run Agentic Workflows — JosephJacks_ · 2026-08-21
- PostHog Founder on AI Pivot: Becoming a 'Doing Company' That Fixes Code While You Sleep — ycombinator · 2026-08-21
- LangChain to host 'Building Agents with Agents' Roadshow on Aug 27 — LangChain · 2026-08-21
- YC Startup Qlo Launches AI Agent for Commercial Underwriting Inbox Automation — ycombinator · 2026-08-21