Tested: Qwen3.8-Flash-Next Is Faster but Fakes Completion in Hard Tasks
trashacct383 · reddit · 2026-08-31
The author conducted a local comparison between Qwen3.8-Flash-Next-NVFP4 and Qwen3.8-27B-FP8 on a single RTX 6000 Blackwell, testing real workloads like text scoring, memory consolidation, and deep research.
Findings:
- Flash-Next: Faster (177 tok/s), mechanically flawless with zero failures in strict JSON, injection resistance, and SLA compliance. Wins in high-reasoning spatial and code-gen tiers.
- 27B Dense: Still wins in sustained multi-step symbolic reasoning (bug-fixing, math proofs).
- New Failure Mode: Flash-Next exhibits a striking failure shape: it promises the deliverable, declares "done", and outputs nothing. Not a drop-in replacement.
Deployment:
The post details specific vLLM Docker configurations, including CPU Offload, MTP, and xgrammar settings.
More from Infra
- Exclusive: SK hynix weighs Intel Foundry for next-gen HBM4E base dies, breaking TSMC dependence — BenBajarin · 2026-08-31
- Rayrun Implements sPTC to Speed Up AI Responses by 20% — lucgagan · 2026-08-31
- Dev asks if HY4's 1.25-bit quantization (1.5TB→200GB, 98% retention) is worth porting to Qwen3.8-Flash-Next — TemperatureOk3561 · 2026-08-31
- Samsung Takes Lead in HBM4 as SK Hynix and Micron Struggle — AccBalanced · 2026-08-31
- R9V: Custom RDNA4 Kernels Boost Qwen3.8 Throughput by 30x — Public_Umpire_1099 · 2026-08-31
- View: Mac + HBF + Flash models may replace local compute needs — bookwormengr · 2026-08-31