Run Qwen3.8-27B NVFP4 on RTX 5090: vLLM, 256K Context, 147 tok/s
AIFlow_ML · x · 2026-08-16
MiaAI-Lab releases a script to serve RadixArk/Qwen3.8-27B-NVFP4 with vLLM 0.27.x on a single RTX 5090 (32GB), with native 256K context, TurboQuant 4-bit KV cache (5.5 GiB), and MTP-3 speculative decoding, exposed as OpenAI-compatible API. Measured 160 tok/s single-stream generation, full 262,144-token context resident, and a patch fixing stock 0.27.1 garbling (0/15 vs 13/15 failures).
More from coding & agent
- Sunil Pai's experiment: LLM-enriched voice transcription that links issues and wiki live — threepointone · 2026-08-16
- Codex intelligently routes tasks to cheaper Luna model to save limits — eyishazyer · 2026-08-16
- SelfMem Paper: AI Agents That Manage Their Own Memory Outperform Fixed Systems — alex_verem · 2026-08-16
- Reddit asks: AI agents are eating API budgets — how is anyone making money? — Nucleif · 2026-08-16
- Claude Code Delivers 5x More Compute Than Codex for $20 — Plenty-Emu3740 · 2026-08-16
- Using AI agents to simplify Apple Search Ads CLI workflows — rudrank · 2026-08-16