Running Qwen3.8-Flash-Next FP8 at 524K context on dual RTX 6000

SpendLucky1273 · reddit · 2026-08-29

The author successfully ran Qwen3.8-Flash-Next FP8 at 524K context on 2x RTX PRO 6000 using vLLM.

Config: TP2 + EP2, MTP3, PLE CPU offload, YaRN 2x, and --max-model-len 524288.

Issue & Fix:

Status: Boots cleanly with 654,980 tokens KV cache and 1.25x concurrency.

Original post →

More from Infra

Infra channel →