Running Qwen 1M context on dual 5090s: custom engine beats vLLM at decode

Littlepharaoh · reddit · 2026-08-30

A developer forked the NInfer engine to add tensor parallelism and YaRN scaling, enabling a 1M token context window for Qwen on dual RTX 5090s (27.4 GB per card, no NVLink).

Performance highlights (500W per GPU):

Original post →

More from Infra

Infra channel →