Running Qwen3.8-Flash-Next with 256K context at 16 tok/s on DDR4 and a Tesla T4

BusTiny207 · reddit · 2026-09-04

A Reddit user runs Unsloth's Qwen3.8-Flash-Next UD-Q4KXL (180B total/6B active, 111GB) on a refurbished Dell R740 with ikllama.cpp, achieving 256K context at 16 tok/s generation on a single Tesla T4 plus 384GB DDR4.

Key details:

The post includes the full launch command and tuning flags — a valuable blueprint for cheap large-context MoE inference on old servers.

Original post →

More from coding & agent

coding & agent channel →