Running Qwen3.8 at 170K Context on a Single 96GB GPU

UltrMgns · reddit · 2026-08-30

A user shared a practical deployment of Qwen3.8-Flash-Next on a single 96GB GPU, achieving a 170K context window and 110 tokens/sec.

Configuration & Optimizations:

Performance:

Related event: Qwen3.8 Runs 170K Context on Single 96GB GPU(2 posts)→

Original post →

More from Infra

Infra channel →