Qwen3.8 Runs 170K Context on Single 96GB GPU
Users deployed Qwen3.8-Flash-Next on a single RTX 6000 Pro with 96GB VRAM via llama.cpp, achieving roughly 170K context length at around 109–110 tokens per second, sharing detailed configurations and quantization options.
2026-08-30 ~ 2026-09-01 · 2 related posts
- Running Qwen3.8 at 170K Context on a Single 96GB GPU — UltrMgns · 2026-08-30
- RTX 6000 Pro Qwen3.8 Benchmark: 96GB VRAM Hits 109 tok/s — FantasticNature7590 · 2026-09-01