RTX 6000 Pro Qwen3.8 Benchmark: 96GB VRAM Hits 109 tok/s

FantasticNature7590 · reddit · 2026-09-01

User performed detailed llama.cpp deployment and performance tests for Qwen3.8-Flash-Next-GGUF (quantized with UD-IQ4XS) on an RTX 6000 Pro with 96GB VRAM.

Key Findings

Benchmark Data (2K Prompt)

| VRAM | Prefill | Decode |

| :--- | :--- | :--- |

| CPU Only | 182.64 | 8.34 |

| 24GB | 260 | 39.01 |

| 48GB | 746.7 | 51.73 |

| 96GB | 1,955 | 109.07 |

The test notes that VRAM limits primarily affect how many expert layers can be loaded, and while the Gated DeltaNet architecture saves memory, it does not make long-context decoding free.

Related event: Qwen3.8 Runs 170K Context on Single 96GB GPU(2 posts)→

Original post →

More from Infra

Infra channel →