RTX 6000 Pro Qwen3.8 Benchmark: 96GB VRAM Hits 109 tok/s
FantasticNature7590 · reddit · 2026-09-01
User performed detailed llama.cpp deployment and performance tests for Qwen3.8-Flash-Next-GGUF (quantized with UD-IQ4XS) on an RTX 6000 Pro with 96GB VRAM.
Key Findings
- Massive Throughput Delta: CPU-only decode achieved 8.34 tok/s; utilizing full 96GB VRAM boosted this to 109.07 tok/s.
- Convergence at Long Context: At a 2K prompt, 96GB was 2.8x faster than 24GB; at 245K context, the advantage shrank to 1.45x (24GB: 14.89 tok/s vs 96GB: 21.61 tok/s).
- PLE Placement Trap: Forcing the 27.2GB PLE (Per-Layer Embedding) table onto CUDA caused a 55.6x slowdown in decode (dropping from 108.5 to 1.95 tok/s); keeping it in system RAM was significantly faster.
Benchmark Data (2K Prompt)
| VRAM | Prefill | Decode |
| :--- | :--- | :--- |
| CPU Only | 182.64 | 8.34 |
| 24GB | 260 | 39.01 |
| 48GB | 746.7 | 51.73 |
| 96GB | 1,955 | 109.07 |
The test notes that VRAM limits primarily affect how many expert layers can be loaded, and while the Gated DeltaNet architecture saves memory, it does not make long-context decoding free.
Related event: Qwen3.8 Runs 170K Context on Single 96GB GPU(2 posts)→
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- mlx-signal-processing brings 10-200x faster signal ops to Apple Silicon — TheMoonMidas · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01