Qwen 3.8 27B NVFP4 benchmarked on 2xV100 with ~400k context

jjusko20 · reddit · 2026-09-17

A Reddit user benchmarks the NVFP4-quantized Qwen 3.8 27B on 2x V100 32GB with NVLink, using dflash2 to reach roughly 400k total context, with a screenshot of results. Useful reference for running long-context local inference on older GPUs.

Original post →

More from Infra

Infra channel →