Running Qwen3.6 27B on Tesla V100: 128K Context Config & Performance

Traditional_Bell8153 · reddit · 2026-08-08

A developer shared their configuration and performance results for running the Qwen3.6 27B model on a Tesla V100 PCIE 32Gb GPU on Reddit.

The setup uses a Q4KM quantized main model paired with a Q80 Multi-Token Prediction (MTP) draft model, tested under a massive 128K context length. The post details the specific llama.cpp parameters used, including Flash Attention, thread counts, and batch sizes, accompanied by a performance screenshot. The author is looking to compare notes with other V100 users.

Original post →

More from Infra

Infra channel →