Ask HN: How Does Qwen 27B Quantized Run on Dual P100 GPUs?

kirisoraa · reddit · 2026-08-21

A user is asking whether anyone has experience running quantized Qwen 27B (q6-q8) across multiple Tesla P100 GPUs (16GB HBM2), with interest in prefill/generation speeds, power consumption, and whether it's a viable budget option for local inference on the secondhand market.

Original post →

More from Infra

Infra channel →