Ask HN: How Does Qwen 27B Quantized Run on Dual P100 GPUs?
kirisoraa · reddit · 2026-08-21
A user is asking whether anyone has experience running quantized Qwen 27B (q6-q8) across multiple Tesla P100 GPUs (16GB HBM2), with interest in prefill/generation speeds, power consumption, and whether it's a viable budget option for local inference on the secondhand market.
More from Infra
- Discussion: Data Privacy in AI Agents and the Case for Private Inference — Many_Audience7660 · 2026-08-21
- Replicating Anthropic Requires 16k Prompt & $250k Synthesis Cost — Ghost_Pilot_MD · 2026-08-21
- Langship: Open Source Tool for Deploying AI Agents Like Terraform — Many_Audience7660 · 2026-08-21
- Qwen3.8-27B now supports NVFP4 and DFlash2 quantization in SGLang — Alibaba_Qwen · 2026-08-21
- Pretraining a Mini Kimi K3 on One H200 for $252: A Complete Worklog — joecole · 2026-08-21
- 130 years, 10^22x more compute per dollar: Kurzweil's graph sparks debate — Singularitarian · 2026-08-21