Running Qwen3.6 Quantized on RTX 5080: Performance Trade-offs
boyeardi · reddit · 2026-07-07
一位用户在 RTX 5080 + 32GB DDR4 上本地运行了 Qwen3.6:27B 在 Q3KM 下约 60 tok/s,35B 在 IQ2 下约 190 tok/s,但精度从 Q6 下调带来的准确性损失明显。该用户另有一张闲置的 1080ti(11GB),考虑用其做 tensor split 分担负载,并询问 16GB 显存用户的实际方案与上限。
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11