Optimizing Local LLMs on Dual 5060 GPUs
MistingFidgets · reddit · 2026-07-11
楼主把第二张 5060 16GB 加进服务器后,总计有 32GB 显存 和 80GB ECC DDR4,想把权重分摊到两张卡上以提升推理能力。
他提到当前环境是 PCIe 3.0,担心 tensor parallel 在这种总线上效果不佳,但实际用 qwen 3.6 35B 仍能跑到大约 3200 tok/s 的 prompt processing 和 100 tok/s 的生成速度;27B 模型在 tensor split 下大约是 600 PP / 23 generation。他现在用 llama.cpp + GGUF,也有 vLLM 的 nvfp4 模型,希望大家给出进一步优化双卡拆权重的设置建议。
More from Infra
- Arbitrum fee simulation shows higher gas capacity but lower L2 revenue under ArbOS61 — tomwanhh · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- SkyPilot exits stealth with $20M to unify fragmented GPU compute across five clouds — skypilot_org · 2026-07-22
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22