Comparing V100 and 5060Ti for Qwen 3.8 inference
ColorsOfCosmos · reddit · 2026-08-31
User discusses upgrading from a 5060Ti 16GB (40t/s, 130k context) for Qwen 3.8 IQ3S. Options include 1x32GB or 2x16GB V100 cards to increase context and improve quantization, constrained by 2x x4 PCIe lanes.
More from Infra
- Local Llama 3.1 install caused slowdown, fixed by uninstall — AiJohnAllen · 2026-08-31
- SimSlim tool fixes AI iOS coding crashes by running more simulators per Mac — bigblueboo · 2026-08-31
- Local 8B Model Document Extraction Demo on iPhone 16 — Better_Comment_7749 · 2026-08-31
- Apple May Scrap 2027 Mobile HBM Plans Due to High Costs — power97992 · 2026-08-31
- Matmul Energy Efficiency Competition Challenges AlphaTensor Algorithm — yaroslavvb · 2026-08-31
- NVIDIA DGX Station offers data-center-class performance — SpendLucky1273 · 2026-08-31