RTX 5090 + 5070 Ti workstation: doubling VRAM wasn't worth it for local LLMs
Lordofwhut · reddit · 2026-10-02
A non-tech Redditor documents building a local LLM workstation around an RTX 5090, then adding a used Threadripper 7960X (96GB ECC DDR5) + 5070 Ti to run bigger models. Findings: Qwen 3.8 27B Q4 fully on the 5090 is fastest; splitting Q8 across both cards trades too much speed for accuracy; going from 32GB to 48GB VRAM mattered far less than expected, leaving the 5070 Ti idle. The post asks for heterogeneous dual-GPU use-case ideas.
More from Infra
- awesome-jev indexes 700 production tools around TypeSafe AI's decision model Jev — Remarkable-Gur719 · 2026-10-02
- smolvm v1.22 ships near-instant VM resume for undoing agent actions, 6.5k stars — LoganGrasby · 2026-10-02
- Broadcom to lend Anthropic up to $42B for AI chips, eyeing top customer slot by 2027 — rohanpaul_ai · 2026-10-02
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- Dev open-sources GPT-2-tools to run original 1.5B GPT-2 XL locally on CPU — MikePFrank · 2026-10-02
- CoreWeave launches serverless GPUs: hourly-billed, no contract, private preview — altryne · 2026-10-02