How much VRAM/RAM is a meaningful upgrade for local LLM inference? Reddit weighs in

MiceLiceandVice · reddit · 2026-09-03

A local inference enthusiast kicked off a Reddit discussion on which hardware jumps actually matter. His main rig is 32GB DDR5 + 16GB GDDR7, and he argues going to 64+16 isn't worthwhile while aligned capacity (e.g. 32+32) is.

He plans to build a DDR4 + HBM2 inference machine as a complement — slower, but cheaper hardware gives much greater total capacity, letting him run a wider diversity of models. He asks the community which config best maximizes model access if the compromise is outright speed: 32+32, 128+32, or 64+64? The underlying thesis: capacity determines which models you can run; speed only determines how fast.

Related event: Reddit Debates RAM vs VRAM Upgrades for Local LLMs(2 posts)→

Original post →

More from Infra

Infra channel →