Budget Inference Dilemma: 24GB GPU for Dense Models vs. RAM for MoE?

Agitated_Camel1886 · reddit · 2026-07-30

A developer posted asking for advice on hardware selection for a local inference machine. With a budget of £500-600, the primary goal is a smooth chat and RAG experience.

They are torn between three options:

They are asking the community for specific hardware advice: what is the minimum memory bandwidth required for CPU-only MoE inference to be viable, and is a budget GPU necessary to accelerate prompt processing?

Original post →

More from Infra

Infra channel →