Discussion: Can 128GB RAM + 16GB VRAM Run Qwen 3.5 122B?
Stock-Union6934 · reddit · 2026-07-06
A user asked in the community about the performance of running Qwen 3.5 122B (MoE architecture) on a setup with 128GB DDR5 RAM and 16GB VRAM, seeking feedback from those with practical experience. This involves hardware configuration selection and feasibility for local deployment of large models.
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11