Local LLM Deployment: Is 128GB or 256GB System RAM Better for 200-300B Models?
Thin_Pollution8843 · reddit · 2026-08-08
The author sparks a technical discussion: when equipped with 128GB of VRAM (using 8-channel DDR4), should you pair it with 128GB or 256GB of system RAM?
- 128GB + 128GB setup: Barely sufficient for q8 quantized models like DeepSeek, keeping total usage under 200GB including context.
- 256GB setup: For higher quantization of 200B-300B parameter models, the VRAM/RAM weight split ratio heavily impacts performance, making larger RAM more secure.
Since the setup uses cheap, used slow memory sticks, the author seeks community feedback on real-world performance impacts of different split ratios.
More from Infra
- Firebird Launches CIS Region's Largest AI Factory with 70,000 NVIDIA GPUs — nvidia · 2026-08-09
- Semiconductor Layer Hits $20T Market Cap Excluding Giants, Analyst Warns of Overheating — pdamodaran · 2026-08-09
- llama.cpp Slower Than Ollama? Local LLM Deployment Config Pitfalls — anshulsingh8326 · 2026-08-09
- Google's Decade Endgame: Challenging Nvidia's AI Chip Supremacy — BorisMPower · 2026-08-08
- Running MiniMax Video Generation on RTX 3060 Ti Takes 30 Minutes, Sparking Discussion — Zestyclose_Mountain6 · 2026-08-08
- Optimizing Small LLMs: Why the Standard Playbook Fails Below 1.5B Params — oli266 · 2026-08-08