Trillion-param models on a PC: hot weights in RAM, experts on NVMe
carrigmat · x · 2026-08-27
carrigmat explains why the weights don't all live in RAM:
- Most parameters in modern LLMs are experts activated only intermittently — those stay on NVMe drives.
- Only the persistently 'hot' weights remain in memory, so 128-192GB RAM plus NVMe can carry a trillion-parameter model.
More from Infra
- Prefix Sliding enables efficient test-time scaling by cutting memory costs — Niklas Muennighoff · 2026-08-27
- LightningAI offers instant H100 access on its self-owned AI cloud — LightningAI · 2026-08-27
- Cohere releases Parse 5, claiming best price-performance for enterprise document parsing — irombie · 2026-08-27
- M7 Ultra may feature native FP8, potentially boosting GLM 5.3-flash performance — Brilliant-Hall1387 · 2026-08-27
- ChronoScale announces 50MW NVIDIA GB300 deployment with Microsoft for AI inference — r_jegaa · 2026-08-27
- Engineering Win: mxfp8 x mxfp4 Matmul Outperforms Standard mxfp8 — zephyr_z9 · 2026-08-27