Prepping for Local LLM Inference: Enthusiast Builds 30TB SSD & 256GB RAM Rig
reto-wyss · reddit · 2026-08-03
As local LLMs grow in size, a developer shared an extreme hardware setup to meet future local inference demands.
- Storage & Bandwidth: Uses 12x Gen 4 3.2TB SSDs (two per GPU) to achieve around 60GB/s bandwidth across 30TB of storage for model loading.
- Memory & Cache: Equips 256GB of DDR4 RAM dedicated to KV cache, utilizing high-endurance drives to write cache to disk.
- Context: The author jokingly references rumored models like "Le Chaton FAT," noting that while affording multiple high-end GPUs is tough, stacking disks and memory is a viable prep for the next era of local inference.
More from Infra
- Cloudflare Kicks Off Agents Week: Building a Cloud Native for AI Agents — ritakozlov · 2026-08-03
- Open Video Models: License Restrictions and VRAM Requirements Compared — Mysterious_Sign_9501 · 2026-08-03
- AMD signs $14B+ deal with Core Scientific for 530MW AI data center capacity — Beth_Kindig · 2026-08-03
- Developer Reports Spending Over $1M on OpenAI Codex This Year — cloneofsimo · 2026-08-03
- 30B Video MoE Quantizations Tested: Most Users Should Wait — EntireBig7258 · 2026-08-03
- AI Data Centers Consume Up to 1.5 Billion Gallons of Water Yearly — AndyMasley · 2026-08-03