Musk says RAM supply is the problem as quantized 25GB MoE local run speculated for 2028-29
elonmusk · x · 2026-10-11
Replying to a discussion about running a large model locally, Elon Musk said "the problem is RAM supply." The original poster speculated the model could be quantized down to 25GB (likely as an MoE), with NAND-based sampling arriving around 2028 and feasibility by late 2028 or early 2029. Musk's terse reply points to memory supply chains as the bottleneck.
Related event: Qualcomm CEO: AI Firms Want Phones Running 100B-Parameter Models by 2028(6 posts)→
More from Infra
- Jevons paradox in AI: falling token prices keep GPU demand tight — AccBalanced · 2026-10-11
- Local AI coding hits only 20-35% of Sonnet's speed in weeks-long app-build tests — julianharris · 2026-10-11
- A 37ms Gap Left GPUs Idle ~40% of the Time; Fixing It Removed the Tiny Idles — HankYeomans · 2026-10-11
- Musk courts chip engineers as Terafab targets 10 chip designs a year — elonmusk · 2026-10-11
- RTX 5090 Hits $5,000 as AI Firms Buy Gaming GPUs by the Pallet — chemist_slime · 2026-10-11
- Qwen3.8-27B Hits 140 tok/s on a Single RTX 3090 with a CUDA Megakernel, KL Divergence 0.0009 — Adorable_Weakness_39 · 2026-10-11