Why Unified Memory Can Run 70B Models
ermantrout · hn · 2026-07-10
The article explains why mini PCs with unified memory can sometimes run 70B models, whereas a "large GPU" might not necessarily be the better fit.
The core ideas are:
- Unified memory allows flexible sharing between CPU and GPU memory, accommodating larger model weights.
- The real bottleneck is often not "capacity," but bandwidth, latency, and actual throughput.
- The article also cautions that while larger models can run, speeds are typically significantly limited, especially in scenarios requiring high throughput.
More from Infra
- NVIDIA launches Vera Rubin with 10x better performance per watt — nvidia · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Arbitrum fee simulation shows higher gas capacity but lower L2 revenue under ArbOS61 — tomwanhh · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- SkyPilot exits stealth with $20M to unify fragmented GPU compute across five clouds — skypilot_org · 2026-07-22
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22