Prediction: flagship-quality local models on 16GB machines within 18 months
julianharris · x · 2026-09-18
The author predicts that within 18 months, machines with 16GB of RAM will comfortably run flagship-quality models at double speed. The extra RAM will instead buy concurrent sessions — envy will shift from model quality to how many agent sessions you can run at once: "128GB? Yeah baby, 8 agents running simultaneously."
More from Infra
- Hyperbolic hires quant researchers to build GPU compute as a tradable asset class — YiMaTweets · 2026-09-18
- Crusoe raises $3.9B Series F at $30.9B valuation to fuel AI energy buildout — beffjezos · 2026-09-18
- Periodic Labs details its stack: 4.1x Megatron throughput, frontier-beating science models on 1,300 H200s — hsu_byron · 2026-09-18
- Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once — julianharris · 2026-09-18
- Jeff Dean: a handful of workloads will dominate world compute, 'crying out' for specialized silicon — AccBalanced · 2026-09-18
- Silicon Data chart shows Nvidia GPUs retaining value well above depreciation schedules — AccBalanced · 2026-09-18