Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context
HankYeomans · x · 2026-10-12
The author runs a 456GB-parameter model across 192GB VRAM, 256GB RAM/CPU, and NVMe, reporting usable performance: 60.1 tok/s decode (C1) and 124.8 tok/s aggregate (C4), 16K prefill at 4,361 tok/s (TTFT 3.7s), 32K prefill at 4,513 tok/s (TTFT 7.1s). With 1M or 4×262K context, they think it could even spawn frontier agents; real-task testing is next.
More from Infra
- Engineer joins NVIDIA's Groq LPU compilers team, working on multi-chip partitioning — blelbach · 2026-10-12
- Linus Ekenstam wants nothing less than a 100B-param model running on your phone — LinusEkenstam · 2026-10-12
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Fireworks: open models plus fine-tuning match closed ones — Cursor gets 13x faster inference — AI Engineer · 2026-10-12
- VitalOps launches agentic inference optimization, 2.6x median speedup — abhijithneil · 2026-10-12
- Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s — human_in_the_looop · 2026-10-12