Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context

HankYeomans · x · 2026-10-12

The author runs a 456GB-parameter model across 192GB VRAM, 256GB RAM/CPU, and NVMe, reporting usable performance: 60.1 tok/s decode (C1) and 124.8 tok/s aggregate (C4), 16K prefill at 4,361 tok/s (TTFT 3.7s), 32K prefill at 4,513 tok/s (TTFT 7.1s). With 1M or 4×262K context, they think it could even spawn frontier agents; real-task testing is next.

Original post →

More from Infra

Infra channel →