Frontier Models on Consumer GPUs: Qwen3.6 on 8GB VRAM at 39 tok/s

berkeley_ai · x · 2026-08-22

Optimizations for memory and compute now allow consumer GPUs to serve frontier-scale models at interactive speeds. Benchmarks show Qwen3.6 35B running at 39 tok/s on an 8GB RTX 4060 laptop, DeepSeek-V4-Flash 284B at 22-25 tok/s on an RTX 5090, and GLM-5.2 753B at 15 tok/s on an RTX PRO 6000. This significantly lowers the barrier to entry for deploying robust agentic workflows, enabling tools like Claude Code to run locally for $0.

Related event: UC Berkeley and MIT Open-Source FreeToken: Frontier MoE Models on Consumer PCs(11 posts)→

Original post →

More from Infra

Infra channel →