Frontier Models on Consumer GPUs: Qwen3.6 on 8GB VRAM at 39 tok/s
berkeley_ai · x · 2026-08-22
Optimizations for memory and compute now allow consumer GPUs to serve frontier-scale models at interactive speeds. Benchmarks show Qwen3.6 35B running at 39 tok/s on an 8GB RTX 4060 laptop, DeepSeek-V4-Flash 284B at 22-25 tok/s on an RTX 5090, and GLM-5.2 753B at 15 tok/s on an RTX PRO 6000. This significantly lowers the barrier to entry for deploying robust agentic workflows, enabling tools like Claude Code to run locally for $0.
More from Infra
- NVIDIA DGX SuperPOD Cuts Drug Discovery from Years to a Week — nvidia · 2026-08-29
- Passive Oculink setup unintentionally creates CUDA task scheduler — TinFoilHat_69 · 2026-08-29
- RX 9070 XT shows huge speed variations with ComfyUI — bosox62 · 2026-08-29
- Open Source Static Performance Model for LLM Inference — stanfordnlp · 2026-08-29
- Prediction: Closed frontier models to become downloadable by 2027 — imjustnewatai · 2026-08-29
- Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading — StartupTim · 2026-08-29