Running a 295B Model on 16GB VRAM
antirez · x · 2026-07-10
A developer demonstrated how to run Tencent's Hy3, a 295B parameter MoE model, on a single RTX 4060 Ti 16GB GPU. Using NeutronStar (based on the author's CUDA branch ds4), attention layers and shared experts are kept in VRAM, while routed experts are streamed from the SSD on a per-token basis. This enables interactive chat with a massive model on a roughly $400 GPU.
The implementation adds a GQA attention path and utilizes a native 2-bit GGUF format for ds4, achieving speeds of about 1.8 tok/s. The author also provided a link to the model weights.
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11