ktransformers: Run LLMs on 24GB VRAM
dosco · x · 2026-07-18
The **ktransformers** team from Tsinghua optimizes MoE model deployment with a straightforward approach: - Keeping frequently used experts on the GPU while offloading less common ones to the CPU, enabling larger models to run on smaller VRAM. - The post claims this allows **DeepSeek-V3 / R1 to support 139K context on 24GB VRAM**, achieving up to **28x** speedups over standard setups. It also makes fine-tuning DeepSeek-V3 on **4x RTX 4090** GPUs possible. - Developed by the **Tsinghua MADSYS Lab**, the project is licensed under **Apache 2.0** and has surpassed **17,000** stars on GitHub.
Related event: ktransformers Enables Local Inference of Massive Models on 24GB VRAM(4 posts)→
More from Infra
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21