AirLLM update runs 2.8T Kimi K3 on under 4GB VRAM

pmttyji · reddit · 2026-08-20

AirLLM released updates significantly reducing inference memory, enabling massive models on consumer hardware. It now supports Qwen3.8-27B (3.33GB VRAM) and the 2.8T parameter Kimi K3 (3.72GB VRAM), and can run DeepSeek-V3 (671B) on 12GB. The tool uses per-expert streaming for MoE models, requiring no quantization, distillation, or pruning.

Original post →

More from Infra

Infra channel →