AirLLM update runs 2.8T Kimi K3 on under 4GB VRAM
pmttyji · reddit · 2026-08-20
AirLLM released updates significantly reducing inference memory, enabling massive models on consumer hardware. It now supports Qwen3.8-27B (3.33GB VRAM) and the 2.8T parameter Kimi K3 (3.72GB VRAM), and can run DeepSeek-V3 (671B) on 12GB. The tool uses per-expert streaming for MoE models, requiring no quantization, distillation, or pruning.
More from Infra
- ComfyUI with Flux 2 Klein and Qwen Image: inference slows ~4x after a few runs — ROBOTTTTT13 · 2026-08-20
- Running DeepSeek V4 on 16x RTX 5060 Ti via PLX switches — Primary_Exchange21 · 2026-08-20
- FrankenGit: Memory Safety Constitution for Pure-Rust Git Hosting Implementation — doodlestein · 2026-08-20
- Snowflake Turns Model Routing Into a Data Governance Feature — shashib · 2026-08-20
- llama.cpp PR Uses AVX2 to Speed Up Large Batch IQ Quantization — pmttyji · 2026-08-20
- NVIDIA Integrates TriAttention: Trigonometric KV Compression for Long Context — 青稞AI · 2026-08-20