AirLLM runs a 70B LLM on a single 4GB GPU by streaming layer weights from disk

techNmak · x · 2026-10-03

The 35K-star open-source project AirLLM runs a 70B LLM on a single 4GB GPU without quantization, distillation, or pruning.

The trick is changing when weights enter VRAM:

For MoE models it goes further: experts are streamed individually, loading only the ones a token actually routes to. The repo claims it can run the 2.8T-parameter Kimi K3 this way.

Original post →

More from Infra

Infra channel →