AirLLM runs a 70B LLM on a single 4GB GPU by streaming layer weights from disk
techNmak · x · 2026-10-03
The 35K-star open-source project AirLLM runs a 70B LLM on a single 4GB GPU without quantization, distillation, or pruning.
The trick is changing when weights enter VRAM:
- The checkpoint is split into per-layer shards on disk, and the full Hugging Face model is created on PyTorch's meta device, so parameters take no real memory
- Hooks load each layer's weights onto the GPU right before it runs and drop them immediately after, while prefetching the next layer in a pipelined fashion
- The full 70B checkpoint never needs to sit in VRAM at once
For MoE models it goes further: experts are streamed individually, loading only the ones a token actually routes to. The repo claims it can run the 2.8T-parameter Kimi K3 this way.
More from Infra
- DGX Spark shortage derails $8,800 donation plan as buyer can't find stock — cyrus_zei · 2026-10-03
- Perplexity to vertically integrate agentic infra on NVIDIA's Vera CPU, ditches x86 — AravSrinivas · 2026-10-03
- 2.3x faster Qwen3.8 27B on RTX 5090: ninfer vs llama.cpp benchmarked across 4 setups — theexile1337 · 2026-10-03
- GPT-6 Astra burns 30 minutes of compute on oversized tasks, then dumps all progress at the limit — CautiousMagazine3591 · 2026-10-03
- SemiAnalysis asks if GPUs are money printers: TPU v7, Vera Rubin vs Blackwell — AccBalanced · 2026-10-03
- gufo Ships Pre-packaged Windows Build for AMD Strix Halo, 40 tps on Agentic Workloads — hiImMate · 2026-10-03