AirLLM runs 70B models on a 4GB GPU via layer-wise inference, scaling to 405B on 8GB

JensHonack · x · 2026-10-11

Open-source AirLLM uses "Layer-wise Inference" — loading, computing, and flushing one layer at a time — so a 70B model fits on a single 4GB GPU and scales up to Llama 3.1 405B on 8GB of VRAM. No quantization needed by default, supports Llama/Qwen/Mistral, works on Linux, Windows, and macOS, fully open source. The main tradeoff likely being latency.

Original post →

More from Infra

Infra channel →