Run 70B Inference on a 4GB GPU

lyogavin · github · 2026-07-18

AirLLM focuses on **running 70B inference on a single 4GB GPU**. The repository covers topics like open-models, open-source-models, lora, and qlora, highlighting its goal to make large models usable under extremely low VRAM conditions.

Original post →

More from Infra

Infra channel →