How much memory for a 30B model? A quick precision-to-VRAM calculation

ashishllm · x · 2026-09-24

A simple interview-worthy calculation: each 1B parameters needs 4GB at FP32, 2GB at FP16, 1GB at int8, and 0.5GB at int4—so a 30B model ranges from 120GB down to 15GB. Quantization cuts memory from 4GB to 500MB per billion parameters. The author notes many aspiring AI engineers can't answer this.

Original post →

More from Infra

Infra channel →