FP8 weights at BF16 memory size, KV cache remains the VRAM bottleneck
AlpinDale · x · 2026-08-28
Discusses the performance trade-off when FP8 checkpoints occupy as much memory as BF16 on hardware supporting FP8 activations (e.g., sm89+). The view is that unless the model is very large, the KV cache remains the biggest VRAM consumer.
More from Infra
- Australia Datacentres Use 3% Power, Set to Hit 13% by 2035 — nordicinst · 2026-08-28
- Cloudflare saved 100TB of memory with 5 changes to 1.1.1.1's DNS cache — ritakozlov · 2026-08-28
- GPT price cuts trigger 13.8x usage surge — scaling01 · 2026-08-28
- Dual-GPU on AM5 delivers 0.1GB/s instead of 8GB/s: Promontory bridge blamed — Ed-2-Zero-9 · 2026-08-28
- DwarfStar Adds GLM 5.3 Flash Support with Q2/Q4 on MacBook — antirez · 2026-08-28
- After NVIDIA's llama.cpp acquisition, are used V100s still a safe cheap-VRAM bet? — OnlineParacosm · 2026-08-28