FP8/FP4 Quantization Delivers Only ~1.5x and 2x Real Speedups, Far Below Theoretical Gains
scaling01 · x · 2026-09-07
The author argues that FP8 and FP4 quantization do not provide the theoretical 2x and 4x speedups over FP16/FP8 — realistic gains are closer to 1.5x and 2x. Also comments that 100 days is a somewhat ridiculous timeline while 30 days is overly optimistic.
More from Infra
- Qdrant calls out Actian's embedded vector DB comparison for its own "not embedded" pick — qdrant_engine · 2026-09-07
- Yandex Makes Pretrained LLMs Interactive by Reorganizing KV Cache, No Retraining Needed — arpit_bhayani · 2026-09-07
- Greenland pays $124/month for 15 Mbps as Starlink offers 200 Mbps in Denmark for $57 — XFreeze · 2026-09-07
- Red Hat and Hugging Face host PyTorch systems night in Bengaluru with 170+ contributors — PyTorch · 2026-09-07
- EU lawmakers push EUROPA consortium: open-source AI in all 24 EU languages plus shared compute — EvaMaydell · 2026-09-07
- Naura demos key etch process for 64-layer 3D DRAM without EUV, selectivity above 500:1 — pstAsiatech · 2026-09-07