Red Hat Releases Quantized Kimi K3: FP4/FP8 for Accelerated Inference with Minimal Quality Loss
_akhaliq · x · 2026-08-01
Red Hat AI has released hardware-optimized quantized versions of Moonshot's Kimi K3 model to accelerate local inference and boost throughput.
- NVFP4 Version: Designed for Blackwell architecture, quantizing MoE layers to 4-bit. Evaluations show minimal quality degradation, with GPQA score dropping slightly from 93.5 to 91.0.
- FP8-Block Version: Optimized for Hopper architecture (H100/H200), leveraging native Tensor Cores for high throughput with Day Zero support via vLLM.
The quantized weights are now open-sourced on Hugging Face and recommended for deployment with vLLM.
More from Infra
- Musk Predicts 99.99% of AI Compute Will Eventually Migrate to Space — XFreeze · 2026-08-01
- MediaTek Targets $12-16B Data Center Revenue by 2027, TPU v8t to Exceed $2B — BenBajarin · 2026-08-01
- Developer Burns 50M Tokens on a Single Task Due to Unmanaged Thread — msg · 2026-08-01
- AI Infrastructure Boom: How Many Datacenters Can Fit in an Airport? — chrisalbon · 2026-08-01
- 7x RTX 3090s Barely Run DeepSeek-V4 Q8 Quantization Locally — _akhaliq · 2026-08-01
- a16z: AI Infra Demand Unabated as Supply Chains Struggle — a16z · 2026-08-01