Red Hat releases an FP8-quantized Kimi-K3 checkpoint tuned for Hopper GPUs
_akhaliq · x · 2026-07-29
Red Hat AI released an FP8BLOCK-quantized checkpoint for Moonshot’s Kimi-K3, tuned for Hopper GPUs.
- The checkpoint is designed to run on NVIDIA H100/H200 with Hopper-native FP8 tensor core throughput.
- Red Hat says it has day-zero support with vllmproject.
- The Hugging Face model card notes the quantization was applied to MLP and expert layers, while attention layers were left unquantized because of shape mismatches.
- The repo is MIT-licensed and includes evaluation results plus the full quantization recipe.
More from Infra
- A proposal to make KV cache portable across machines, data centers, and the WAN — knowrohit07 · 2026-07-29
- Bittensor raises q from 0.61 to 0.75, easing its emission gate for mid-ranked subnets — markjeffrey · 2026-07-29
- OpenRouter’s moat comes from routing data and tooling it can refine multiple times a day — mmurph · 2026-07-29
- Hyperscaler credit spreads may be overpricing risk as GPU spot rents run 2x contract rates — GavinSBaker · 2026-07-29
- OpenAI adds GPT-Live-Transcribe and GPT-Transcribe to its API — OpenAIDevs · 2026-07-29
- A new MCP server lets agents buy and manage proxies through plain-language chat — Key_Economy3728 · 2026-07-29