Kimi K3 Full Fine-Tuning Launches with Minimal MXFP4 Loss

burny_tech · x · 2026-08-30

Applied Compute launched full fine-tuning for Kimi K3 on the AC2 platform, reducing GPU requirements per training replica by 40% through memory optimizations. Tests show the KL divergence between bf16 training and MXFP4 inference is approximately the same as between bf16 training and bf16 inference, indicating high efficiency with minimal precision loss.

Related event: Kimi K3 Full Fine-Tuning Launches on AC2 with 40% Fewer GPUs(3 posts)→

Original post →

More from Infra

Infra channel →