Kimi K3 Full Fine-Tuning Launches with Minimal MXFP4 Loss
burny_tech · x · 2026-08-30
Applied Compute launched full fine-tuning for Kimi K3 on the AC2 platform, reducing GPU requirements per training replica by 40% through memory optimizations. Tests show the KL divergence between bf16 training and MXFP4 inference is approximately the same as between bf16 training and bf16 inference, indicating high efficiency with minimal precision loss.
Related event: Kimi K3 Full Fine-Tuning Launches on AC2 with 40% Fewer GPUs(3 posts)→
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01