StepFun co-founder proposes KITE: PD-separation-inspired training for scaling agentic LLMs

teortaxesTex · x · 2026-09-29

Following the Step 5 Preview release, StepFun co-founder Yibo Zhu published a technical deep dive on Zhihu exploring whether training can benefit from the same "PD separation" idea used in inference, proposing KITE (KV-Invariant Transformer Expansion).

The core challenge is jointly optimizing model quality, training cost, and inference cost. MoE is the classic three-way win, but approaches like Upcycling and YOCO tend to trade one dimension against another: scaling them to recover model quality raises training or inference cost.

KITE combines scaling with upcycling while keeping KV computation invariant: a smaller prefiller is first trained to generate the KV cache, allowing model expansion without changing KV computation — a potential new path for scaling agentic LLMs.

Original post →

More from coding & agent

coding & agent channel →