StepFun co-founder proposes KITE: PD-separation-inspired training for scaling agentic LLMs
teortaxesTex · x · 2026-09-29
Following the Step 5 Preview release, StepFun co-founder Yibo Zhu published a technical deep dive on Zhihu exploring whether training can benefit from the same "PD separation" idea used in inference, proposing KITE (KV-Invariant Transformer Expansion).
The core challenge is jointly optimizing model quality, training cost, and inference cost. MoE is the classic three-way win, but approaches like Upcycling and YOCO tend to trade one dimension against another: scaling them to recover model quality raises training or inference cost.
KITE combines scaling with upcycling while keeping KV computation invariant: a smaller prefiller is first trained to generate the KV cache, allowing model expansion without changing KV computation — a potential new path for scaling agentic LLMs.
More from coding & agent
- A naive video transcription fixer pipeline: extract audio+frames, ASR, then correct with a frontier model — capetorch · 2026-09-29
- bdsqlsz is vibe-coding DLSS5 weight training for his in-development 3D game — bdsqlsz · 2026-09-29
- We Built an On-Call Agent That Failed the Right Way — Memory Can Learn the Wrong Lesson — Similar-Split7292 · 2026-09-29
- Extending Jev Mode to Images: Constrained llama.cpp Outputs as Image Selections — opUserZero · 2026-09-29
- Scraping Xiaohongshu hit posts with Codex + a wired Android phone — huangyun_122 · 2026-09-29
- Reverse-Engineering MW2, Minecraft and Skate 3 With Claude and DeepSeek Into One Playable Game — ericcalyborn · 2026-09-29