Distilling an LLM down to 2B to run AI NPCs on consumer GPUs in near-real-time
Kangwook_Lee · x · 2026-09-29
- Why run AI NPCs on device? Prototyping with APIs is fine, but APIs can't deliver near-real-time interaction for in-game characters.
- The team distilled their model via SFT and off-policy KD, followed by agentic on-policy distillation (OPD), down to a 2B LLM.
- The result runs on consumer GPUs, enabling near-real-time AI NPC behavior.
More from Research
- Ramez's new piece tests intelligence explosion claims against current AI scaling data — tobyordoxford · 2026-09-29
- Paper Puts a Number on AI-Human Attention Gap: Context Windows Up 3,906x, Focus Down — alex_verem · 2026-09-29
- Max Welling: AI Is Searching 10^60 Possible Materials and Accelerating Its Own Invention — charlieharris01 · 2026-09-29
- StepFun co-founder proposes KITE: PD-separation-inspired training for scaling agentic LLMs — teortaxesTex · 2026-09-29
- Alignment researcher points to 'self-undermining unilateral optimization' classics — edelwax · 2026-09-29
- World's first live AI-assisted brain surgery removes tumour using real-time DINOv3 inference — TimDarcet · 2026-09-29