Kangwook Lee's Team Rented a Korean PC Bang, Recruited 1000+ Players to Train On-Device Game AI Ally Distilled to 2B
Kangwook Lee's team built a pipeline that runs game AI NPCs in real time on local consumer-grade GPUs, and revealed in detail how the data for the StarCraft AI "Ally" was collected and compressed. The core idea is to combine large-scale play sessions with human players and iterative distillation so the model can achieve near-real-time interaction on-device.
Confirmed
- The authors note that prototyping AI NPCs via APIs is feasible, but APIs struggle to deliver near-real-time interaction, which motivated the local deployment.
- Compression pipeline: SFT and off-policy KD first, followed by agentic training, and finally distilling the model down to 2B scale to run on consumer GPUs.
- For data collection, the team borrowed the DAgger idea: collect trajectories, apply corrections to the trajectories and continue training, then deploy the new model back to the PC bang for the next round of collection, iterating repeatedly.
- The most striking detail: the team rented a Korean PC bang and recruited over 1,000 human players to actually play against early versions of Ally in order to collect on-policy rollout data.
- Since everything runs on edge devices with a very tight KV cache budget, the team developed a dedicated context compaction algorithm and released full examples of the reasoning/tool-calling/observation traces.
Why it matters
- This case demonstrates a complete engineering path for compressing large models to run on-device: distillation, on-policy human data, iterative correction training, and context compensation all fit together, offering direct reference value for teams aiming to build low-latency agents on local devices.
- Renting a PC bang and recruiting a thousand players also illustrates a practical, scalable approach to collecting high-quality on-policy data.
2026-09-29 ~ 2026-09-29 · 5 related posts
Primary sources
- StarCraft AI Ally team rented a PC bang and recruited 1,000+ humans to collect on-policy rollout data — Kangwook_Lee ·
- Distilling an Agent Down to 2B for Consumer GPUs via DAgger-Style Loops — Kangwook_Lee ·
- Edge-run agent uses context compaction algorithm to fit tight KV cache budget — Kangwook_Lee ·
- [source] Edge-run agent uses context compaction algorithm to fit tight KV cache budget — Kangwook_Lee · 2026-09-29
- [source] StarCraft AI Ally team rented a PC bang and recruited 1,000+ humans to collect on-policy rollout data — Kangwook_Lee · 2026-09-29
- Team Rented a PC Bang and Recruited 1,000+ People to Train Game AI — Kangwook_Lee · 2026-09-29
- [source] Distilling an Agent Down to 2B for Consumer GPUs via DAgger-Style Loops — Kangwook_Lee · 2026-09-29
- Distilling an LLM down to 2B to run AI NPCs on consumer GPUs in near-real-time — Kangwook_Lee · 2026-09-29