Samsung, Oxford and PKU propose TrOPD to distill frontier-model reasoning into on-device small models
jiqizhixin · x · 2026-09-26
Samsung, Oxford and Peking University present TrOPD
- Context: Frontier models keep scaling, but rising inference costs are the bottleneck to adoption. For Samsung's hundreds of millions of edge devices, on-device models must be smart enough within tight memory, compute, power and latency budgets.
- Method: Trust Region On-Policy Distillation (TrOPD) builds on On-Policy Distillation, which unlike GRPO's outcome-reward exploration uses a teacher LLM for fine-grained supervision so the on-device model directly learns the stronger model's reasoning.
- Key idea: TrOPD identifies the teacher's trustworthy supervision regions, letting the small model inherit large-model capabilities more stably and effectively while avoiding unreliable supervision.
More from Infra
- Vpipe vs Draw Things on M5 Pro: 24% faster at 1K, finishes 2K where Draw Things crashes — TgoAI · 2026-09-26
- AMD publishes educational GEMM optimization ladder for Helios MI455X GPUs with HipKittens — salykova_ · 2026-09-26
- Terafab starts hiring: 1 TW/year chip output and orbital AI compute in its sights — seanmcdonaldxyz · 2026-09-26
- Blog: Scaling LLM Inference from a Single Node to Millions — abhijithneil · 2026-09-26
- Pay-as-you-go vs committed LLM API volume: real procurement questions from a scaling team — LeviYagami · 2026-09-26
- Germany and the Netherlands put €40m into a challenge to design AI chips with AI — VraserX · 2026-09-26