Supervising a student's future actions with a large teacher: an early form of on-policy distillation

abursuc · x · 2026-09-16

At the #ssad2026 workshop, the author points to a way of thinking about LLMs and think-ahead: use a large teacher model to supervise a small student's future actions — squint a bit and it's an early form of on-policy distillation from the driving world.

A follow-up note covers mid-level slot representations: they worked well, but progress elsewhere (e.g., SAM2's implicit tracking) pushed researchers toward other problems — "that's research: looking at things that don't work yet."

Related event: SSAD 2026 Explores LLM-Style Think-Ahead Supervision for Autonomous Driving(3 posts)→

Original post →

More from Research

Research channel →