Supervising a student's future actions with a large teacher: an early form of on-policy distillation
abursuc · x · 2026-09-16
At the #ssad2026 workshop, the author points to a way of thinking about LLMs and think-ahead: use a large teacher model to supervise a small student's future actions — squint a bit and it's an early form of on-policy distillation from the driving world.
A follow-up note covers mid-level slot representations: they worked well, but progress elsewhere (e.g., SAM2's implicit tracking) pushed researchers toward other problems — "that's research: looking at things that don't work yet."
Related event: SSAD 2026 Explores LLM-Style Think-Ahead Supervision for Autonomous Driving(3 posts)→
More from Research
- AI is not yet driving drug development, Axios reports — polymute · 2026-09-16
- MIT's xvr aligns 2D X-rays with 3D scans in seconds at sub-millimeter precision — nordicinst · 2026-09-16
- Critch & Russell's 2023 'production web' AI-risk taxonomy gets a fresh look — seanwbren · 2026-09-16
- AI researchers predicted 2054 for a Millennium Problem; reality arrived far sooner — ben_j_todd · 2026-09-16
- Paper2Agent turns research papers into interactive AI agents (Nature) — EricTopol · 2026-09-16
- CPAL 2027 heads to Tokyo: small conference on parsimony and learning — YiMaTweets · 2026-09-16