Does Imitation Learning Really Outperform RL?
srchvrs · x · 2026-07-15
This discussion questions: since everyone wants to add RL on top of imitation learning to address distribution shift, why do imitation learning agents still often perform better in public benchmarks?
The original reply's core argument:
- Many assume 'distribution shift' is a core problem that RL must solve;
- But on some open benchmarks, imitation learning agents like TFv5/6, DrivoR dominate;
- Therefore, it is worth re-examining whether the distribution shift problem is overemphasized.
Related event: Why imitation learning still tops RL on benchmarks(2 posts)→
More from Research
- Style-similarity analysis puts Kimi K3 closer to Claude Fable 5 than to K2.6 — soumitrashukla9 · 2026-07-21
- A GLP1R variant may explain stronger Ozempic weight loss, and the team built an agent workflow — julia_kiseleva · 2026-07-21
- Proceedings for the second geometry-grounded representation learning workshop are now online — erikjbekkers · 2026-07-21
- New survey maps how agentic systems are learning to improve themselves — SchmidhuberAI · 2026-07-21
- A curated TTS list for voice agents tracks latency, cancellation, and evals — mahimairaja · 2026-07-21
- Jacob Tsimerman interview frames LLMs as a turning point for mathematical discovery — stevenstrogatz · 2026-07-21