Paper: When Does On-Policy Interaction Help in Value-Based Imitation Learning?
canondetortugas · x · 2026-08-12
This paper investigates the nature of performance improvements brought by expert interaction (like DAgger) in value-based imitation learning (IL).
Key findings:
- Relaxed Representational Demands: Expert interaction relaxes the representational requirements on the learner. The learner only needs a model capable of realizing the expert's value function, bypassing the stricter need to realize the expert's policy itself.
- OVI Algorithm: The authors introduce OVI, an interactive on-policy IL algorithm that is statistically and computationally efficient given access to a linear maximization oracle.
Related event: Study Reveals Key Role of Interaction in Value-Based Imitation Learning(2 posts)→
More from Research
- Ensemble-Conditioned Guidance Reframes Molecular Design Around Conformational Ensembles — _onionesque · 2026-09-23
- ICLR author proposes submission caps and exhaustive appendices to fight AI paper flood — algo_diver · 2026-09-23
- Programmable Si photonic circuit hits 29 fW static power per pi phase shift — jwt0625 · 2026-09-23
- Yoav Goldberg: some tasks just need deterministic rules — agents can write them — yoavgo · 2026-09-23
- Yoav Goldberg: shape predictor variables and decisions as a decision tree — yoavgo · 2026-09-23
- Yoav Goldberg: For Recurring Tasks, Tune Bespoke Predictors Instead of Always Using Reasoning LLMs — yoavgo · 2026-09-23