NeurIPS 2026 paper AVIC: RL policy learns when and how much to imagine

mohitban47 · x · 2026-09-30

Paper AVIC has been accepted to NeurIPS 2026 (the AC metareview strongly recommended it as a Spotlight). Its core finding: for spatial reasoning, more imagination is not always better — the key is knowing when to imagine and how much.

AVIC learns an RL policy that adaptively invokes a video world model and controls how much to imagine, balancing answer accuracy against imagination cost for more accurate and efficient reasoning.

Original post →

More from Models

Models channel →