VisionCoach: RL framework rewards correct visual attention for grounded video reasoning, SOTA zero-shot

mohitban47 · x · 2026-09-10

VisionCoach, presented at ECCV 2026 (Poster #148, Sep 10), is an RL framework for grounded video reasoning. Video reasoning models often miss where to look or lean on language priors, so instead of supervising only final answers, VisionCoach rewards correct visual attention, using dynamic visual prompting as a training-time coach for spatio-temporal grounding. Self-distillation keeps inference simple and tool-free. It achieves state-of-the-art zero-shot results across video reasoning, understanding, and temporal grounding benchmarks such as V-STAR.

Original post →

More from Research

Research channel →