VisionCoach: RL framework rewards correct visual attention for grounded video reasoning, SOTA zero-shot
mohitban47 · x · 2026-09-10
VisionCoach, presented at ECCV 2026 (Poster #148, Sep 10), is an RL framework for grounded video reasoning. Video reasoning models often miss where to look or lean on language priors, so instead of supervising only final answers, VisionCoach rewards correct visual attention, using dynamic visual prompting as a training-time coach for spatio-temporal grounding. Self-distillation keeps inference simple and tool-free. It achieves state-of-the-art zero-shot results across video reasoning, understanding, and temporal grounding benchmarks such as V-STAR.
More from Research
- CoopEval: a framework for comparing cooperation mechanisms in multi-agent systems — conitzer · 2026-09-10
- Open Yap 1K: 1,000 hours of natural two-speaker conversations, free for commercial use — realmrfakename · 2026-09-10
- Hank Yang: AI Excels at Well-Defined Problems, So the Real Skill Is Defining New Ones — hankyang94 · 2026-09-10
- Jacobian conjecture drama: Anthropic's Alpoge responds to leaked BGV paper concerns — suchenzang · 2026-09-10
- NNsight 0.8 pre-release ships faster engine, MoE and near-native vLLM support — davidbau · 2026-09-10
- ECCV talk outlines three pillars for embodied AI: motion prediction, evidence, streaming — CSProfKGD · 2026-09-10