EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026

jiqizhixin · x · 2026-09-11

Harbin Institute of Technology and Zhongguancun College present EASE (Evidence-Anchored Spatial Attention), accepted to EMNLP 2026 Findings. Outcome-only reward in multimodal RLVR raises accuracy but leaves a systematic mismatch between where the model looks and where the evidence is. EASE adds a process-level signal: an annotation pipeline produces bounding boxes for answer-supporting regions (training metadata only), each converted to a 2D Gaussian (2σ covering half the box), equal-weight mixing of multiple boxes, and 10% background smoothing. Gains: +2.9 on Qwen2.5-VL-7B, +3.1 on Qwen3-VL-4B, +2.5 on Qwen3-VL-8B over the DAPO baseline.

Original post →

More from Research

Research channel →