EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026
jiqizhixin · x · 2026-09-11
Harbin Institute of Technology and Zhongguancun College present EASE (Evidence-Anchored Spatial Attention), accepted to EMNLP 2026 Findings. Outcome-only reward in multimodal RLVR raises accuracy but leaves a systematic mismatch between where the model looks and where the evidence is. EASE adds a process-level signal: an annotation pipeline produces bounding boxes for answer-supporting regions (training metadata only), each converted to a 2D Gaussian (2σ covering half the box), equal-weight mixing of multiple boxes, and 10% background smoothing. Gains: +2.9 on Qwen2.5-VL-7B, +3.1 on Qwen3-VL-4B, +2.5 on Qwen3-VL-8B over the DAPO baseline.
More from Research
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11
- MutexaGPT: LLM agents plus MD simulations hit 40% on enzyme design, 4x the baseline — bravo_abad · 2026-09-11