MoCA Paper Decouples Perception and Reasoning in VLMs

The spotlight paper MoCA proposes decoupling perception and reasoning in VLM training. The resulting 7B model achieves significant benchmark improvements, addressing distinct failure modes where VLMs either fail to 'see' or fail to 'reason'.

2026-07-09 ~ 2026-07-09 · 2 related posts