MoCA Decouples Perception and Reasoning in Training
VectorInst · x · 2026-07-09
Wenhu Chen's spotlight paper MoCA decouples perception and reasoning during training, resulting in a 7B model. The post notes that the model shows improvements on both perception and reasoning benchmarks, outperforming GPT-4o in several tests.
Related event: MoCA Paper Decouples Perception and Reasoning in VLMs(2 posts)→
More from Multimodal
- NVIDIA ships Nemotron audio-native open weights in 2B and 30B sizes — victormustar · 2026-07-21
- HOMIE pairs Qwen3-VL-2B with Wan2.1 for human-object-centric video personalization — switch2stock · 2026-07-21
- Video Models Cut Ad Production Costs by 90-99%: Runway Enterprise Data — c_valenzuelab · 2026-07-21
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21