Tencent BAC’s ProLaViT pushes multimodal reasoning into a progressive latent chain

机器之心 · wechat · 2026-07-21

What ProLaViT changes

Tencent BAC proposes ProLaViT (Progressive Latent Visual Thought) for multimodal reasoning tasks where models can see an image but still struggle to reason correctly over space, geometry, and logic.

Results

Built on Qwen2.5-VL-7B-Instruct, the full method reports a 75.11% average accuracy across MMVP, VisPuzzle, VStar, ChartQA, BLINK, and CV-Bench.

Overall, the work argues that structured latent reasoning can bridge the gap between expensive explicit image generation and unstable implicit one-step inference.

Original post →

More from Multimodal

Multimodal channel →