Paper proposes Internalized Visual Thinking for efficient video reasoning

koustuvsinha · x · 2026-08-19

The paper "Beyond Visual CoT" introduces Internalized Visual Thinking (IVT), a method for proactive video reasoning. Unlike Visual CoT which generates intermediate images, IVT trains models to predict latent embeddings of future frames and removes the image generation pathway at inference. Results show IVT outperforms Visual CoT in 4 of 6 settings while reducing inference latency from 6.56s to 1.22s, cutting costs by over 5x.

Original post →

More from Research

Research channel →