Apple's Internalized Visual Thinking Drops the Paint-the-Future Pipeline for ~5x Faster Video Reasoning

jiqizhixin · x · 2026-09-12

Apple proposes Internalized Visual Thinking (IVT). Visual chain-of-thought generates future frames then decodes/re-encodes them, adding latency fatal for proactive video reasoning. IVT instead trains the model to learn both the text answer and latent-space representations of future frames, then removes the entire future-visual-prediction pipeline at inference — the model answers directly from the current video. It improves results across 6 dataset-task settings on Ego-Exo4D and Ego4D while reasoning 5x faster.

Original post →

More from Multimodal

Multimodal channel →