Apple's Internalized Visual Thinking matches VisualCoT accuracy with ~5x faster video reasoning

机器之心 · wechat · 2026-09-03

Apple researchers propose Internalized Visual Thinking (IVT), a post-training framework for proactive video reasoning that learns to predict the future internally instead of generating future frames at inference time.

Motivation

Method

Results

Limits: no gains at longer hop=3 horizons; multi-branch futures remain hard to evaluate. IVT opens a middle ground between explicit visual chain-of-thought and pure text reasoning: think during training, answer directly at inference.

Original post →

More from Research

Research channel →