Foresight lets streaming VLMs plan future perception without retraining, beating baselines by 9.5%
Ashok Prasad Neupane · hf · 2026-10-07
Foresight is a training-free method that lets streaming vision-language models dynamically plan future computation.
- Key insight: streaming VLMs inherently anticipate the immediate future, which can guide computation without retraining.
- A dual-stream Siamese LLM architecture (shared weights, encoders, KV cache) runs one LLM on the incoming stream while a second anticipates future context and plans when to reason, what to check, and how densely to sample.
- Online reconfiguration uses schema-guided decoding and lightweight diff updates with minimal overhead.
- With a frozen Qwen3-VL-8B backbone, it hits 23.0 mean joint F1 on OmniPro Online (beating the strongest trained baseline by 9.5%), plus +6.7 on StreamingBench and +15.4 on OVO-Bench. Code will be open-sourced.
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07