COLM paper asks whether VLMs can internalize tool calls in latent space instead of calling them

PMinervini · x · 2026-10-07

A study being presented at COLM explores whether vision-language models, which benefit from tool calls enabling "thinking with images," can internalize the effects of calling the correct tools directly in latent space — eliminating the need for explicit tool calls at inference time.

The author shared a thread with details ahead of the 11 AM COLM talk, positioning the work in the emerging area of VLM latent reasoning.

Original post →

More from Research

Research channel →