New preprint finds VLM OCR attention heads that verbalize far more than text, enabling a logit lens for image tokens

gsarti_ · x · 2026-09-29

A new preprint from Sheridan Feucht examines how images in VLMs align with words. The team identified a set of attention heads responsible for OCR, but surprisingly found these heads verbalize much more than just text in the image. They leveraged the discovery to build a simple logit lens for image tokens, offering a new interpretability window into multimodal representations.

Original post →

More from Research

Research channel →