New preprint finds VLM OCR attention heads that verbalize far more than text, enabling a logit lens for image tokens
gsarti_ · x · 2026-09-29
A new preprint from Sheridan Feucht examines how images in VLMs align with words. The team identified a set of attention heads responsible for OCR, but surprisingly found these heads verbalize much more than just text in the image. They leveraged the discovery to build a simple logit lens for image tokens, offering a new interpretability window into multimodal representations.
More from Research
- AI agents claim Collatz breakthrough: positive proportion of numbers reach 1, Lean-formalized — AlexKontorovich · 2026-09-29
- Averaging OCR attention heads across layers accidentally yields a strong interpretability lens — benno_krojer · 2026-09-29
- Self-supervised multisensory pretraining robot RL paper wins best student paper at IROS — GeorgiaChal · 2026-09-29
- saezlab open-sources mc-ASTRA, a Python framework for studying tissue remodeling across patients and perturbations — anshulkundaje · 2026-09-29
- Sampling the SK Model at β<1: New Polynomial-Time Algorithm Closes a 2022 Open Problem — canondetortugas · 2026-09-29
- Testing Jev on Reasoning-Intensive Regression: Wicked Fast but Classification-Focused — dbreunig · 2026-09-29