Visualizing Gemma 4 Image Patch Predictions

matthen2 · x · 2026-07-19

The image displays image patch predictions made by Gemma 4 12B on a scene from the National Galleries of Scotland in Edinburgh. The model generates the most salient next-word predictions for different regions of the image, mapping elements like buildings, signs, people, and skies to specific vocabulary.

This visualization essentially demonstrates how the model "deconstructs" an image to make local semantic associations, rather than just performing whole-image classification. The author aims to highlight that predicting at the raw image patch level reveals how foundation models understand scene details.

Related event: Visualizing Image Patch Predictions in Gemma 4(4 posts)→

Original post →

More from Multimodal

Multimodal channel →