Gemma 4 12B Image Patch Predictions

matthen2 · x · 2026-07-19

This image visualizes the "next token prediction" at the image patch level for Gemma 4 12B. Different image regions correspond to various high-frequency predicted words, indicating the model's distinct local preferences for scene, object, material, weather, and orientation information.

By overlaying these predicted words directly onto the original image, it's evident that the model generates specific semantic associations for areas like buildings, roads, vehicles, and people. Overall, it serves as an observational diagram of a multimodal model's internal representations rather than a standard generation result.

Related event: Visualizing Image Patch Predictions in Gemma 4(4 posts)→

Original post →

More from Multimodal

Multimodal channel →