What Gemma 4 Predicts When Watching Videos
matthen2 · x · 2026-07-19
The author asks: What exactly is a multimodal LLM "thinking" about when watching a video? The post mentions that Gemma 4 12B directly reads raw image patches and processes them like tokens. Although it hasn't been trained to make predictions on these "patch tokens," the video demonstrates what happens if you forcibly sample from its next-token head to see the predicted next step. This post is more of an exploration of the internal representations and prediction mechanisms of multimodal models rather than a standard product introduction.
Related event: Visualizing Image Patch Predictions in Gemma 4(3 posts)→
More from Multimodal
- Creator says he directed a full short film with Adobe Firefly AI Assistant — LudovicCreator · 2026-07-21
- Reddit users say Krea 2 Turbo regains strong facial expressions with bypass LoRAs — YentaMagenta · 2026-07-21
- AI music demo blends Suno v5.5, Reason Studios and Grok Imagine 1.5 — Kyrannio · 2026-07-21
- AI anime workflow article breaks storytelling into repeatable prompt steps — Aiden_Tech_Ai · 2026-07-21
- Claude is being pitched as a free workflow for viral YouTube Shorts scripts — Aiden_Tech_Ai · 2026-07-21
- A Bittensor game demo claims a two-person team built a playable 3D world in 30 days — markjeffrey · 2026-07-21