What Gemma 4 Predicts When Watching Videos
matthen2 · x · 2026-07-19
The author asks: What exactly is a multimodal LLM "thinking" about when watching a video?
The post mentions that Gemma 4 12B directly reads raw image patches and processes them like tokens. Although it hasn't been trained to make predictions on these "patch tokens," the video demonstrates what happens if you forcibly sample from its next-token head to see the predicted next step. This post is more of an exploration of the internal representations and prediction mechanisms of multimodal models rather than a standard product introduction.
Related event: Visualizing Image Patch Predictions in Gemma 4(4 posts)→
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11