What Gemma 4 Predicts When Watching Videos
matthen2 · x · 2026-07-19
The author asks: What exactly is a multimodal LLM "thinking" about when watching a video?
The post mentions that Gemma 4 12B directly reads raw image patches and processes them like tokens. Although it hasn't been trained to make predictions on these "patch tokens," the video demonstrates what happens if you forcibly sample from its next-token head to see the predicted next step. This post is more of an exploration of the internal representations and prediction mechanisms of multimodal models rather than a standard product introduction.
Related event: Visualizing Image Patch Predictions in Gemma 4(4 posts)→
More from Multimodal
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11