Symbolic supervision gives video models equal reasoning at ~1/30 the compute, Rubik's Cube study finds
niloofar_mire · x · 2026-10-03
- A new arXiv paper tests hidden-state reasoning in video generation models with a 2x2x2 Rubik's Cube benchmark: given a fixed view of three faces, models must predict the sticker configuration after nine prescribed moves.
- Key finding 1: validation MSE follows an approximate power law, but lower loss does not reliably indicate downstream reasoning. A 70M autoregressive model hits 44.6% frame accuracy at 0.1 PF-days, while a 1B model sits at 0.3% and only reaches 83.7% at 3.14 PF-days.
- Key finding 2: adding symbolic state supervision lifts a 20M model's frame accuracy from 31.1% to 67.3% at the same data budget — matching video-only scaling with roughly 1/30 the compute. Learning scalable, generalizable symbolic representations may be a crucial complement to pure scaling.
More from Multimodal
- GemPix 2.5 Flash string spotted in Gemini iOS update, hinting at Nano Banana 2.5 Flash — lyraxana · 2026-10-03
- PotionUI 0.0.14 Drops the GPU Requirement: Cloud Models via OpenRouter, Built-in Editor, Video Director — 0roborus_ · 2026-10-03
- Researcher trains a diffusion model from scratch in two weeks — and finds its art beautiful — pbaylies · 2026-10-03
- Static: a sci-fi fake trailer made entirely with Kling — SightsFilms · 2026-10-03
- Redditor Releases SITCOM, an AI-Generated Psychological Horror Short Film — Sufficient_Flow_415 · 2026-10-03
- Glitched Gotham: a Midjourney --sref code for glitch aesthetics — egeberkina · 2026-10-03