PAWBench Results: Even Told the Target Outcome, Video Models Hit It Only 38–58% of the Time
RisingSayak · x · 2026-08-28
Key results from the PAWBench thread: the authors tried three fixes—clearer outcome instructions, more diverse sampling, and additional training—each helping a different part of the problem, none solving it end to end. Even when given the desired outcome, models produced it only 38–58% of the time. Models also react to the wrong cues: changing actual physics moves predictions too little, while irrelevant visual or text cues move them too much.
Related event: PAWBench Reveals Video Generators Fail Probabilistic Alignment Tests(5 posts)→
More from Research
- Google DeepMind brings AI Co-Scientist from simulation to real-world experiments — omarsar0 · 2026-08-28
- Shunyu Yao: RL finally works, marking the start of AI's 'Second Half' — dejavucoder · 2026-08-28
- Zero-WAM: robots learn unseen tasks from a single human demo video, no fine-tuning — deepakpathak · 2026-08-28
- LeVJEPA: Video Pretraining Without Complex Components — randall_balestr · 2026-08-28
- Study: Task Diversity Inhibits Continual Reinforcement Learning — RubenEVillegas · 2026-08-28
- AI reconstruction reveals plasma jet physics in galaxy — bravo_abad · 2026-08-28