PAWBench Results: Even Told the Target Outcome, Video Models Hit It Only 38–58% of the Time

RisingSayak · x · 2026-08-28

Key results from the PAWBench thread: the authors tried three fixes—clearer outcome instructions, more diverse sampling, and additional training—each helping a different part of the problem, none solving it end to end. Even when given the desired outcome, models produced it only 38–58% of the time. Models also react to the wrong cues: changing actual physics moves predictions too little, while irrelevant visual or text cues move them too much.

Related event: PAWBench Reveals Video Generators Fail Probabilistic Alignment Tests(5 posts)→

Original post →

More from Research

Research channel →