PROWBench: 170 Programmable Episodes Test Whether Video World Models Render What the Program Specifies

AlayaLab · hf · 2026-10-02

PROWBench from AlayaLab evaluates visual adherence of programmable world models—the promising foundation for next-gen game engines—to fine-grained, program-specified world events, a dimension existing benchmarks rarely test. It comprises 170 programmatically constructed episodes and 600 proxy videos, logging entity states and timestamped events (including off-camera ones) as replayable world records from which synchronized views and proxy representations (coarse 3D, bounding boxes) are rendered. Generated videos are checked against the observable consequences of program execution via entity control, long-horizon memory, and two VLM-based metrics: Logic-Render Alignment and Interaction Success Rate.

Original post →

More from Multimodal

Multimodal channel →