PROWBench: 170 Programmable Episodes Test Whether Video World Models Render What the Program Specifies
AlayaLab · hf · 2026-10-02
PROWBench from AlayaLab evaluates visual adherence of programmable world models—the promising foundation for next-gen game engines—to fine-grained, program-specified world events, a dimension existing benchmarks rarely test. It comprises 170 programmatically constructed episodes and 600 proxy videos, logging entity states and timestamped events (including off-camera ones) as replayable world records from which synchronized views and proxy representations (coarse 3D, bounding boxes) are rendered. Generated videos are checked against the observable consequences of program execution via entity control, long-horizon memory, and two VLM-based metrics: Logic-Render Alignment and Interaction Success Rate.
More from Multimodal
- Image editing demo with Nano Banana — tkasasagi · 2026-10-02
- Creator: Opus 5.5 edits videos well, but forcing AI to clip without real need yields garbage — AlchainHust · 2026-10-02
- Reddit user's VEC concept car AI video shows startlingly realistic motion — Vashukanni · 2026-10-02
- Fable 5.5 rumored: webpage morphs its art style to match passing images — lxfater · 2026-10-02
- Oxford VGG Unveils SynCity 3000, Generating Globally Coherent Scene-Scale 3D Worlds — rsasaki0109 · 2026-10-02
- Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights — _reachsumit · 2026-10-02