WorldReward: a vision-language reward model for evaluating camera-conditioned world models
Yibin Wang · hf · 2026-09-04
WorldReward is a vision-language reward model for evaluating camera-conditioned world models.
- Evaluation: aligns generated video chunks with actions and aggregates preferences along two dimensions — execution consistency and visual quality
- Significance: provides a quantifiable reward signal for world model quality, usable in the reward-modeling stage of world model training, relevant to embodied AI and video generation
More from Research
- FlashRender: Few-step camera-controlled generative rendering via MeanFlow distillation — everex · 2026-09-04
- LatentStream: progressive latent memory evolution for streaming video understanding — Hongyu Qu · 2026-09-04
- Temporal Context Routing aligns script timing in joint audio-video generation — Yichen Liu · 2026-09-04
- Puffin-World scales unified multimodal model with native 3D world states — Kang Liao · 2026-09-04
- Agent reliability degrades with trajectory length, but no benchmark isolates it, dev finds — rio_ARC · 2026-09-04
- The Bayes Bandit: A Mathematical Take on Curiosity in Reinforcement Learning — CatAstro_Piyush · 2026-09-04