TOPReward reads a VLM’s token probabilities and works with zero training

DJiafei · x · 2026-07-26

A robotics reward-model project called TOPReward claims to read a pretrained VLM’s internal belief directly from token probabilities.

The post is a reply thanking others for featuring the project and noting that more follow-up work is using their reward model to improve policies.

Original post →

More from Embodied

Embodied channel →