TOPReward reads a VLM’s token probabilities and works with zero training
DJiafei · x · 2026-07-26
A robotics reward-model project called TOPReward claims to read a pretrained VLM’s internal belief directly from token probabilities.
- It needs zero training and no fine-tuning.
- It works on open-source models.
- The author says it beats previous methods by a wide margin and is already running in the real world.
- The work has also attracted follow-up research using the same reward model to refine policies.
The post is a reply thanking others for featuring the project and noting that more follow-up work is using their reward model to improve policies.
More from Embodied
- Vivix A1 accepts voice, text and images mid-interaction, demo says — Scobleizer · 2026-07-26
- Vivix A1 demo shows synchronized speech, gaze, facial motion and body movement — Scobleizer · 2026-07-26
- Travis Kalanick weighs humanoid robots against specialized machines — MarwaEldiwiny · 2026-07-26
- Reservoir Farms says 20+ startups are now paying to test agtech on its farms — DynamicWebPaige · 2026-07-26
- Nomagic is hiring senior robotics and engineering roles for its physical AI team — m_wulfmeier · 2026-07-26
- A robotic elephant-trunk gripper uses an internal camera to sense touch — rvp · 2026-07-26