Freeform preference learning lets robots take natural-language reward axes over binary feedback

chris_j_paxton · x · 2026-09-09

Specifying reward functions is one of the hardest parts of robot RL: binary success/failure labels and binary trajectory preferences are too sparse for long-horizon tasks.

Marcel Torne and collaborators propose freeform preference learning:

The team discussed the method on the RoboPapers podcast.

Original post →

More from Embodied

Embodied channel →