Freeform Preferences: Multi-Dimensional Robot Supervision

chelseabfinn · x · 2026-07-03

The author introduces "freeform preferences": supervisors first define relevant evaluation dimensions and then provide preferences along them. Dimensions can be fixed scoring criteria or natural language descriptions. This approach eliminates ambiguity, comprehensively covers all aspects, and provides denser supervision signals (method diagram included). The post serves as the core method explanation for their reward model research thread.

Original post →

More from Research

Research channel →