RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier

Carles Gelada's team designs multi-effort RL rewards as R = S − λe·C, arguing geometrically that λe must set the iso-reward slope tangent to the Pareto curve to lift the entire cost-performance frontier.

2026-09-12 ~ 2026-09-12 · 3 related posts