RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier
Carles Gelada's team designs multi-effort RL rewards as R = S − λe·C, arguing geometrically that λe must set the iso-reward slope tangent to the Pareto curve to lift the entire cost-performance frontier.
2026-09-12 ~ 2026-09-12 · 3 related posts
- Multi-effort RL reward design: R = S − λ_e·C to push up the whole Pareto frontier — carlesgelada · 2026-09-12
- Multi-effort RL: set cost penalty to the Pareto curve slope to lift the whole frontier — carlesgelada · 2026-09-12
- Why the λ_e penalty must tangent the Pareto curve to lift the whole frontier in RL — carlesgelada · 2026-09-12