Multi-effort RL: set cost penalty to the Pareto curve slope to lift the whole frontier

carlesgelada · x · 2026-09-12

Carles Gelada details a multi-effort RL reward design aimed at pushing up the entire Pareto frontier: reward R = S − λe·C, where S is solve/fail, C the rollout cost, and λe an effort-specific penalty. The key result: the correct λe equals the slope of the base model's Pareto curve at each effort level, approximated once and held constant during RL. This works because λe sets the slope of iso-reward lines; only when those lines are tangent to the Pareto frontier does improving reward guarantee frontier improvement.

Related event: RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier(3 posts)→

Original post →

More from Research

Research channel →